Methods and apparatus for deep learning-based audio object extraction
The DNN-based audio object estimation method efficiently extracts multiple audio signals from noisy mixtures, addressing the limitations of existing speech-focused techniques by enhancing audio experiences in UGC and short video content.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-24
- Publication Date
- 2026-04-02
AI Technical Summary
Existing deep learning-based speech enhancement techniques are limited to processing speech signals and fail to effectively handle diverse audio content in user-generated content (UGC), which often includes various audio signals such as music, ambient sounds, and sound effects, necessitating advanced methods for improved audio object extraction.
A method and apparatus using a deep neural network (DNN) for audio object estimation that processes time-frequency representations of audio mixtures to extract multiple audio signals, employing two-stage mask estimation modules and a hybrid loss function for training, enabling efficient and flexible extraction of audio objects like speech and music.
Enables effective extraction of multiple audio objects from noisy mixtures, enhancing the audio experience in UGC and short video content scenarios by improving the handling of diverse audio sources beyond speech.
Smart Images

Figure US2025047728_02042026_PF_FP_ABST
Abstract
Description
[0001] PCT / US25 / 47728 24 September 2025 (24.09.2025)
[0002] D24140W001
[0003] METHODS AND APPARATUS FOR DEEP LEARNING-BASED AUDIO OBJECT EXTRACTION
[0004] CROSS-REFERENCE TO RELATED APPLICATIONS
[0005] 5 This application claims the benefit of priority from International Patent Application No. PCT / CN2024 / 121955, filed on 27 September 2024, and United States Provisional Patent Application No. 63 / 809,979, filed on 21 May 2025, each of which is incorporated herein in their entirety.
[0006] TECHNICAL FIELD
[0007] The present disclosure generally relates to the technical field of signal processing, and more particularly to deep learning-based audio object extraction (or estimation) methods and apparatus.
[0008] 15 BACKGROUND
[0009] In a broad sense, deep learning-based speech enhancement related techniques may be generally designed to remove unwanted artifacts such as background noise, reverberation, and other distortions, for example with the potential objective of extracting clean speech from a mixed audio signal. These techniques have shown promising performance in isolating speech
[0010] 20 content; however, they typically tend to be tailored specifically for speech signals, but may usually not account for other types of audio signals, such as music.
[0011] In parallel with the rapid growth of the short video domain, user generated content (UGC) has become increasingly popular across various platforms. Generally speaking, such UGC may commonly contain a wide range of audio signals beyond speech, including for example music, ambient sounds, and sound effects. As a result, the diversity of audio content in UGC may be generally understood to pose new challenges for audio enhancement systems that are limited to speech-specific processing.
[0012] To provide a more immersive and engaging audio experience in modem multimedia applications, e.g., in the context of UGC or the like, there appears to be a need for advanced
[0013] 30 audio enhancement techniques that are generally capable of handling and improving multiple types of audio sources, or in other words, not limited to speech alone. PCT / US25 / 47728 24 September 2025 (24.09.2025)
[0014] D24140W001
[0015] SUMMARY
[0016] In view of the above, the present disclosure provides methods, apparatus, and programs, as well as computer-readable storage media for audio object estimation (or extraction), having
[0017] 5 the features of the respective independent claims. The dependent claims relate to preferred embodiments.
[0018] According to a first aspect of the present disclosure, there is provided a method for audio object estimation (or, in some possible cases, also audio object extraction). For instance, the provided method may be used to estimate (or extract) one or more audio objects (e.g., speech,
[0019] 10 music, etc.) from an audio mixture. In particular, the method may comprise obtaining a timefrequency representation (in a time-frequency domain) of an audio signal that comprises a mixture of at least a first signal and a second signal of different types. The method may further comprise obtaining (or determining) a set of bin features from the time-frequency representation, wherein each bin feature in the set may correspond to a respective time¬
[0020] 15 frequency bin of the time-frequency representation. The bin features may be obtained (or determined) from the time-frequency representation by using any suitable feature extraction technique, for example. The method may further comprise feeding the set of bin features to a (e.g., pre-trained) neural network model to estimate (or determine) a first mask and a second mask respectively corresponding to the first signal and the second signal. In other words, the
[0021] 20 neural network model may be used for determining a first mask that is associated with the first signal; and also for determining a second mask that is associated with the second signal. Accordingly, the method may further comprise respectively applying the first mask and the second mask to the time-frequency representation, in order to obtain an estimated (or extracted) time-frequency representation of the first signal and an estimated (or extracted) time-frequency representation of the second signal. It may be worth noting that, as will also be understood and appreciated by the skilled person, the first / second audio signal (or the respective estimation / extract! on thereof) and / or the first / second mask should not be understood to constitute a limitation of any kind. For instance, with suitable adaption or the like, the provided method may be extended so as to be enabled to support estimation or
[0022] 30 extraction of more than two audio signals (objects) from an input audio mixture.
[0023] Configured as proposed, broadly speaking, this provided method generally enables supporting of estimation (or extraction) of more than one audio signal (object / content) from noisy audio mixtures in an efficient yet flexible manner, which in turn helps improve audio experience in PCT / US25 / 47728 24 September 2025 (24.09.2025)
[0024] D24140W001 the overall audio application system, e.g., in the context of UGC and / or short video content scenarios, or the like.
[0025] In some example embodiments, the first signal may be a speech signal and the second signal may be a music signal. Of course, any other suitable audio signal, for instance of any suitable
[0026] 5 type, may be applied as well. For instance, in some possible cases, it may also be possible to further split (categorize) the speech (or music, or the like) into, for instance, speech from male / female / child speakers, or the like. Therefore, this should not be understood to constitute a limitation of any kind.
[0027] In some example embodiments, the first mask may comprise values (e.g., between 0 and 1)
[0028] 10 indicative of ratios of the first signal in the audio signal for respective time-frequency bins, and the second mask may comprise values (e.g., between 0 and 1) indicative of ratios of the second signal in the audio signal for respective time-frequency bins. For instance, in some possible examples, a value of zero (or close to zero) for the first / second mask may be understood to indicate no (or close to none) first / second signal being present in the respective
[0029] 15 time-frequency bin. Accordingly, in some possible examples, applying the first / second mask to the time-frequency representation thereby obtaining the estimated time-frequency representation of the first / second signal may involve multiplying, for each time-frequency bin, the respective value of the first / second mask with the time-frequency representation. Of course, as will be understood and appreciated by the skilled person, the first / second mask may
[0030] 20 be implemented by using any other suitable representation as well.
[0031] In some example embodiments, the method may further comprise: obtaining the audio signal and transforming the audio signal by using a transform function, to thereby obtain the timefrequency representation of the audio signal. The audio signal may be obtained by using any suitable means, for example from a recording device, from a file, or the like, depending on various implementations.
[0032] In some example embodiments, the method may further comprise: applying a corresponding inverse transform function (that is associated with the above transform function) to the estimated time-frequency representation of the first signal, to thereby obtain an estimated (extracted) first signal from the audio signal; and similarly, applying the inverse transform
[0033] 30 function to the estimated time-frequency representation of the second signal, to thereby obtain an estimated (extracted) second signal from the audio signal.
[0034] In some example embodiments, the transform function may comprise one of: a short time Fourier transform (STFT), a modified discrete X transform (MDXT), or a filterbank based PCT / US25 / 47728 24 September 2025 (24.09.2025)
[0035] D24140W001 transform such as complex quadrature mirror filterbank (CQMF). Of course, as will be understood and appreciated by the skilled person, any other suitable transform function and correspondingly also the associated inverse transform function may be applied as well, depending on various implementations.
[0036] 5 In some example embodiments, the neural network model is a deep neural network (DNN) based model. Of course, as will be understood and appreciated by the skilled person, any other suitable neural network model (or deep learning based model) may be applied as well, depending on various implementations.
[0037] In some example embodiments, the DNN based model may comprise a first two-stage mask estimation module and a second two-stage mask estimation module configured for respectively estimating the first mask and the second mask.
[0038] In some example embodiments, at least one (e.g., any one or both) of the first and second two- stage mask estimation modules may be configured for estimating, at a first stage, a magnitude mask, and at a second stage, a complex mask. Accordingly, in some possible cases, both the
[0039] 15 first and second two-stage mask estimation modules may be used, each being configured for estimating, at the (respective) first stage, a respective magnitude mask, and at the (respective) second stage, a respective complex mask.
[0040] In some example embodiments, applying the first mask may involve: applying, at the first stage, the magnitude mask to the time-frequency representation to obtain a modified time¬
[0041] 20 frequency representation; and applying, at the second stage, the complex mask to the modified time-frequency representation, in order to obtain the estimated time-frequency representation of the first signal.
[0042] In some example embodiments, applying the magnitude mask to the time-frequency representation to obtain the modified time-frequency representation may comprise: applying the magnitude mask to magnitude values of the time-frequency representation to obtain a modified magnitude spectrum; and obtaining the modified time-frequency representation having a modified complex-valued spectrum based on the modified magnitude spectrum and a phase spectrum corresponding to the time-frequency representation. In other words, the modified time-frequency representation having the modified complex-valued spectrum may
[0043] 30 be obtained based on the modified magnitude spectrum of the original phase (spectrum) of the input audio signal. PCT / US25 / 47728 24 September 2025 (24.09.2025)
[0044] D24140W001
[0045] In some example embodiments, applying the magnitude mask and / or the complex mask may comprise deep filtering that involves weighted averaging among neighboring time-frequency bins.
[0046] In some example embodiments, the deep filtering used for applying the magnitude mask at the
[0047] 5 first stage may be implemented according to: f) denotes the modified magnitude spectrum, MS1(t, f) denotes the magnitude mask, |f (t, / ) | denotes the magnitude spectrum corresponding to the time-frequency representation, t and f represent time and frequency indices, and U , U2, V and V2are predetermined values indicative of a to-be-estimated size of the magnitude mask. Of course, as will be understood and appreciated by the skilled person, the deep filtering used for applying the magnitude mask may be implemented in any other suitable means as well.
[0048] In some example embodiments, the modified complex-valued spectrum f51(t, ) may be obtained according to: where operator <p is used for calculating an argument of a complex number.
[0049] In some example embodiments, the deep filtering used for applying the complex mask at the second stage may be implemented according to:
[0050] 20 where Ys'2(t, f) denotes the estimated time-frequency representation of the first signal, M52(t, ) denotes the complex mask, and M , M2, and N2are predetermined values indicative of a to-be-estimated size of the complex mask. Of course, as will be understood and appreciated by the skilled person, the deep filtering used for applying the complex mask may be implemented in any other suitable means as well.
[0051] In some example embodiments, at least one (e.g., any one or both) of the first and second two- stage mask estimation modules may compnse: a first activation layer configured to, at the first stage, limit the magnitude mask to a value range of 0 to 1 (or any other suitable value range); and a second activation layer configured to, at the second stage, limit the complex mask to a value range of -1 to 1 (or any other suitable value range). In some possible examples, the first
[0052] 30 may be implemented as a sigmoid function and / or the second activation layers may be implemented as a tanh function, or the like, depending on various circumstances. PCT / US25 / 47728 24 September 2025 (24.09.2025)
[0053] D24140W001
[0054] In some example embodiments, the DNN based model may comprise a processing chain comprising, in this order, a feature extraction module, followed by an encoder module, followed by a decoder module, and followed by the first and second two-stage mask estimation modules. This is merely one possible example implementation and should not be
[0055] 5 understood to constitute a limitation of any kind. Of course, as will be understood and appreciated by the skilled person, any other suitable processing chain, or more generally, the DNN based model, may be implemented as well, depending on various circumstances.
[0056] In some example embodiments, the neural network may have been trained based on at least one training pair. In particular, each training pair may comprise, on the one hand, a first reference signal and a second reference signal as ground truth, and on the other hand, a corresponding audio mixture signal. The audio mixture signal (or sometimes also referred to as a noisy mixture) may comprise, among others (e.g., background noises, etc.), at least a first signal (corresponding to the first reference signal) and a second signal (corresponding to the second reference signal) of different types.
[0057] 15 In some example embodiments, the neural network has been trained based on a loss function. As will be described in great detail below, the loss function may also be referred to as a hybrid loss function in the sense that more than one component has been taken into account when designing such (hybrid) loss function.
[0058] According to a second aspect of the disclosure, there is provided a method of training a neural
[0059] 20 network for use in audio object estimation (or, in some possible cases, also audio object extraction). In particular, the method may comprise obtaining an audio mixture signal, and corresponding first and second reference signals of different types that are used as ground truth. In other words, for the purpose of training, one or more training pairs may be first obtained, each of which may comprise, on the one hand, an audio mixture signal, and on the other hand, a corresponding first reference signal and a corresponding second reference signal that are used as ground truth. The audio mixture signal (or sometimes also referred to as noisy mixture) may comprise, among others (e.g., background noises, etc.) at least a first signal (corresponding to the first reference signal) and a second signal (corresponding to the second reference signal) of different types. The method may further comprise obtaining (e.g., by
[0060] 30 using any suitable feature extraction means, or the like) a first set of bin features from the audio mixture signal, a second set of bin features from the first reference signal, and a third set of bin features from the second reference signal, wherein each bin feature in the first to third sets corresponds to a respective time-frequency bin of the audio mixture signal, the first PCT / US25 / 47728 24 September 2025 (24.09.2025)
[0061] D24140W001 reference signal and the second reference signal. The method may further comprise feeding the first set of bin features to the neural network to estimate a first mask and a second mask respectively corresponding to the first reference signal, and the second reference signal. In other words, the neural network model may be used for determining, based on the first set of
[0062] 5 bin features, the first mask that is associated with the first signal, and also the second mask that is associated with the second signal. The method may yet further comprise obtaining a fourth set of bin features based on the first mask and the first set of bin features and also obtaining a fifth set of bin features based on the second mask and the first set of bin features. Finally, the method may comprise evaluating a loss function based on the first to fifth sets of
[0063] 10 bin features; and updating parameters (e.g., weights, etc.) of the neural network based on a result of the evaluation.
[0064] Configured as proposed, broadly speaking, this provided method enables efficient, flexible and yet reliable training for supporting estimation (or extraction) of more than one audio signal (object / content) from noisy audio mixtures, which in turn helps improve audio experience in the overall audio application system, e.g., in the context of UGC and / or short video content scenarios, or the like.
[0065] In some example embodiments, the first reference signal may be a reference speech signal and the second reference signal may be a reference music signal. Of course, any other suitable audio signal, for instance of any suitable type, may be applied as well.
[0066] 20 In some example embodiments, obtaining the first set of bin features from the audio mixture signal may comprise: transforming the audio mixture signal by using a transform function to obtain a transformed audio signal; and extracting a respective bin feature from the transformed audio signal for each time-frequency bin to obtain the first set of bin features. Similarly, obtaining the second set of bin features from the first reference signal may comprise: transforming the first reference signal by using the transform function to obtain a transformed first reference signal; and extracting a respective bin feature from the transformed first reference signal for each time-frequency bin to obtain the second set of bin features. Further, obtaining the third set of bin features from the second reference signal may comprise: transforming the second reference signal by using the transform function to obtain a
[0067] 30 transformed second reference signal; and extracting a respective bin feature from the transformed second reference signal for each time-frequency bin to obtain the third set of bin features. PCT / US25 / 47728 24 September 2025 (24.09.2025)
[0068] D24140W001
[0069] In some example embodiments, the transform function may comprise one of: a short time Fourier transform (STFT), a modified discrete X transform (MDXT), or a filterbank based transform such as complex quadrature mirror filterbank (CQMF). Of course, as will be understood and appreciated by the skilled person, any other suitable transform function may¬
[0070] 5 be applied as well, depending on various implementations.
[0071] In some example embodiments, the first mask may comprise values (e.g., between 0 and 1) indicative of ratios of the first signal in the audio signal for respective time-frequency bins, and the second mask may comprise values (e.g., between 0 and 1) indicative of ratios of the second signal in the audio signal for respective time-frequency bins. For instance, in some possible examples, a value of zero or close to zero for the first / second mask may be understood to indicate no (or close to none) first / second signal being present in the respective time-frequency bin. Of course, as will be understood and appreciated by the skilled person, the first / second mask may be implemented by using any other suitable representation as well.
[0072] In some example embodiments, the neural network model is a deep neural network (DNN)
[0073] 15 based model. Of course, as will be understood and appreciated by the skilled person, any other suitable neural network model (or deep learning based model) may be applied as well, depending on various implementations.
[0074] In some example embodiments, the DNN based model may comprise a first two-stage mask estimation module and a second two-stage mask estimation module configured for
[0075] 20 respectively estimating the first mask and the second mask.
[0076] In some example embodiments, at least one (e.g., any one or both) of the first and second two- stage mask estimation modules may be configured for estimating, at a first stage, a magnitude mask, and at a second stage, a complex mask. Accordingly, in some possible cases, both the first and second two-stage mask estimation modules are each configured for estimating, at the (respective) first stage, a respective magnitude mask, and at the (respective) second stage, a respective complex mask.
[0077] In some example embodiments, obtaining the fourth set of bin features based on the first mask and the first set of bin features may involve: applying, at the first stage, the magnitude mask to the first set of bin features to obtain a modified first set of bin features; and applying, at the
[0078] 30 second stage, the complex mask to the modified first set of bin features, to obtain the fourth set of bin features. PCT / US25 / 47728 24 September 2025 (24.09.2025)
[0079] D24140W001
[0080] In some example embodiments, applying the magnitude mask to the first set of bin features to obtain the modified first set of bin features may comprise: applying the magnitude mask to magnitude values of the first set of bin features to obtain a modified magnitude spectrum; and obtaining the modified first set of bin features having a modified complex-valued spectrum
[0081] 5 based on the modified magnitude spectrum and a phase spectrum corresponding to the first set of bin features. In other words, the modified first set of bin features having the modified complex-valued spectrum may be obtained based on the modified magnitude spectrum of the original phase (spectrum) of the (input) audio mixture signal.
[0082] In some example embodiments, applying the magnitude mask and / or the complex mask may comprise deep filtering that involves weighted averaging among neighboring time-frequency bins.
[0083] In some example embodiments, the deep filtering used for applying the magnitude mask at the first stage may be implemented according to:
[0084] 15 where Y^' (t, f) denotes the modified magnitude spectrum, MS1(t, f) denotes the magnitude mask, | Y(t, f) | denotes the magnitude spectrum corresponding to the first set of bin features, t and f represent time and frequency indices, and U , U2, V and k2are predetermined values indicative of a to-be-estimated size of the magnitude mask. Of course, as will be understood and appreciated by the skilled person, the deep filtering used for applying the magnitude mask
[0085] 20 may be implemented in any other suitable means as well.
[0086] In some example embodiments, the modified complex-valued spectrum K51(t, / ) may be obtained according to: where operator (p is used for calculating an argument of a complex number.
[0087] In some example embodiments, the deep filtering used for applying the complex mask at the second stage may be implemented according to: where YS1(t, f) denotes a complex-valued spectrum corresponding to the fourth set of bin features, MS2(t, f) denotes the complex mask, and N and N2are predetermined
[0088] 30 values indicative of a to-be-estimated size of the complex mask. Of course, as will be PCT / US25 / 47728 24 September 2025 (24.09.2025)
[0089] D24140W001 understood and appreciated by the skilled person, the deep filtering used for applying the complex mask may be implemented in any other suitable means as well.
[0090] In some example embodiments, at least one (e.g., any one or both) of the first and second two- stage mask estimation modules may comprise: a first activation layer configured to, at the first
[0091] 5 stage, limit the magnitude mask to a value range of 0 to 1 (or any other suitable value range); and a second activation layer configured to, at the second stage, limit the complex mask to a value range of -1 to 1 (or any other suitable value range). In some possible examples, the first may be implemented as a sigmoid function and / or the second activation layers may be implemented as a tanh function, or the like, depending on various circumstances
[0092] In some example embodiments, the DNN based model may comprise a processing chain comprising, in this order, a feature extraction module, followed by an encoder module, followed by a decoder module, and followed by the first and second two-stage mask estimation modules. This is merely one possible example implementation and should not be understood to constitute a limitation of any kind. Of course, as will be understood and
[0093] 15 appreciated by the skilled person, any other suitable processing chain, or more generally, the DNN based model, may be implemented as well, depending on various circumstances.
[0094] In some example embodiments, the loss function may comprise a weighted sum of a first loss function corresponding to the first reference signal and a second loss function corresponding to the first reference signal. In this sense, such loss function (by taking various parts into
[0095] 20 consideration) may also be referred to as a hybrid loss function. Of course, as will be understood and appreciated by the skilled person, the loss function may be designed in any other suitable means, depending on various implementations.
[0096] In some example embodiments, the loss function Loss may be implemented according to: where Lossp(S, S) denotes the first loss function, S denotes a spectrum of the first reference signal, S denotes an estimated spectrum after the first mask has been applied, Lossp(M, M) denotes the second loss function, M denotes a spectrum of the second reference signal, M denotes an estimated spectrum after the second mask has been applied, and a is a predefined weighting coefficient between the first and second loss functions.
[0097] 30 In some example embodiments, at least one of the first and second loss functions Losspmay comprise a weighted sum of a respective perceptual loss function and a respective Mean Squared Error (MSE) loss function implemented according to: PCT / US25 / 47728 24 September 2025 (24.09.2025)
[0098] D24140W001 where Lossmdenotes the perceptual loss function, Losscdenotes the MSE loss function, and f> is a predefined weighting coefficient.
[0099] In some example embodiments, for the first loss function, the perceptual loss function
[0100] 5 Lossm(S, S) may be implemented according to: where the MSE loss function Lossc(S,S) is implemented according to: where S denotes a spectrum of the first reference signal, S denotes an estimated spectrum after the first mask has been applied, p is a predefined spectral compression factor, operator (p is used for calculating an argument of a complex number, and m is a predefined tuning parameter.
[0101] 15 In some example embodiments, the loss function may further comprise a residual loss function (in addition to the above-illustrated first / second loss function).
[0102] In some example embodiments, the loss function Loss may be implemented according to: where Lossp(S, S) denotes the first loss function, Lossp(M, M) denotes the second loss function, and Lossp(N, N) denotes the residual loss function, wherein
[0103] N = 7 - S - M, and
[0104] N = I — S — M, where S denotes a spectrum of the first reference signal, S denotes an estimated spectrum after
[0105] 25 the first mask has been applied, M denotes a spectrum of the second reference signal, M denotes an estimated spectrum after the second mask has been applied, N denotes a spectrum of a residual signal, N denotes a spectrum of an estimated residual signal, 7 denotes a PCT / US25 / 47728 24 September 2025 (24.09.2025)
[0106] D24140W001 complex-valued spectrum of the audio mixture signal, and a, / ? and y are predefined weighting coefficients.
[0107] In some example embodiments, updating the parameters of the neural network based on the evaluation may comprise updating weights in the neural network.
[0108] 5 According to a third aspect of the present disclosure, there is provided an apparatus including a processor and a memory coupled to the processor. The processor may be adapted to cause the apparatus to carry out all steps according to any of the example methods described in the foregoing aspects.
[0109] According to a fourth aspect of the present disclosure, there is provided a computer program.
[0110] 10 The computer program may include instructions that, when executed by a processor, cause the processor to carry out all steps of the example methods described throughout the present disclosure.
[0111] According to a fifth aspect of the present disclosure, there is provided a computer-readable storage medium. The computer-readable storage medium may store the aforementioned computer program.
[0112] It will be appreciated that apparatus features and method steps may be interchanged in many ways. In particular, the details of the disclosed method(s) can be realized by the corresponding apparatus (or system), and vice versa, as the skilled person will appreciate. Moreover, any of the above statements made with respect to the method(s) are understood to likewise apply to
[0113] 20 the corresponding apparatus (or system), and vice versa.
[0114] BRIEF DESCRIPTION OF DRAWINGS
[0115] Example embodiments of the disclosure are explained below with reference to the accompanying drawings, wherein
[0116] 25 Fig. 1 schematically illustrates an example framework for training a neural network for use in audio object estimation according to some example embodiments of the present disclosure,
[0117] Fig. 2 schematically illustrates an example neural network,
[0118] Fig. 3 schematically illustrates an example neural network for use in audio object estimation according to some example embodiments of the present disclosure, PCT / US25 / 47728 24 September 2025 (24.09.2025)
[0119] D24140W001
[0120] Fig. 4 schematically illustrates an example of an implementation of a two stage mask estimation module for use in a neural network according to some example embodiments of the present disclosure,
[0121] Fig- 5 is a flowchart illustrating an example of a method of training a neural network for use
[0122] 5 in audio object estimation according to some example embodiments of the present disclosure,
[0123] Fig. 6 schematically illustrates an example framework for using a neural network for audio object estimation according to some example embodiments of the present disclosure,
[0124] Fig. 7 is a flowchart illustrating an example of a method of using a neural network for audio object estimation according to some example embodiments of the present disclosure,
[0125] Fig. 8 schematically illustrates another example of a method of using a neural network for audio object estimation, and
[0126] Fig. 9 is a schematic block diagram of an example electronic device or architecture.
[0127] DETAILED DESCRIPTION
[0128] 15 The Figures (Figs.) and the following description relate to preferred embodiments by way of illustration only. It should be noted that from the following discussion, alternative embodiments of the structures and methods disclosed herein will be readily recognized as viable alternatives that may be employed without departing from the principles of what is claimed.
[0129] 20 As briefly mentioned above, deep learning-based speech enhancement techniques may be understood to target removing unwanted artifacts such as noise, reverberation, or the like, and extracting clean speech from the mixture. Usually, these techniques only focus on speech signal and ignore other types of audio signals. However, with the rapid development of short video field, the user generated content (UGC) has become more and more popular, and the UGC typically contains various ty pes of audio signals. To bring a more immersive audio experience, there appears to be a general need to develop methods that can handle multiple types of audio sources in addition to speech signals.
[0130] References are now made to the figures. In particular, it is to be noted that identical or tike reference numbers used in the figures of the present disclosure may, unless indicated
[0131] 30 otherwise, indicate identical or tike elements, such that repeated description thereof may be omitted for reasons of conciseness. It is also to be noted that while some of the examples PCT / US25 / 47728 24 September 2025 (24.09.2025)
[0132] D24140W001 shown in the figures and described below may seem to explicitly make reference to speech and / or music, they are merely provided as possible examples for easy illustration purposes only, and thus should not be understood to constitute limitations of any kind. As will be understood and appreciated by the skilled person, any other suitable type of audio signal may
[0133] 5 be applied as well, depending on various implementations.
[0134] Fig. 1 schematically illustrates an example framework 100 for training a neural network for use in audio object estimation according to some example embodiments of the present disclosure.
[0135] In particular, as will be understood and appreciated by the skilled person, in the training stage, for each training step, it would be necessary to first collect the clean speech, the clean music (which would be used as the ground truth), room impulse response, noise, or the like.
[0136] Then, suitable techniques, such as data augmentation or the like, may be performed to obtain the training pair(s) (namely, the noisy mixture 101, the clean speech 102, and the clean music 103). The noisy mixture 101, the clean speech 102, and the clean music 103 are fed into a
[0137] 15 suitable feature extraction module 110 (or, in some possible cases, undergo a suitable feature extraction process, or the like) to obtain complex- valued (or simply complex for short) bin features 111, 112, and 113, respectively. In some possible cases, the feature extraction process / module may involve determining respective complex spectrum value for each time / frequency bin or the like, for instance.
[0138] 20 More specifically, in some possible examples, the respective waveform may be fed into a suitable transform function to obtain time-frequency bin features. Depending on various implementations, the transform can be short-time Fourier transform (STFT), modified- discrete cosine transform (MDCT), modified discrete X transform (MDXT), a filterbank based transform such as complex quadrature mirror filter (CQMF) transform, or any other suitable time-frequency transform.
[0139] Then the complex bin features of the noisy mixture 111 may be fed into an audio object estimation / extraction model 120, for example, a neural network (or deep learning) based model, to obtain the predicted speech and music bin masks 122 and 123, respectively. The predicted bin masks 122, 123, and the spectrum of the noisy mixture 111 may be used to
[0140] 30 obtain the predicted speech and music spectrums 132 and 133, respectively.
[0141] Finally, the predicted speech and music spectrums 132, 133, the corresponding clean speech and music spectrums 112, 113, and the noisy mixture spectrum 111 may be used to calculate PCT / US25 / 47728 24 September 2025 (24.09.2025)
[0142] D24140W001
[0143] (or evaluate) a (hybrid) loss function 140. which would in turn be used to update the model parameters (e.g., weights of the neural network).
[0144] As the aim of the final framework is to estimate / extract a speech signal and a music signal that are as close as possible to the respective reference / clean speech and music signals, in this final
[0145] 5 step, the estimated / predicted speech and music spectrums 132, 133 and the respective spectrums (or bin features) of the reference / clean speech and music signals have to be compared, for example with the help of the (hybrid) loss function 140. Any suitable loss function may be used to evaluate the performance of the performed audio object estimation / extraction, such as a mean square error (MSE) based loss function or the like. Notably, the present disclosure generally proposes a hybrid loss function to further improve the performance of the overall audio object estimation / extraction system, which will be described in more detail below.
[0146] The result of the (hybrid) loss function may be evaluated. For instance, in some possible cases, the evaluation may be performed on the result of the (hybrid) loss function over
[0147] 15 multiple training pairs, each of which comprises, on the one hand, a reference speech signal and a reference music signal as ground truth, and on the other hand, a corresponding (noisy) audio mixture signal. Accordingly, depending on the evaluation result(s), parameters of the neural network may be updated. Updating the parameters may include updating of weights in the neural network. The neural network may be trained, for example, until the result of the
[0148] 20 loss function meets a (predetermined) threshold (or condition), or the like.
[0149] As described above, a neural network would be needed in order to be able to eventually (e.g., after proper training) estimate respective speech and music masks based on the (timefrequency) bin features of the input noisy mixture. Such neural network may be implemented in any suitable means, depending on various circumstances. For instance, Fig. 2 schematically illustrates a possible example implementation of a neural network 200. Therein, the exemplary neural network model structure 200 may be seen to comprise a processing chain that includes: a feature extraction module 220 (which may take, for example band (or bin) based features 210 as input), an encoder module 230, a decoder module 240 and a final CNN layer 250. In some possible cases, the encoder module 230 may, optionally, have one or more
[0150] 30 downsample layers and other CNN layers and dense connections together. Similarly, in some possible cases, the decoder module 240 may, optionally, have one or more up-sample layers and other CNN layers and dense connections together. The final CNN layer 250 may be understood to be used to estimate a suitable mask 260, such as a speech mask or the like. PCT / US25 / 47728 24 September 2025 (24.09.2025)
[0151] D24140W001
[0152] As illustrated earlier, the present disclosure generally seeks to be able to (simultaneously) estimate or extract multiple audio contents / objects (e.g., speech and music) from an input audio mixture signal. In view of this, the neural network (e.g., the exemplary neural network model structure 200 shown in Fig. 2) may need appropriate extension. An example neural
[0153] 5 network 300 for use in audio object estimation according to some example embodiments of the present disclosure is schematically illustrated in Fig. 3. As indicated above, identical or like reference numbers used in Fig. 3 may, unless indicated otherwise, indicate identical or like elements in Fig. 2, such that repeated description thereof may be omitted for reasons of conciseness.
[0154] In particular, as shown in Fig. 3, compared to the model structure 200 of Fig. 2, the model 300 generally uses complex-valued input (bin) features 310 as input and proposes to use two so-called two-stage mask estimation modules 351 and 352 instead of a simple CNN layer (e.g., the CNN layer 250 in the example of Fig. 2) to estimate speech and music bin masks 361 and 362, respectively.
[0155] 15 Now, with reference to Fig. 4, a possible example of an implementation of a two stage mask estimation module 400 for use in a neural network according to some example embodiments of the present disclosure will be illustrated in detail. Such two stage mask estimation module 400 may be used to implement one or both of the two two-stage mask estimation modules 351 and 352 as exemplarily shown in Fig. 3, for example.
[0156] 20 In particular, as exemplarily shown in Fig. 4, at the first stage, the output of module 420 is fed into a (first) CNN layer 431 in order to be able to estimate magnitude masks 451. Here, it may be understood that this module 420 schematically corresponds to the feature extraction module 320, the encoder module 330, and the decoder module 340 of the example shown in Fig- 3 in combination, such that the output of this module 420 may be simply considered to schematically correspond to the output of the decoder module 340. In some possible cases, the CNN layer 431 may be followed by an activation layer (e.g., a sigmoid function or the like) 441 to limit the mask values between 0 and 1 (or any other suitable values / range). Then, the magnitude masks 451 and the input bin features 410 are fed into a deep filter module 461 in order to obtain a modified magnitude spectrum. As can be understood and appreciated by the
[0157] 30 skilled person, such deep filtering may be implemented in any suitable means. For instance, in some possible examples, the form of the deep filter may be seen as follows: PCT / US25 / 47728 24 September 2025 (24.09.2025)
[0158] D24140W001 where Y^ (t, / ) denotes the modified magnitude spectrum, MS1(t, f) denotes the estimated magnitude mask, | Y (t, / ) | represents the magnitude of the input spectrum, and t and f represent time and frequency indices.
[0159] It may be worth mentioning that, in this illustrated example, more particularly in the first
[0160] 5 stage, the mask is estimated for a size of (1^ + V2+ 1, U1+ U2+ 1) and then applied to the magnitude of the input spectrum in the form of deep filter. However, as other suitable implementation may be applicable as well, as will be understood and appreciated by the skilled person.
[0161] Then, the modified complex spectrum may be obtained by using such
[0162] 10 phase from the input spectrum, which is shown as: where operator <p may be understood to be used to calculate the argument of a complex number.
[0163] Subsequently, in the second stage, the output of module 420 is also fed into another (second) CNN layer 432 to estimate complex masks 452. Similar to above, the CNN layer 432 may be
[0164] 15 usually followed by an activation layer (e.g., a tanh function or the like) 442 to limit the mask values between -1 and 1 (or any other suitable values / range). Then the complex masks 452 and the modified complex spectrum from the first stage are fed into a (second) deep filter module 462 to obtain the final extracted complex spectrum 470. In some possible cases, the form of the (second) deep filter 462 in the second stage may be implemented as follows (or in
[0165] 20 any other suitable manner): where Ys' (t, / ) represents the final extracted complex spectrum, MS2(t, f) represents the estimated complex mask, and YS1(t, f) represents the modified spectrum from the first stage. In this exemplary implementation of the second stage, the mask of size (M + M2+ 1, N1+ N2+ 1) is estimated and applied to the spectrum from the first stage in the form of deep filter. PCT / US25 / 47728 24 September 2025 (24.09.2025)
[0166] D24140W001
[0167] Notably, in some experiments, it may be generally observed that using a two-stage mask estimation method may appear to be able to obtain cleaner and higher quality output, and at the same time prevent noise leakage. Moreover, computing deep filter for output reconstruction appears to be able to help the model to leverage neighbor time-frequency bins
[0168] 5 in a learnable way.
[0169] It may also be worth mentioning that, as can also be understood and appreciated by the skilled person, parameters, namely, and V2, and M , M2, N and N2, may be predetermined (or predefined) for indicating (controlling) the to-be-estimated sizes of the magnitude mask and the complex mask respectively. For instance, in some possible examples, U1may be 0, 1,
[0170] 10 2, etc., U2may be 0, 1, etc., V and V2may be 1, 2, etc., may be 0, 1, 2, etc., M2may be 0, 1, etc., and and N2may be 1, 2, etc. In some possible cases, if U2and M2are set to zero, then the two-stage mask estimation may be considered causal. Of course, any other suitable predetermined (or predefined) values for these parameters may be possible as well, depending on various implementations.
[0171] 15 In order to mitigate for example the over-suppression issue of conventional deep learningbased speech enhancement techniques, and possibly also taking the complex domain knowledge during training into consideration, a (hybrid) loss function that combines a perceptual loss function in the magnitude spectrum domain and an MSE based loss function in the complex spectrum domain may be proposed.
[0172] 20 In some possible examples, the combined (hy brid) perceptual loss function may be implemented as follows: where Lossm(S,S) represents the perceptual loss function in the magnitude spectrum domain, Lossc(S, S') represents the MSE loss function in the complex spectrum domain, / ? is a (e.g., predetermined) weighting coefficient between the magnitude spectrum domain loss function
[0173] 25 and the complex spectrum domain loss function, m is a (e.g., predetermined) tuning parameter that may be understood to control the shape of the asymmetric penalty, S and S represent the reference (i.e., clean) complex spectrum and the estimated complex spectrum, p is a (e.g., PCT / US25 / 47728 24 September 2025 (24.09.2025)
[0174] D24140W001 predetermined) spectral compression factor, operator (p may be understood to calculate the argument of a complex number.
[0175] Accordingly, in order to be able to perform (simultaneous) speech and music estimation / extract on, in some possible examples, a straightforward loss function may be
[0176] 5 implemented as: where Lossp(S, S) denotes the combined perceptual loss function for speech signal, S and S are the clean speech spectrum and the estimated speech spectrum, Lossp(M, M) denotes the combined perceptual loss function for music signal, M and M are the clean music spectrum and the estimated music spectrum, and a is a (e.g., predetermined) weighting coefficient
[0177] 10 between the speech related loss and the music related loss.
[0178] In some experimental evaluations, some over-suppressions in the extracted speech and music objects may still be observed, even though the above-proposed perceptual loss is used. In other words, the residual signal (namely, the mixture minus the extracted speech and music) appears to still contain some (leftover) speech and music content that should not be contained.
[0179] 15 Accordingly, in order to possibly alleviate the over-suppression issue and improve the speech and music extraction performance, in the present disclosure, a hybrid loss function that combines the object-specific loss function with a residual loss function is further proposed. Based on some experimental evaluations, the proposed hybrid loss function appears to be able to largely alleviate the over-suppression issue and improve the speech and music extraction
[0180] 20 performance.
[0181] In some possible examples, the hybrid loss function may be implemented as follows:
[0182] N = l — S — M (11)
[0183] N = 1 - S - M (12) denote the combined perceptual loss functions for speech and music signals, respectively, LosspN, N) denotes the residual loss between the noise N and the estimated residual signal N, 1 denotes the complex spectrum of input (noisy)
[0184] 25 audio signal, and y is a (e.g., predetermined) weighting coefficient between the speech related PCT / US25 / 47728 24 September 2025 (24.09.2025)
[0185] D24140W001 loss and the residual loss. In some possible examples, parameter a may be 0.8, 1, 1.2, etc., and parameter y can be 0.5, 1.0, etc, or any other suitable value, depending on various implementations.
[0186] Generally speaking, the purpose of the residual loss is to remove the speech and music
[0187] 5 contents in the residual signal and put them back into the extracted speech and music signal. It may be worth mentioning that here the residual loss may appear to be calculated by using the combined perceptual loss function. Of course, as will be understood and appreciated by the skilled person, any other suitable loss functions, such as MSE based loss, mean absolute error (MAE) based loss can also be used as the residual loss function as well, depending on various implementations.
[0188] Reference is now made to Fig. 5, which is a flowchart schematically illustrating an example of a method 500 of training a neural network for use in audio object estimation according to some example embodiments of the present disclosure. The method 500 may correspond to the steps / procedures performed according to the training of the audio object estimation / extraction
[0189] 15 framework 100 as exemplarily shown in Fig. 1. Depending on various implementations, blocks of the method 500 may be performed by a suitable apparatus, or the like.
[0190] In particular, method 500 may comprise, at step S510, obtaining an audio mixture signal, and corresponding first and second reference signals of different types that are used as ground truth. In other words, for the purpose of training, one or more training pairs may be first
[0191] 20 obtained, each of which may comprise, on the one hand, an audio mixture signal, and on the other hand, a corresponding first reference signal and a corresponding second reference signal that are used as ground truth. The audio mixture signal (or noisy mixture) may comprise, among others (e.g., background noises, etc.), at least a first signal (corresponding to the first reference signal) and a second signal (corresponding to the second reference signal) of different types. The first and second signals may be speech and music signals respectively, as schematically illustrated above. Of course, as will be understood and appreciated by the skilled person, any other suitable audio signal may be applied as well, depending on various implementations.
[0192] Method 500 may further comprise, at step S520, obtaining (e.g., by using any suitable feature
[0193] 30 extraction means, or the like) a first set of bin features from the audio mixture signal, a second set of bin features from the first reference signal, and a third set of bin features from the second reference signal, wherein each bin feature in the first to third sets corresponds to a PCT / US25 / 47728 24 September 2025 (24.09.2025)
[0194] D24140W001 respective time-frequency bin of the audio mixture signal, the first reference signal and the second reference signal.
[0195] Method 500 may further comprise, at step S530, feeding the first set of bin features to the neural network to estimate a first mask and a second mask respectively corresponding to the
[0196] 5 first reference signal, and the second reference signal. In other words, the neural network model may be used for determining, based on the first set of bin features, the first mask that is associated with the first signal, and also the second mask that is associated with the second signal.
[0197] Method 500 may yet further comprise, at step S540, obtaining a fourth set of bin features based on the first mask and the first set of bin features; and, at step S550, also obtaining a fifth set of bin features based on the second mask and the first set of bin features.
[0198] Finally, the method may comprise, at step S560, evaluating a loss function based on the first to fifth sets of bin features; and, at step S570, updating parameters (e.g., weights, etc.) of the neural network based on a result of the evaluation.
[0199] 15 After training of the audio object estimation / extract! on framework, such (trained) audio object estimation / extraction framework may be ready to be used for practice audio object estimation / extraction (sometimes also referred to as the inference stage, in comparison to the training stage as described with reference to Fig. 1). Fig. 6 schematically illustrates an example framework 600 for using such (trained) neural network for audio object estimation
[0200] 20 according to some example embodiments of the present disclosure. As indicated above, identical or like reference numbers used in Fig. 6 may, unless indicated otherwise, indicate identical or like elements in Fig. 1, such that repeated description thereof may be omitted for reasons of conciseness.
[0201] More particularly, in the inference stage, an input audio (test) signal 601 may be fed into the feature extraction module / process 610 to obtain the complex- valued bin features 611. The test signal 601 may comprise an audio signal (e.g., input audio signal) with a mixture of different types of audio, including speech, music, background noise, or any combination thereof. The test signal 601 may also include one or more audio objects, one or more audio channels, static or dynamic metadata, or any combination thereof.
[0202] 30 Then, the complex bin features 611 may be fed into the trained audio object estimation / extraction model 620 to obtain the predicted speech and music bin masks 622 and 623, respectively. PCT / US25 / 47728 24 September 2025 (24.09.2025)
[0203] D24140W001
[0204] Subsequently, the predicted bin masks 622, 623 and the spectrum of the input test signal 601 are used to obtain the predicted spectrums 632 and 633 respectively, which are then used to obtain the estimated / extracted speech and music waveforms 662 and 663, respectively, for example through transform inverse 650 (corresponding to the previously applied transform
[0205] 5 function).
[0206] Fig. 7 is a flowchart schematically illustrating an example of a method 700 of using a neural network for audio object estimation according to some example embodiments of the present disclosure. The method 700 may correspond to the steps / procedures performed according to the audio object estimation / extraction framework 600 as exemplarily shown in Fig. 6. Depending on various implementations, blocks of the method 700 may be performed by a suitable apparatus, or the like.
[0207] In particular, method 700 may comprise, at step S710, obtaining a time-frequency representation (in a time-frequency domain) of an audio signal that comprises a mixture of at least a first signal and a second signal of different types.
[0208] 15 Method 700 may further comprise, at step S720, obtaining (or determining) a set of bin features from the time-frequency representation, wherein each bin feature in the set corresponds to a respective time-frequency bin of the time-frequency representation. The bin features may be obtained (or determined) from the time-frequency representation by using any suitable feature extraction technique, for example.
[0209] 20 Method 700 may further comprise, at step S730, feeding the set of bin features to a neural network model (e.g., the neural network that has been trained according to the illustration with reference to Figs. 1 to 5) to estimate (determine) a first mask and a second mask respectively corresponding to the first signal and the second signal. In other words, the neural network model may be used for determining the first mask that is associated with the first signal; and for determining the second mask that is associated with the second signal.
[0210] Accordingly, method 700 may yet further comprise, at step S740, respectively applying the first mask and the second mask to the time-frequency representation, in order to obtain an estimated (or extracted) time-frequency representation of the first signal and an estimated (or extracted) time-frequency representation of the second signal.
[0211] 30 To summarize the above, broadly speaking, the present disclosure may be seen to aim at developing a deep learning (or neural network) based audio object estimation / extraction PCT / US25 / 47728 24 September 2025 (24.09.2025)
[0212] D24140W001 system that can simultaneously extract speech and music (or the like) contents from noisy mixtures.
[0213] Accordingly, in a broad sense, some of the example embodiments of the present disclosure may be understood to relate generally to a deep learning (or neural network) based audio
[0214] 5 object estimation / extract! on that includes a bin-based neural network model with a two-stage mask estimation method and a hybrid loss function. In simple terms, the techniques proposed in the present disclosure may be understood to involve for example: a bin-based audio object extraction model based on a deep neural network (DNN)-based model that uses a two-stage mask estimation method to extract speech and music content from noisy mixtures, resulting in
[0215] 10 cleaner and higher quality spectrums; and a hybrid loss function that combines the objectspecific loss function with a residual loss function, which can mitigate over-suppression issues and extract more content from the mixture to speech and music objects.
[0216] In some examples, the present disclosure may provide a deep learning-based audio object extraction system with bin-based model and a hybrid loss function including:
[0217] 15 a bin-based model that estimates complex bin masks using two two-stage mask estimation modules, each corresponding to speech and music signal, respectively; and a hybrid loss function that combines the object-specific perceptual loss with a residual loss.
[0218] In some examples, the bin-based model may comprise two two-stage mask estimation
[0219] 20 modules after the decoder module of the LensNet model.
[0220] In some examples, the hybrid loss function may include: two perceptual loss functions for speech and music content respectively; and a residual loss function to calculate the loss between the residual and noise.
[0221] In some examples, the residual may be obtained by subtracting the extracted speech and music
[0222] 25 from the mixture.
[0223] In some examples, the two-stage mask estimation module may comprise two CNN layers, each CNN layer followed by an activation layer and a deep filter module.
[0224] In some examples, the first CNN layer may estimate magnitude masks and the second CNN layer may estimate complex masks. PCT / US25 / 47728 24 September 2025 (24.09.2025)
[0225] D24140W001
[0226] In some examples, the deep filter in the first stage may be used to obtain modified magnitude spectrums, and the deep filter in the second stage may be used to obtain the final extracted complex spectrums.
[0227] In some examples, the techniques described herein may relate to a deep learning-based audio
[0228] 5 extraction method for extracting two or more types of audio from an input audio signal, the method including: extracting features from the input audio signal; estimating, via a deep neural network-based model, a mask of a first type of audio and a mask of a second type of audio; determining a first complex spectrum based on the estimated mask of the first type and a complex spectrum of the first type; determining a second complex spectrum based on the estimated mask of the second type and a complex spectrum of the second type; and determining an audio signal of the first type based on the complex spectrum of the first type and an audio signal of the second type based on the complex spectrum of the second type.
[0229] In some examples, the techniques described herein may relate to a method, wherein the first type of audio corresponds to music and the second type of audio corresponds to speech.
[0230] 15 In some examples, the techniques described herein may relate to a method, wherein the input audio signal corresponds to a mixed audio signal including a plurality of types of audio.
[0231] In some examples, the techniques described herein may relate to a method, wherein extracting features from the input audio signal includes determining, via a feature extraction module, complex bin features.
[0232] 20 In some examples, the techniques described herein may relate to a method, wherein the complex bin features include time-frequency bin features.
[0233] In some examples, the techniques described herein may relate to a deep learning-based audio extraction system for extracting two or more types of audio from an input audio signal, the system including: a feature extraction module configured to extract features from the input audio signals; one or more multi-stage mask estimation modules configured to determine a mask of a type of audio, the multi-stage mask estimation module including: a first stage configured to determine a modified magnitude spectrum based on a magnitude mask and the extracted features; a second stage configured to determine a complex spectrum based on a complex mask and the modified magnitude spectrum; and determining, from the complex
[0234] 30 spectrum, the mask of the type of audio.
[0235] In some examples, the techniques described herein may relate to a system, including at least two multi-stage mask estimation modules, wherein a first multi-stage mask estimation module PCT / US25 / 47728 24 September 2025 (24.09.2025)
[0236] D24140W001 is configured to determine a speech mask, and a second multi-stage mask estimation module is configured to determine a music mask.
[0237] In some examples, the techniques described herein may relate to a system, further including an encoder module and a decoder module.
[0238] 5 In some examples, the techniques described herein may relate to a system, further including: a first convolutional neural network layer configured to estimate the magnitude mask and a second convolutional neural network layer configured to estimate the complex mask.
[0239] In some examples, the techniques described herein may relate to a system, wherein the first stage includes a first deep filter module configured to take, as input, the magnitude mask and the extracted features and output, the modified magnitude spectrum.
[0240] In some examples, the techniques described herein may relate to a system, wherein the second stage includes a second deep filter module configured to take, as input, the complex mask and the modified magnitude spectrum, and output, the complex spectrum.
[0241] In some examples, the techniques described herein may relate to a system, wherein the system
[0242] 15 is trained to minimize a hybrid loss function including a perceptual loss function for speech content, a perceptual loss function for music content, and a residual loss function.
[0243] In some examples, the techniques described herein may relate to an apparatus including a processor and a memory, configured to perform the method described above.
[0244] In some examples, the techniques described herein may relate to a computer program product
[0245] 20 including instructions which, when the program is executed by a computer, cause the computer to carry out the method described above.
[0246] In some examples, the techniques described herein may relate to a computer-readable storage medium storing the computer program product.
[0247] Further, Fig. 8 schematically illustrates another example of a method 800 of using a neural network for estimating / extracting two or more types of audio from an input audio signal. The method may be performed by an electronic processor (for example, the electronic apparatus 900 of Fig. 9), which may be configured to perform the method via execution of machineexecutable instructions. The various process blocks illustrated in Fig. 8 provide examples of various methods disclosed herein, and it is understood that some blocks may be removed,
[0248] 30 added, combined, or modified without departing from the spirit of the present disclosure. PCT / US25 / 47728 24 September 2025 (24.09.2025)
[0249] D24140W001
[0250] At step S810, the method 800 may extract features from the input audio signal. For example, the electronic processor 901 may extract features from the input audio signal via a timefrequency transform.
[0251] At step S820, the method 800 may estimate, via a deep neural network-based model, a mask
[0252] 5 of a first type of audio and a mask of a second type of audio. For example, the electronic processor 901 may estimate, via a DNN-based model, a speech mask and a music mask.
[0253] At step S830, the method 800 may determine a first complex spectrum based on the estimated mask of the first type and a complex spectrum of the first type.
[0254] At step S840, the method 800 may determine a second complex spectrum of a second type based on the estimated mask of the second type and a complex spectrum of the second type.
[0255] Finally, at step S850, the method 800 may determine an audio signal of the first type based on the complex spectrum of the first type and an audio signal of the second type based on the complex spectrum of the second type.
[0256] Moreover, Fig. 9 is a schematic block diagram of an example electronic device or architecture
[0257] 15 900 suitable for implementing example embodiments of the present disclosure. Architecture 900 may include, but is not limited to servers and client devices, systems, modules and methods as described in reference to Figs. 1 to 8.
[0258] As shown, the architecture 900 includes central processing unit (CPU) 901 which is capable of performing various processes in accordance with a program stored in, for example, read
[0259] 20 only memory (ROM) 902 or a program loaded from, for example, storage unit 908 to random access memory (RAM) 903. The CPU 901 may be, for example, an electronic processor 901, which may include one or more processor cores, and in some examples, the processor 901 may be multiple processors. In RAM 903, the data used when CPU 901 performs the various processes is also stored, as required. CPU 901, ROM 902, and RAM 903 are connected to one another via bus 904. Input / output (I / O) interface 905 is also connected to bus 904.
[0260] The following components are connected to I / O interface 905: input unit 906, that may include a keyboard, a mouse, or the like; output unit 907 that may include a display such as a liquid crystal display (LCD) and one or more speakers; storage unit 908 including a hard disk, or another suitable storage device; and communication unit 909 which may include a network
[0261] 30 interface card such as a network card (e.g., wired or wireless). PCT / US25 / 47728 24 September 2025 (24.09.2025)
[0262] D24140W001
[0263] In some implementations, input unit 906 includes one or more microphones in different positions (depending on the host device) enabling capture of audio signals in various formats (e.g., mono, stereo, spatial, immersive, and other suitable formats).
[0264] In some implementations, output unit 909 includes systems with various numbers of speakers.
[0265] 5 Output unit 907 (depending on the capabilities of the host device) can render audio signals in various formats (e.g., mono, stereo, immersive, binaural, and other suitable formats).
[0266] In some embodiments, communication unit 909 is configured to communicate with other devices (e.g., via a network). Drive 910 is also connected to I / O interface 905, as required. Removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, a flash drive or another suitable removable medium is mounted on drive 910, so that a computer program read therefrom is installed into storage unit 908, as required. A person skilled in the art would understand that although apparatus 900 is described as including the abovedescribed components, in real applications, it is possible to add, remove, and / or replace some of these components and all these modifications or alterations all fall within the scope of the
[0267] 15 present disclosure.
[0268] In accordance with example embodiments of the present disclosure, the processes described above may be implemented as computer software programs or on a computer-readable storage medium. For example, embodiments of the present disclosure include a computer program product including a computer program tangibly embodied on a machine readable medium, the
[0269] 20 computer program including program code for performing methods. In such embodiments, the computer program may be downloaded and mounted from the network via the communication unit 909, and / or installed from the removable medium 911, as shown in Fig. 9.
[0270] Generally, various example embodiments of the present disclosure may be implemented in hardware or special purpose circuits (e.g., control circuitry), software, logic or any combination thereof. For example, the units discussed above can be executed by control circuitry (e.g., CPU 901 in combination with other components of Fig. 9), thus, the control circuitry may be performing the actions described in this disclosure. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software which may be executed by a controller, a processor and / or other computing device(s), which
[0271] 30 may include control circuitry. While various aspects of the example embodiments of the present disclosure are illustrated and described as block diagrams, flowcharts, or using some other pictorial representation, it will be appreciated that the blocks, apparatus, systems, techniques, or methods described herein may be implemented in, as non-limiting examples, PCT / US25 / 47728 24 September 2025 (24.09.2025)
[0272] D24140W001 hardware, software, firmware, special purpose circuits or logic, general purpose hardware or controller or other computing devices, or some combination thereof.
[0273] Interpretation
[0274] A computing device implementing the techniques described above can have the following
[0275] 5 example architecture. Other architectures are possible, including architectures with more or fewer components. In some implementations, the example architecture includes one or more processors (e.g., dual -core Intel® Xeon® Processors), one or more output devices (e.g., LCD), one or more network interfaces, one or more input devices (e.g., mouse, keyboard, touch-sensitive display) and one or more computer-readable mediums (e.g., RAM, ROM, SDRAM, hard disk, optical disk, flash memory, etc.). These components can exchange communications and data over one or more communication channels (e.g., buses), which can utilize various hardware and software for facilitating the transfer of data and control signals between components.
[0276] The term “computer-readable medium” refers to a medium that participates in providing
[0277] 15 instructions to processor for execution, including without limitation, non-volatile media (e.g., optical or magnetic disks), volatile media (e.g., memory) and transmission media. Transmission media includes, without limitation, coaxial cables, copper wire and fiber optics.
[0278] Computer-readable medium can further include operating system (e.g., a Linux® operating system), network communication module, audio interface manager, audio processing manager
[0279] 20 and live content distributor. Operating system can be multi-user, multiprocessing, multitasking, multithreading, real time, etc. Operating system performs basic tasks, including but not limited to: recognizing input from and providing output to network interfaces and / or devices; keeping track and managing files and directories on computer-readable mediums (e.g., memory or a storage device); controlling peripheral devices; and managing traffic on the one or more communication channels. Network communications module includes various components for establishing and maintaining network connections (e.g., software for implementing communication protocols, such as TCP / IP, HTTP, etc.).
[0280] Architecture can be implemented in a parallel processing or peer-to-peer infrastructure or on a single device with one or more processors. Software can include multiple software
[0281] 30 components or can be a single body of code.
[0282] The described features can be implemented advantageously in one or more computer programs that are executable on a programmable system including at least one programmable PCT / US25 / 47728 24 September 2025 (24.09.2025)
[0283] D24140W001 processor coupled to receive data and instructions from, and to transmit data and instructions to, a data storage system, at least one input device, and at least one output device. A computer program is a set of instructions that can be used, directly or indirectly, in a computer to perform a certain activity or bring about a certain result. A computer program can be written
[0284] 5 in any form of programming language (e.g., Objective-C, Java), including compiled or interpreted languages, and it can be deployed in any form, including as a stand-alone program or as a module, component, subroutine, a browser-based web application, or other unit suitable for use in a computing environment.
[0285] Suitable processors for the execution of a program of instructions include, by way of example,
[0286] 10 both general and special purpose microprocessors, and the sole processor or one of multiple processors or cores, of any kind of computer. Generally, a processor will receive instructions and data from a read-only memory or a random access memory or both. The essential elements of a computer are a processor for executing instructions and one or more memories for storing instructions and data. Generally, a computer will also include, or be operatively
[0287] 15 coupled to communicate with, one or more mass storage devices for storing data files; such devices include magnetic disks, such as internal hard disks and removable disks; magnetooptical disks; and optical disks. Storage devices suitable for tangibly embodying computer program instructions and data include all forms of non-volatile memory, including by way of example semiconductor memory devices, such as EPROM, EEPROM, and flash memory
[0288] 20 devices; magnetic disks such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The processor and the memory can be supplemented by, or incorporated in, ASICs (application-specific integrated circuits).
[0289] To provide for interaction with a user, the features can be implemented on a computer having a display device such as a CRT (cathode ray tube) or LCD (liquid crystal display) monitor or a retina display device for displaying information to the user. The computer can have a touch surface input device (e.g., a touch screen) or a keyboard and a pointing device such as a mouse or a trackball by which the user can provide input to the computer. The computer can have a voice input device for receiving voice commands from the user.
[0290] The features can be implemented in a computer system that includes a back-end component,
[0291] 30 such as a data server, or that includes a middleware component, such as an application server or an Internet server, or that includes a front-end component, such as a client computer having a graphical user interface or an Internet browser, or any combination of them. The components of the system can be connected by any form or medium of digital data PCT / US25 / 47728 24 September 2025 (24.09.2025)
[0292] D24140W001 communication such as a communication network. Examples of communication networks include, e.g., a LAN, a WAN, and the computers and networks forming the Internet.
[0293] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The
[0294] 5 relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. In some embodiments, a server transmits data (e.g., an HTML page) to a client device (e g., for purposes of displaying data to and receiving user input from a user interacting with the client device). Data generated at the client device (e.g., a result of the user interaction) can be received from the client device at the server.
[0295] A system of one or more computers can be configured to perform particular actions by virtue of having software, firmware, hardware, or a combination of them installed on the system that in operation causes or cause the system to perform the actions. One or more computer programs can be configured to perform particular actions by virtue of including instructions
[0296] 15 that, when executed by data processing apparatus, cause the apparatus to perform the actions.
[0297] While this specification contains many specific implementation details, these should not be construed as limitations on the scope of any inventions or of what may be claimed, but rather as descriptions of features specific to particular embodiments of particular inventions. Certain features that are described in this specification in the context of separate embodiments can
[0298] 20 also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable subcombination. Moreover, although features may be described above as acting in certain combinations and even initially claimed as such, one or more features from a claimed combination can in some cases be excised from the combination, and the claimed combination may be directed to a subcombination or variation of a subcombination.
[0299] Similarly, while operations are depicted in the drawings in a particular order, this should not be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed, to achieve desirable results.
[0300] 30 In certain circumstances, multitasking and parallel processing may be advantageous.
[0301] Moreover, the separation of various system components in the embodiments described above should not be understood as requiring such separation in all embodiments, and it should be PCT / US25 / 47728 24 September 2025 (24.09.2025)
[0302] D24140W001 understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0303] Unless specifically stated otherwise, as apparent from the following discussions, it is appreciated that throughout the present disclosure discussions utilizing terms such as
[0304] 5 “processing”, “computing”, “calculating”, “determining”, “analyzing” or the like, refer to the action and / or processes of a computer or computing system, or similar electronic computing devices, that manipulate and / or transform data represented as physical, such as electronic, quantities into other data similarly represented as physical quantities.
[0305] Reference throughout this disclosure to “one example embodiment”, “some example embodiments” or “an example embodiment” means that a particular feature, structure or characteristic described in connection with the example embodiment is included in at least one example embodiment of the present disclosure. Thus, appearances of the phrases “in one example embodiment”, “in some example embodiments” or “in an example embodiment” in various places throughout this disclosure are not necessarily all referring to the same example
[0306] 15 embodiment. Furthermore, the particular features, structures or characteristics may be combined in any suitable manner, as would be apparent to one of ordinary skill in the art from this disclosure, in one or more example embodiments.
[0307] As used herein, unless otherwise specified the use of the ordinal adjectives “first”, “second”, “third”, etc., to describe a common object, merely indicates that different instances of like
[0308] 20 objects are being referred to and are not intended to imply that the objects so described must be in a given sequence, either temporally, spatially, in ranking, or in any other manner.
[0309] Also, it is to be understood that the phraseology and terminology used herein are for the purpose of description and should not be regarded as limiting. The use of “including,” “comprising,” or “having” and variations thereof are meant to encompass the items listed thereafter and equivalents thereof as well as additional items. Unless specified or limited otherwise, the terms “mounted”, “connected”, “supported”, and “coupled” and variations thereof are used broadly and encompass both direct and indirect mountings, connections, supports, and couplings.
[0310] In the claims below and the description herein, any one of the terms comprising, comprised of
[0311] 30 or which comprises is an open term that means including at least the elements / features that follow, but not excluding others. Thus, the term comprising, when used in the claims, should not be interpreted as being limitative to the means or elements or steps listed thereafter. For example, the scope of the expression a device comprising A and B should not be limited to PCT / US25 / 47728 24 September 2025 (24.09.2025)
[0312] D24140W001 devices consisting only of elements A and B. Any one of the terms including or which includes or that includes as used herein is also an open term that also means including at least the elements / features that follow the term, but not excluding others. Thus, including is synonymous with and means comprising.
[0313] 5 It should be appreciated that in the above description of example embodiments of the present disclosure, various features of the present disclosure are sometimes grouped together in a single example embodiment, Fig., or description thereof for the purpose of streamlining the present disclosure and aiding in the understanding of one or more of the various inventive aspects. This method of disclosure, however, is not to be interpreted as reflecting an intention that the claims require more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects he in less than all features of a single foregoing disclosed example embodiment. Thus, the claims following the Description are hereby expressly incorporated into this Description, with each claim standing on its own as a separate example embodiment of this disclosure.
[0314] 15 Furthermore, while some example embodiments described herein include some but not other features included in other example embodiments, combinations of features of different example embodiments are meant to be within the scope of the present disclosure, and form different example embodiments, as would be understood by those skilled in the art. For example, in the following claims, any of the claimed example embodiments can be used in
[0315] 20 any combination.
[0316] In the description provided herein, numerous specific details are set forth. However, it is understood that example embodiments of the present disclosure may be practiced without these specific details. In other instances, w ell-known methods, structures and techniques have not been shown in detail in order not to obscure an understanding of this description.
[0317] Thus, while there has been described what are believed to be the best modes of the present disclosure, those skilled in the art will recognize that other and further modifications may be made thereto without departing from the spirit of the present disclosure, and it is intended to claim all such changes and modifications as fall within the scope of the present disclosure. For example, any formulas given above are merely representative of procedures that may be used.
[0318] 30 Functionality may be added or deleted from the block diagrams and operations may be interchanged among functional blocks. Steps may be added or deleted to methods described within the scope of the present disclosure. PCT / US25 / 47728 24 September 2025 (24.09.2025)
[0319] D24140W001
[0320] Enumerated Example Embodiments
[0321] Various aspects and implementations of the present disclosure may also be appreciated from the following enumerated example embodiments (EEEs), which are not claims.
[0322] EEE 1. A deep learning-based audio object extraction system with bin-based LensNet
[0323] 5 model and a hybrid loss function comprising of a. A bin-based LensNet model that estimates complex bin masks using two two-stage mask estimation modules, each corresponding to speech and music signal, respectively. b. A hybrid loss function that combines the object-specific perceptual loss with a
[0324] 10 residual loss.
[0325] EEE 2. As in EEE 1, where the bin-based LensNet model comprising of two two-stage mask estimation modules after the decoder module of the LensNet model.
[0326] EEE 3. As in EEE 1, where the hybrid loss function comprising of a. Two perceptual loss functions for speech and music content respectively. b. A residual loss function to calculate the loss between the residual and noise.
[0327] EEE 4. As in EEE 3, the residual is obtained by subtracting the extracted speech and music from the mixture.
[0328] EEE 5. As in EEE 1 or 2, where the two-stage mask estimation module comprising two CNN layers, each CNN layer follows an activation layer and a deep filter module.
[0329] 20 EEE 6. As in EEE 5, the first CNN layer estimates magnitude masks and the second CNN layer estimates complex masks.
[0330] EEE 7. As in EEE 6, the deep filter in the first stage is used to obtain modified magnitude spectrums, and the deep filter in the second stage is used to obtain the final extracted complex spectrums.
[0331] 25 EEE 8. A deep learning-based audio extraction method for extracting two or more types of audio from an input audio signal, the method comprising: extracting features from the input audio signal; estimating, via a deep neural network-based model, a mask of a first type of audio and a mask of a second type of audio; PCT / US25 / 47728 24 September 2025 (24.09.2025)
[0332] D24140W001 determining a first complex spectrum based on the estimated mask of the first type and a complex spectrum of the first type; determining a second complex spectrum based on the estimated mask of the second type and a complex spectrum of the second type; and
[0333] 5 determining an audio signal of the first type based on the complex spectrum of the first type and an audio signal of the second type based on the complex spectrum of the second type.
[0334] EEE 9. The method of EEE 8, wherein the first type of audio corresponds to music and the second type of audio corresponds to speech.
[0335] EEE 10. The method of EEE 8 or 9, wherein the input audio signal corresponds to a mixed audio signal comprising a plurality of types of audio.
[0336] EEE 11. The method of any one of EEE 8 to 10, wherein extracting features from the input audio signal comprises determining, via a feature extraction module, complex bin features.
[0337] 15 EEE 12. The method of EEE 11, wherein the complex bin features comprise timefrequency bin features.
[0338] EEE 13. A deep learning-based audio extraction system for extracting two or more types of audio from an input audio signal, the system comprising: a feature extraction module configured to extract features from the input audio signals;
[0339] 20 one or more multi-stage mask estimation modules configured to determine a mask of a type of audio, the multi-stage mask estimation module comprising: a first stage configured to determine a modified magnitude spectrum based on a magnitude mask and the extracted features; a second stage configured to determine a complex spectrum based on a complex mask and the modified magnitude spectrum; and determining, from the complex spectrum, the mask of the type of audio.
[0340] EEE 14. The system of EEE 13, comprising at least two multi-stage mask estimation modules, wherein a first multi-stage mask estimation module is configured to determine a speech mask, and a second multi-stage mask estimation module is configured to determine a
[0341] 30 music mask. PCT / US25 / 47728 24 September 2025 (24.09.2025)
[0342] D24140W001
[0343] EEE 15. The system of EEE 13 or 14, further comprising an encoder module and a decoder module.
[0344] EEE 16. The system of any one of EEEs 13 to 15, further comprising: a first convolutional neural network layer configured to estimate the magnitude mask
[0345] 5 and a second convolutional neural network layer configured to estimate the complex mask.
[0346] EEE 17. The system of any one of EEEs 13 to 16, wherein the first stage comprises a first deep filter module configured to take, as input, the magnitude mask and the extracted features and output, the modified magnitude spectrum.
[0347] EEE 18. The system of any one of EEEs 13 to 17, wherein the second stage comprises
[0348] 10 a second deep filter module configured to take, as input, the complex mask and the modified magnitude spectrum, and output, the complex spectrum.
[0349] EEE 19. The system of any one of EEEs 13 to 18, wherein the system is trained to minimize a hybrid loss function comprising a perceptual loss function for speech content, a perceptual loss function for music content, and a residual loss function.
[0350] EEE 20. An apparatus comprising a processor and a memory, configured to perform the method according to any of EEEs 8 to 12.
[0351] EEE 21. A computer program product comprising instructions which, when the program is executed by a computer, causes the computer to carry out the method according to any of EEEs 8 to 12.
[0352] 20 EEE 22. A computer-readable storage medium storing the computer program product according to EEE 21.
Claims
PCT / US25 / 47728 24 September 2025 (24.09.2025)D24140W001CLAIMS1. A method for audio object estimation, comprising: obtaining a time-frequency representation of an audio signal that comprises a mixture5 of at least a first signal and a second signal of different types; obtaining a set of bin features from the time-frequency representation, wherein each bin feature in the set corresponds to a respective time-frequency bin of the time-frequency representation; feeding the set of bin features to a neural network model to estimate a first mask and a10 second mask respectively corresponding to the first signal and the second signal; and respectively applying the first mask and the second mask to the time-frequency representation, in order to obtain an estimated time-frequency representation of the first signal and an estimated time-frequency representation of the second signal.15 2. The method according to claim 1, wherein the first signal is a speech signal and the second signal is a music signal.
3. The method according to claim 1 or 2, wherein the first mask comprises values indicative of ratios of the first signal in the audio signal for respective time-frequency bins,20 and the second mask comprises values indicative of ratios of the second signal in the audio signal for respective time-frequency bins.
4. The method according to any one of the preceding claims, wherein the method further comprises: obtaining the audio signal; and transforming the audio signal by using a transform function, to thereby obtain the time-frequency representation of the audio signal.
5. The method according to claim 4, wherein the method further comprises:30 applying a corresponding inverse transform function to the estimated time-frequency representation of the first signal, to thereby obtain an estimated first signal from the audio signal; and applying the inverse transform function to the estimated time-frequency representation of the second signal, to thereby obtain an estimated second signal from the audio signal.PCT / US25 / 47728 24 September 2025 (24.09.2025)D24140W0016. The method according to claim 4 or 5, wherein the transform function comprises one of: a short time Fourier transform, STFT, a modified discrete X transform, MDXT, or a filterbank based transform such as complex quadrature mirror filterbank, CQMF.
57. The method according to any one of the preceding claims, wherein the neural network model is a deep neural network, DNN, based model.
8. The method according to claim 7, wherein the DNN based model comprises a first10 two-stage mask estimation modules and a second two-stage mask estimation module configured for respectively estimating the first mask and the second mask.
9. The method according to claim 8, wherein at least one of the first and second two- stage mask estimation modules is configured for estimating, at a first stage, a magnitude mask, and at a second stage, a complex mask.
10. The method according to claim 9, wherein applying the first mask involves: applying, at the first stage, the magnitude mask to the time-frequency representation to obtain a modified time-frequency representation; and20 applying, at the second stage, the complex mask to the modified time-frequency representation, in order to obtain the estimated time-frequency representation of the first signal.
11. The method according to claim 10, wherein applying the magnitude mask to the time-frequency representation to obtain the modified time-frequency representation comprises: applying the magnitude mask to magnitude values of the time-frequency representation to obtain a modified magnitude spectrum; and obtaining the modified time-frequency representation having a modified complex¬30 valued spectrum based on the modified magnitude spectrum and a phase spectrum corresponding to the time-frequency representation.PCT / US25 / 47728 24 September 2025 (24.09.2025)D24140W00112. The method according to claim 10 or 11, wherein applying the magnitude mask and / or the complex mask comprises deep filtering that involves weighted averaging among neighboring time-frequency bins.5 13. The method according to claim 12 when depending on claim 11, wherein the deep filtering used for applying the magnitude mask at the first stage is implemented according to:where Y^ (t, f) denotes the modified magnitude spectrum, MS1(t, denotes the magnitude mask, | Y (t, ) | denotes the magnitude spectrum corresponding to the timefrequency representation, t and f represent time and frequency indices, and U1, U2, and 72are predetermined values indicative of a to-be-estimated size of the magnitude mask.
14. The method according to claim 13, wherein the modified complex-valued spectrum YS1(t, ) is obtained according to:where operator (p is used for calculating an argument of a complex number.
15. The method according to claim 14, wherein the deep filtering used for applying the complex mask at the second stage is implemented according to:where YS2it, f) denotes the estimated time-frequency representation of the first signal, MS2it, f) denotes the complex mask, and M , M2, N and N2are predetermined values indicative of a to-be-estimated size of the complex mask.
16. The method according to any one of claims 10 to 15, wherein at least one of the first and second two-stage mask estimation modules comprises: a first activation layer configured to, at the first stage, limit the magnitude mask to a value range of 0 to 1 ; and a second activation layer configured to, at the second stage, limit the complex mask to30 a value range of -1 to 1.
17. The method according to any one of claims 8 to 16, wherein the DNN based model comprises a processing chain comprising, in this order, a feature extraction module, followedPCT / US25 / 47728 24 September 2025 (24.09.2025)D24140W001 by an encoder module, followed by a decoder module, and followed by the first and second two-stage mask estimation modules.
18. The method according to any one of the preceding claims, wherein the neural5 network has been trained based on at least one training pair, each training pair comprising, on the one hand, a first reference signal and a second reference signal as ground truth, and on the other hand, a corresponding audio mixture signal.
19. The method according to any one of the preceding claims, wherein the neural10 network has been trained based on a loss function.
20. A method of training a neural network for use in audio object estimation, the method comprising: obtaining an audio mixture signal, and corresponding first and second reference signals of different types that are used as ground truth; obtaining a first set of bin features from the audio mixture signal, a second set of bin features from the first reference signal, and a third set of bin features from the second reference signal, wherein each bin feature in the first to third sets corresponds to a respective time-frequency bin of the audio mixture signal, the first reference signal and the second20 reference signal; feeding the first set of bin features to the neural network to estimate a first mask and a second mask respectively corresponding to the first reference signal, and the second reference signal; obtaining a fourth set of bin features based on the first mask and the first set of bin features; obtaining a fifth set of bin features based on the second mask and the first set of bin features; evaluating a loss function based on the first to fifth sets of bin features; and updating parameters of the neural network based on a result of the evaluation.3021. The method according to claim 20, wherein the first reference signal is a reference speech signal and the second reference signal is a reference music signal.PCT / US25 / 47728 24 September 2025 (24.09.2025)D24140W00122. The method according to claims 20 or 21, wherein obtaining the first set of bin features from the audio mixture signal comprises: transforming the audio mixture signal by using a transform function to obtain a transformed audio signal; and5 extracting a respective bin feature from the transformed audio signal for each timefrequency bin to obtain the first set of bin features; wherein obtaining the second set of bin features from the first reference signal comprises: transforming the first reference signal by using the transform function to obtain a10 transformed first reference signal; and extracting a respective bin feature from the transformed first reference signal for each time-frequency bin to obtain the second set of bin features; and wherein obtaining the third set of bin features from the second reference signal comprises: transforming the second reference signal by using the transform function to obtain a transformed second reference signal; and extracting a respective bin feature from the transformed second reference signal for each time-frequency bin to obtain the third set of bin features.20 23. The method according to claim 22, wherein the transform function comprises one of: a short time Fourier transform, STFT, a modified discrete X transform, MDXT, or a filterbank based transform such as complex quadrature mirror filterbank, CQMF.
24. The method according to any one of claims 20 to 23, wherein the first mask comprises values indicative of ratios of the first reference signal in the audio mixture signal for respective time-frequency bins, and the second mask comprises values indicative of ratios of the second reference signal in the audio mixture signal for respective time-frequency bins.
25. The method according to any one of claims 20 to 24, wherein the neural network30 model is a deep neural network, DNN, based model.
26. The method according to claim 25, wherein the DNN based model comprises a first two-stage mask estimation module and a second two-stage mask estimation module configured for respectively estimating the first mask and the second mask.PCT / US25 / 47728 24 September 2025 (24.09.2025)D24140W00127. The method according to claim 26, wherein at least one of the first and second two-stage mask estimation modules is configured for estimating, at a first stage, a magnitude mask, and at a second stage, a complex mask.
528. The method according to claim 27, wherein obtaining the fourth set of bin features based on the first mask and the first set of bin features involves: applying, at the first stage, the magnitude mask to the first set of bin features to obtain a modified first set of bin features; and10 applying, at the second stage, the complex mask to the modified first set of bin features, to obtain the fourth set of bin features.
29. The method according to claim 28, wherein applying the magnitude mask to the first set of bin features to obtain the modified first set of bin features comprises: applying the magnitude mask to magnitude values of the first set of bin features to obtain a modified magnitude spectrum; and obtaining the modified first set of bin features having a modified complex-valued spectrum based on the modified magnitude spectrum and a phase spectrum corresponding to the first set of bin features.2030. The method according to claim 28 or 29, wherein apply ing the magnitude mask and / or the complex mask comprises deep filtering that involves weighted averaging among neighboring time-frequency bins.
31. The method according to claim 30 when depending on claim 29, wherein the deep filtering used for applying the magnitude mask at the first stage is implemented according to:where Y (t, ) denotes the modified magnitude spectrum, MS1(t, ) denotes the magnitude mask, | Y(t, f) | denotes the magnitude spectrum corresponding to the first set of30 bin features, t and f represent time and frequency indices, andU2, V1and V are predetermined values indicative of a to-be-estimated size of the magnitude mask.
32. The method according to claim 31, wherein the modified complex-valued spectrum YS1(t, f) is obtained according to:PCT / US25 / 47728 24 September 2025 (24.09.2025)D24140W001YS1(t,f) = where operator <p is used for calculating an argument of a complex number.
33. The method according to claim 32, wherein the deep filtering used for applying the5 complex mask at the second stage is implemented according to:where Ys'2(t, ) denotes a complex-valued spectrum corresponding to the fourth set of bin features, MS2(t, f) denotes the complex mask, and M15M2,and N2are predetermined values indicative of a to-be-estimated size of the complex mask.1034. The method according to any one of claims 28 to 33, wherein at least one of the first and second two-stage mask estimation modules comprises: a first activation layer configured to, at the first stage, limit the magnitude mask to a value range of 0 to 1 ; and15 a second activation layer configured to, at the second stage, limit the complex mask to a value range of -1 to 1.
35. The method according to any one of claims 26 to 34, wherein the DNN based model comprises a processing chain comprising, in this order, a feature extraction module,20 followed by an encoder module, followed by a decoder module, and followed by the first and second two-stage mask estimation modules.
36. The method according to any one of claims 20 to 35, wherein the loss function comprises a weighted sum of a first loss function corresponding to the first reference signal and a second loss function corresponding to the first reference signal.
37. The method according to claim 36, wherein the loss function Loss is implemented according to:30 where Lossp(S, S) denotes the first loss function, S denotes a spectrum of the first reference signal, S denotes an estimated spectrum after the first mask has been applied, Lossp(M, M) denotes the second loss function, M denotes a spectrum of the second referencePCT / US25 / 47728 24 September 2025 (24.09.2025)D24140W001 signal, M denotes an estimated spectrum after the second mask has been applied, and a is a predefined weighting coefficient between the first and second loss functions.
38. The method according to claim 36 or 37, wherein at least one of the first and5 second loss functions Losspcomprises a weighted sum of a respective perceptual loss function and a respective Mean Squared Error, MSE, loss function implemented according to:where Lossmdenotes the perceptual loss function, Losscdenotes the MSE loss function, and / ? is a predefined weighting coefficient.
39. The method according to claim 38, wherein for the first loss function, the perceptual loss function Lossm(S, S) is implemented according to:Lossm(S, S)~ = mdlff - diff - 1, wherethe MSE loss function LosscS,S) is implemented according to:where S denotes a spectrum of the first reference signal, S denotes an estimated spectrum after the first mask has been applied, p is a predefined spectral compression factor,20 operator <p is used for calculating an argument of a complex number, and m is a predefined tuning parameter.
40. The method according to any one of claims 36 to 39, wherein the loss function further comprises a residual loss function.
41. The method according to claim 40, wherein the loss function Loss is implemented according to:where Lossp(S, S) denotes the first loss function, Lossp(M, M) denotes the second30 loss function, and Lossp(^N, N^ denotes the residual loss function, whereinPCT / US25 / 47728 24 September 2025 (24.09.2025)D24140W001N = I - S - M, andN = I — S — M, where S denotes a spectrum of the first reference signal, S denotes an estimated spectrum after the first mask has been applied, M denotes a spectrum of the second reference5 signal, M denotes an estimated spectrum after the second mask has been applied, N denotes a spectrum of a residual signal, N denotes a spectrum of an estimated residual signal, / denotes a complex-valued spectrum of the audio mixture signal, and a, / 3 and y are predefined weighting coefficients.10 42. The method according to any one of claims 20 to 41, wherein updating the parameters of the neural network based on the evaluation comprises updating weights in the neural network.
43. An apparatus, comprising a processor and a memory coupled to the processor, wherein the processor is adapted to carry out the method according to any one of claims 1 to 42.
44. A program comprising instructions that, when executed by a processor, cause the processor to carry out the method according to any one of claims 1 to 42.2045. A computer-readable storage medium storing the program according to claim 44.
Citation Information
Patent Citations
Content-based audio stream separation
US20190206417A1