Method for separation of audio sources with different time-frequency characteristics

The neural network system addresses the challenge of poor audio source separation by using a feature extractor and separation blocks to generate accurate main and residual source masks, resulting in improved detection of residual audio sources.

WO2025136697A1PCT designated stage expired Publication Date: 2025-06-26DOLBY LABORATORIES LICENSING CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/058940
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-14
Filing Date
2024-12-06
Publication Date
2025-06-26

AI Technical Summary

Technical Problem

Existing audio source separation methods often struggle with poor separation accuracy, especially for audio sources with indefinite or time-varying frequency characteristics, leading to inaccurate classification of residual audio sources.

Method used

A computer-implemented neural network system that includes a feature extractor, separation blocks for generating main source masks, a residual mask extractor, and residual classifiers to predict residual source activity confidence metrics, thereby improving source separation accuracy.

Benefits of technology

The system achieves more accurate predictions of residual source activity by reducing information complexity and leveraging a shared feature extractor for efficient and scalable implementation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024058940_26062025_PF_FP_ABST
    Figure US2024058940_26062025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to a neural network system for residual source activity detection in an input audio signal. The neural network system comprising a feature extractor trained to predict a set of latent variables based on an input audio frame of the input audio signal and at least two separation blocks, configured to generate a main source mask for separating a respective main audio source. The neural network system further comprises a residual mask extractor, configured to form a combined source mask by combining all main source masks and determine a residual mask that complements the combined source mask and at least one residual classifier, comprising a neural network trained to predict a residual source activity confidence metric based on the residual mask, the residual source activity confidence metric indicating a likelihood of a residual audio source being active.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD FOR SEPARATION OF AUDIO SOURCES WITH DIFFERENT TIMEFREQUENCY CHARACTERISTICSCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims the benefit of priority from PCT Application No. PCT / CN2023 / 139771 filed on 19 December 2023, and U.S. Provisional Application No. 63 / 565,483, filed on 14 March 2024. each of which is incorporated by reference herein in its entirety.TECHNICAL FIELD

[0002] The present disclosure relates to a computer implemented neural network for performing source separation and / or classification and a method for training the computer implemented neural network.BACKGROUND

[0003] In the field of audio processing, audio source separation is a challenging problem encountered in many implementations. Audio source separation relates to separating one or more target audio sources from a mixed audio signal comprising a mix of different audio sources. Commonly, the target audio source is human speech and the mixed audio signal is a recorded audio signal containing audio content associated with a large variety of audio sources (e.g. wind noise, traffic noise, birdsong, or the sound of airplanes) in addition to speech. By separating the speech from the other audio sources a clean speech audio signal can be formed without interference wherein the clean speech audio signal can be used in isolation or boosted to increase speech intelligibility.

[0004] A basic form of speech separation can be realized by filtering the mixed audio signal with a bandpass filter tuned to t pical speech frequencies, such as frequencies from 150 Hz to 8 kHz. Recently, more sophisticated speech separation processing techniques have been developed relying on neural networks trained to predict a time-frequency gain mask for separating the speech from other audio content on a frame-by-frame basis.

[0005] Using similar methods, other types of audio sources can also be extracted from a mixed audio signal. For example, the bandpass filter can be tuned to, or the neural can be trained to, separate music, birdsong or other audio sources with distinguishable time frequency characteristics if desired.SUMMARY

[0006] A drawback with existing solutions for source separation is that separation accuracy often times can be poor. For example, a passband filter will allow any audio content falling withing the same frequency band as speech to pass through the filter which may have detrimental effects on the speech intelligibility. Some neural network based approaches solve this by predicting a more accurate and fine granularity gain mask for each frame. However, for certain types of audio sources, and especially audio sources having an indefinite or time-varying frequency characteristics, previous neural network based source separation methods still fail to provide accurate source separation, and even fail to provide accurate classification of whether certain types of audio sources are active or not.

[0007] It is a purpose of the present disclosure to a provide an improved computer implemented neural network, and a method for training the computer implemented neural network, that overcomes at least some of the drawbacks of existing solutions.

[0008] According to a first aspect, there is provided a computer implemented neural network system for residual source activity detection in an input audio signal comprising a mix of residual and main audio sources. The neural network system comprising a feature extractor, comprising a neural network trained to predict a set of latent variables based on an input audio frame of the input audio signal, the latent variables representing latent features of the input audio frame and at least two separation blocks, wherein each separation block is configured to generate a main source mask for separating a respective main audio source using a trained neural network configured to receive the set of latent variables. The neural network further comprising a residual mask extractor, configured to receive the respective main source mask of each separation block, form a combined source mask by combining all main source masks and determine a residual mask that complements the combined source mask and at least one residual classifier, comprising a neural network trained to predict a residual source activity confidence metric based on the residual mask, the residual source activity confidence metric indicating a likelihood of a residual audio source being active. The neural network system further comprises an output stage, configured to output the at least one residual source activity confidence metric.

[0009] By predicting main source separation masks and forming the residual mask using the main source masks the at least one residual classifier is presented with a reduced amount information enabling the residual classifier to perform more accurate predictions of whether a residual source is active. Additionally, the neural network system is efficiently implemented and scalable since the at least two source separators share a common feature extractor. Accordingly, additional source separators and / or residual classifiers may be added as needed.

[0010] A main audio source is an audio associated with a physical audio source that features at least one of a specific time-frequency pattern, a specific time or frequency envelope, a specific energy distribution or a specific frequency distribution in terms of e.g. bandwidth and band center frequency. For example, a main audio source may exhibit characteristic harmonics, specific dominating frequencies or a specific energy distribution. Examples of main audio sources include speech or birdsong.

[0011] A residual audio source is audio associated with a physical audio source which does not feature a specific time-frequency pattern, a specific time or frequency envelope, a specific energy distribution or a specific frequency distribution. Examples of residual audio sources include the sounds associated with a car or sound associated with chewing.

[0012] According to some implementations, the neural network in each separation block is trained to output a candidate mask for separating the main audio source, wherein each separation block further comprises a main classifier comprising a neural network trained to predict a main source activity confidence metric based on the candidate mask, and a source mask modifier configured to modify the candidate mask based on the main source activity confidence metric to form the source mask.

[0013] With a main classifier operating in each separation block, a main source activity' metric can be extracted which in turn is used to modify the candidate mask to make the resulting main source mask more accurate. For example, even when the main audio source associated with a separation block is inactive in an input audio frame the predicted candidate gain mask may still indicate, in error, that some audio content belongs to main audio source. By utilizing a predicted main source activity metric to adjust the candidate mask the consequences of this type of false positives can be mitigated.

[0014] For example, the source mask modifier is configured to make the candidate mask less inclusive in response to the main source activity confidence metric being below a predetermined threshold level.

[0015] Making an audio source separation mask less inclusive is understood as modifying the separation mask to include less audio content (in terms of spectral energy) when it is applied to an audio frame.

[0016] According to a second aspect of the invention there is provided a method for training the computer implemented neural network system according to the first aspect. The method comprises obtaining a plurality of training audio frames, each training audio frame comprising one of: a main audio source, a residual audio source, and a mix of a main and residual audio source and obtaining, for each training audio frame, ground truth data indicating main audio source frequency characteristics and a residual source activity confidence metric and providingthe training frame to the computer implemented neural network. The method further comprises, for each training audio frame, updating the learnable parameters of the at least two separation blocks using a first loss function based on the main source mask generated by each separation block and the frequency characteristics of the ground truth data; and updating the learnable parameters of the at least one residual classifier using a second loss function based on the source activity confidence metric predicted by the at least one residual classifier and the residual source activity confidence metric of the ground truth data.

[0017] According to a third aspect of the invention there is provided a method for training the computer implemented neural network system according to the first aspect. The method comprises obtaining a first set of training audio frames, each training audio frame in the first set comprising a main audio source and obtaining, for each training audio frame in the first set, first ground truth data indicating the frequency characteristics of the main audio source. The method further comprises providing the training frame of the first set to the computer implemented neural netw ork and updating, for each training audio frame, the learnable parameters of each separation block, w hile the learnable parameters of the at least one residual classifier is kept constant, using a first loss function based on the source mask generated by each separation block and the frequency characteristics of the first ground truth data. The method further comprises obtaining a second set of training audio frame, each training audio frame in the first set comprising a residual audio source, obtaining, for each training audio frame of the second set, second ground truth data indicating a residual source activity confidence metric, and providing each training audio frame of the second set to the computer implemented neural netw ork. The method further comprises updating the learnable parameters of the at least one residual classifier, while the learnable parameters of each separation block is kept constant, using a second loss function based on the source activity confidence metric predicted by at least one residual classifier and the residual source activity7confidence metric of the second ground truth data.

[0018] That is, the neural network system may be trained with the classifiers and separation blocks trained separately or trained together.

[0019] The frequency characteristics ground truth data indicates the time-frequency characteristics of each respective main audio source included in the training audio frame. The goal of the training is partly to enable each source separation block to predict a source separation mask that matches the respective frequency characteristics of the ground truth data. The residual source activity confidence metric of the ground truth data may be binary value indicating, for each training frame whether the residual audio source is active. The goal of the training is partly to enable each residual classifier to predict a residual source activity metric that matches the residual source activity metric of the ground truth data.DESCRIPTION OF THE DRAWINGS

[0020] Aspects of the present disclosure will be described in more detail with reference to the appended drawings, showing exemplary embodiments.

[0021] Figure 1 is a block diagram illustrating a neural network system according to some implementations.

[0022] Figure 2 is a block diagram illustrating a source separation block of the neural network system according to some implementations.

[0023] Figure 3a is an illustrative example of a series of candidate masks according to some implementations.

[0024] Figure 3b is an illustrative example of a senes of main source masks that have been formed by making the respective candidate mask less inclusive, according to some implementations.

[0025] Figure 4 is a block diagram illustrating a neural network system with feedforward of the input audio frame to the residual classifier and the source separation blocks, according to some implementations.

[0026] Figure 5 is a block diagram illustrating how the neural network system may be used together with an audio processor configured to process input audio frames based on the output provided by the neural network system, according to some implementations.

[0027] Figure 6 is a block diagram illustrating how training frames may be prepared for training the neural network system, according to some implementations.

[0028] Figure 7a is a flowchart illustrating a method for training the neural network system with the source separation blocks trained together with the residual classifier.

[0029] Figure 7b is a flowchart illustrating an alternative method for training the neural netw ork system, with the source separation blocks trained separately from the residual classifier.DETAILED DESCRIPTION

[0030] Systems and methods disclosed in the present application may be implemented as software, firmware, hardware or a combination thereof. In a hardware implementation, the division of tasks does not necessarily correspond to the division into physical units; to the contrary, one physical component may have multiple functionalities, and one task may be carried out by several physical components in cooperation.

[0031] The computer hardware may for example be a server computer, a client computer, a personal computer (PC), a tablet PC. a set-top box (STB), a personal digital assistant (PDA), a cellular telephone, a smartphone, an AR / VR wearable, automotive infotainment system, a web appliance, a netw ork router, switch or bridge, or any machine capable of executing instructions(sequential or otherwise) that specify actions to be taken by that computer hardware. Further, the present disclosure shall relate to any collection of computer hardware that individually or jointly execute instructions to perform any one or more of the concepts discussed herein.

[0032] Certain or all components may be implemented by one or more processors that accept computer-readable (also called machine-readable) code containing a set of instructions that when executed by one or more of the processors carry out at least one of the methods described herein. Any processor capable of executing a set of instructions (sequential or otherwise) that specify actions to be taken are included. Thus, one example is a typical processing system (e g., computer hardware) that includes one or more processors. Each processor may include one or more of a CPU, a graphics processing unit, and a programmable DSP unit. The processing system further may include a memory subsystem including a hard drive, SSD, RAM and / or ROM. A bus subsystem may be included for communicating between the components. The software may reside in the memory subsystem and / or within the processor during execution thereof by the computer system.

[0033] The one or more processors may operate as a standalone device or may be connected, e.g., networked to other processor(s). Such a network may be built on various different network protocols, and may be the Internet, a Wide Area Network (WAN), a Local Area Network (LAN), or any combination thereof.

[0034] The software may be distributed on computer readable media, which may comprise computer storage media (or non-transitory media) and communication media (or transitory7media). As is well know n to a person skilled in the art, the term computer storage media includes both volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Computer storage media includes, but is not limited to, physical (non-transitory) storage media in various forms, such as EEPROM, flash memory' or other memory7technology7, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store the desired information and which can be accessed by a computer. Further, it is well known to the skilled person that communication media (transitory ) ty pically embodies computer readable instructions, data structures, program modules or other data in a modulated data signal such as a carrier wave or other transport mechanism and includes any information delivery media.

[0035] Fig. 1 show s schematically' the architecture of a computer implemented neural net 'ork system 10 for source classification and / or source separation according to some implementations.The neural network system 10 obtains as an input audio frame representing a duration of audio content.

[0036] In general, any audio signal may be represented with a temporal sequence of frames, each frame representing a short duration of audio content. The duration of audio content represented by a single frame may vary’ from one implementation to another implementation and the frames may be non-overlapping or partially overlapping. For example, a frame comprises a waveform domain representation of a short duration of audio content or a frame comprises a plurality of Modified Discrete Cosine Transform (MDCT) samples which represent the original time domain representation of the audio signal in the frame. A frame can also be represented with a plurality of frequency bands, wherein each frequency band is associated with at least one value indicating the spectral energy of the respective frequency band. A sequence of such frames forms a timefrequency representation indicating the spectral magnitude in each frequency’ band over time (i.e. over multiple frames). A sequence of two or more frames will be referred to as a segment.

[0037] For example, a segment may comprise 200 consecutive frames with 50% overlap, wherein each frame comprises 64 (e.g. octave) frequency bands for a total of 200*64=12800 bins in the time-frequency representation of the segment. In some implementations, a 2048 point FFT is used to convert 48 kHz sampled audio content into the time-frequency representation.

[0038] At least one input frame is used as input to the neural network system 10. It is however also investigated that a segment, i.e. two or more subsequent frames are used as input. In the following it will be assumed that the input to the neural network segment is one input audio frame, however, it is understood that the neural network system 10 may be configured to operate on two or more input audio frames (i.e. an input audio segment) instead. For example, a frame represents the (log) spectral magnitude of its audio content in each of a plurality of frequency bands.

[0039] Each frame may therefore form a one dimensional time-frequency representation spanning a frame duration of time and a plurality of frequency bands. In general, the timefrequency representation may be continuous in both time and / or frequency but in many applications the time-frequency representation is discrete and represented with a plurality of bins spanning the different frequency bands. Each bin represents a portion of the frequency (e.g. a single frequency band) at a time portion associated with the frame.

[0040] In one implementation, each frame is originally represented with 1024 frequency bins whereby the 1024 frequency bins are grouped into 64 bands by matrix multiplication with a banding matrix of dimensions 1024*64 prior to providing the frame to neural network system 10.

[0041] The above segment and frame details are merely exemplary, and it is understood that many different types of frames and segments may be used, having e.g. a different number of number of frequency bands and spanning longer or shorter temporal durations.

[0042] The input audio frame is provided to a feature extractor 1 comprising a plurality of neural network layers trained to generate a set of latent variables based on the input audio frame. The latent variables represent so called “hidden” or latent features of the input audio frame and, in general, the latent variables cannot be interpreted or analyzed outside the neural network system 10. The latent variables, and the features they represent, are defined by the training of the neural network system 10 whereby the neural network of the feature extractor 1 has learned to generate latent variables that are beneficial for subsequent neural network layers in the neural network system 10 and that enable the neural network system 10 as a whole to achieve good accuracy.

[0043] In some implementations, the input audio frame comprises a plurality7of bins spanning a frequency dimensions (e.g. 64 frequency bands). In general, feature extractor 1 may be configured to output latent features having any dimensions wherein the latent features have a higher dimensionality, lower dimensionality or the same dimensionality7as the input audio frame. In one example, the feature extractor 1 is configured to output latent variables having a smaller dimension in frequency compared to the input audio frame. The latent variables may e.g. consist of one vector for each frame wherein each vector has a number of elements that is smaller than the number of frequency bands in the input audio frame. In this way, the temporal accuracy is maintained while the feature extractor 1 learns to predict, for each frame, lower dimension frequency features.

[0044] The latent variables are provided to each of N source separation blocks 2a, 2b, wherein N > 2. Each source separation block 2a, 2b is associated with a respective main audio source.

[0045] A main audio source is a type of audio source which features expected, predetermined and / or well-defined time-frequency characteristics. A main audio source may alternatively be referred to as being distinguishable, as opposed to indistinguishable, meaning that it is possible for a human listener, or a trained neural network, to identify the type of main audio source even if a very short sample of the main audio source is presented to the listener or neural network.

[0046] For example, a main audio source is an audio source associated with a specific spectral energy distribution, specific harmonics or specific frequencies.

[0047] Human speech is one example of a main audio source. Human speech features a well defined spectral energy distribution and specific harmonics that makes it possible for a human listener, or a trained neural network, to identify human speech even if a very short audio clip (for example a single frame) containing speech is presented to the human listener or neural network.

[0048] Another example of a main audio source is animal sounds such as birdsong, or the sounds made by other animals such as cicadas, cats, dogs or various other animals. Like human speech, animal sounds often feature as specific spectral energy distribution and harmonics meaning that it is possible for a human listener, or a trained neural network, to identify the animal sound even if a very short audio clip containing speech is presented to the human listener or neural network. Especially, cicada sound is very characteristic and easy to identify due to the specific mix of frequencies involved.

[0049] Yet another example of a main audio source is the sounds associated with various musical instruments, such as guitar, drums, bass or violin. Each musical instrument generates a sound with well defined harmonics or specific frequencies making it possible for a human listener, or trained neural network, to identify the musical instrument even if only presented with a short audio clip with a musical instrument.

[0050] Yet another example of a main audio source is the sound associated with wind noise. Wind noise, while exhibiting some random behavior, often features a characteristic energy distribution over lower frequencies, such as from about 50 Hz to about 2 kHz. Another example of a main audio source is the sound of rain.

[0051] Each separation block 2a, 2b comprises a trained neural network that obtains the latent variables as an input and, using the output of the trained neural network, a main source mask Maski, Mask2 for separating the main audio source is predicted and outputted by each source separation block 2a, 2b. That is, the first source separation block 2a outputs a first main source mask Maski for separating a first main audio source type and the second source separation block 2b outputs a second main source mask Mask2 for separating a second, different, main audio source type. More than two source separation blocks 2a, 2b may be provided and in general N source separation blocks may be provided, each of the N source separation blocks predicting a respective main source mask Maski, Mask2, . .. , M ask N for a respective, different main audio source.

[0052] That is, the main audio source associated with each source separation block is a different, and unique, compared to the other main audio sources. As an example, the first main audio source associated with the first source separation block la is speech and the second main audio source associated with the second source separation block lb is birdsong. As another example, the first main audio source associated with the first source separation block la is female speech and the second main audio source associated with the second source separation block lb is male speech.

[0053] Each main source mask Maski, Mask2, . . . , Maskx may be a frequency gain mask carrying a gain value for each of a plurality frequency bins representing the input audio frame.That is. each predicted main source mask Maski, Mask2, , MaskN may feature the same frequency resolution as the input audio frame whereby each predicted main source mask Maski, Mask2, . . . , MaskN can be applied directly to the input audio frame. Alternatively, it is also envisaged that at least one, e g. each, predicted main source mask Maski, Mask2, . .. , MaskN may feature a frequency resolution that is different from that of the input audio frame wherein the at least one predicted main source mask Maski, Mask2, ... , MaskN is upsampled or downsampled to match the frequency resolution of the input audio frame. In some implementations, each main source mask Maski, Mask2, ... , MaskN is one dimensional, spanning multiple frequency bins wherein each frequency bin is associated with the same time stamp. In some implementations, when the input to the neural network system is two or more audio frames, each main source mask Maski, Mask2, . . . , MaskN is two dimensional, spanning multiple frequency bins across two or more time stamps (one time stamp for each input audio frame).

[0054] Each gain mask is one out of two types of gain masks. Either the gain mask features high gain values for frequency bins associated with the main audio source (type A) or the gain mask features high gain values for all frequency bins associated with audio content which is different from the main audio source (type B). While there are differences regarding how the different types of gain masks are to be implemented to achieve source separation, both types are said to separate the main audio source type from other background audio content which does not belong to the main audio source type.

[0055] A ty pe A gain mask includes higher gain values for frequency bins associated with the main audio source and lower gain values for frequency bins associated with audio content which does not belong to the main audio source. Accordingly, to perform source separation with the goal of isolating the main audio source the gain values of the type A gain mask may be applied to the input audio frame, whereby the main audio source content will be amplified whereas audio which does not belong to main audio source the is attenuated or silenced.

[0056] A ty pe B gain mask includes lower gain values for frequency bins associated with the main audio source and higher gain values for frequency bins associated with audio content which does not belong to the main audio source. Accordingly, to perform source separation with the goal of isolating the main audio source the gain values of the type B gain mask can be individually inverted and then applied to a frequency bin representation of the input audio frame, whereby the main audio source content will be amplified and audio content which does not belong to the main audio source is attenuated or silenced. Alternatively, the type B gam mask is applied directly the input audio frame, whereby the main audio source content will be attenuated or silenced and the audio content which does not belong to the main audio source is amplified. This forms a background audio frame which includes all audio content which does not belong tothe main audio source. By subtracting the background audio frame from the input audio frame an audio frame with the main audio source isolated can be obtained.

[0057] That is, irrespective of the type of gain mask predicted by the source separation blocks each gain mask is configured to separate the main audio source from any background audio content that does not belong to the main audio source.

[0058] Assuming that each main source mask is a ty pe A mask the main source masks Maski, Mask2, . . . , Maskx are provided to a residual mask extractor 3 which collects the main source masks Maski, Mask2, .. . , Maskx and forms a combined source mask Maskcom by combining all main source masks Maski, Mask2, ... , Maskx. For example, the residual mask extractor 3 performs elementwise (i.e. bin-wise) summation of all main source masks Maski, Mask2, ... , Maskx to form the combined mask Maskcom. The combined mask Maskcom will therefore be configured to separate all of the main audio sources from the input audio frame from any audio content not belonging to any of the main audio sources.

[0059] As an example, main source mask Maski separates speech and main mask Mask2 separates birdsong. Accordingly, applying main source mask Maski to the input audio frame will keep speech but suppress birdsong as well as any additional background audio content (if present) and applying main source mask Mask2 to the input audio frame will keep birdsong but suppress speech as well as any additional background audio content (if present). The combined mask Maskcom will in this scenario capture both speech and birdsong while any background audio content, if present, is excluded.

[0060] Using the combined mask Maskcom, the residual mask extractor 3 extracts a residual mask Mask being the complement to the combined mask Maskcom. The complement of a gain mask may be defined as a gain mask that includes all audio content excluded by the original gain mask, and vice versa. For example, for a binary gain mask where each bin is associated with either “1”, indicating inclusion of the bin, or “0”, indicating exclusion of the bin the, the complement to this mask is a mask with the values in each bin inverted. As another example, for a gain mask where each bin is associated with a value ranging from 0 to 1 with 1 indicating complete inclusion of a bin, 0 indicating complete exclusion of a bin and values between 0 and indicating partial inclusion of a bin the complement to this mask is a mask with 1 - X in each bin, with X being the original value of the bin.

[0061] If each main source mask is a type B main source mask, the residual mask MaskRmay be formed by combining the type B main source masks directly.

[0062] The residual mask MaskR, being the complement to the combined mask Maskcom, is provided to at least one classifier 4a, 4b comprising a neural network trained to predict a sourceactivity metric SAMN+I indicating the confidence of a respective residual audio source being active.

[0063] Each residual audio source is different from the main audio sources associated with each of the separation blocks 2a, 2b. The residual audio source types are associated with audio sources which are not easily distinguishable by a human listener or neural network due to the residual audio sources being associated with a time-varying or non-specific audio characteristics. That is, the residual audio sources do not exhibit an easily definable time-frequency pattern in terms of spectral energy, any clearly distinguishable harmonics or any specific frequencies.

[0064] One example of a residual audio source is motor vehicle sounds, such as sounds associated with a car. The sound a car makes will vary depending on the type of propulsion system (e.g. petrol, diesel, electric or hybrid) used by the car, whether the car is moving or parked and, if the car is moving, the shape of the car which influences the air turbulence caused by the car, the types of car tires used and the level of strain of the engine. For example, if a petrol or diesel car is idling it makes a mechanical engine sound that may be indistinguishable from a lawnmower engine sound or the sound of some other mechanical machine, such as a compressor or pump. On the other hand, when a car passes by at high speed the engine sound is usually different compared to when the car is idling and there is also the addition of noise caused by air turbulence created by the car as it goes by and the sound of the tires contacting the ground.

[0065] Additional examples of residual audio sources are sounds associated with chewing, sounds associated with footsteps, sounds associated with aircrafts, sounds associated with printers, and sounds associated with a door opening or closing. For example, the sound of chewing could vary dramatically depending on the type of food chewed and the rate of the chewing. As another example, the sound of footsteps could vary dramatically depending on the step pace, the type of shoe and the properties of the ground (which e g. could be dry, wet, covered with leaves or covered with snow).

[0066] Accordingly, it is challenging to predict source separation masks for residual audio sources since the physical audio source responsible for generating the sounds may emit a wide variety of different sounds that are not easily capturable by a single source separation mask.

[0067] With the neural network system of fig. 1 the main source masks Maski, Mask2, ... , Maskx for the main audio sources have been eliminated in the residual mask MaskR which is provided to the at least one residual classifier 4a, 4b. This means that information relating to the audio content which with high accuracy has been determined to belong to a main audio source, has been eliminated to present the residual classifier(s) 4a, 4b with less information. In turn, this enables the residual classifier(s) 4a, 4b to make a more accurate prediction of the likelihood ofthe residual audio source being active since less information will obscure the residual audio source.

[0068] Each residual classifier 4a, 4b has been trained to output a residual audio source confidence metric which may be a value between 0 and 1, with 0 indicating that the residual audio source is not active (i.e. likelihood of being active is 0%). 1 indicating that the residual audio source is active (i.e. likelihood of being active is 100%) and values between 0 and 1 indicating the likelihood of the residual audio source being active. Each residual classifier 4a, 4b comprises a plurality of neural network layers wherein the first layer is configured to receive the residual gain mask Mask as an input and the last layer is configured to output a single parameter, the residual source activity metric SAMN+I for the associated residual audio source.

[0069] In some implementations, the residual classifier 4a, 4b is configured to obtain multiple consecutive residual masks Maska, each representing a respective input audio frame, and predict a single source activity metric SAM for all consecutive residual masks Maska. That is, the residual classifier 4a. 4b may take more than one residual Maska into account when determining the source activity metric SAM whereby the source activity metric is said to be segment based since it is based on more than one frame that together form a segment. A segment based source activity metric SAM may still be determined for each new' frame that is input to the neural network. For example, the residual classifier 4a. 4b may be configured to receive a current residual mask Maska, and a predetermined number of past residual masks Maska, whereby when a new' residual mask Maska becomes availible, the new residual mask Maska becomes the current residual mask Maska.

[0070] Considering more than one residual mask Maska may be used in combination when the residual classifier 4a. 4b is implemented as a convolutional neural network. CNN.

[0071] In some implementations, each residual classifier 4a, 4b comprises a recurrent neural network, RNN. In an RNN the output from at least one node in the network may influence the subsequent input to the same node. This means the RNN features a type of memory' which in effect increases the temporal context window the residual classifier is aware of. With RNN implementations it may be sufficient that the residual classifier 4a, 4b obtains only one current residual mask MaskR, since the inherent memory properties of the RNN may increase the temporal context of the residual classifier is aw are of 4a, 4b. However, it is also envisaged that more than one residual mask MaskR is input to an RNN implementation of the residual classifier 4a, 4b.

[0072] In some implementations, the last layer of the plurality of neural netw ork layers that constitutes the residual classifier 4a, 4b has sigmoid activation.

[0073] In some implementations, the residual source activity metric SAMN+I of each, at least one residual classifier 4a, 4b is the output of the neural network system 10. Accordingly, the neural network system 10 accomplishes accurate activity detection for residual audio sources which otherwise are difficult to detect accurately. In some implementations, at least two residual classifiers 4a, 4b are used, each predicting a respective residual source activity metric SAMN+I, SAMN+2. In general, M > I residual classifiers 4a, 4b may be present each predicting an individual residual source activity metric SAMN+I, SAMN+2, ... SAMN+M for an individual residual audio source.

[0074] Optionally, the neural network system 10 outputs additional information as well, such as each of the predicted source masks Maski, Mask2, . .. , Mask\. The outputs of the neural network system 10 may be used in a wide variety7of audio processing tasks as will be described below, in connection to fig. 5.

[0075] Turning to fig. 2 the details of source separation block 2a is shown, according to some implementations. It is envisaged that each source separation block 2a, 2b in the neural network system 10 of fig. 1 comprises the corresponding conceptual parts as source separation block 2a albeit tuned for a different main audio source.

[0076] As mentioned above, the source separation block 2a comprises a neural network 21 comprising a plurality of neural network layers. The neural network 21 is trained to obtain as an input the latent variables associated with an input audio frame, and predict, based on the latent variables, a candidate mask Maskcan for separating the main audio source in the input audio frame.

[0077] The candidate mask Maskcan is provided to a main source classifier 22 which, similar to the residual classifier described above, comprises a neural network trained to predict a source activity metric SAMi. Specifically the main source classifier 22 is trained to predict a main source activity metric SAMi indicating the likelihood of the main source associated with the separation block 2a being active. The main source classifier 22 predicts the main source activity metric SAMi based at least on the candidate mask Maskcan output by the neural network 21.

[0078] Besides the fact that the main source classifier 22 is trained to predict a main source activity7metric SAMi associated with a main audio source, instead of a residual audio source, the main source classifier 22 may otherwise be equivalent to the residual classifier. That is, the main source classifier 22 may according to some implementations comprises an RNN and / or is configured to output a single parameter being a value between 0 and 1.

[0079] The main source activity7metric SAMi is provided to a mask modifier 23 alongside the candidate mask Maskcan. The mask modifier 23 is configured to modify the candidate mask Maskcan based on the main source activity metric SAMi to form the main source mask Maski.

[0080] In some implementations, the mask modifier 23 is configured to make the candidate mask Maskcan less inclusive for the main audio source, when the main source activity metric SAMi is below a predetermined threshold.

[0081] The process of making a (gain) mask more or less “inclusive’' is understood by the person skilled in the art as adjusting the gain values of each bin such that, when the mask is applied to audio content, more or less audio content (e.g. in terms of spectral energy) is included.

[0082] Turning to fig. 3a a sequence of four exemplary candidate masks Maskcan is show n as columns with each mask comprising four frequency bands. In practice, the candidate mask Maskcan has many more bins representing frequency bands (such as ten or more) and the 4 bin gam masks of fig. 3a are mainly for illustration purposes.

[0083] Each bin in in each candidate mask Maskcan has a gain value and in this example the gain value is expressed in linear units w ith 0 indicating a complete silencing of the bin and 1 indicating that a linear gain of unity is applied. It is understood that other gain values can also be used in the same manner (e.g. a scale from 0 - 100, or decibels).

[0084] As seen in the exemplary candidate mask Maskcan of the first column in fig. 3a, the bins have mostly zeros. For a type A mask this indicates that these bins should be silenced since they do not contain audio content that belongs to the main audio source. On the other hand, the lower frequency bins of the third column mask are associated with high values, such as 1, and for a type A mask this indicates that these bins should be kept since they do contain audio content that is predicted to be associated w ith the main audio source. For a type B mask the situation is reversed wherein low' values indicate bins associated with the main source and vice versa for the high values.

[0085] In general, even if the main audio source is not active in an input audio frame, the associated candidate gain mask Maskcan will not be all zeros (or ones for a type B masks) meaning that for some input audio frames the candidate gain mask Maskcan is inaccurate since it indicates the main audio source is present in some bins despite this not being the case.

[0086] To this end, the mask modifier 23 takes each candidate mask and makes it less inclusive if the main source activity metric SAMi is below a predetermined threshold. For a type A mask the process of making the candidate mask Maskcan less inclusive involves decreasing the predicted gain value of each bin. In fig. 3b the result after decreasing the gain value of each bin in each candidate mask Maskcan to form the main source mask Maskl is shown. In this this example, the gain value in each bin of each bin has been decreased by subtracting 0.50 from each predicted gain value (and capping the minimum gain to 0). When each main source mask Maski is applied to the corresponding input audio frame less audio content (less spectral energy ) will be left meaning that the mask has been made less inclusive.

[0087] For a type B mask the mask modified may instead increase the gain value of each bin whereby the type B mask becomes less inclusive when used to separate the main audio source. Tn fact, the type B mask becomes more inclusive for capturing background audio content which has as a consequence that it becomes less inclusive for the main audio source.

[0088] The above processing of the mask modifier is merely exemplary. It is envisaged that the mask modifier may e.g. decrease / increase the gain value of each bin in accordance with a function (e.g. linear function, polynomial function or exponential) of the main source activity metric SAMi. Alternatively, in some implementations, if the main source activity metric SAMi is below a predetermined, optionally source specific, threshold (such as 0.5 or 0.3) for a candidate mask Maskcan the mask modifier 23 sets the gain of all bins of the mask to 0, effectively silencing the mask.

[0089] Alternatively or additionally, the process of making the masks less inclusive is controlled based on a respective main source signal-to-noise ratio, SNR. As also shown in fig. 2 the mask modifier 23 may receive an SNR indicating the (log) spectral power of the main audio source as it appears after it has been separated by the candidate mask Maskcan, corresponding to a signal component, with respect to the (log) spectral power of the input audio frame, with the main audio source as it appears after it has been separated by the candidate mask Maskcan removed, corresponding to a noise component. If the signal component is much smaller than the noise component this may be taken as an indicator that the corresponding mask should be silenced.

[0090] A global SNR threshold may be set for all sources, or an individual SNR threshold may be set for each source individually wherein, if the SNR is below the SNR threshold for the respective source, the associated candidate mask Maskcan is silenced. For example, when the SNR is determined for log spectral power levels examples of appropriate SNR thresholds are - 5dB for birdsong, 10 dB for other animal sounds, 0 dB for wind sounds and 0 dB for rain sounds.

[0091] In some implementations, the main source activity metric SAMi of each source separator is output by the neural network system alongside the residual source activity metric(s) SAMN+I, SAMN+2. Hereby, an audio processing system may utilize the source activity metrics of both main and residual audio objects to perform or select an audio processing to be applied. For example, the source activity metrics SAMi, SAM2, ... SAMN, SAMN+I, ... SAMN+M of both main and residual audio sources may be used to determine an acoustic scene.

[0092] In some implementations, the candidate mask Maskcan is used as the main source mask Maski, whereby the classifier 22 and / or the mask modifier 23 may be omitted. In such implementations, the processing in each separation block 2a is made less complex at the cost of some non-silent main source masks being outputted even when the associated main source is not active.

[0093] As also is shown in fig. 2 the main source classifier 22 according to some implementations may be configured and trained to predict the main source activity metric SAMi based on the input audio frame in addition to the predicted candidate mask Maskcan. Accordingly, in some cases when the neural network 21 in the source separation block 2a is strained, and the predicted gain mask is less reliable, the main source classifier 22 may still achieve accurate prediction of the main source activity metric SAMi by also taking the input audio frame into account. In some implementations, the input audio frame and the candidate mask Maskcan have the same frequency resolution whereby the two can be summed and the sum provided to the neural network 21.

[0094] In fact, the input audio frame may alternatively, or in addition, be provided to each of the residual classifiers 4a, 4b as well as shown in fig. 4 illustrating an alternative neural network system 10’ where the input audio signal is fed forward and used as input to each residual classifer 4a, 4b and / or each main classifier in each source separation block 2a, 2b. In this implementation, the accuracy of also the residual source activity metric SAMN+I is further enhanced, especially for cases where one or more of the predicted main source masks Maski, Mask2, . . . , Maskw is not reliable.

[0095] Fig. 5 is a block diagram illustrating how the neural network system 10 can be used together with an audio processor 5 to process an input audio frame based on the output of the neural network system 10, wherein the output comprises at least one of: at least one main source mask Maski, Mask2, . . . , Mask at least one main source activity metric SAMi, SAM2, . . . , SAMN and at least one residual source activity' metric SAMN+I, SAMN+2,... , SAMN+M.

[0096] An input audio frame is provided to the neural network system 10 which in turn generates the above described output whereby the output is provided to the audio processor 5.

[0097] The audio processor 5 is configured to perform audio processing on the input audio frame to form a processed audio frame, wherein the audio processing is controlled by the output of the neural network system 10.

[0098] In one exemplary implementation, the audio processor 5 obtains at least one main source mask Maski, Mask2,... , Maskw and applies the main source mask Maski, Mask2,... , Maskw to the input audio frame to form the processed audio frame. For example, the at least one main source mask Maski, Mask2,.. . , Maskw comprises a speech mask whereby the intelligibility' of speech is enhanced in the processed audio frame by application of the speech mask to isolate speech. As another example, the at least one source mask Maski, Mask2,... , Maskw comprises a birdsong mask whereby the audio processor is configured to output a multi-channel, or object based, processed audio frame wherein the audio processor extracts the birdsong and places it in apredetermined channel or at a predetermined object location (such as in a height channel or at an object location above the listening position).

[0099] In another exemplary' implementation, the audio processor 5 obtains at least one main audio source activity metric SAMi, SAM2, , SAMN and / or residual audio source metric SAMN+I, SAM N+2, . .. , SAM N+M and determines, based on the main / residual audio source metric SAMi, . . . , SAMN+M an acoustic scene. The audio processor 5 then selects an audio processing scheme based on the acoustic scene and uses the selected audio processing scheme to process the input audio frame to form the processed audio frame. For example, if source activity metric associated with birdsong and wind noise (two main audio sources) and car sounds (a residual audio source) are relatively high the acoustic scene may be established to be an outdoor urban scene. On the other hand if the source activity metric of birdsong and wind noise are relatively high and the source activity' metric of car sounds is relatively low the acoustic scene may be determined as an outdoor nature scene. Each acoustic scene ty pe may be associated with its own individual audio processing scheme. For example, the aggressiveness of a noise suppression algorithm may be tuned based on the acoustic scene with less aggressive noise suppression in an outdoor nature scene compared to an outdoor urban scene.

[0100] The neural network system 10 is not limited to use cases with an audio processor and it is envisaged that the neural network system 10 can be used in many other applications as well. For example, source activity metrics SAMi, ... , SAMN+M of the neural network system 10 can be used to label audio frames for the purposes of generating training data for neural networks or machine learning purposes.

[0101] With reference to fig. 6 and fig. 7a a method for training the neural network system will now be described.

[0102] Firstly, since the neural network system may learn to predict a mask (and optionally source activity7metric) for an arbitrary' number N of main audio sources and a source activity' metric of an arbitrary number M of residual audio sources the training data used for training the neural network system should comprise examples spanning all N + M main and residual audio sources. The training data may e.g. be extracted from publicly available data sets, created synthetically or recorded.

[0103] In some implementations, multiple audio frames, each containing audio representing at least one main or residual source, are collected from a public dataset or recording. A portion of the multiple audio frames are labeled manually indicating which mam or residual audio source is active in which audio frame. Optionally, an audio file comprising multiple audio frames is labeled and using timestamps it is indicated which audio source is active at what times of the audio file which in turn can be mapped to individual audio frames. Additionally, a manualdenoising process is performed such that each audio frames comprises a mix of main and residual audio sources with as little background noise as possible.

[0104] Since a large number of training frames are needed the manually labeled audio frames are used to train a source classifier. To increase the robustness of the source classifier, the manually labeled audio frames are augmented to increase the amount of training data availible for training the source classifier.

[0105] The source classifier is then used to classify the remaining portion of the multiple audio frame to obtain a large number of labeled audio frames, where some are manually labeled, and some are labeled by the source classifier. In some implementations, a wiener filter is also designed to denoise the audio frames labeled by the source classifier.

[0106] After this process, a large number of audio frames may be available with each frame comprising at least one clean example of a residual or main audio source. With reference to fig. 6 it will now be described how these training audio frames can be combined to form training audio frames suitable for training the neural network system.

[0107] In some implementations, one residual audio source or main audio source is the most important audio source to accurately detect and / or predict an accurate separation mask for. Commonly, human speech is prioritized but it is understood that human speech may be replaced with any other audio source depending on the application. In fig. 6, human speech is prioritized and in a first database 6a audio frames containing human speech is stored.

[0108] To create a training audio frame a first audio frame from the first database 6a is selected (e.g. randomly) and provided to random leveling unit 7a which applies a random gain to the first audio frame. A second audio frame is selected (e.g. randomly) from the second database 6b and provided to random leveling unit 7b which applies a random gain to the second audio frame. The second database 6b stores audio frames associated with any audio source (main or residual) other than the prioritized audio source in the first database 6a. For example, the first database 6a stores audio frames with human speech and the second database 6b stores audio frames comprising at least one of animal sounds, wind noise, car sounds, chewing and sounds that are all examples of main and residual audio sources.

[0109] The randomly gain adjusted first audio frame and the randomly gain adjusted second audio frame are provided to a first mixer 8a which mixes the randomly gain adjusted first audio frame with the randomly gain adjusted second audio frame. This results in a mixed audio frame carrying a mix of the prioritized audio source and at least one other audio source with the SNR of the prioritized audio source with respect to the at least one other audio source being randomized (due to random gains applied in the random levelling units 7a, 7b).

[0110] The mixed frame is then optionally provided to a second mixer 8b, together with a randomly leveled third audio frame selected (e.g. randomly) from a third database 6c. Alternatively, the output of the first mixer 8a is used as the training audio frame.[OHl] The third database stores audio frames comprising audio not associated with any main or residual audio source. For example, the third database stores audio frames containing noise, such as white noise or pink noise. As another example, if e.g. audio associated with aircraft is not used as a main or residual audio source aircraft sounds may be included in the audio frames of the third database.

[0112] Accordingly, at the second mixer 8b the mixed audio frame is mixed with the third audio frame such that the SNR of the mixed audio frame with respect to the third audio frame is randomized. The result of the mixing at the second mixer 8b is used as a training frame which carries a mixture of the prioritized audio source, one or more main and / or residual audio sources and additional audio content not associated with any of the main or residual audio sources.

[0113] This training data can therefore be used to train the neural network system such that it leams to distinguish between different audio sources even when the relative level of the different audio sources is randomized. Additionally, since the training data comprises additional audio content not associated with a main or residual audio source the neural network will leam to not merely divide the input audio into different complementary parts but also to disregard audio content not associated with a main or residual audio source.

[0114] For each training frame, ground truth data is also collected indicating which audio source is active in each training frame and, if a main audio source is active, the frequency distribution of the spectral energy of the main audio source(s) is recorded as frequency characteristics data. When training the neural network system, this ground truth data will be compared to the predicted output of the neural network system to determine a loss function which is used to update the learnable parameters of the neural network system. The ground truth source activity measure is compared to the predicted source activity metric of the neural network system and the predicted main source masks are compared to the frequency characteristics of the ground truth data. The purpose of the training is to obtain a neural network system that predicts main source masks that match the time-frequency distribution of the spectral energy and source activity7metrics that match the ground truth activity7.

[0115] In fig. 6 the randomized SNR levels are achieved using separate conceptual parts 7a, 7b, 7c for performing random levelling prior to each mixer 8a, 8b. It is however envisaged that this functionality can be integrated into each mixer 8a, 8b using e.g. a randomized mixing ratio in each mixer 8a, 8b.

[0116] Fig. 7a and fig. 7b are flowcharts illustrating alternative methods for training the neural network system 10 of fig. 1 . The method of fig. 7a involves training the residual classifier(s) 4a, 4b and the separation blocks la, lb together and the method of fig. 7b involves training the residual classifier(s) 4a, 4b and the separation blocks la, lb in sequence, with the separation blocks la, lb trained first and the residual classifier(s) 4a. 4b trained second.

[0117] With reference to fig. 1 and fig. 7a the method for training the residual classifier(s) 4a, 4b and the separation blocks 2a, 2b together will now be described. At step SI a training audio frame is obtained and at step S2 the training audio frame is input to the neural network system 10 via the feature extractor 1.

[0118] The source separation blocks 2a. 2b each predicts a mam source mask Maski, Mask2, ... , MaskN and at step S31 these main source masks Maski, Mask2, ... , Masks are accessed, and a first loss function is evaluated. Evaluating the first loss function comprises using the frequency characteristics of the ground truth data and the respective main source mask Maski, Mask2, .... MaskN to evaluate the loss function.

[0119] For example, the frequency characteristics of the ground truth data indicates, for each main source, a frequency energy distribution of the respective source and the first loss function is based on the difference between the predicted main source mask (or the predicted main source mask applied to the training audio frame) and the frequency energy distribution of the frequency characteristic. If the neural network is perfectly trained, the predicted main source mask Maski, Mask2, . . . , MaskN, when applied to the training frame, will yield the respective main audio source exactly as included in the training audio frame. In general, however, the neural network system 10 will not perform perfectly and there are always some differences between the main audio source as predicted by application of the predicted main source mask to the training audio frame and the frequency characteristics of the ground truth data whereby the learnable parameters of the at least the source separation blocks 2a, 2b can be updated at step S41 to reduce the errors.

[0120] The updating of the learnable parameters is governed by the first loss function, wherein if the loss function indicates a larger loss (larger error) the update step is increased compared to if the loss function indicates a smaller loss (smaller error). As the frequency characteristics and main source mask are frequency representations the first loss function evaluated at step S31 is a loss function suitable for describing errors distributed across multiple frequency bins. As an example, the first loss function is the mean absolute error. MAE across all frequency bins, the mean squared error, MSE, across all frequency bins or tuned extensions thereof.

[0121] In some implementations, only the learnable parameters of the associated separation block 2a, 2b (and optionally the feature extractor) are updated. Since each separation block 2a,2b is tasked with predicting a separate main source mask a large loss at one separation module 2a which has failed to predict an accurate main source mask Maski for its main audio source should not influence the learnable parameters of a second separation module 2b which is tasked with predicting a main source mask Maskz for some other type of main audio source.

[0122] The neural network system 10 will further predict at least one residual source activity metric SAMN+I, SAMN+2, SAMN+M, and at step S32 the predicted at least one residual source activity metric SAMN+I, SAMN+2, SAMN+M and a residual source activity metric of the ground truth data are used to evaluate a second loss function. The residual source activity metric of the ground truth data may be a binary value (e.g. obtained by manual labelling) indicating whether the residual audio source is active or not. Since the at least one residual source activity metric SAMN+I, SAM +2, SAMN+M is a single parameter (e.g. a value between 0 and 1) and the ground truth residual source activity7metric is a binary7parameter (e.g. either 0 or 1) the second loss function is any loss function suitable for describing the error between two such parameters. For example, the second loss function is binary cross entropy.

[0123] Based on the loss obtained from the second loss function, the learnable parameters of at least the residual classifier are updated to decrease the loss at step S42. If more than one residual classifiers 4a. 4b are used the second loss function is evaluated independently for each residual classifier, based on the residual source activity metric SAMN+I, SAMN+2, SAMN+M of each residual classifier and the corresponding ground truth source activity metric. The learnable parameters of at least each classifier 4a, 4b are then updated based on the respective loss determined with the second loss function.

[0124] In some implementations, it is only the learnable parameters of the respective residual classifier 4a, 4b which are updated at step S42. Alternatively, it is envisaged that the error can be propagated further back in the neural network system 10 to also update the internal weights of one or more separators 2a, 2b and / or the feature extractor 1.

[0125] Notably steps S31, S41 and steps S32, S42 may be performed in parallel all performed for the same training audio frame.

[0126] Additionally, it is noted that each training audio frame does not necessarily contain all main and residual audio sources as long as the total set of training audio frames used to train the neural network system spans all main and residual audio source types. For each training audio frame that misses one or more main audio source in the mix, the frequency characteristics of the ground truth data may indicate complete silence for these main audio sources since it is preferable if each separation block 2a, 2b leams to output completely silent main source mask (i.e. a minimum gain for each bin in a type A mask and a maximum gain for each bin in a type B mask) when its associated main audio source is not present in mix of the training audio frame.

[0127] With reference to fig. 1 and fig. 7b an alternative method for training the neural network system 10 will now be described. In a first training epoch El , first training frames are obtained at step SI ’ and at step S2’ the first training frames are provided to the neural network system 10. The purpose of the first training frames is to train the separation blocks 2a, 2b and, to this end, the first training frames need not include any examples of residual audio sources and the associated ground truth data may only indicate the frequency characteristics of the main audio source(s) in the mix. The method then goes to step S3 and step S4 which correspond to steps S3 and S4 in figure 7a. Accordingly, a first loss function is evaluated using predicted main source masks at step S3 and the learnable parameters of the separation blocks are updated at step S4.

[0128] During the first training epoch El the residual mask extractor 3 and the at least one residual classifier 4a, 4b may be omitted from the neural network system 10 since they are not used. A plurality of first training audio frames are used to train the separation blocks of the during the first training epoch El and, after the first training epoch El has been completed, the separation blocks 2a. 2b offer satisfactory accuracy for determining the main source masks Maski, Mask2, . . . , Maskv

[0129] The method then goes to step S5 being the first step of the second training epoch E2 wherein step S5 comprises obtaining second training audio frames.

[0130] The purpose of the second training frames is to train the residual classifier(s) 4a, 4b and. to this end, the second training frames need not include any examples of main audio sources and the associated ground truth data may only indicate the ground truth residual source activity of the residual audio source(s) included the mix. At step S6 the second training audio frames are provided to the neural network.

[0131] If the residual mask extractor and / or the at least one residual classifier 4a. 4b is not present during the first training epoch El these components are now added to the neural network system during the second training epoch E2.

[0132] For each inputted second training audio frame, the neural network system 10 outputs at least one predicted residual source activity metric SAMN+I, SAMN+2, . . . , SAMN+M and at step S7 the second loss function is evaluated to determine a loss whereby at step S8 the learnable parameters of the residual classifier 4a, 4b is updated based on the loss determined by the second loss function. Notably, during the second training epoch the learnable parameters of the separation blocks 2a. 2b are kept constant or “frozen” whereby only the learnable parameters of the residual classifiers 4a, 4b are updated.

[0133] In some implementations, the neural network system further comprises one or more main classifiers (see e.g. fig. 2) and it is understood that these main classifiers may be trained in amanner analogous to the residual classifiers, either together with separation blocks 2a, 2b or after the separation blocks 2a, 2b in the second training epoch E2.

[0134] Unless specifically stated otherwise, as apparent from the following discussions, it is appreciated that throughout the disclosure discussions utilizing terms such as “processing’", “computing"’, “calculating”, “determining”, “analyzing” or the like, refer to the action and / or processes of a computer hardware or computing system, or similar electronic computing devices, that manipulate and / or transform data represented as physical, such as electronic, quantities into other data similarly represented as physical quantities.

[0135] It should be appreciated that in the above description of exemplary embodiments of the disclosure, various features are sometimes grouped together in a single embodiment, figure, or description thereof for the purpose of streamlining the disclosure and aiding in the understanding of one or more of the various inventive aspects. This method of disclosure, however, is not to be interpreted as reflecting an intention that the claimed disclosure requires more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects lie in less than all features of a single foregoing disclosed embodiment. Thus, the claims following the Detailed Description are hereby expressly incorporated into this Detailed Description, with each claim standing on its own as a separate embodiment of this disclosure. Furthermore, while some embodiments described herein include some but not other features included in other embodiments, combinations of features of different embodiments are meant to be within the scope of the disclosure, and form different embodiments, as would be understood by those skilled in the art. For example, in the following claims, any of the claimed embodiments can be used in any combination.

[0136] Furthermore, some of the embodiments are described herein as a method or combination of elements of a method that can be implemented by a processor of a computer system or by other means of carrying out the function. Thus, a processor with instructions for carrying out such a method or element of a method forms a means for carrying out the method or element of a method. Note that when the method includes several elements, e.g.. several steps, no ordering of such elements is implied, unless specifically stated. Furthermore, an element described herein of an apparatus embodiment is an example of a means for carrying out the function performed by the element for the purpose of carry ing out the embodiments of the disclosure. In the description provided herein, numerous specific details are set forth. However, it is understood that embodiments of the disclosure may be practiced without these specific details. In other instances, well-known methods, structures and techniques have not been shown in detail in order not to obscure an understanding of this description. Thus, while there has been described specific embodiments of the disclosure, those skilled in the art will recognize that other and furthermodifications may be made thereto without departing from the spirit of the disclosure, and it is intended to claim all such changes and modifications as falling within the scope of the disclosure.

[0137] Various aspects of the present disclosure may be appreciated from the following Enumerated Example Embodiments (EEEs):EEE1. A computer implemented neural network system for residual source activity detection in an input audio signal comprising a mix of residual and main audio sources, the neural network system comprising: a feature extractor, comprising a neural network trained to predict a set of latent variables based on an input audio frame of the input audio signal, the latent variables representing latent features of the input audio frame, at least two separation blocks, wherein each separation block is configured to generate a main source mask for separating a respective main audio source using a trained neural network configured to receive the set of latent variables, a residual mask extractor, configured to receive the respective main source mask of each separation block, form a combined source mask by combining all main source masks and determine a residual mask that complements the combined source mask, at least one residual classifier, comprising a neural network trained to predict a residual source activity confidence metric based on the residual mask, the residual source activity confidence metric indicating a likelihood of a residual audio source being active, and an output stage, configured to output the at least one residual source activity confidence metric.EEE2. The computer implemented neural network system according to EEE1, wherein the neural network in each separation block is trained to output a candidate mask for separating the main audio source, and wherein each separation block further comprises: a main classifier comprising a neural network trained to predict a main source activity confidence metric based on the candidate mask, and a source mask modifier configured to modify' the candidate mask based on the main source activity confidence metric to form the source mask.EEE3. The computer implemented neural network system according to EEE2, wherein the source mask modifier is configured to make the candidate mask less inclusive in response to the main source activity confidence metric being below a predetermined threshold level.EEE4. The computer implemented neural network system according to EEE2 or EEE3, wherein the output stage is further configured to output the main source activity confidence metric of each separation block.EEE5. The computer implemented neural network system according to any one of EEE2-EEE4, wherein each main classifier is trained to predict a main source activity confidence metric based on the candidate mask and the input audio frame.EEE6. The computer implemented neural network system according to any one of the preceding EEEs, wherein the neural network of the at least one residual classifier is trained to predict a residual source activity confidence metric based on the residual mask and the input audio frame.EEE7. The computer implemented neural network system according to any one of the preceding EEEs, comprising: at least two residual classifiers, each comprising a neural network trained to predict a respective residual source activity confidence metric based on the residual mask, wherein the output stage is configured to output the at least two residual source activity confidence metrics.EEE8. The computer implemented neural network system according to any one of the preceding EEEs, wherein the at least one residual classifier comprises a recurrent neural network. RNN.EEE9. The computer implemented neural network system according to any one of the preceding EEEs, wherein the output stage is further configured to output the main source mask of each separation block.EEE10. The computer implemented neural network system according to any one of the preceding EEEs, wherein each main audio source is an audio source associated with at least one of a predetermined harmonic pattern, predetermined frequency distribution, predetermined dominating frequency and predetermined energy distribution across frequencies.EEE11. The computer implemented neural network system according to EEE10, wherein the main audio source of the at least two separation blocks are selected from a group comprising speech, animal sound and wind noise.EEE12. The computer implemented neural network system according to any one of the preceding EEEs, wherein the residual audio source is selected from a group comprising motor vehicle sounds and waterfall sounds.EEE13. An audio processing system, comprising the computer implemented neural network system according to any one of the preceding EEEs, further comprising: an audio processor, configured to process the input audio frame with a processing scheme, wherein the audio processing scheme is controlled based on the at least one residual source activity confidence metric.EEE14. The audio processing system according to EEE13, wherein the audio processor is configured to process the input audio frame with one of a plurality of audio processing schemes, wherein each audio processing scheme is associated with an audio scene, wherein the audio processor is further configured to: determine the acoustic scene of the input audio frame based on the at least one residual source activity confidence metric, select an audio processing scheme associated with the acoustic scene, and process the input audio frame with the selected audio processing scheme.EEE15. A method for training the computer implemented neural network system according to any one of EEE1-EEE12, comprising: obtaining a plurality of training audio frames, each training audio frame comprising one of: a main audio source, a residual audio source, and a mix of a main and residual audio source; obtaining, for each training audio frame, ground truth data indicating main audio source frequency characteristics and a residual source activity confidence metric; providing the training frame to the computer implemented neural network; for each training audio frame: updating the learnable parameters of the at least two separation blocks using a first loss function based on the main source mask generated by each separation block and the frequency characteristics of the ground truth data; and updating the learnable parameters of the at least one residual classifier using a second loss function based on the source activity confidence metric predicted by the at least one residual classifier and the residual source activity7confidence metric of the ground truth data.EEE16. A method for training the computer implemented neural network system according to any one of EEE1 -EEE12, comprising: obtaining a first set of training audio frames, each training audio frames in the first set comprising a main audio source; obtaining, for each training audio frame of the first set, first ground truth data indicating the frequency characteristics of the main audio source; providing each training frame of the first set to the computer implemented neural network; updating, for each training audio frame, the learnable parameters of each separation block, while the learnable parameters of the at least one residual classifier is kept constant, using a first loss function based on the source mask generated by each separation block and the frequency characteristics of the first ground truth data; obtaining a second set of training audio frames, each training audio frame in the second set comprising a residual audio source; obtaining, for each training audio frame of the second set, second ground truth data indicating a residual source activity confidence metric; providing each training audio frame of the second set to the computer implemented neural network; and updating the learnable parameters of the at least one residual classifier, while the learnable parameters of each separation block is kept constant, using a second loss function based on the source activity confidence metric predicted by at least one residual classifier and the residual source activity confidence metric of the second ground truth data.EEE17. The method according to EEE15 or EEE16, wherein the first loss function is selected from a group comprising mean absolute error, MAE, mean squared error, MSE, and tuned extensions thereof.EEE18. The method according to any one of EEE15-EEE17, wherein the second loss function is binary' cross-entropy.

Claims

CLAIMS1 . A computer implemented neural network system for residual source activity detection in an input audio signal comprising a mix of residual and main audio sources, the neural network system comprising: a feature extractor, comprising a neural network trained to predict a set of latent variables based on an input audio frame of the input audio signal, the latent variables representing latent features of the input audio frame, at least two separation blocks, wherein each separation block is configured to generate a main source mask for separating a respective main audio source using a trained neural network configured to receive the set of latent variables, a residual mask extractor, configured to receive the respective main source mask of each separation block, form a combined source mask by combining all main source masks and determine a residual mask that complements the combined source mask, at least one residual classifier, comprising a neural network trained to predict a residual source activity confidence metric based on the residual mask, the residual source activity confidence metric indicating a likelihood of a residual audio source being active, and an output stage, configured to output the at least one residual source activity confidence metric.

2. The computer implemented neural network system according to claim 1, wherein the neural network in each separation block is trained to output a candidate mask for separating the main audio source, and wherein each separation block further comprises: a main classifier comprising a neural network trained to predict a main source activity confidence metric based on the candidate mask, and a source mask modifier configured to modify the candidate mask based on the main source activity confidence metric to form the source mask.

3. The computer implemented neural network system according to claim 2, wherein the source mask modifier is configured to make the candidate mask less inclusive in response to the main source activity confidence metric being below a predetermined threshold level.

4. The computer implemented neural network system according to claim 2 or claim 3. wherein the output stage is further configured to output the main source activity confidence metric of each separation block.

5. The computer implemented neural network system according to any one of claims 2-4, wherein each main classifier is trained to predict a main source activity confidence metric based on the candidate mask and the input audio frame.

6. The computer implemented neural network system according to any one of the preceding claims, wherein the neural network of the at least one residual classifier is trained to predict a residual source activity confidence metric based on the residual mask and the input audio frame.

7. The computer implemented neural network system according to any one of the preceding claims, comprising: at least two residual classifiers, each comprising a neural network trained to predict a respective residual source activity confidence metric based on the residual mask, wherein the output stage is configured to output the at least two residual source activity confidence metrics.

8. The computer implemented neural network system according to any one of the preceding claims, wherein the at least one residual classifier comprises a recurrent neural network, RNN.

9. The computer implemented neural network system according to any one of the preceding claims, wherein the output stage is further configured to output the main source mask of each separation block.

10. The computer implemented neural network system according to any one of the preceding claims, wherein each main audio source is an audio source associated with at least one of a predetermined harmonic pattern, predetermined frequency distribution, predetermined dominating frequency and predetermined energy distribution across frequencies.

11. The computer implemented neural network system according to claim 10, wherein the main audio source of the at least two separation blocks are selected from a group comprising speech, animal sound and wind noise.

12. The computer implemented neural network system according to any one of the preceding claims, wherein the residual audio source is selected from a group comprising motor vehicle sounds and waterfall sounds.

13. An audio processing system, comprising the computer implemented neural network system according to any one of the preceding claims, further comprising: an audio processor, configured to process the input audio frame with a processing scheme, wherein the audio processing scheme is controlled based on the at least one residual source activity confidence metric.

14. The audio processing system according to claim 13, wherein the audio processor is configured to process the input audio frame with one of a plurality of audio processing schemes, wherein each audio processing scheme is associated with an audio scene, wherein the audio processor is further configured to: determine the acoustic scene of the input audio frame based on the at least one residual source activity confidence metric, select an audio processing scheme associated with the acoustic scene, and process the input audio frame with the selected audio processing scheme.

15. A method for training the computer implemented neural network system according to any one of claims 1-12. comprising: obtaining a plurality of training audio frames, each training audio frame comprising one of: a main audio source, a residual audio source, and a mix of a main and residual audio source; obtaining, for each training audio frame, ground truth data indicating main audio source frequency characteristics and a residual source activity confidence metric; providing the training frame to the computer implemented neural network; for each training audio frame: updating the learnable parameters of the at least two separation blocks using a first loss function based on the main source mask generated by each separation block and the frequency characteristics of the ground truth data; and updating the learnable parameters of the at least one residual classifier using a second loss function based on the source activity confidence metric predicted by the at least one residual classifier and the residual source activity7confidence metric of the ground truth data.

16. A method for training the computer implemented neural network system according to any one of claims 1-12, comprising: obtaining a first set of training audio frames, each training audio frames in the first set comprising a main audio source;obtaining, for each training audio frame of the first set, first ground truth data indicating the frequency characteristics of the main audio source; providing each training frame of the first set to the computer implemented neural network; updating, for each training audio frame, the learnable parameters of each separation block, while the learnable parameters of the at least one residual classifier is kept constant, using a first loss function based on the source mask generated by each separation block and the frequency characteristics of the first ground truth data; obtaining a second set of training audio frames, each training audio frame in the second set comprising a residual audio source; obtaining, for each training audio frame of the second set, second ground truth data indicating a residual source activity confidence metric; providing each training audio frame of the second set to the computer implemented neural network; and updating the learnable parameters of the at least one residual classifier, while the learnable parameters of each separation block is kept constant, using a second loss function based on the source activity confidence metric predicted by at least one residual classifier and the residual source activity confidence metric of the second ground truth data.

17. The method according to claim 15 or claim 16, wherein the first loss function is selected from a group comprising mean absolute error, MAE, mean squared error, MSE, and tuned extensions thereof.

18. The method according to any one of claims 15-17, wherein the second loss function is binary cross-entropy.

Citation Information

Patent Citations

  • Systems and methods for speech separation and neural decoding of attentional selection in multi-speaker environments

    US20190066713A1

  • Multi-channel speech separation

    US20190139563A1