Source separation and remix in signal processing

JP2024540567A5Pending Publication Date: 2025-10-30DOLBY LABORATORIES LICENSING CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024529682
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-05-04
Filing Date
2022-10-26
Publication Date
2025-10-30

AI Technical Summary

Technical Problem

Existing neural network models for audio signal processing struggle to effectively separate and enhance speech intelligibility while preserving desired audio components, as they are trained for specific types of noise and fail when encountering different noise definitions.

Method used

A method and system for audio processing that separates audio signals into stationary and non-stationary noise components, using trained models to determine and adjust weighting factors for each component, allowing for enhanced speech intelligibility and preservation of ambient noise.

Benefits of technology

The method achieves improved speech intelligibility and ambient sound preservation by accurately separating and remixing audio signals, addressing the limitations of existing models that fail with different noise definitions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2023091276000001
    Figure 2023091276000001
  • Figure 2023091276000002
    Figure 2023091276000002
  • Figure 2023091276000003
    Figure 2023091276000003
Patent Text Reader

Abstract

The present disclosure relates to a method and an audio processing system (1) for performing source separation. The method comprises the steps of: in ) from the audio signal; JPEG2024540567000189.jpg75), stationary noise content ( JPEG2024540567000190.jpg75), and non-audio content ( JPEG2024540567000191.jpg75) (S2a, S2b, S2c). JPEG2024540567000192.jpg75) is the non-audio content ( The method further comprises: JPEG2024540567000194.jpg75) and the non-audio content ( JPEG2024540567000195.jpg75) based on the difference between non-stationary noise content ( JPEG2024540567000196.jpg76) (S3), obtaining a set of weighting factors (S5), and weighting the audio content ( JPEG2024540567000197.jpg75), the stationary noise content ( JPEG2024540567000198.jpg75), and the non-stationary noise content ( and forming (S6) a processed audio signal based on a combination of the encoded image data (JPEG2024540567000199.jpg76).
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to the following priority applications: International Application PCT / CN2021 / 131462 (Reference: D21131WO), filed November 18, 2021, U.S. Provisional Application 63 / 288,996 (Reference: D21131USP1), filed December 13, 2021, U.S. Provisional Application 63 / 336,824 (Reference: D21131USP2), filed April 29, 2022, and European Patent Application 22171560.0, filed May 04, 2022, each of which is incorporated by reference in its entirety.

[0002] TECHNICAL FIELD OF THEINVENTION The present invention relates to a method and audio processing system for source separation and remix. [Background technology]

[0003] 2. Background of the Invention The recorded audio signal may contain representations of one or more audio sources in addition to the noise component. In particular for User Generated Content (UGC), it can generally be said that when recording audio, in addition to the noise audio component (such as white noise), many individual audio sources are picked up.

[0004] For example, consider a user using a headset or smartphone to record the audio track of a video, record a podcast, or make a phone call from the sidewalk of a busy street or in a forest under windy conditions. A recorded audio signal from a busy street may include, for example, the user's voice, as well as other nearby pedestrians, nearby pedestrians' mobile phone ringtones, passing cars and buses, sounds from a nearby construction site, emergency vehicle sirens, and noise components. Similarly, a recorded audio signal from a forest may include, for example, the user's voice, birds chirping, airplanes passing overhead, wind rustling leaves, and noise.

[0005] Since the recorded audio signal is composed of audio coming from all of these recorded sources, the desired audio signal, e.g., the voice of a user recording a video or making a phone call, will be less intelligible. For this reason, neural network models for speech separation have been proposed. These neural network models can take as input an audio signal containing recorded speech along with other audio sources and noise, and output either a processed audio signal with enhanced speech intelligibility, or a speech separation filter (often called a "mask") to suppress non-speech audio components of the audio signal. Thus, the use of neural network models can enhance the intelligibility of the speech contained in the audio signal, making it possible for users to record the audio signal in many locations.

[0006] In other situations, especially in Professionally Generated Content (PGC), such as the recording of a movie audio track, all audio sources (at least additional audio sources in addition to the recorded voice) may be of interest. For example, in a movie audio track recorded in a forest under windy conditions, the voice, rustling leaves, and birdsong are the desired audio signal components, whereas the sound of an airplane passing overhead is the undesired audio signal components. Thus, by enhancing the voice intelligibility using a neural network for voice separation, the separately recorded audio signals containing only the birdsong and only the rustling leaves can be mixed with the intelligibility-enhanced voice to achieve the desired mix of audio sources for the movie audio track. Here, the final mix has enhanced voice intelligibility but also contains the birdsong and rustling leaves, but does not contain the sound of the airplane passing by, providing a desired and believable ambience effect. Summary of the Invention [Problem to be solved by the invention]

[0007] Summary of the Invention The drawback of the previous solutions is that many neural network models perform well in terms of removing noise components, but each model is trained to remove a specific type of predetermined noise. Due to different definitions of noise, one neural network model works well when the noise definition used to train the model overlaps with the unwanted noise to be removed. However, as soon as the trained model is applied to remove a noise with a different definition from the noise definition used during training, the noise suppression performance will degrade.

[0008] For example, a trained speech separation model may be aggressive and trained to treat all audio signal components that are not speech as noise. Using such a speech separation on, say, a movie audio track where speech, birds chirping, and rustling leaves are all desirable audio signals will suppress the birds chirping and rustling leaves, and only the speech will be separated. On the other hand, using a less aggressive speech separation model, trained to predict and remove only stationary background noise, will only suppress stationary background noise, and not suppress unwanted sounds such as, say, a plane flashing overhead (which is not an example of stationary background noise).

[0009] It is therefore an object of the present disclosure to provide an enhanced method for audio processing that mitigates at least some of the drawbacks of the existing solutions mentioned above. [Means for solving the problem]

[0010] A first aspect of the present invention relates to a method of processing audio for source separation, comprising the steps of: obtaining an audio signal comprising a mixture of speech content and noise content; determining speech content from the audio signal; determining stationary noise content from the audio signal; and determining non-speech content from the audio signal, the stationary noise content being a proper subset of the non-speech content. The method further comprises determining non-stationary noise content based on a difference between the stationary noise content and the non-speech content, obtaining a set of weighting coefficients comprising weighting coefficients respectively corresponding to each of the speech content, the stationary noise content, and the non-stationary noise content, and forming a processed audio signal based on a combination of the speech content, the stationary noise content, and the non-stationary noise content weighted by the respective weighting coefficients.

[0011] Stationary noise content means noise content that is constant in time and does not carry interpretable information. White noise and thermal noise are both examples of stationary noise. Further examples of stationary noise are pink noise, Gaussian noise, noise introduced by audio amplifiers, and noise with any time-independent distribution.

[0012] Non-speech can be defined as the difference between a clean speech audio signal (e.g. a speech signal recorded in an anechoic chamber with stationary noise removed) and a clean speech audio signal with added disturbances (stationary noise, birds chirping, etc.) i.e. non-speech content includes not only stationary noise but also other kinds of non-stationary noise such as birds chirping or rain.

[0013] The first aspect of the present invention is based at least in part on the realization that by extracting non-stationary noise as the difference between non-speech content and stationary noise content, two independent noise content types are obtained in addition to the independent speech content. This facilitates remixing since the relative magnitudes of the three content types are adjusted by selecting a desired set of weighting coefficients. For example, by adjusting the three weighting coefficients, the stationary noise content is omitted entirely, the non-stationary noise is attenuated but not omitted entirely, and the speech content is amplified. The result is a processed audio signal that enhances speech intelligibility while also providing some ambience (since at least a portion of the non-stationary noise content is preserved).

[0014] In some aspects, determining the stationary noise content includes subjecting the audio signal to a stationary noise separator model trained to predict a stationary noise mask for removing the stationary noise content from the audio signal, and determining the stationary noise content based on the stationary noise mask and the audio signal.

[0015] Thus, an accurate trained model (e.g., implemented with a neural network) can be used to determine the stationary noise content given a representation of an audio signal. The stationary noise content can be precisely defined and large amounts of training data are readily available, recorded or synthetically created, which means that the stationary noise separator model can be trained to be very accurate.

[0016] Similarly, in some aspects, determining the non-speech content includes subjecting the audio signal to a speech separator model trained to predict a noise mask for removing the non-speech content from the audio signal, and determining the non-speech content based on the noise mask and the audio signal.

[0017] The process of separating speech from any audio signal can be performed accurately using a model (e.g., implemented with a neural network) trained to predict a mask for separating speech content given a representation of the audio signal. Furthermore, the same mask used to extract speech content can also be used to extract non-speech content. This means that the same trained model can be used to determine both speech and non-speech content.

[0018] It is difficult to train a model to separate different types of noise, such as stationary and non-stationary noise content. However, some aspects of the first aspect of the present invention utilize trained models adapted to more clearly separate different types of audio content, such as speech and stationary noise, and subsequent manipulation of the separated audio content to more accurately separate the different types of noise. This manipulation includes determining the difference between stationary noise and non-speech content.

[0019] In some aspects, the method further comprises band-pass filtering the non-stationary noise content with a band-pass filter configured to isolate noise objects in the non-stationary noise.

[0020] That is, the non-stationary noise may contain audio content associated with multiple non-stationary noise objects, but application of an appropriate band-pass filter isolates at least one desired noise object. The advantage of applying a band-pass filter to the non-stationary noise content is that the filter does not pass speech content or stationary noise content (as these are not present in the non-stationary noise content).

[0021] In some aspects, the bandpass filter is obtained by analyzing an example audio signal, the method further comprising the steps of collecting an example audio signal including at least one example of a noise object, determining a frequency distribution of the example audio signal, and defining the bandpass filter based on the frequency distribution of the example audio signal.

[0022] For this purpose, the frequency distribution of any non-stationary object(s) may be determined and used to generate a bandpass filter for filtering the non-stationary noise.

[0023] According to a second aspect of the present invention, there is provided an audio processing system comprising an audio content separation unit configured to obtain an audio signal comprising a mixture of speech content and noise content and to determine speech content, stationary noise content and non-speech content from the audio signal, the stationary noise content being a proper subset of the non-speech content, the audio content separation unit further configured to determine non-stationary noise content based on a difference between the stationary noise content and the non-speech content, and the audio processing system further comprises a mixing unit configured to obtain a set of weighting coefficients including weighting coefficients respectively corresponding to each of the speech content, the stationary noise content and the non-stationary noise content, and to form a processed audio signal based on a combination of the speech content, the stationary noise content and the non-stationary noise content weighted with the respective weighting coefficients.

[0024] According to a third aspect of the present invention, there is provided a non-transitory computer-readable medium storing instructions which, when executed by one or more processors, cause the one or more processors to perform a method according to the first aspect of the present invention. [Brief description of the drawings]

[0025] BRIEF DESCRIPTION OF THE DRAWINGS Aspects of the present invention will now be described in more detail with reference to the accompanying drawings showing presently preferred embodiments.

[0026] [Figure 1A] FIG. 1a illustrates an audio signal that is separated into non-speech content, speech content, stationary noise content, and residual content according to some aspects. [Figure 1B] FIG. 1b illustrates an audio signal that is separated into non-speech content, speech content, stationary noise content, and residual content according to some aspects.

[0027] [Diagram 2] FIG. 2 illustrates different types of non-speech content that an audio processing system separates from an audio signal, according to some aspects.

[0028] [Figure 3A] FIG. 3a is a block diagram illustrating different audio processing systems for source separation, according to some aspects. [Figure 3B] FIG. 3b is a block diagram illustrating a different audio processing system for source separation, according to some aspects. [Figure 3C] FIG. 3c is a block diagram illustrating a different audio processing system for source separation, according to some aspects.

[0029] [Figure 4] FIG. 4 is a flow chart illustrating a method according to some aspects.

[0030] [Diagram 5] FIG. 5 is a block diagram illustrating an audio processing system according to some aspects that includes an audio separator model for separating at least two different types of audio content.

[0031] [Figure 6A] FIG. 6a illustrates different options for an audio processing system with a classifier and a selector according to some aspects. [Figure 6B] FIG. 6b illustrates different options for an audio processing system with a classifier and a selector according to some aspects. [Figure 6C] FIG. 6c illustrates different options of an audio processing system with a classifier and a selector according to some aspects.

[0032] [Figure 7]FIG. 7 illustrates an example setup for training stationary noise separator and speech separator models according to some aspects. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0033] DETAILED DESCRIPTION OF THE PRESENTLY PREFERRED EMBODIMENTS The systems and methods disclosed herein can be implemented as software, firmware, hardware, or a combination thereof. In a hardware implementation, the division of tasks does not necessarily correspond to a division into physical units. Rather, one physical component may have multiple functions, and one task may be performed cooperatively by multiple physical components.

[0034] The computer hardware may be, for example, a server computer, a client computer, a personal computer (PC), a tablet PC, a set-top box (STB), a personal digital assistant (PDA), a mobile phone, a smart phone, a web appliance, a network router, a switch or bridge, or any device capable of executing instructions (sequential or otherwise) that specify actions to be performed by the computer hardware. Furthermore, the present disclosure relates to any collection of computer hardware that individually or collectively executes instructions to perform any one or more of the concepts described herein.

[0035] Some or all of the components may be implemented by one or more processors that accept computer-readable (also called machine-readable) code. The computer-readable (machine-readable) code includes a set of instructions that, when executed by one or more of the processors, performs at least one of the methods described herein. Any processor capable of executing a set of instructions (sequential or otherwise) that specify actions to be performed is included. Thus, one example is a typical processing system (i.e., computer hardware) that includes one or more processors. Each processor may include one or more of a CPU, a graphics processing unit (GPU), and a programmable DSP unit. The processing system may further include a memory subsystem that includes a hard drive, SSD, RAM, and / or ROM. A bus subsystem may be included for communication between the components. The software may reside in the memory subsystem and / or in the processor during its execution by the computer system.

[0036] One or more processors may operate as stand-alone devices or may be, for example, networked to other processor(s). Such a network may be built on a variety of different network protocols and may be the Internet, a wide area network (WAN), a local area network (LAN), or any combination thereof.

[0037] Software may be distributed on computer readable media, which may consist of computer storage media (or non-transitory media) and communication media (or transitory media). As is well known to those skilled in the art, the term computer storage media includes both volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information, such as computer readable instructions, data structures, program modules or other data. Computer storage media includes various forms of physical (non-transitory) storage media, such as, but not limited to, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVDs) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and that can be accessed by a computer. Furthermore, communication media (transitory) typically embody computer readable instructions, data structures, program modules, or other data in the form of a modulated data signal, such as a carrier wave or other transport mechanism, and is well known to those skilled in the art to include any information delivery media. Figure 1a illustrates the transmission of an audio signal S in The audio signal S in is a mixture of a desired source s and noise n, where the desired source s is, for example, speech content. in may be a mono audio signal, a stereo audio signal, or a multi-channel audio signal with more than two channels (eg the audio signal is a 5.1 or 7.1.2 audio signal).

[0038] Figure 1a shows an audio signal S in The audio signal S in is a mixture of a desired source s and noise n, where the desired source s is, for example, speech content. inmay be a mono audio signal, a stereo audio signal, or a multi-channel audio signal with more than two channels (eg the audio signal is a 5.1 or 7.1.2 audio signal).

[0039] An audio signal S consisting of a mixture of speech and noise content in is sometimes referred to as x(k) in the time domain, where k is the time sample index. Thus, x(k) can be expressed in the time domain as:

number

number

[0040] Audio signal S in may be fed to a trained model that has been trained to output masks M1, M2 to suppress certain noises, where the masks M1, M2 are typically masks that suppress the desired speech S for each time frame and frequency bin. m,f and audio signal mixture X m,f That is, the mask M is defined as follows:

number

[0041] Depending on the type and training of the mask prediction model, the masks M1, M2 may suppress different types of noise. inare shown separated by masks M1, M2, this is merely a simple illustration and the illustration should not be construed as merely showing, for example, time and frequency frames. As is evident from Equation 3, masks M1, M2 are made up of multiple mask values, one for each time and frequency bin, typically real numbers between zero and one that describe the degree to which each time and frequency bin should be suppressed.

[0042] With further reference to FIG. 1b, the audio signal S in Of the non-voice A first trained model 11 is trained to output a first mask M1 that suppresses all audio components that are JPEG2024540567000005.jpg75. in The manner in which the audio signal S in When mask M1 is applied to the speech, the trained model Only what is considered to be JPEG2024540567000006.jpg75 remains. Audio This first trained model 11 may be used to perform aggressive speech intelligibility enhancement, since the mask M1 removes all sounds that are not considered to be JPEG2024540567000007.jpg75. While this may be appropriate in some cases, this type of speech intelligibility enhancement may not be appropriate in other cases. For example, if a character is speaking in a busy street in the audio track of a video, aggressive speech intelligibility enhancement may remove traffic sounds from the street that are important for context and immersion.

[0043] Audio signal S in Stationary noise content Output a mask M2 that suppresses only JPEG2024540567000008.jpg75 and excludes all audio content that is not stationary noise content (residual content The second trained model 12 is trained to apply the mask M2 to the audio signal S in By applying it to the signal, stationary noise (i.e., noise with a constant probability distribution over time) is effectively removed, while other types of potentially undesirable noise (e.g., the revving of a nearby car engine) are left unaffected.

[0044] By using these two trained models simultaneously, a first model 11, which is a speech separator model trained to output a first mask M1, and a second model 12, which is a stationary noise separator model trained to output a second mask M2, where the first mask M1 is a non-speech mask. The second mask M2 is for suppressing stationary noise. JPEG2024540567000011.jpg75 is suppressed, and the audio signal S in The first model 11 estimates the speech content of JPEG2024540567000012.jpg75 and non-audio content JPEG2024540567000013.jpg75 (i.e. noise such as birds chirping and stationary noise) can be calculated as follows:

number

number

[0045] Output audio signal S out Then, from equations 4, 5, 6 and 7, JPEG2024540567000018.jpg721 and This can be determined by combining JPEG2024540567000019.jpg75 as follows:

number

number

[0046] The above audio signal components from Eq. JPEG2024540567000028.jpg736 is not independent, for example audio content JPEG2024540567000029.jpg75 is partially or entirely residual content JPEG2024540567000030.jpg75. This means that it may not be possible to achieve the desired mix of components from Equation 8. For this reason, non-speech content JPEG2024540567000031.jpg75 with stationary noise content Non-stationary noise content using JPEG2024540567000032.jpg75 We define a new type of noise content, called JPEG2024540567000033.jpg76 or object noise content, as follows:

number

number

[0047] Stationary Noise Content JPEG2024540567000038.jpg74 and non-stationary noise content JPEG2024540567000039.jpg76 is the audio signal S in is an independent part of (and conversely JPEG2024540567000040.jpg717 is dependent). Stationary noise content JPEG2024540567000041.jpg74, for example, captures white noise and non-stationary noise content. JPEG2024540567000042.jpg76 captures all content that is neither stationary noise content nor audio content. Examples of noises in JPEG2024540567000043.jpg76 include birds chirping, leaves rustling, cars, airplanes, helicopters, sirens, wind gusts, rain and thunder. These examples (as well as others not mentioned) are the respective noise objects. JPEG2024540567000044.jpg723 is formed. Here, each noise object JPEG2024540567000045.jpg723 is non-stationary noise content It is a proper subset of JPEG2024540567000046.jpg76 and is associated with audio content of a particular type or from a particular audio source (e.g., a machine, an animal, or a vehicle).

[0048] Therefore, the audio signal component JPEG2024540567000047.jpg735 is combined in a manner similar to Equation 8 as follows:

number

[0049] The output signal S calculated using Equation 12 using JPEG2024540567000054.jpg729 out From Equation 8, It can also be expressed as JPEG2024540567000055.jpg729, and from Equation 9 Note that the weighting coefficients can also be expressed as JPEG2024540567000056.jpg719. Thus, there is a mapping between all three sets of weighting coefficients, i.e., weighting coefficients α2, β2, γ2, μ2, weighting coefficients α1, β1, γ1, μ1, and weighting coefficients c1, c2, c3. However, the representation from Equation 12 is not supported for three independent content types (if JPEG2024540567000057.jpg75 is omitted) to produce an output audio signal S out This has the advantage that it makes it easier to remix more precisely.

[0050] JPEG2024540567000058.jpg75 is audio content JPEG2024540567000059.jpg75 and non-audio content predicted by the first trained model11 Since the residual content of JPEG2024540567000060.jpg75 includes overlap between the two, in some aspects, β1 or β2 are set to zero or the residual content is JPEG2024540567000061.jpg75 is omitted from Equation 8 and Equation 12.

[0051] Referring to Figure 2, different types of non-audio content JPEG2024540567000062.jpg75 is a schematic diagram showing non-audio content. JPEG2024540567000063.jpg75 is stationary noise content JPEG2024540567000064.jpg75, where the stationary noise content is white noise N w This includes various forms of stationary noise content, such as: JPEG2024540567000065.jpg714 with non-speech audio content The difference from JPEG2024540567000066.jpg75 is the non-stationary noise content Define JPEG2024540567000067.jpg76. Non-stationary noise content JPEG2024540567000068.jpg76 contains one or more noise objects Contains JPEG2024540567000069.jpg723. Noise object JPEG2024540567000070.jpg723 is neither speech nor stationary noise (e.g. birds chirping).

[0052] Figure 3a shows a block diagram of an audio processing system 1. With further reference to the flow chart of figure 4, a method for performing audio processing for source separation according to several aspects will now be described in detail.

[0053] In step S1, an audio signal containing a mix of speech and noise content is obtained and applied to an audio separation unit 10. The audio separation unit 10 separates non-speech content in the audio signal. Audio content from JPEG2024540567000071.jpg75 The speech separator model 11 is trained to predict a mask M1 for separating the speech content JPEG2024540567000072.jpg75. The mask M1 can be applied to the audio signal, for example according to Equations 4 and 5 above, to obtain the speech content. JPEG2024540567000073.jpg75 and non-audio content JPEG2024540567000074.jpg75 are determined in steps S2a and S2c, respectively.

[0054] Similarly, an audio signal may contain stationary noise content. Residual audio content from JPEG2024540567000075.jpg75 JPEG2024540567000076.jpg75. The mask M2 is applied to the audio signal, for example according to Equation 7 above, to obtain a stationary noise separator model 12 trained to predict a mask M2 for isolating the stationary noise content JPEG2024540567000076.jpg75. JPEG2024540567000077.jpg75 is determined in step S2b.

[0055] In step S3, non-stationary noise content JPEG2024540567000078.jpg76 is the non-speech noise content predicted by the speech separator model 11 JPEG2024540567000079.jpg75 and the stationary noise predicted by the stationary noise separator model 12 JPEG2024540567000080.jpg75. Alternatively, the audio separation unit 10 may determine the audio content JPEG2024540567000081.jpg75, non-audio content JPEG2024540567000082.jpg75 with stationary noise content Non-stationary noise content is removed by outputting JPEG2024540567000083.jpg75. JPEG2024540567000084.jpg76 is determined by the auxiliary calculation unit.

[0056] The method then proceeds to step S5. Step S5 includes JPEG2024540567000085.jpg75, stationary noise content JPEG2024540567000086.jpg714, with non-stationary noise content determining at least one weighting factor for each of the JPEG2024540567000087.jpg76, the weighting factor being based on an independent audio content in the output audio signal; JPEG2024540567000088.jpg75, stationary noise content JPEG2024540567000089.jpg714, with non-stationary noise content The weighting coefficients may be predetermined or set by a user / mixing engineer, for example, to obtain a desired mix of JPEG2024540567000090.jpg76. Furthermore, as described below, a selector may select or suggest a set of weighting coefficients based on detected noise objects present in the audio signal.

[0057] In step S6, the audio content JPEG2024540567000091.jpg75, stationary noise content JPEG2024540567000092.jpg714, with non-stationary noise content JPEG2024540567000093.jpg76 is combined with respective weighting factors by mixer unit 14 to form a processed audio signal, for example according to equation 12 above, i.e. different independent content types of the audio signal are remixed to form a processed output audio signal.

[0058] Optionally, stationary noise content, as seen in the exemplary embodiment of FIG. 3b. JPEG2024540567000094.jpg75 and residual content JPEG2024540567000095.jpg75 are both determined in step S2b, for example by using Equations 6 and 7 above, whereby the stationary noise content JPEG2024540567000096.jpg75 and residual content JPEG2024540567000097.jpg75 are both used in the mixer unit 14 in combination with their respective weighting factors.

[0059] FIG. 3c shows another optional embodiment, In step S3, non-stationary noise is removed before being applied to JPEG2024540567000098.jpg72414. JPEG2024540567000099.jpg76 is processed with a bandpass filter 13. Furthermore, the filtered non-stationary noise may be smoothed with a smoothing kernel or filter (not shown) before being provided to the mixer unit 14. The embodiment of Fig. 3c can be combined with other embodiments, such as, for example, the embodiment shown in Fig. 3b. Furthermore, the non-stationary noise It is also envisaged that both JPEG2024540567000100.jpg76 and the non-stationary noise processed by filter 13 may be provided to a mixing unit 14, as shown in FIG. 6a.

[0060] The filter 13 may then be determined by collecting an exemplary audio signal and determining a frequency distribution of the exemplary audio signal, which includes at least one example of a (non-stationary) target noise object, such as a bird chirp, or a group of target noise objects, such as traffic sounds, and the frequency distribution of the exemplary audio signal reveals the energy distribution of the audio signal, which allows defining a suitable bandpass filter 13 having a passband that passes at least a predetermined portion of the exemplary audio signal. For example, the bandpass filter 13 is defined as narrow as possible, but still characterized by a passband that passes at least 50%, preferably at least 70%, and most preferably at least 90% of the energy of the test signal. That is, the bandpass filter 13 filters and attenuates noise objects that are different from the target noise object(s).

[0061] To obtain a more accurate bandpass filter 13, the example audio signal should be composed of clean examples of the target noise object or group of noise objects. For this purpose, the target audio signal is manually cleaned to remove audio components or noise that are not examples of the target noise object(s) or cleaned by a reliable automatic process. Furthermore, to avoid averaging errors, a longer example audio signal with more / longer examples of the target noise object(s) is preferred. For example, the example audio signal is composed of at least 1 hour, preferably at least 5 hours, most preferably at least 10 hours of audio content of the noise object.

[0062] As an illustrative example, the noise object of interest is bird songs, which results in an exemplary audio signal containing 10 hours of clean bird songs, and a frequency distribution is determined. The frequency distribution reveals that most of the signal energy is contained between 3 kHz and 7 kHz, which allows a bandpass filter 13 with a passband between 3 kHz and 7 kHz and a stopband starting at 1 kHz and 9 kHz, respectively, to filter the bird songs from non-stationary noise. It is defined to separate it from other noise objects present in JPEG2024540567000101.jpg76.

[0063] The audio processing system 1 shown in Fig. 5 is identical to the audio processing system described in relation to Fig. 3a, except for the presence of a different type of speech separator model 11'. The speech separator model 11' in Fig. 5 takes an audio signal and is trained to predict at least two masks to separate at least two different types of speech present in the audio signal. In the illustrated embodiment, the speech separator model 11' is trained to predict a speech signal without reverberation, called dry speech, JPEG2024540567000102.jpg74, Early reverberation Dry audio with early and late reverberation, JPEG2024540567000103.jpg74 Predicting three masks to isolate dry audio with JPEG2024540567000104.jpg74. Different audio types JPEG2024540567000105.jpg717 is fed to the mixing unit 14, where it is mixed with stationary noise content JPEG2024540567000106.jpg74 and non-stationary noise JPEG2024540567000107.jpg76 with weighting coefficients according to the respective audio types. Thus, the output audio signal S out Equation 12 (residual content JPEG2024540567000108.jpg75 or not) as audio content JPEG2024540567000109.jpg75 You can change it by replacing

number

[0064] Therefore, for example, by setting α2 and α3 to small values ​​relative to α1, the output audio signal S out In the output audio signal S out In this mode, dry sound with early reverberation is emphasized.

[0065] Late reverberation refers to speech reverberation that has a reverberation time above a certain threshold, and early reverberation refers to speech reverberation that has a time constant below a certain threshold.

[0066] The Audio Separator Model 11' is a device that can be used to separate different audio types. Alternatively, the speech separator may be constructed by training a model for each of the different types of speech. The image may consist of a single separator model 11' trained to predict one mask for separating each of the JPEG2024540567000116.jpg716.

[0067] In the embodiment of the audio processing system 1 of FIG. 5, sound types different in terms of reverberation are extracted, but sound types different in other respects are also extracted as sound types with different reverberation characteristics. It is contemplated that such a configuration may be used as an alternative to, or in addition to, JPEG2024540567000117.jpg716. For example, the speech separator model 11' may be configured (trained) to separate at least two types of speech that differ in at least one of the gender of the voice making the speech, the age of the voice making the speech, and the language of the speech.

[0068] 6a, 6b and 6c each show a block diagram of an audio processing system 1 including a classifier 15 according to several aspects which will now be described in further detail.

[0069] In Fig. 6a, the classifier 15 receives an audio signal, and the classifier 15 is trained to predict the presence of at least a noise object in the audio signal. The classifier 15 may further be trained to predict the presence of at least a noise object in the audio signal, where the at least one noise object is at least one noise object of a predetermined set of noise objects. For example, the classifier 15 may be trained to predict the presence of at least one of bird chirps, traffic sounds, wind sounds, rain sounds, thunder sounds, sirens sounds, airplane sounds, helicopter sounds, and machine sounds (such as washing machine, drill, lawnmower sounds, etc.) in the audio signal. The selector 16 selects filter data 172a, 172b, 172c associated with the predicted noise object based on the at least one noise object predicted to be present in the audio signal, and applies a filter 13' as described by the selected filter data 172a, 172b, 172c to the non-stationary noise. For example, the classifier 15 may predict that birdsong is present in the audio signal, so that the birdsong filter 13' selected by the selector 16 may filter out non-stationary noise content. applied to JPEG2024540567000118.jpg76.

[0070] To this end, the classifier 15 may be a neural network trained to predict the presence of at least one noise object given a representation of an audio signal, where the neural network predicts the likelihood that the audio signal contains one or more predefined noise objects, and the noise object associated with the greatest likelihood is assumed to be the predicted noise object.

[0071] The selector 16 can retrieve the filter 13' from a database 171 of different sets of filter data 172a, 172b, 172c, each set of filter data being associated with a noise object and describing the filter 13' to be applied. For example, for each noise object present in a given set of noise objects that are possible outputs of the classifier 15, there is a corresponding set of filter data 172a, 172b, 172c in the database 171. Furthermore, as shown in FIG. 6a, in addition to the filtered non-stationary noise, a non-stationary noise JPEG2024540567000119.jpg76 may be provided to the mixing unit 14. JPEG2024540567000120.jpg76 and filtered non-stationary noise are Respective weighting coefficients are provided that allow the relative signal strengths of JPEG2024540567000121.jpg76 to be altered as desired (e.g., by a user or mixing engineer).

[0072] In the exemplary embodiment shown in Figure 6a, classifier 15 predicts birdsong as one noise object present in the audio signal and provides an indication of birdsong to selector 16. Selector 16 accesses database 171 and finds that filter data 172b describes a filter 13' associated with birdsong (e.g., a filter having a passband between 3kHz and 7kHz as described above), which causes selector 16 to select filter data 172b and select bird song filter 13' as a filter for non-stationary noise content. This allows it to be applied to JPEG2024540567000122.jpg76.

[0073] 6b illustrates another audio processing system 1 including a classifier 15 according to some aspects. The classifier 15 predicts the presence of at least one noise object (e.g., the presence of at least one noise object among a predetermined set of noise objects) and provides the predicted noise object(s) to a selector 16. The selector 16 accesses a database 173 of trained noise object separation models 174a, 174b, 174c and provides the at least one predicted noise object(s). At least one trained noise object separation model 174a is selected that is trained to predict a mask for separating JPEG2024540567000123.jpg710. The predicted mask of the selected noise object separation model 174a is applied to the audio signal to obtain JPEG2024540567000124.jpg710. And the noise object JPEG2024540567000125.jpg710 is fed to the mixing unit 14, where non-stationary noise is removed. JPEG2024540567000126.jpg76, Stationary noise JPEG2024540567000127.jpg74 and audio content JPEG2024540567000128.jpg75. Here, each content type has its own weighting factor. Thus, a user or a mixing engineer can set the weighting factor as desired, for example, to prioritize stationary and non-stationary noise. JPEG2024540567000129.jpg713 suppressed non-stationary noise noise object JPEG2024540567000130.jpg710 and audio content Amplify only JPEG2024540567000131.jpg75.

[0074] Although the audio processing system 1 of FIG. 6a and FIG. 6b uses the classifier 15 and the selector 16 to select the appropriate filter data 172a, 172b, 172c or the noise object separator model 174a, 174b, 174c, it is envisioned that the classifier 15 and the selector 16 may select more than one (e.g., two or more) filter or noise object separator model if the classifier 15 detects that two or more noise objects are present in the audio signal. Furthermore, a filter or noise object separator model may be associated with a group of noise objects rather than a single noise object. For example, there may be a trained natural object separator model or natural filter that is selected if the classifier 15 detects at least one of birds chirping, rustling leaves, and rain.

[0075] In connection with Figures 6a and 6b above, we explain how the classifier 15 and selector 16 are used to dynamically and based on the content of the audio signal, change the filter 13' applied to the non-stationary noise or change which object noise separator model 174a, 174b, 174c is used. Thus, the number of audio content types provided to the mixing unit 14 can be changed depending on the content of the audio signal, allowing a user or mixing engineer to select the desired relative signal strength for each component by manually selecting the weighting coefficients. However, as shown in Figure 6c, the weighting coefficients may be determined automatically, e.g., selected by the selector 16 from a database 175 of weighting coefficient sets 176a, 176b, 176c based on the noise object(s) that the classifier 15 predicts to be present in the audio signal. Each set of weighting coefficients 176a, 176b, 176c in the database is composed of at least one value of each of α2, γ2 and μ2.

[0076] For example, if the classifier 15 predicts the presence of birds singing, the selector 16 may select a set of weighting coefficients 176c that suppresses stationary noise, amplifies non-stationary noise, and amplifies audio content, since birds singing is believed to add a pleasant ambience while not interfering with speech intelligibility. On the other hand, if the classifier 15 predicts the presence of wind noise, the selector 16 may select a different set of weighting coefficients 176a that suppresses stationary and non-stationary noise (including wind noise) while amplifying audio content, since wind noise is believed to be a non-useful disturbance.

[0077] In this way, the selector 16 automatically selects the appropriate set of weighting coefficients 176a, 176b, 176c for every audio signal according to a predefined set of rules. The user or the mixing engineer optionally provides preferences to modify the rules. The preferences may indicate, for example, a desire to suppress some noise objects more than others (e.g., suppress all artificial noise objects such as machinery and traffic sounds, but leave all natural sounds such as birds singing, rain and thunder). Alternatively or additionally, the preferences may indicate, for example, a desire to enhance speech intelligibility at the expense of less ambience, such as completely omitting reverberation and stationary noises and attenuating all noise objects.

[0078] In some embodiments (not shown), the classifier 15 uses the outputs of the stationary noise separator model 12 and the speech separator model 11 to extract non-stationary noise content. JPEG2024540567000132.jpg76 (instead of the entire audio signal). The noise object contains non-stationary noise content. Since the image contains the noise object JPEG2024540567000133.jpg76, the classifier 15 is still able to correctly predict the presence of at least one noise object. Since JPEG2024540567000134.jpg76 contains only audio content, which is a proper subset of the audio signal content, the classification will be more accurate.

[0079] FIG. 7 shows how the stationary noise separator model 12 and the speech separator model 11 are trained to predict the corresponding masks M1, M2. Training data in speech form is obtained from a speech database 179. The speech database 179 is composed of audio signals with clean speech audio signals corresponding to a number of different speakers, languages, and signal bit rates. Similarly, noise training data is obtained from a noise database 177. The noise is composed of a number of non-speech sounds, such as different types of stationary noise (e.g., white noise) and different types of non-stationary noise (e.g., rain and barking dogs). The training speech and noise data are combined in a mixer and given to the stationary noise separator model 12 and the speech separator model 11, respectively, for training.

[0080] During training, the internal weights and / or parameters of the separation models 11, 12 are adjusted to predict a mask M1 that accurately separates speech and a mask M2 that accurately separates stationary noise. To achieve this, the audio signal obtained after applying the mask M1 is compared to a ground truth signal containing clean speech from the speech database 179, and the audio signal obtained after applying the mask M2 is compared to a ground truth signal containing only stationary noise added from the noise database 177. By modifying the internal weights and / or parameters of the separation models 11, 12 to minimize the mismatch between the audio signal with the respective mask applied and the ground truth signal, the models 11, 12 gradually learn to predict the masks M1, M2 for accurate speech and stationary noise separation.

[0081] One or more noise object separator models 174a, 174b, 174c of the database 173 described in relation to Fig. 6b may be obtained by a similar training setup, except that for the noise object separator models 174a, 174b, 174c, the ground truth signal is a clean signal representative of a noise object (such as the exemplary audio signal described above) and the training signal is the clean signal representative of the noise object mixed with at least one of other noise objects, speech and stationary noise.

[0082] Unless otherwise noted, and as will be apparent from the description that follows, throughout the description of this disclosure, terms such as "processing," "computing," "calculating," "determining," "analyzing," and the like are used and are understood to refer to the actions and / or processes of computer hardware or computing systems, or similar electronic computing devices, that manipulate and / or transform data represented as physical quantities, such as electronic quantities, into other data also represented as physical quantities.

[0083] In the above description of exemplary embodiments of the present invention, it should be understood that various features of the invention may be grouped together in a single embodiment, figure, or description thereof for the purpose of streamlining the disclosure and aiding in the understanding of one or more of the various inventive aspects. However, this method of disclosure should not be interpreted as reflecting an intention that the embodiments of the present invention utilize more features than are expressly recited in each claim. Rather, as the following claims reflect, inventive aspects lie in fewer than all features of a single disclosed embodiment as set forth above. Accordingly, the claims following the detailed description are expressly incorporated herein, with each claim standing on its own as a separate embodiment of the present invention. Furthermore, although some embodiments described herein include some but not other features included in other embodiments, combinations of features of different embodiments are meant to form different embodiments within the scope of the present invention as understood by those skilled in the art. For example, in the following claims, any of the claimed embodiments may be used in any combination.

[0084] Furthermore, some of the embodiments are described herein as a method or a combination of elements of a method that can be implemented by a processor of a computer system or other means for performing the function. Thus, a processor having instructions for implementing such a method or elements of a method forms a means for implementing the method or elements of a method. It should be noted that when a method includes several elements, e.g. several steps, no ordering of such elements is implied unless otherwise specified. Furthermore, an element of an embodiment as an apparatus described herein is an example of a means for performing the functions performed by that element for implementing an embodiment of the present invention. In the description provided herein, numerous specific details are specified. However, it is understood that an embodiment of the present invention can be implemented without these specific details. In other instances, well-known methods, structures and techniques have not been shown in detail in order not to obscure an understanding of this specification.

[0085] Those skilled in the art will understand that the aspects of the present invention are in no way limited to the above-mentioned embodiment. On the contrary, many modifications and variations are possible within the scope of the appended claims. For example, the classifier 15 and selector 16 in the manner shown in Fig. 6a, Fig. 6b, Fig. 6c are used to select filter data 172a, 172b, 172c, noise object separator model 174a, 174b, 174c, or weighting coefficient set 176a, 176b, 176c, but it is envisioned that the classifier and selector may simultaneously select two or all three of the filter(s), noise object separator model(s), or weighting coefficient set. For example, the noise object separator model 174a, 174b, 174c may be sufficient to separate noise objects, while the filter 13' may be used to further improve the quality of the separation of noise objects.

[0086] Various aspects of the present invention will be understood from the following enumerated example embodiments (EEE): EEE1. A method for processing audio, comprising: receiving an audio signal comprising a mixture of speech content and noise content; determining background noise and object noise from said noise content; enhancing the audio content to generate audio-enhanced audio, where enhancing the audio content includes applying one or more first gains to the audio content, one or more second gains to the background noise, and one or more third gains to the object noise; providing the voice-enhanced audio to a downstream device; A method comprising: EEE2. The method according to EEE1, wherein the step of determining background noise and object noise comprises combining and remixing type 1 noise and type 2 noise, the type 1 noise and type 2 noise being defined in a noise database and corresponding to respective models for generating respective masks for enhancing speech under each type of noise. EEE3. The method according to EEE2, wherein said background noise corresponds to said type 1 noise and said object noise corresponds to the difference between said type 1 noise and said type 2 noise. EEE4. The method according to EEE2 or 3, wherein at least one of the one or more second gains or the one or more third gains is different from the gains corresponding to the Type 1 noise and Type 2 noise defined in the respective models. EEE5. A method for processing audio, comprising: receiving an audio mixture; separating and remixing said audio mixture based on a particular source type; Includes, a method. EEE6. The method of EEE5, wherein the type of source includes at least one of noise or equipment sound. EEE7. solving the problem of overlap between source types by providing type definitions, where the difference information between types is used in the remix process; The method according to EEE5 or 6, comprising: EEE8. The method of any one of EEE5 to 7, comprising performing post-processing including extending from said particular source type to other source types. EEE9. representing new source types by combining the source type classifiers; performing separation and mixing using said new source type; The method according to any one of EEE5 to 8, comprising: EEE10. one or more processors; a non-transitory computer-readable medium storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations 1-9; The system has: EEE11. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations 1-9.

Claims

1. 1. A method of processing audio for source separation, comprising: An audio signal (S) containing a mixture of speech and noise content in (S1) obtaining a The audio signal is extracted from the audio content ( ) determining the Extract stationary noise content ( ) (S2b), Extracting non-speech content (e.g., ), wherein the stationary noise content ( ) is the non-audio content ( ) is a proper subset of (S2c); The stationary noise content ( ) and the non-audio content ( ) and non-stationary noise content ( ) determining step (S3); The audio content ( ), the stationary noise content ( ), and the non-stationary noise content ( obtaining a set of weighting factors (S5) including weighting factors respectively corresponding to each of the The audio content weighted by each weighting coefficient ( ), the stationary noise content ( ), and the non-stationary noise content ( forming a processed audio signal based on the combination of the A method comprising:

2. The stationary noise content ( ) determining The audio signal (S in ) to the stationary noise content ( ) to remove stationary noise mask (M 2 ) applying the audio signal to a stationary noise separator model (12) trained to predict The stationary noise mask (M 2 ) and the audio signal (S in ) based on the stationary noise content ( and determining The method of claim 1.

3. The non-audio content ( ) determining The audio signal is then extracted from the non-speech content ( ) to remove noise mask (M 1 ) to a speech separator model (11) trained to predict the audio signal (S in ) and The noise mask (M 1 ) and the audio signal (S in ) based on the non-audio content ( and determining 3. The method according to claim 1 or 2.

4. The non-stationary noise content ( ) noise object The non-stationary noise content ( ) further comprising a step (S4) of band-pass filtering the 3. The method according to claim 1 or 2.

5. The non-stationary noise content ( ), wherein each bandpass filter (13) is configured to filter out non-stationary noise ( ) different noise objects and separating the The method of claim 4.

6. The audio signal (S in ) to the audio signal (S in ) noise objects present in a noise object classifier model (15) trained to output a prediction of Each of them is the non-stationary noise ( ) different noise objects providing a plurality of bandpass filters (13) configured to separate selecting the bandpass filter (13') associated with the predicted noise object; The method of claim 4 further comprising:

7. Each bandpass filter (13) Noise Object collecting an example audio signal that includes at least one example of determining a frequency distribution of said exemplary audio signal; defining the bandpass filter based on the frequency distribution of the exemplary audio signal; The method of claim 4, wherein the polymer is obtained by

8. smoothing the filtered non-stationary noise with a smoothing filter. The method of claim 4.

9. The weighting factors are determined based on the stationary noise content ( ) to the non-stationary noise content ( 3. The method of claim 1 or 2, wherein the method is shown to enhance the activity of the IL-1 receptor agonist.

10. providing at least two sets of weighting factors, each set of weighting factors associated with a respective audio source type; The audio signal (S in ) noise objects present in subjecting the audio signal to a classifier model trained to output a prediction of further comprising The step of obtaining a set of weighting coefficients comprises: the predicted noise object from the at least two sets selecting a set associated with 3. The method according to claim 1 or 2.

11. The audio signal (S in ) based on the non-stationary noise content ( ) at least one noise object forming a proper subset of and determining the set of weighting factors includes a noise object weighting factor for each noise object, and the combination is further based on the noise objects being weighted by the noise object weighting factor.

3. The method according to claim 1 or 2.

12. The step of determining at least one noise object comprises: The audio signal (S in ) to the noise object providing the audio signal to an object separation model (172') trained to predict a mask for separating The audio signal (S in ) and the audio signal (S in ) to the noise object based on a mask to separate the noise object determining The method of claim 11 , comprising:

13. Each is a different noise object providing a plurality of trained object separation models (172') trained to predict a mask for separating a The audio signal (S in ) to a classifier model (15) trained to output predicted noise objects present in the audio signal (S in ) and selecting, from the plurality of trained object separation models, the trained object separation model associated with the predicted noise object; The audio signal (S in ) to the object separation model selected to predict a mask for separating the predicted noise object from the audio signal (S in ) and The method of claim 11 further comprising:

14. - an audio signal (S) containing a mixture of speech and noise content in ) and - extracting audio content (e.g., ) is determined, - extracting stationary noise content (e.g., ) is determined, - extracting from the audio signal the stationary noise content ( ) is non-audio content ( ) such that the non-audio content ( ) is determined, - the stationary noise content ( ) and the non-audio content ( ) and non-stationary noise content ( ) to determine An audio processing system (1) comprising an audio content separation unit (10) configured as follows: The audio processing system (1) further comprises a mixing unit (14), which comprises: - the audio content ( ), the stationary noise content ( ), and the non-stationary noise content ( a set of weighting factors, each of which includes a weighting factor corresponding to each of the - the audio content weighted by the respective weighting factors ( ), the stationary noise content ( ), and the non-stationary noise content ( ) to form a processed audio signal based on a combination of 1. An audio processing system configured as follows:

15. A non-transitory computer readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to perform the method of claim 1 or 2.