System and method for separating audio tracks and matching audio tracks with text subtitles

Through neural network separation and enhancement of audio, the problem of specific sound separation in single-channel audio is solved, and precise audio track matching and control based on text description is achieved, improving audio quality.

CN120390955APending Publication Date: 2025-07-29GOOGLE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202280102894.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2022-12-30
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The prior art is difficult to effectively separate and enhance a particular sound in a single-channel audio input, especially in audio mixtures, and it is difficult for computers to accurately match and control audio content based on text descriptions.

Method used

Using a neural network-based method, the audio separation network and the text embedding network are used to separate the input audio waveform into multiple audio tracks, and a classifier is used to determine whether the audio track matches the input text description to achieve enhancement or suppression of the audio content.

Benefits of technology

Improves the accuracy and flexibility of audio separation, and users can selectively enhance or suppress specific audio tracks based on text descriptions, improving the actual and perceived quality of the audio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120390955A_ABST
    Figure CN120390955A_ABST
Patent Text Reader

Abstract

A computer-implemented method is provided. The method includes receiving, by a computing device, an input audio waveform and an input text description. The method further includes separating the input audio waveform into a plurality of audio tracks by a neural network. The method further includes determining, by a neural network, whether the input text description describes a track of the plurality of separated tracks. The method further includes, upon determining that the input textual description describes a track of the plurality of tracks, providing, by the computing device, the track corresponding to the input textual description to the interactive user interface.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0001] Many modern computing devices, including mobile phones, personal computers, and tablet computers, include audio capture devices such as microphones and video cameras. The audio capture devices can capture audio from various audio sources such as people, animals, music, background sounds, etc. The captured video may include audio that can correspond to entities such as people, animals, landscapes, and / or objects.

[0002] Some audio capture devices and / or computing devices can correct or otherwise enhance audio content. For example, some audio capture devices can provide corrections to eliminate artifacts such as speech distortion, bandwidth reduction, elimination and / or suppression of certain frequency bands. After the captured audio has been corrected, the corrected audio can be saved, played, transmitted, and / or otherwise utilized. Summary of the Invention

[0003] In one aspect, a computing device can be configured to isolate any sound in the audio. Thus, sounds from certain sources can be enhanced while sounds from other sources can be suppressed. With the support of a system composed of machine learning components, the audio capture device can be configured to enable a user to enhance audio content.

[0004] In some aspects, a mobile device can be configured with these features such that audio can be enhanced substantially in real time. In some cases, the mobile device can automatically enhance the audio. In other aspects, a mobile phone user can non-destructively enhance the audio to match their preferences. Additionally, for example, pre-existing audio in a user's library can be enhanced based on the techniques described herein.

[0005] Accordingly, the present disclosure provides a neural network that separates tracks in an input audio waveform and determines whether the tracks match an input text description or caption. In some cases, the neural network can estimate the degree of overlap of each separated track and use the probability of this overlap to calculate an estimated track corresponding to the input text description. In some aspects, a joint embedding network (sometimes also referred to as SoundWords) can be trained to embed audio and text into an embedding in a shared representation that indicates the proximity of a particular audio embedding to a particular text embedding. In some aspects, SoundWords can be a text-audio embedding model trained by contrast. As described herein, the separation model can be trained to be controlled by a natural language input, such as, for example, "remove the dog barking in the distance". When such an input is received, the trained model can separate background audio including the dog barking, identify the track corresponding to the dog barking, remove or suppress the track corresponding to the dog barking, and provide background audio without the dog barking.

[0006] In one aspect, a computer-implemented method is provided. The method includes receiving, by a computing device, an input audio waveform and an input text description. The method further includes separating, by a neural network, the input audio waveform into a plurality of tracks. The method also includes determining, by the neural network, whether the input text description describes a track among the plurality of separated tracks. The method also includes, when it is determined that the input text description describes a track among the plurality of tracks, providing, by the computing device, the track corresponding to the input text description to an interactive user interface.

[0007] In another aspect, a computing device is provided. The computing device includes one or more processors and a data storage device. Computer-executable instructions are stored on the data storage device that, when executed by the one or more processors, cause the computing device to perform operations. The operations include receiving, by the computing device, an input audio waveform and an input text description. The operations further include separating, by a neural network, the input audio waveform into a plurality of tracks. The operations also include determining, by the neural network, whether the input text description describes a track among the plurality of separated tracks. The operations also include, when it is determined that the input text description describes a track among the plurality of tracks, providing, by the computing device, the track corresponding to the input text description to an interactive user interface.

[0008] In another aspect, an article is provided. The article includes one or more computer-readable media having computer-readable instructions stored thereon that, when executed by one or more processors of a computing device, cause the computing device to perform operations. The operations include receiving, by the computing device, an input audio waveform and an input text description. The operations further include separating the input audio waveform into a plurality of audio tracks by a neural network. The operations also include determining, by the neural network, whether the input text description describes an audio track among the plurality of separated audio tracks. The operations also include, when it is determined that the input text description describes an audio track among the plurality of audio tracks, providing, by the computing device, the audio track corresponding to the input text description to an interactive user interface.

[0009] In another aspect, a system is provided. The system includes means for receiving, by a computing device, an input audio waveform and an input text description; means for separating the input audio waveform into a plurality of audio tracks by a neural network; means for determining, by the neural network, whether the input text description describes an audio track among the plurality of separated audio tracks; and means for providing, by the computing device, the audio track corresponding to the input text description to an interactive user interface when it is determined that the input text description describes an audio track among the plurality of audio tracks.

[0010] In another aspect, a computer-implemented method is provided. The method includes receiving, by a computing device, training data including a plurality of audio segments and a plurality of audio text descriptions. The method further includes training a neural network based on the training data to perform operations of: receiving (i) an input audio waveform, and (ii) a text description of an audio track, separating the input audio waveform into a plurality of audio tracks, and determining whether the input text description describes an audio track among the plurality of separated audio tracks. The method also includes providing, by the computing device, the trained neural network.

[0011] In another aspect, a computing device is provided. The computing device includes one or more processors and a data storage device. The data storage device has computer-executable instructions stored thereon that, when executed by the one or more processors, cause the computing device to perform operations. The operations include receiving, by the computing device, training data including a plurality of audio segments and a plurality of audio text descriptions. The operations also include training a neural network based on the training data to perform operations of: receiving (i) an input audio waveform, and (ii) a text description of an audio track, separating the input audio waveform into a plurality of audio tracks, and determining whether the input text description describes an audio track among the plurality of separated audio tracks. The operations also include providing, by the computing device, the trained neural network.

[0012] In another aspect, an article is provided. The article includes one or more computer-readable media having computer-readable instructions stored thereon that, when executed by one or more processors of a computing device, cause the computing device to perform operations. The operations include receiving, by the computing device, training data including a plurality of audio segments and a plurality of audio text descriptions. The operations also include training a neural network based on the training data to perform the following operations: receiving (i) an input audio waveform, and (ii) a text description of a sound track, separating the input audio waveform into a plurality of sound tracks, and determining whether the input text description describes a sound track among the plurality of separated sound tracks. The operations also include providing, by the computing device, the trained neural network.

[0013] In another aspect, a system is provided. The system includes means for receiving, by a computing device, training data including a plurality of audio segments and a plurality of audio text descriptions; means for training a neural network based on the training data to perform the following operations: receiving (i) an input audio waveform, and (ii) a text description of a sound track, separating the input audio waveform into a plurality of sound tracks, and determining whether the input text description describes a sound track among the plurality of separated sound tracks; and means for providing, by the computing device, the trained neural network.

[0014] The foregoing summary is illustrative only and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments, and features described above, additional aspects, embodiments, and features will become apparent by reference to the figures and the following detailed description and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 An example inference phase of a neural network for separating audio corresponding to a text description according to an example embodiment is shown.

[0016] Figure 2 An example training phase of a neural network for separating audio corresponding to a text description according to an example embodiment is shown.

[0017] Figure 3 is a diagram showing a training phase and an inference phase of a machine learning model according to an example embodiment.

[0018] Figure 4 depicts a distributed computing architecture according to an example embodiment.

[0019] Figure 5 is a block diagram of a computing device according to an example embodiment.

[0020] Figure 6 depicts a network of a computing cluster arranged as a cloud-based server system according to an example embodiment.

[0021] Figure 7 is a flowchart of a method according to an example embodiment.

[0022] Figure 8 is another flowchart of a method according to an example embodiment. Detailed Description

[0023] In recent years, the progress of deep learning has driven some progress in audio enhancement technology. A particularly interesting topic is the problem of extracting one or more specific sounds from a single-channel (mono) audio input that may contain various different sounds. The specific sound can be described via a text input.

[0024] This document describes a method for separating a desired sound source from one or more audio mixtures based on a text description or caption. In some embodiments, this can be achieved by combining two models. A first sound separation model can be used to separate the audio tracks in a single-channel mixture. A second model can be trained to jointly embed both the separated audio tracks and the text description into the same embedding in a shared representation and determine a matching score. Some traditional models perform audio separation conditioned on captions. However, captions do not describe the audio exhaustively, resulting in noise in the conditioning information. Therefore, separating the audio tracks and then performing caption matching helps reduce and / or eliminate such errors generated in the conditioning network.

[0025] In some aspects, to enable a user to control audio enhancement features, the techniques described herein apply a neural network-based model to adjust audio content. The techniques described herein include receiving input audio and predicting output audio that enhances specific audio tracks and / or suppresses specific audio tracks. In some examples, the trained neural network model can operate on a variety of computing devices, including but not limited to mobile computing devices (e.g., smartphones, tablet computers, cellular phones, laptops), fixed computing devices (e.g., desktop computers), and server computing devices.

[0026] A training data set consisting of pairs of audio segments and text descriptions of the audio segments can be used to train a neural network to perform one or more aspects described herein. In some examples, the neural network can be arranged as an encoder / decoder neural network.

[0027] In one example, a copy of the trained neural network can reside on a mobile computing device. The mobile computing device can include a microphone that can capture input audio. The input audio can be provided to the trained neural network residing on the mobile computing device. In response, the trained neural network can generate predicted output audio that enhances and / or suppresses one or more audio tracks in the input audio. Then, the user of the mobile computing device can listen to the output audio.

[0028] Thus, the mobile computing device can enhance the input audio and then output the output audio (e.g., provide the output audio via an audio output device of the mobile computing device). In other examples, the trained neural network does not reside on the mobile computing device; instead, the mobile computing device (e.g., via the Internet or another data network) provides the input audio to a trained neural network located remotely. The remotely located neural network can process the input audio and provide the output audio to the mobile computing device. In other examples, non-mobile computing devices can also use the trained neural network to modify audio.

[0029] Thus, the techniques described herein can improve audio by applying more desirable and / or selectable audio enhancements, thereby enhancing the actual and / or perceived quality of the audio. Thus, enhancing the actual and / or perceived quality of the audio can provide benefits by making the audio clearer and less affected by certain background sounds. These techniques are very flexible and can thus be applied to a variety of audio enhancements including any sound source.

[0030] Introduction and Overview

[0031] Hybrid audio tracks are common in various environments. For example, common street sounds may include multiple independent tracks. Additionally, for example, the commentary in a stadium and / or the rulings made by referees, umpires, etc. may be drowned out by a cheering crowd. Although a person may be able to separate the various sounds, a computer may have difficulty performing this task, especially when the audio is provided only as a single-channel input and the goal is to focus on any sound (e.g., commentary or a specific street sound).

[0032] Typically, when the type of audio to be focused on is known in advance, a dedicated enhancement algorithm can be trained to separate that audio. In some cases, a system can be built to enhance speech and / or extract the audio corresponding to individual instruments. There are also models that can solve the general sound separation problem, which aims to separate a single-channel audio mixture into individual component signals regardless of their categories. These methods can be utilized, followed by a subsequent selection step that picks the component signals the user might want to focus on. Although this approach may be slightly more general than having a predefined set of target categories, it typically involves estimating each individual sound in the mixture, making the user interface more complex and consuming more computational resources compared to only enhancing the target sound of one type or category.

[0033] Some existing methods are based on imposing conditional constraints on the behavior of neural sound separation models to utilize additional information present in modalities other than audio. Specifically, audiovisual methods can be used for specific tasks (such as speech enhancement) and separating a limited set of categories (such as instruments). An audiovisual general separation model can also be trained to separate all sounds originating from visible objects on the screen. Other examples of conditional constraint inputs can include accelerometer data for improving speech enhancement, target speaker embeddings for speech separation, or sound categories for general separation, using a vector indicating one or more desired source categories or an audio snippet from the desired sound category. Conditional constraints can generally be performed by injecting the embeddings extracted from the conditional constraint inputs into the sound separation network.

[0034] General sound separation (such as the task of separating all sounds from a sound mixture regardless of their categories) can be achieved through supervised data. Some weakly supervised methods can use sound categories as weak labels. In some cases, fully unsupervised methods can learn directly from the raw sound mixture. Some example general sound separation models can be conditioned on images, audio, and / or text. However, such conditional constraints may lead to ambiguity, i.e., how the model will specifically interpret the provided conditional constraints. For example, if an image of an owl hooting is provided to the model, it may not be clear whether it can focus on the sounds of all birds, only the sounds made by the owl, or the sounds made by this particular individual owl. Additionally, for example, it may not be clear whether the model should focus on the hooting sound made by the owl rather than other sounds, such as the sound of the owl flapping its wings. Similar types of ambiguity can occur when the conditional constraint is based only on audio. There may be additional complexities in audio conditional constraints because some types of sounds can change significantly over time. For example, the noise when an engine starts is different from the continuous rumbling sound when the engine is running.

[0035] Textual conditional constraints of separation models typically rely on the text transcription of the target audio track. For example, text-based non-negative matrix factorization for speech separation can be performed. As another example, lyric-based vocal separation can also be performed. Some computer vision models involve image and natural language joint embedding models. Such models do not rely on a fixed ontology, but rather provide a natural interface to support cross-modal retrieval and zero-shot classification applications. Additionally, for example, some joint embedding models are built on separately trained and fixed neural audio embeddings and text embeddings. Although such joint embeddings have been evaluated in cross-modal retrieval tasks, these models have not been considered in separation or generative audio applications.

[0036] In some cases, separating arbitrary sounds from a mixture (referred to as "general sound separation") can be achieved for a fixed number of sounds. Conditional information about which sound classes are present can improve separation performance. The availability of the Free Universal Sound Separation (FUSS) dataset has extended the scope to separating a variable number of sounds, which can then be used to process more realistic data. Additionally, for example, specific sound classes can be extracted from an input sound mixture. Such methods typically use curated data containing isolated sounds for training, which limits their application in truly open-domain data and poses challenges such as, for example, annotation costs, accurate simulation of real acoustic mixtures, and / or biased datasets. Some of these challenges can be overcome by replacing the strong supervision of the reference source signal with weak supervision labels from related modalities such as sound classes, visual inputs, or spatial locations from multi-microphone recordings.

[0037] Generally, the general sound separation process is more effective when the audio track is first separated before determining the match with the text description. Therefore, a general sound separation model based on text-driven machine learning is needed, where a single-channel sound is first separated and the separated track is matched with the text description. For example, given any input audio, one or more audio tracks in the input mixture can be separated, and a probability score indicating the audio correspondence of each separated track can be estimated. A higher probability score indicates that the separated source corresponds to the text description, while a lower probability score indicates that the separated source does not correspond to the text description. The separated tracks are weighted according to their estimated probabilities and can then be added together to reconstruct the audio mixture. Since real-world audio may contain an unknown number of tracks belonging to an undefined class ontology, the machine learning model disclosed herein provides an effective solution to the text-to-audio matching problem. In some embodiments, an embedding model can be trained to co-embed an audio segment and its text description tightly in the same embedding space. The machine learning model disclosed herein does not limit the domain of the audio, such as, for example, musical instruments or human speakers.

[0038] Network Architecture

[0039] Figure 1 Illustrates an example inference stage 100 of a neural network for separating audio corresponding to a text description according to an example embodiment. In some embodiments, a computing device may receive an input audio waveform 105 and an input text description 130. The input audio waveform 105 may be input into an audio separation network 110. In some embodiments, the audio separation network 110 identifies one or more estimated audio tracks in the input audio waveform 105. For example, the input audio waveform 105 may be a mixture of multiple audio tracks. The audio separation network 110 estimates multiple audio tracks in the input audio waveform 105.

[0040] An example architecture of the audio separation network 110 may include learnable convolutional encoder layers and decoder layers. For example, each input audio waveform 105 may be 2.5 milliseconds (ms) of audio and may be represented by multiple coefficients that capture the audio features in the encoded mixture of the input audio waveform 105. Some embodiments may use the short-time Fourier transform as the representation. In some embodiments, the audio separation network 110 may process these coefficients and predict a mask, which is another neural representation of the input audio waveform 105. The mask may be multiplied with the coefficients, and an inverse transform may be applied to generate multiple estimated audio tracks or tracks 115.

[0041] In some implementations, mixture consistency projection may be applied to constrain the separated sources to add up to the input mixture. The audio separation network 110 may process an input mixture waveform of T samples and output estimated audio tracks 115. For example, M estimated audio tracks may be output as vectors of length , where . The masking network estimates M masks, which are multiplied with the activations of the encoded version of the input audio waveform 105. The final time-domain signal may be calculated by applying a decoder (e.g., a transposed convolutional layer) to the masked coefficients.

[0042] In some embodiments, the neural network includes an audio embedding network 120 to generate an audio embedding that includes a representation of the audio features in the estimated audio tracks 115. For each separated source m, time-domain audio samples Therefore, a neural network can be used to generate corresponding audio embeddings. For example, the audio embedding network 120 can operate on a mel spectrogram (e.g., a 1-second patch) to obtain a time-varying audio embedding of the audio track. In some embodiments, the audio embedding network 120 can use the MobileNet v1 architecture. This architecture can include stacked two-dimensional (2D) separable dilated convolutional blocks with a dense layer at the end. In some embodiments, the audio embedding network 120 can use the ResNet architecture. Additional and / or alternative audio embedding architectures can be used.

[0043] The audio embedding network 120 generates one or more audio embeddings 125. The input text description 130 is input into the text embedding network to generate a text embedding 135. In some embodiments, the input text description 130 can be a natural language description of the audio clip. For example, an audio clip of an owl's hooting may be associated with text descriptions such as "hooting of an owl", "an owl hooting", "owl hoot", etc. It should be noted that there may be multiple audio clips for each type of sound. For example, an audio clip of an owl hooting can include audio clips of different lengths from a single audio recording, audio clips corresponding to the same owl but at different times and / or locations, audio clips corresponding to different owls, and so on.

[0044] Furthermore, for example, the text description "owl hoot" can be associated with one or more audio clips corresponding to an owl hooting. However, the text description "hooting horn" may be associated with one or more audio clips of the sound made by a car horn, such as, for example, honking, rumbling, tooting, etc. As another example, "hoot with laughter" may be associated with one or more audio clips of human voices corresponding to scorn, disapproval, joy, etc.

[0045] The classifier 140 can match the text embedding 135 with one or more audio embeddings 125 to determine if a match exists. For example, the classifier 140 outputs a probability score for each estimated audio track 115 individually and assigns a value between 0 and 1. The probability score indicates the likelihood that a given estimated audio track corresponds to the input text description 130. A higher probability indicates a higher likelihood that a given estimated audio track corresponds to the input text description 130, while a lower probability indicates a lower likelihood that a given estimated audio track corresponds to the input text description 130. Thus, the classifier 140 can generate a label or ranking 145 indicating whether a given estimated audio track corresponds to the input text description 130.

[0046] In some implementations, a threshold probability score may indicate whether a given estimated audio track corresponds to the input text description 130. For example, a probability score exceeding the threshold probability score indicates a match with the input text description 130, while a probability score not exceeding the threshold probability score does not indicate a match with the input text description 130. Generally, the threshold probability score may depend on various factors, such as, for example, audio features such as sound type, source type, audio quality, environment type, and so on. Additionally, for example, in some implementations, a separated audio track in a given audio waveform may be associated with a threshold probability score. In some other implementations, more than one threshold probability score may be used for a given audio waveform. Additionally, for example, the threshold probability score may be a learnable parameter of a neural network.

[0047] In some embodiments, based on whether a particular audio track in one or more estimated audio tracks corresponds to the input text description 130, the computing device may modify the audio content of the particular audio track associated with the input text description 130. Additionally, for example, when determining that a separated audio track does not correspond to the input text description 130, the computing device may suppress the audio track. In some embodiments, when determining that a separated audio track does not correspond to the input text description 130, the computing device may provide the audio track to the user via an interactive user interface. Additionally, for example, the computing device may provide user-selectable options to enhance or suppress such an audio track. In some embodiments, when determining that a separated audio track does not correspond to the input text description 130, the computing device may provide the user with a notification via the interactive user interface that the audio track does not correspond to the input text description 130. Once the audio track is enhanced or suppressed, it can be mixed to recreate the audio track mixture, and the computing device may use the computing device to provide the modified audio content to the interactive user interface.

[0048] In some embodiments, a computing device may receive two or more text descriptions and identify the audio tracks corresponding to each of the two or more text descriptions. Generally, this correspondence may be one-to-many. For example, a first text description may correspond to a first set of separated audio tracks, and a second text description may correspond to a second set of separated audio tracks. The computing device may separate these audio tracks and label them. For example, the computing device may label the first set of separated audio tracks with the first text description, and / or label the second set of separated audio tracks with the second text description. Additionally and / or alternatively, the computing device may label any remaining separated audio tracks as not corresponding to the first and second text descriptions. An interactive user interface may be configured to provide user-friendly tools that enable a user to select the separated audio tracks and indicate which tracks (if any) need to be enhanced or suppressed. As used herein, the terms enhance or suppress generally may refer to modifying the signal corresponding to an audio track to adjust one or more frequencies (e.g., treble, bass, etc.), modifying the volume, removing noise from the signal, adding a time delay or offset to the signal, repeating or deleting certain portions of the signal, and so on. In some embodiments, the user may be provided with the ability to input whether the audio tracks corresponding to the text descriptions have been correctly identified. Such user feedback may be used as labeled data to train the neural networks (e.g., classifiers) described herein.

[0049] In some embodiments, a computing device (e.g., a mobile device) may generate a request to determine whether an input text description describes an audio track in an input waveform. The computing device (e.g., a mobile device) may then send the request from the computing device (e.g., a mobile device) to a second computing device (e.g., a cloud server). The second computing device (e.g., a cloud server) may include a trained version of a neural network. Subsequently, the computing device (e.g., a mobile device) may receive from the second computing device (e.g., a cloud server) a determination of whether the input text description describes an audio track in the input waveform. The computing device (e.g., a mobile device) may then output a modified version of the waveform (e.g., by enhancing the audio track corresponding to the input text description and / or suppressing the audio track corresponding to the input text description). In some embodiments, the computing device may include a microphone and use the microphone to generate audio content.

[0050] In some embodiments, providing an audio track involves providing user-selectable options via an interactive user interface to enhance or suppress an audio track corresponding to an input text description. For example, an interactive user experience can be provided where the user can adjust the audio content of an input audio waveform. This provides additional creative flexibility to the user, enabling them to find their own balance between sounds from different audio tracks based on a text description that describes one or more of the audio tracks in the audio track. For example, when the audio captures a conversation between two people, the computing device can separate the audio tracks (e.g., each person as an audio track). Options to enhance or suppress the separated audio track that matches a given text description (e.g., the name of the speaker) can be provided to the user via an interactive graphical user interface. In some embodiments, speech recognition technology can be applied to identify the speaker. The user can indicate the selection of a person, and the computing device can enhance the audio track corresponding to the selected person, and / or suppress the audio track corresponding to the other person. Additionally, for example, audio enhancement can be applied to existing audio files in the user's library. Although this example illustrates the enhancement and / or suppression of audio tracks corresponding to people, similar techniques can also be applied to other audio tracks. In some embodiments, for example, a musical composition can include audio tracks corresponding to different musical instruments. The user can input a text description corresponding to the name of the musical instrument (e.g., piano, viola, lead guitar, etc.). The techniques described herein can be applied to separate a musical composition into audio tracks and can identify the audio track corresponding to the text description. Subsequently, options to enhance or suppress the separated audio track that matches the text description can be presented to the user via an interactive user interface.

[0051] In some embodiments, receiving an input audio waveform and an input text description involves receiving the input text description from the user via an interactive user interface. For example, the interactive user interface can be an audio interface, and the user can input a text description of the audio track (e.g., using a microphone of the computing device). Additionally, for example, the interactive user interface can be a text input interface, and the user can input a text description of the audio track (e.g., using a text input medium such as a keyboard).

[0052] In some embodiments, receiving an input audio waveform and an input text description involves receiving the input audio waveform via an interactive user interface. For example, the interactive user interface can be an audio interface communicatively linked to a microphone of the computing device, and the microphone can be used to record the input audio waveform. In some embodiments, the recorded audio can be saved in a library. Audio separation can be performed in real-time when the input audio waveform is recorded, or audio separation can be performed on the recorded audio file.

[0053] Figure 2Shows an example training phase 200 of a neural network for separating audio corresponding to a text description according to an example embodiment. In some embodiments, the neural network can be a fully convolutional neural network as described herein. During training, the neural network can receive one or more input training datasets as input. The neural network can include node layers for processing the input audio and input text. Example layers can include, but are not limited to, an input layer, a convolutional layer, an activation layer, a pooling layer, and an output layer. The input layer can store the input data, such as the audio features of the input audio and the inputs from other layers of the neural network. The convolutional layer can compute the outputs of neurons connected to local regions in the input. In some examples, the predicted output can be fed back into the neural network again as input to perform iterative refinement. The activation layer can determine whether the output of the previous layer is "activated" or actually provided (e.g., provided to the next layer). The pooling layer can downsample the input. For example, the neural network can involve one or more pooling layers for downsampling the input by a predetermined factor (e.g., factor 2) in the horizontal dimension and / or the vertical dimension. The output layer can provide the output of the neural network to software and / or hardware that interfaces with the neural network; for example, to hardware and / or software for displaying, printing, transmitting, and / or otherwise providing the enhanced audio.

[0054] For illustrative purposes, the audio waveform 205 can be a mixture of one or more audio tracks. In some embodiments, the audio waveform 205 can include multiple audio mixes, audio mix 1, audio mix 2, and so on. During preprocessing and the like, some audio tracks may be filtered out. Additionally, for example, a mixture of mixtures (MoM) 205a of audio mixes can be generated by mixing multiple audio mixes (audio mix 1, audio mix 2, etc.). The audio MoM 205a can be input into the audio separation network 210. The audio separation network 210 can share one or more common aspects with the audio separation network 110.

[0055] The audio separation network 210 can separate the audio MoM 205a into estimated audio tracks 215. As shown, the estimated audio tracks 215 can include the separated audio tracks: separated audio 1, separated audio 2, …, separated audio M. In some embodiments, each separated audio track from the estimated audio tracks 215 can be input into the audio embedding network. The audio embedding network can share one or more common aspects with the audio embedding network 120. The audio embedding network can generate one or more audio embeddings, where each audio track (separated audio 1, separated audio 2, …, separated audio M) from the estimated audio tracks 215 corresponds to a respective separate audio embedding. For illustrative purposes, separated audio 1, separated audio 2, …, separated audio M can represent the separated audio tracks and / or the audio embeddings.

[0056] The input text description 220 can be subtitles for audio mixing 2. In some embodiments, the input text description 220 can be input into a text embedding network to generate a text embedding. In some embodiments, the input text description 220 can be a natural language description of an audio track. For illustrative purposes, the input text description 220 can represent the text input and / or the text embedding of the text input. The matching network 225 can match the input text description 220 (such as subtitles for audio mixing 2) with the separated audio tracks 215 (such as, separated audio 1, separated audio 2, …, separated audio M) to determine if there is a match. In some embodiments, the text embedding can be matched with one or more audio embeddings to determine if there is a match. For example, the matching network 225 can output a probability score for each estimated audio track 215 individually and assign a value between 0 and 1. The probability score indicates the likelihood that a given estimated audio track (such as separated audio 1, separated audio 2, …, separated audio M) corresponds to the input text description 220 (such as subtitles for audio mixing 2). A higher probability indicates a higher likelihood that the given estimated audio track corresponds to the input text description 220, while a lower probability indicates a lower likelihood that the given estimated audio track corresponds to the input text description 220. Thus, the matching network 225 can generate a label or ranking indicating whether a given estimated audio track corresponds to the input text description 220. In some embodiments, additional and / or alternative subtitles can be input. For example, the input text description 220 can be subtitles for audio mixing 1, and the model can be trained on such subtitles. In some embodiments, more than one subtitle can be input. For example, a first subtitle for audio mixing 1 and a second subtitle for audio mixing 2 can be provided.

[0057] For example, separated audio 1 can be associated with a relatively high probability score. Similarly, separated audio 2 can be associated with a relatively low probability score. Thus, the neural network predicts that separated audio 1 is likely to correspond to the input text description 220, while separated audio 2 is likely not to correspond to the input text description 220.

[0058] In some embodiments, the training of the neural network can include unsupervised Mixing Invariant Training (MixIT). Specifically, the training can involve determining a MixIT loss 230 that depends only on audio information. For example, the MixIT loss 230 can involve comparing the estimated audio tracks 215 with the input audio waveform 205 to minimize the loss. For example, the sum of the estimated audio tracks 215 can be determined and compared with the input audio waveform 205 to determine if all the audio tracks in the input audio waveform 205 have been accounted for.

[0059] For example, a neural network can be trained with a MixIT loss 230 that measures the fidelity between the sum of the separated sources (e.g., estimated track 225) assigned by MixIT (e.g., MixIT assignment 235 of audio mix 2) and the reference mixture (e.g., input audio waveform 205). For example, the MixIT separation loss 230 can be used to optimize the assignment of the estimated tracks to the sum of two reference mixtures as follows:

[0060]

[0061] where the mixing matrix is constrained to a set of binary matrices where the sum of each column is 1. Due to the constraint on each source can only be assigned to one reference mixture. In some embodiments, the audio separation network 210 can be pre-trained using mixture invariant training (MIxIT) on a large amount of data, which is a method of training a separation model on the original audio mixtures.

[0062] In some embodiments, the neural network can be trained according to a caption matching loss 245 based on the difference between one or more estimated tracks 215 and the input text description 220. For example, various subsets of the separated tracks 215 can be determined. For each such subset, the sum of the respective tracks can be input into an audio embedding network to generate an audio embedding for each track subset. Subsequently, the distance between the text embedding of the caption and the corresponding audio embedding of each track subset can be determined, for example, in the shared representation generated by the joint embedding network. At block 240, a subset of the separated tracks can be selected from the estimated tracks 215, where the subset minimizes the distance between the text embedding of the caption and the corresponding audio embedding corresponding to the selected subset.

[0063] Then the input text description 220 can be compared with the subset of the separated tracks to determine the caption matching loss 245. In some embodiments, the caption matching loss 245 can consider a text caption matching loss. Generally, training the neural network involves minimizing the caption matching loss 245 such that the caption can provide a more accurate match to the separated tracks.

[0064] Generally, the caption matching loss 245 cannot be computationally intensive. For example, since there are M separated tracks, the number of possible subset combinations will be on the order of 2^M. In some embodiments, a smaller joint embedding network can be trained to identify the subset that minimizes the distance. For example, the trained model can identify one or more relationships between the audio embeddings of the separated tracks and the audio embedding of the audio mixture including the separated tracks. This relationship can be utilized when training the smaller joint embedding network.

[0065] In some embodiments, the audio mix 1 can also be associated with subtitles. Accordingly, a subtitle matching loss 245 can be applied based on the subtitles for the audio mix 1 to ensure that there is no overlap between the corresponding assignments of subtitle matching for the audio mix 1 and the subtitle matching for the audio mix 2. This can provide additional training examples for training the matching network 225.

[0066] In some embodiments, the neural network can be trained according to a classification loss 260. For example, one or more subtitle matching assignments 250 can be compared with the predicted subtitle matching assignments 255. For example, one or more subtitle matching assignments 250 can be a subset of the MixIT assignment 235 for the audio mix 2. Additionally, for example, the predicted subtitle matching assignments 255 can be a set of probabilities generated by the matching network 225, where these probabilities indicate the likelihood that the separated audio tracks correspond to the subtitles for the audio mix 2. The classification loss 260 can indicate the degree of match between the subtitle matching assignment 250 and the predicted subtitle matching assignment 255. Generally, training of the neural network involves minimizing the classification loss 260 such that the predicted matches are more consistent with the subtitle matching assignments.

[0067] In some embodiments, training data including multiple audio mixes and multiple audio text descriptions can be received. For example, a dataset including pairs of audio mixes and text descriptions can be received. The audio mixes can be audio samples of sounds such as, for example, animal sounds, bird calls, sounds emitted by musical instruments, engine sounds of various types of vehicles, water sounds in different environments, human speech, and other human voices emitted by humans, natural sounds, and so on. As a labeled example, the audio tracks in each audio mix can be associated with a text description of the audio.

[0068] In some embodiments, audio extraction can be performed by a model including one or more components. In some embodiments, the audio embedding network and the text embedding network can be components of (or integrated into) a joint embedding network that generates a shared representation including joint embeddings, where the audio embedding of a given audio track is within a threshold distance of the text embedding of the text description of the given audio track. In some embodiments, the joint embedding network can be a frozen network. For example, the sound extraction model can include a joint embedding network that obtains text- or audio-based input vectors and computes corresponding embeddings. For example, the joint embedding network can be trained on pairs (audio, text) that can include various scales, supervision qualities, and language usage patterns. As an illustrative example, a first set of approximately 50 million 10-second sound segments extracted from randomly selected internet videos (1 segment / video) can be utilized. The associated text can be the names and natural language descriptions of knowledge graph entities associated with each video.

[0069] As another illustrative example, the second set can be samples from the AudioSet dataset, which includes approximately two million 10-second audio clips. The dataset covers approximately 527 sound event categories. For example, AudioSet consists of: an extended ontology containing 332 audio event categories, and a collection of 2,084,320 human-labeled 10-second sound clips extracted from YouTube videos. The ontology is specified as a hierarchical graph of event categories, covering a variety of human and animal sounds, musical instruments and genres, and common everyday environmental sounds. In some aspects, the natural language diversity can be restricted here to one or two label variants per category. The associated text can include (i) standard human-readable labels associated with each audio clip in the dataset; (ii) a collection of human-provided natural language sound descriptions that can be gathered (e.g., approximately 50K AudioSet clips); and (iii) the AudioCaps dataset, which includes natural language descriptions of approximately 46K AudioSet clips. AudioCaps is a large-scale dataset containing approximately 46K audio clip and human-written text pairs collected via crowdsourcing on the AudioSet dataset.

[0070] As another example, the third set can include approximately 110K sound events and / or scene recordings from a library (such as a professional sound effects library). Each recording can be paired with a text description. In some embodiments, the text description can be an unstructured list of labels, a multi-sentence natural language description, etc.

[0071] Generally, the two models, the audio separation network 210 and the joint embedding network (e.g., including an audio embedding network and a text embedding network), can be trained separately because they can involve different types of training data. For example, the joint embedding network can be trained on audio samples paired with corresponding text descriptions, and the audio separation network 210 can be trained using single-channel audio mixtures.

[0072] In some embodiments, a neural network can be trained based on training data to perform the following operations: receive (i) an input audio waveform 205, and (ii) a text description 220 of an audio track, separate the input audio waveform into multiple audio tracks, and determine whether the input text description describes an audio track among the multiple separated audio tracks. For example, the neural network can receive an input audio waveform 205 from which one or more audio tracks are to be separated or filtered out. In some embodiments, the input audio waveform 205 can be a single-channel mixed audio. For example, the input audio waveform 205 can be a single-channel audio including one or more audio tracks such as a car horn sound, an owl's call, human voices in a nearby store, and the sound of a car engine. Additionally, for example, the neural network can receive a text description 220 of an audio track. For example, the neural network can receive a text input such as "car engine". In some embodiments, the audio separation network 210 can separate the audio track into estimated audio tracks 215 and can embed the text input 220 into a shared representation (e.g., by using a joint embedding network).

[0073] In some embodiments, the training of the neural network can involve training the neural network to generate a first audio corresponding to the text description from the input audio waveform (e.g., as an enhanced output), and generate a second audio not corresponding to the text description from the input audio waveform (e.g., as a suppressed output).

[0074] In some embodiments, the generation of the shared representation includes training a joint embedding network. The training of the joint embedding network can be based on training data including multiple pairs, each pair including an audio segment and a text description of the audio segment. For example, an audio segment and a text description of the audio segment can be received. The audio segment can be an audio sample of a sound such as, for example, an animal sound, a bird call, a sound emitted by a musical instrument, engine sounds of various types of vehicles, water sounds in different environments, human voices, and other human sounds emitted by humans, natural sounds, and so on. Each audio segment can be associated with a text description. Thus, the multiple pairs can be "(audio segment, text description)". Such pairings are typically many-to-many mappings. In other words, each audio segment can be paired with multiple text descriptions, and each text description can correspond to multiple audio segments.

[0075] In some embodiments, the text description can be a natural language description of an audio clip. For example, an audio clip of an owl's hooting may be associated with text descriptions such as "hooting of an owl", "an owl hooting", "owl hoot", etc. Thus, multiple pairs can include "(audio clip of an owl hooting, hooting of an owl)" or "(audio clip of an owl hooting, owl hoot)", and so on. It should be noted that there may be multiple audio clips of each type of sound. For example, the audio clips of an owl hooting can include audio clips of different lengths from a single audio recording, audio clips corresponding to the same owl but at different times and / or locations, audio clips corresponding to different owls, and so on. Additionally, for example, the text description can describe multiple audio clips.

[0076] In some embodiments, the joint embedding network can be a joint embedding model of natural language and sound, which includes dedicated towers for each, but can be trained to terminate in the same target space. Similar to the prior art, contrastive learning can be performed using a large set of audio clips paired with their associated natural language descriptions. By applying a contrastive multi-view coding loss function, the learned embeddings can co-localize the audio and its underlying semantic categories expressed via free-form natural language in the same target space.

[0077] In some embodiments, the training of the joint embedding network can be warm-started from separate audio encoder models and text encoder models. For example, the joint embedding network can include dedicated towers, such as an audio embedding tower for generating audio embeddings and a text embedding tower for generating text embeddings. As an illustrative example, the audio embedding tower can include a modified Resnet-50 (e.g., Resnetish-50) architecture trained on multiple 64-channel log mel spectrograms. Mel encoding can be performed by converting the signal to a mel-scale magnitude spectrogram and then converting back using the original input phase. One or more variants of mel encoding can be used, such as narrowband (where mel bins cover frequencies in the range of 60 to 5000 Hz) and wideband (where mel bins cover frequencies in the range of 0 to 16,000 Hz). Additional and / or alternative frequency bands can be utilized. In some embodiments, the audio embedding tower can be pre-trained. In some embodiments, the final classifier layer of the Resnetish-50 architecture can be replaced with a fully-connected final layer having a number of units corresponding to the dimension of the shared representation. For example, the fully-connected final layer can map to a shared embedding space (e.g., a 64-dimensional space). As another illustrative example, the audio embedding tower can take short segments (e.g., 10-second segments) as input and apply average pooling to the frame-level outputs to produce a single segment-level embedding.

[0078] In some embodiments, for the text embedding tower, an embedding model based on the publicly available case-insensitive Transformer-based Bidirectional Encoder Representations from Transformers (BERT) can be used. As an illustrative example, a transformer architecture based on special classification token (CLS-token) pooling, along with a fully-connected final layer for mapping to the shared representation, can be used. CLS is the last hidden state of BERT. For example, the fully-connected final layer can map to a shared embedding space (e.g., a 64-dimensional space).

[0079] In some embodiments, after model initialization, the parameters in the audio embedding tower and the text embedding tower can be updated during training. Each batch can be constructed with a preset mixture of training data sources. For illustrative purposes, the preset mixture can include 30% from the 50M video dataset, 5% from the 50K collected caption dataset, 25% from the AudioCaps dataset, 10% from the AudioSet label dataset, and 30% from the professional sound dataset. Other combinations are also considered within the general scope of this specification. In some embodiments, these ratios may not be optimized for training. The training loss can be calculated on examples across multiple accelerators, resulting in an effective batch size (e.g., 6144) for the target audio-text pairs. In some embodiments, the Adam optimizer with a learning rate of 5e-5 can be used, and its trainable softmax temperature has an initial value (e.g., 0.1).

[0080] In some embodiments, the joint embedding network includes a symmetric encoder-decoder U-net neural network with skip-connections. For example, the architecture can include a symmetric encoder-decoder U-net network with skip-connections operating in the waveform domain, where the architecture of the decoder layer mirrors the structure of the encoder, and the skip-connections run between each encoder block and its mirrored decoder block.

[0081] Training a machine learning model for generating inferences / predictions

[0082] Figure 3 FIG. 300 according to an example embodiment is shown, which shows the training phase 302 (e.g., as Figure 2 shown) and the inference phase 304 (e.g., as Figure 1 shown) of the trained machine learning model 332. Some machine learning techniques involve training one or more machine learning algorithms on a set of input training data to identify patterns in the training data and provide output inferences and / or predictions regarding the patterns in the training data. The resulting trained machine learning algorithm can be referred to as a trained machine learning model. For example, Figure 3 FIG. 302 shows the training phase, where one or more machine learning algorithms 320 are being trained on training data 310 to become the trained machine learning model 332. Then, during the inference phase 304, the trained machine learning model 332 can receive input data 330 (e.g., input audio waveform, input text description) and one or more inference / prediction requests 340 (possibly as part of the input data 330), and correspondingly provide one or more inferences and / or predictions 350 as output (e.g., predicting whether the audio track in the input audio waveform corresponds to the input text description).

[0083] Thus, the trained machine learning model 332 can include one or more models of one or more machine learning algorithms 320. The machine learning algorithms 320 can include, but are not limited to: artificial neural networks (e.g., convolutional neural networks, recurrent neural networks, Bayesian networks, hidden Markov models, Markov decision processes, logistic regression functions, support vector machines, suitable statistical machine learning algorithms, and / or heuristic machine learning systems as described herein). The machine learning algorithms 320 can be supervised or unsupervised and can implement any suitable combination of online learning and offline learning.

[0084] In some examples, the machine learning algorithms 320 and / or the trained machine learning model 332 can be accelerated using an on-device co-processor such as a graphics processing unit (GPU), a tensor processing unit (TPU), a digital signal processor (DSP), and / or an application specific integrated circuit (ASIC). Such an on-device co-processor can be used to accelerate the machine learning algorithms 320 and / or the trained machine learning model 332. In some examples, the trained machine learning model 332 can be trained, resident, and executed to provide inferences on a specific computing device, and / or can otherwise make inferences for a specific computing device.

[0085] During the training phase 302, the machine learning algorithms 320 can be trained by using unsupervised, supervised, semi-supervised, and / or weakly supervised learning techniques by at least providing training data 310 as training input. Unsupervised learning involves providing a portion (or all) of the training data 310 to the machine learning algorithms 320, and the machine learning algorithms 320 determining one or more output inferences based on the provided portion (or all) of the training data 310. Supervised learning involves providing a portion of the training data 310 to the machine learning algorithms 320, where the machine learning algorithms 320 determine one or more output inferences based on the provided portion of the training data 310 and accept or correct the output inferences based on the correct results associated with the training data 310. In some examples, the supervised learning of the machine learning algorithms 320 can be managed by a set of rules and / or a set of labels for the training input, and the set of rules and / or the set of labels can be used to correct the inferences of the machine learning algorithms 320.

[0086] Semi-supervised learning involves having correct labels for a portion (but not all) of the training data 310. During semi-supervised learning, supervised learning is used for a portion of the training data 310 that has correct results, and unsupervised learning is used for a portion of the training data 310 that does not have correct results. In some examples, the machine learning algorithms 320 and / or the trained machine learning model 332 can be trained using other machine learning techniques, including but not limited to incremental learning and curriculum learning.

[0087] In some examples, the machine learning algorithm 320 and / or the trained machine learning model 332 can use transfer learning techniques. For example, transfer learning techniques can involve pre-training the trained machine learning model 332 on one data set and additionally training it using the training data 310. More specifically, the machine learning algorithm 320 can be pre-trained on data from one or more computing devices, and the resulting trained machine learning model is provided to a specific computing device, where the specific computing device is intended to execute the trained machine learning model during the inference phase 304. Then, during the training phase 302, the pre-trained machine learning model can be additionally trained using the training data 310, where the training data 310 can be derived from the kernel data and non-kernel data of the specific computing device. This further training of the machine learning algorithm 320 and / or the pre-trained machine learning model using the training data 310 of the specific computing device data can be performed using supervised learning or unsupervised learning. Once the machine learning algorithm 320 and / or the pre-trained machine learning model has been trained at least on the training data 310, the training phase 302 can be completed. The resulting trained machine learning model can be used as at least one of the trained machine learning models 332.

[0088] In particular, once the training phase 302 has been completed, the trained machine learning model 332 can be provided to the computing device (if not already on the computing device). The inference phase 304 can begin after the trained machine learning model 332 has been provided to the specific computing device.

[0089] During the inference phase 304, the trained machine learning model 332 can receive the input data 330 and generate and output one or more corresponding inferences and / or predictions 350 regarding the input data 330. Thus, the input data 330 can be used as an input to the trained machine learning model 332 for providing the corresponding inferences and / or predictions 350 to the kernel components and non-kernel components. For example, the trained machine learning model 332 can generate inferences and / or predictions 350 in response to one or more inference / prediction requests 340. In some examples, the trained machine learning model 332 can be executed as part of other software. For example, the trained machine learning model 332 can be executed by an inference or prediction daemon to be readily available for providing inferences and / or predictions upon request. The input data 330 can include data from the specific computing device on which the trained machine learning model 332 is executed and / or input data from one or more computing devices other than the specific computing device.

[0090] The input data 330 can include an audio waveform and a text description of the audio. The audio waveform can include sounds corresponding to various objects.

[0091] The inference and / or prediction 350 may include audio embeddings and / or text embeddings, predictions, estimated audio tracks, and / or other output data generated by a trained machine learning model 332 operating on input data 330 (and training data 310). In some examples, the trained machine learning model 332 may use the output inference and / or prediction 350 as input feedback 360. The trained machine learning model 332 may also rely on past inferences as input for generating new inferences.

[0092] An audio separation network, an audio embedding network, a text embedding network, a joint embedding network, a classifier, etc. may be examples of the machine learning algorithm 320. After training, a trained version of such a neural network may be an example of the trained machine learning model 332. In such an approach, an example of an inference / prediction request 340 may be a request to predict whether an audio track in a single-channel audio waveform corresponds to an input text description, and a corresponding example of the inference and / or prediction 350 may be an output indicating that the audio track in the single-channel audio waveform corresponds to the input text description. In some examples, a given computing device may include a trained neural network 300, possibly after training a neural network. Then, the given computing device may receive a request to predict whether an audio track in a single-channel audio waveform corresponds to an input text description and use the trained neural network to generate a prediction.

[0093] In some examples, two or more computing devices may be used to provide a prediction; for example, a first computing device may generate and send a request to predict whether an audio track in a monophonic audio waveform corresponds to an input text description. Then, a second computing device may use a trained version of the neural network (possibly after training) to generate a prediction and respond to the request from the first computing device. When receiving a response to the request, the first computing device may provide the requested output (e.g., using a user interface and / or display, a printed copy, an electronic communication, etc.).

[0094] Example data network

[0095] Figure 4Depicts a distributed computing architecture 400 according to an example embodiment. The distributed computing architecture 400 includes server devices 708, 410 configured to communicate with programmable devices 404a, 404b, 404c, 404d, 404e via a network 406. The network 406 can correspond to a local area network (LAN), a wide area network (WAN), a wireless local area network (WLAN), a wireless wide area network (WWAN), a corporate intranet, the public Internet, or any other type of network configured to provide a communication path between networked computing devices. The network 406 can also correspond to a combination of one or more LANs, WANs, corporate intranets, and / or the public Internet.

[0096] Although Figure 4 only five programmable devices are shown, the distributed application architecture can serve dozens, hundreds, or thousands of programmable devices. Additionally, the programmable devices 404a, 404b, 404c, 404d, 404e (or any additional programmable devices) can be any kind of computing device, such as a mobile computing device, a desktop computer, a wearable computing device, a head-mounted device (HMD), a network terminal, a mobile computing device, and so on. In some examples, as shown for programmable devices 404a, 404b, 404c, 404e, the programmable devices can be directly connected to the network 406. In other examples, as shown for programmable device 404d, the programmable device can be indirectly connected to the network 406 via an associated computing device (such as programmable device 404c). In this example, the programmable device 404c can act as the associated computing device for relaying electronic communications between the programmable device 404d and the network 406. In other examples, as shown for programmable device 404e, the computing device can be part of and / or located inside a vehicle (such as a car, a truck, a bus, a boat or ship, an airplane, etc.). In Figure 4 other examples not shown, the programmable device can be both directly and indirectly connected to the network 406.

[0097] The server devices 408, 410 can be configured to perform one or more services requested by the programmable devices 404a - 404e. For example, the server devices 408 and / or 410 can provide content to the programmable devices 404a - 404e. The content can include, but is not limited to, web pages, hypertext, scripts, binary data (such as compiled software), images, audio, and / or video. The content can include compressed and / or uncompressed content. The content can be encrypted and / or decrypted. Other types of content are also possible.

[0098] As another example, server devices 408 and / or 410 can provide access to software for databases, search, computing, graphics, audio, video, World Wide Web / Internet utilization, and / or other functions to programmable devices 404a - 404e. Many other examples of server devices are also possible.

[0099] Computing device architecture

[0100] Figure 5 is a block diagram of an example computing device 500 according to an example embodiment. Specifically, Figure 5 the illustrated computing device 500 can be configured to perform at least one function and / or method 700, 800 related to and / or associated with the neural network described herein.

[0101] The computing device 500 can include a user interface module 501, a network communication module 502, one or more processors 503, a data storage device 504, one or more cameras 518, one or more sensors 520, and a power system 522, all of which can be linked together via a system bus, network, or other connection mechanism 505.

[0102] The user interface module 501 may be operable to send data to and / or receive data from an external user input / output device. For example, the user interface module 501 can be configured to send data to and / or receive data from a user input device such as a touch screen, computer mouse, keyboard, keypad, touchpad, trackball, joystick, voice recognition module, and / or other similar devices. The user interface module 501 can also be configured to provide output to a user display device such as one or more cathode ray tubes (CRTs), liquid crystal displays, light - emitting diodes (LEDs), displays using digital light processing (DLP) technology, printers, light bulbs, and / or other similar devices whether now known or later developed. The user interface module 501 can also be configured to generate an audible output using devices such as speakers, speaker jacks, audio output ports, audio output devices, headphones, and / or other similar devices. The user interface module 501 can further be configured with one or more haptic devices that can generate haptic output such as vibration and / or other output detectable by touch and / or physical contact with the computing device 500. In some examples, the user interface module 501 can be used to provide a graphical user interface (GUI) for utilizing the computing device 500. For example, the user interface module 501 can be used to provide selectable audio tracks, where the selectable audio tracks are separated from the input audio waveform. Additionally, for example, the user interface module 501 can be used to receive a user's selection of an audio track. In some embodiments, the selected audio track can be enhanced or suppressed.

[0103] The network communication module 502 may include one or more devices that provide one or more wireless interfaces 507 and / or one or more wired interfaces 508 that are configurable to communicate via a network. The wireless interface 507 may include one or more wireless transmitters, receivers, and / or transceivers, such as Bluetooth™ transceivers, Zigbee® transceivers, Wi-Fi™ transceivers, WiMAX™ transceivers, LTE™ transceivers, and / or other types of wireless transceivers that can be configured to communicate via a wireless network. The wired interface 508 may include one or more wired transmitters, receivers, and / or transceivers, such as Ethernet transceivers, Universal Serial Bus (USB) transceivers, or similar transceivers that can be configured to communicate via twisted pair, coaxial cable, fiber optic link, or similar physical connections to a wired network.

[0104] In some examples, the network communication module 502 may be configured to provide reliable, protected, and / or authenticated communication. For each communication described herein, information for facilitating reliable communication (e.g., ensuring message delivery) may be provided, which may be part of a message header and / or message tail (e.g., packet / message sequencing information, encapsulation headers and / or encapsulation tails, size / time information, and transmission verification information such as cyclic redundancy check (CRC) and / or parity values). One or more cryptographic protocols and / or algorithms may be used to secure (e.g., encode or encrypt) the communication and / or decrypt / decode the communication, such as but not limited to the Data Encryption Standard (DES), Advanced Encryption Standard (AES), Rivest-Shamir-Adelman (RSA) algorithm, Diffie-Hellman algorithm, Secure Sockets Protocol (such as Secure Sockets Layer (SSL) or Transport Layer Security (TLS)), and / or Digital Signature Algorithm (DSA). Other cryptographic protocols and / or algorithms may be used to protect (and then decrypt / decode) the communication in addition to those listed herein.

[0105] The one or more processors 503 may include one or more general-purpose processors and / or one or more special-purpose processors (e.g., digital signal processors, tensor processing units (TPU), graphics processing units (GPU), application-specific integrated circuits, etc.). The one or more processors 503 may be configured to execute computer-readable instructions 506 contained in the data storage device 504 and / or other instructions as described herein.

[0106] The data storage device 504 may include one or more non-transitory computer-readable storage media that can be read and / or accessed by at least one of the one or more processors 503. The one or more computer-readable storage media may include volatile and / or non-volatile storage components that are integrated, in whole or in part, with at least one of the one or more processors 503, such as optical, magnetic, organic, or other memory or disk storage devices. In some examples, the data storage device 504 may be implemented using a single physical device (e.g., one optical, magnetic, organic, or other memory or disk storage unit), while in other examples, the data storage device 504 may be implemented using two or more physical devices.

[0107] The data storage device 504 may include computer-readable instructions 506 and possibly additional data. In some examples, the data storage device 504 may include storage required to execute at least a portion of the methods, scenarios, and techniques described herein and / or at least a portion of the functionality of the devices and networks described herein. In some examples, the data storage device 504 may include storage for a trained neural network model 512 (e.g., a model of a trained neural network such as the network described in the inference phase 100). Specifically, in these examples, the computer-readable instructions 506 may include instructions that, when executed by the processor 503, cause the computing device 500 to provide some or all of the functionality of the trained neural network model 512.

[0108] In some examples, the computing device 500 may include one or more cameras 518. The cameras 518 may include one or more image capture devices equipped to capture video, such as still cameras and / or video cameras. The one or more images may be one or more of the images used in the video footage. The cameras 518 may capture light and / or electromagnetic radiation emitted as visible light, infrared radiation, ultraviolet light, and / or as light at one or more other frequencies.

[0109] In some examples, computing device 500 may include one or more sensors 520. The sensors 520 may be configured to measure conditions within the computing device 500 and / or conditions in the environment of the computing device 500 and provide data regarding such conditions. For example, the sensors 520 may include one or more of the following: (i) sensors for obtaining data regarding the computing device 500, such as but not limited to a thermometer for measuring the temperature of the computing device 500, a battery sensor for measuring the charge of one or more batteries of the power system 522, and / or other sensors for measuring conditions of the computing device 500; (ii) identification sensors for identifying other objects and / or devices, such as but not limited to radio frequency identification (RFID) readers, proximity sensors, one-dimensional barcode readers, two-dimensional barcode (e.g., quick response (QR) code) readers, and laser trackers, where the identification sensors may be configured to read identifiers (such as RFID tags, barcodes, QR codes) and / or other devices and / or objects configured to be read and provide at least identification information; (iii) sensors for measuring the position and / or movement of the computing device 500, such as but not limited to tilt sensors, gyroscopes, accelerometers, Doppler sensors, GPS devices, sonar sensors, radar devices, laser displacement sensors, and compasses; (iv) environmental sensors for obtaining data indicative of the environment of the computing device 500, such as but not limited to infrared sensors, optical sensors, light sensors, biosensors, capacitance sensors, touch sensors, temperature sensors, wireless sensors, radio sensors, motion sensors, microphones, sound sensors, ultrasonic sensors, and / or smoke sensors; and / or (v) force sensors for measuring one or more forces acting on the computing device 500 (e.g., inertial forces and / or gravity), such as but not limited to one or more sensors for measuring force, torque, ground force, friction in one or more dimensions, and / or a zero moment point (ZMP) sensor for identifying the ZMP and / or the location of the ZMP. Many other examples of the sensors 520 are possible.

[0110] The power supply system 522 may include one or more batteries 524 and / or one or more external power interfaces 526 for supplying power to the computing device 500. Each of the one or more batteries 524 may act as a stored power source for the computing device 500 when electrically coupled to the computing device 500. The one or more batteries 524 of the power supply system 522 may be configured to be portable. Some or all of the one or more batteries 524 may be removably from the computing device 500 easily. In other examples, some or all of the one or more batteries 524 may be located inside the computing device 500 and thus may not be removably from the computing device 500 easily. Some or all of the one or more batteries 524 may be rechargeable. For example, a rechargeable battery may be charged via a wired connection between the battery and another power supply device (such as via one or more power supply devices located outside the computing device 500 and connected to the computing device 500 via one or more external power interfaces). In other examples, some or all of the one or more batteries 524 may be non-rechargeable batteries.

[0111] One or more external power interfaces 526 of the power supply system 522 may include one or more wired power interfaces that implement a wired power connection to one or more power supply devices located outside the computing device 500, such as a USB cable and / or a power cord. One or more external power interfaces 526 may include one or more wireless power interfaces that implement a wireless power connection to one or more external power supply devices, such as a Qi wireless charger. Once a power connection to an external power source is established using one or more external power interfaces 526, the computing device 500 may draw power from the external power source through the established power connection. In some examples, the power supply system 522 may include associated sensors, such as battery sensors associated with one or more batteries or other types of power sensors.

[0112] Cloud-based server

[0113] Figure 6 A network 406 of computing clusters 609a, 609b, 609c arranged as a cloud-based server system according to an example embodiment is depicted. The computing clusters 609a, 609b, 609c may be cloud-based devices that store program logic and / or data for cloud-based applications and / or services; for example, performing at least one function and / or method 700, 800 related to and / or associated with a neural network.

[0114] In some embodiments, the computing clusters 609a, 609b, 609c can be a single computing device residing in a single computing center. In other embodiments, the computing clusters 609a, 609b, 609c can include multiple computing devices in a single computing center, or even multiple computing devices in multiple computing centers located in different geographical locations. For example, Figure 6 depicts each of the computing clusters 609a, 609b, and 609c residing in different physical locations.

[0115] In some embodiments, the data and services at the computing clusters 609a, 609b, 609c can be encoded as computer-readable information stored in a non-transitory tangible computer-readable medium (or computer-readable storage medium) and accessible by other computing devices. In some embodiments, the computing clusters 609a, 609b, 609c can be stored on a single disk drive or other tangible storage medium, or can be implemented on multiple disk drives or other tangible storage media located at one or more different geographical locations.

[0116] Figure 6 depicts a cloud-based server system according to an example embodiment. In Figure 6 it, the functionality of the neural network and / or computing devices can be distributed among the computing clusters 609a, 609b, 609c. The computing cluster 609a can include one or more computing devices 600a, a cluster storage array 610a, and a cluster router 611a connected via a local cluster network 612a. Similarly, the computing cluster 609b can include one or more computing devices 600b, a cluster storage array 610b, and a cluster router 611b connected via a local cluster network 612b. Likewise, the computing cluster 609c can include one or more computing devices 600c, a cluster storage array 610c, and a cluster router 611c connected via a local cluster network 612c.

[0117] In some embodiments, each of the computing clusters 609a, 609b, and 609c can have an equal number of computing devices, an equal number of cluster storage arrays, and an equal number of cluster routers. However, in other embodiments, each computing cluster can have a different number of computing devices, a different number of cluster storage arrays, and a different number of cluster routers. The number of computing devices, cluster storage arrays, and cluster routers in each computing cluster can depend on one or more computing tasks assigned to each computing cluster.

[0118] For example, in computing cluster 609a, computing device 600a may be configured to perform various computing tasks of neural networks, text embedding towers, audio embedding towers, joint embedding networks, audio separation networks, classifiers, and / or computing devices. In one embodiment, the various functionalities of neural networks, text embedding towers, audio embedding towers, joint embedding networks, audio separation networks, classifiers, and / or computing devices may be distributed among one or more of computing devices 600a, 600b, 600c. Computing devices 600b and 600c in corresponding computing clusters 609b and 609c may be configured similarly to computing device 600a in computing cluster 609a. On the other hand, in some embodiments, computing devices 600a, 600b, and 600c may be configured to perform different functions.

[0119] In some embodiments, the computing tasks and stored data associated with neural networks, text embedding towers, audio embedding towers, joint embedding networks, audio separation networks, classifiers, and / or computing devices may be distributed across computing devices 600a, 600b, and 600c at least in part based on: the processing requirements of neural networks, text embedding towers, audio embedding towers, joint embedding networks, audio separation networks, classifiers, and / or computing devices; the processing capabilities of computing devices 600a, 600b, 600c; the latency of network links between computing devices within each computing cluster and between the computing clusters themselves; and / or other factors such as cost, speed, fault tolerance, resiliency, efficiency, and / or other design goals that may affect the overall system architecture.

[0120] The cluster storage arrays 610a, 610b, 610c of computing clusters 609a, 609b, 609c may be data storage arrays that include disk array controllers configured to manage read and write access to groups of hard disk drives. The disk array controllers, either individually or in conjunction with their respective computing devices, may also be configured to manage backup or redundant copies of data stored in the cluster storage arrays to guard against disk drive or other cluster storage array failures and / or network failures that prevent one or more computing devices from accessing one or more cluster storage arrays.

[0121] The functionality similar to that of neural networks, text embedding towers, audio embedding towers, joint embedding networks, audio separation networks, classifiers, and / or computing devices can be distributed across the computing devices 600a, 600b, 600c of computing clusters 609a, 609b, 609c. In a similar manner, the respective active and / or backup portions of these components can be distributed across the cluster storage arrays 610a, 610b, 610c. For example, some cluster storage arrays can be configured to store a portion of the data of neural networks, text embedding towers, audio embedding towers, joint embedding networks, audio separation networks, classifiers, and / or computing devices, while other cluster storage arrays can store other portions of the data of neural networks, text embedding towers, audio embedding towers, joint embedding networks, audio separation networks, classifiers, and / or computing devices. Additionally, for example, some cluster storage arrays can be configured to store the data of a first neural network, while other cluster storage arrays can store the data of a second neural network and / or a third neural network. Further, some cluster storage arrays can be configured to store backup versions of the data stored in other cluster storage arrays.

[0122] The cluster routers 611a, 611b, 611c in the computing clusters 609a, 609b, 609c can include networking equipment configured to provide internal and external communication for the computing clusters. For example, the cluster router 611a in the computing cluster 609a can include one or more Internet switching and routing devices configured to (i) provide local area network communication between the computing device 600a and the cluster storage array 610a via the local cluster network 612a, and (ii) provide wide area network communication between the computing cluster 609a and the computing clusters 609b and 609c via the wide area network link 613a to the network 406. The cluster routers 611b and 611c can include networking equipment similar to that of the cluster router 611a, and the cluster routers 611b and 611c can perform networking functions similar to the networking functions performed by the cluster router 611a for the computing cluster 609a for the computing clusters 609b and 609b.

[0123] In some embodiments, the configuration of the cluster routers 611a, 611b, 611c can be at least partially based on the data communication requirements of the computing devices and the cluster storage arrays, the data communication capabilities of the networking equipment in the cluster routers 611a, 611b, 611c, the latency and throughput of the local cluster networks 612a, 612b, 612c, the latency, throughput, and cost of the wide area network links 613a, 613b, 613c, and / or other factors that can affect the cost, speed, fault tolerance, resilience, efficiency, and / or other design criteria of the regulatory system architecture.

[0124] Example operating methods

[0125] Figure 7 It is a flowchart of a method 700 according to an exemplary embodiment. The method 700 can be executed by a computing device (such as, the computing device 500). The method 700 can start at block 710, where the computing device can receive an input audio waveform and an input text description.

[0126] At block 720, the computing device can separate the input audio waveform into multiple audio tracks through a neural network.

[0127] At block 730, the computing device can determine, through a neural network, whether the input text description describes an audio track among the multiple separated audio tracks.

[0128] At block 740, when the computing device determines that the input text description describes an audio track among the multiple audio tracks, the computing device can provide the audio track corresponding to the input text description to an interactive user interface.

[0129] In some embodiments, the neural network includes a symmetric encoder-decoder U-net neural network with skip connections.

[0130] In some embodiments, the method involves generating, through a neural network, one or more audio embeddings corresponding to one or more of the separated audio tracks, where the one or more audio embeddings include a representation of audio features of the one or more separated audio tracks. Such embodiments also involve generating, through a neural network, a text embedding corresponding to the input text description, where the text embedding includes a representation of text features of the input text description, and where determining whether the input text description describes an audio track among the multiple audio tracks includes applying a classifier to match the one or more audio embeddings with the text embedding.

[0131] In some embodiments, the neural network includes an audio separation network for separating the input audio waveform into multiple audio tracks, an audio embedding network for generating one or more audio embeddings, and a text embedding network for generating a text embedding.

[0132] In some embodiments, the neural network includes a joint embedding network, and the method involves generating, through the joint embedding network, a text embedding of the input text description and corresponding audio embeddings of the multiple audio tracks into a shared representation including a joint embedding, where a particular audio embedding is within a threshold distance of a particular text embedding that describes the particular audio embedding. In some embodiments, the joint embedding network includes an audio embedding tower for generating audio embeddings and a text embedding tower for generating text embeddings.

[0133] In some embodiments, the audio embedding tower includes (a) a modified Resnet-50 architecture trained on multiple 64-channel log Mel spectrograms, and (b) a fully connected final layer having multiple units corresponding to the dimension of the shared representation.

[0134] In some embodiments, the text embedding tower includes (a) a Transformer architecture based on Bidirectional Encoder Representations from Transformers (BERT) pooled according to special classification tokens (CLS-tokens), and (b) a fully connected final layer for mapping to a shared representation.

[0135] Some embodiments relate to receiving training data including a plurality of audio tracks, where each of the plurality of audio tracks is labeled with a text caption associated with the audio track. Such embodiments relate to training a neural network based on the training data.

[0136] In some embodiments, the neural network includes an audio separation network for performing separation of an input audio waveform into a plurality of audio tracks, and wherein training of the audio separation network includes unsupervised mixture invariant training.

[0137] In some embodiments, the neural network includes a classifier configured to determine whether an input text description describes an audio track among a plurality of separated audio tracks, and wherein training of the neural network is based on a matching loss of the classifier, and wherein the matching loss indicates a degree of match between the text description and a given audio track in a given audio waveform.

[0138] In some embodiments, training of the neural network is performed at a computing device.

[0139] In some embodiments, receiving an input audio waveform and an input text description involves receiving the input text description from a user via an interactive user interface.

[0140] In some embodiments, providing an audio track involves providing user-selectable options via an interactive user interface to enhance or suppress an audio track corresponding to the input text description.

[0141] Some embodiments relate to a computing device generating a request to determine whether an input text description describes an audio track among a plurality of audio tracks. Such embodiments relate to sending the request from the computing device to a second computing device, the second computing device including a trained version of the neural network. Such embodiments further relate to, after sending the request, the computing device receiving from the second computing device a determination as to whether the input text description describes an audio track among a plurality of audio tracks. Such embodiments further relate to outputting an audio track corresponding to the input text description based on the received determination that the input text description describes an audio track among a plurality of audio tracks.

[0142] Figure 8 is a flowchart of a method 800 according to an example embodiment. The method 800 may be performed by a computing device (such as, computing device 500). The method 800 may begin at block 810, where the computing device may receive training data including a plurality of audio segments and a plurality of audio text descriptions.

[0143] At block 820, a computing device may train a neural network based on training data to perform the following operations: receive (i) an input audio waveform, and (ii) a text description of a sound track, separate the input audio waveform into multiple sound tracks, and determine whether the input text description describes a sound track among the multiple separated sound tracks.

[0144] At block 830, the computing device may provide the trained neural network.

[0145] In some embodiments, the training of the neural network is performed at the computing device.

[0146] In some embodiments, the neural network includes an audio separation network for performing separating an input audio waveform into multiple sound tracks, and a classifier configured to determine whether the input text description describes a sound track among the multiple separated sound tracks.

[0147] In some embodiments, the training of the neural network involves unsupervised permutation invariant training of the audio separation network.

[0148] In some embodiments, the training of the neural network is based on a matching loss for training the classifier, where the matching loss indicates the degree of match between the text description and a given sound track in a given audio waveform.

[0149] The present disclosure is not limited in terms of the specific embodiments described in this application, and these specific embodiments are intended to be illustrative of various aspects. As will be apparent to those skilled in the art, many modifications and variations can be made without departing from its essence and scope. In addition to the methods and devices enumerated herein, functionally equivalent methods and devices within the scope of the present disclosure will be apparent to those skilled in the art from the foregoing description. Such modifications and variations are intended to fall within the scope of the appended claims.

[0150] The above detailed description has described various features and functions of the disclosed systems, apparatuses, and methods with reference to the accompanying drawings. In the drawings, like symbols generally identify like components unless the context dictates otherwise. The illustrative embodiments described in the detailed description, the drawings, and the claims are not intended to be limiting. Other embodiments can be utilized and other changes can be made without departing from the spirit or scope of the subject matter presented herein. It will be readily understood that the aspects of the present disclosure, as generally described herein and illustrated in the drawings, can be arranged, substituted, combined, separated, and designed in a wide variety of different configurations, all of which are explicitly contemplated herein.

[0151] For any and all ladder diagrams, scenarios, and flowcharts shown in the figures and discussed herein, each box and / or communication can represent information processing and / or information transmission according to example embodiments. Alternative embodiments are included within the scope of these example embodiments. In these alternative embodiments, for example, functions described as boxes, transmissions, communications, requests, responses, and / or messages may not be executed in the order shown or discussed, including being executed substantially simultaneously or in the reverse order, depending on the functions involved. Additionally, more or fewer boxes and / or functions may be used with any of the ladder diagrams, scenarios, and flowcharts discussed herein, and these ladder diagrams, scenarios, and flowcharts may be combined with each other in part or in whole.

[0152] A box representing the processing of information can correspond to circuitry configured to perform a particular logical function of the methods or techniques described herein. Alternatively or additionally, a block representing the processing of information can correspond to a module, segment, or portion of program code, including associated data. The program code can include one or more instructions that can be executed by a processor to implement a particular logical function or action in a method or technique. The program code and / or associated data can be stored on any type of computer-readable medium, such as a storage device including a magnetic disk or hard drive or other storage medium.

[0153] The computer-readable medium can also include non-transitory computer-readable media, such as non-transitory computer-readable media that store data for a short period of time, such as register memory, processor cache, and random access memory (RAM). The computer-readable medium can also include non-transitory computer-readable media that store program code and / or data for a longer period of time, such as secondary or persistent long-term storage devices, such as read-only memory (ROM), optical disks or magnetic disks, and compact disc read-only memory (CD-ROM). The computer-readable medium can also be any other volatile or non-volatile storage system. The computer-readable medium can be considered, for example, a computer-readable storage medium or a tangible storage device.

[0154] Additionally, a box representing one or more information transmissions can correspond to an information transmission between software modules and / or hardware modules within the same physical device. However, other information transmissions can be between software modules and / or hardware modules in different physical devices.

[0155] Although various aspects and embodiments have been disclosed herein, other aspects and embodiments will be apparent to those skilled in the art. The various aspects and embodiments disclosed herein are provided for purposes of explanation and not limitation, and the true scope is indicated by the claims.

Claims

1. A computer-implemented method, comprising: receiving, by a computing device, an input audio waveform and an input text description; separating, by a neural network, the input audio waveform into a plurality of audio tracks; determining, by the neural network, whether the input text description describes an audio track among the plurality of separated audio tracks; and when it is determined that the input text description describes an audio track among the plurality of audio tracks, providing, by the computing device, the audio track corresponding to the input text description to an interactive user interface.

2. The computer-implemented method according to claim 1, wherein the neural network comprises a symmetric encoder-decoder U-net neural network having skip connections.

3. The computer-implemented method according to claim 1, further comprising: generating, by the neural network, one or more audio embeddings corresponding to the one or more separated audio tracks, wherein the one or more audio embeddings comprise a representation of audio features of the one or more separated audio tracks; generating, by the neural network, a text embedding corresponding to the input text description, wherein the text embedding comprises a representation of text features of the input text description, and wherein determining whether the input text description describes an audio track among the plurality of audio tracks comprises applying a classifier to match the one or more audio embeddings with the text embedding.

4. The computer-implemented method according to claim 3, wherein the neural network comprises: an audio separation network configured to perform separating the input audio waveform into the plurality of audio tracks; an audio embedding network configured to generate the one or more audio embeddings; and a text embedding network configured to generate the text embedding.

5. The computer-implemented method according to claim 1, wherein the neural network comprises a joint embedding network, and further comprising: generating, by the joint embedding network, the text embedding of the input text description and the corresponding audio embeddings of the plurality of audio tracks into a shared representation comprising a joint embedding, wherein a particular audio embedding is within a threshold distance of a particular text embedding that describes the particular audio embedding.

6. The computer-implemented method according to claim 5, wherein the joint embedding network comprises an audio embedding tower for generating the audio embeddings and a text embedding tower for generating the text embeddings.

7. The computer-implemented method according to claim 6, wherein the audio embedding tower comprises (a) a modified Resnet-50 architecture trained on a plurality of 64-channel log Mel spectrograms, and (b) a fully-connected final layer having a plurality of units corresponding to the dimension of the shared representation.

8. The computer-implemented method according to claim 6, wherein the text embedding tower comprises (a) a transformer architecture based on Bidirectional Encoder Representations from Transformers (BERT) pooled according to a special classification token (CLS-token), and (b) a fully-connected final layer for mapping to the shared representation.

9. The computer-implemented method according to claim 1, further comprising: Receiving training data including a plurality of audio tracks, wherein each of the plurality of audio tracks is labeled with a text caption associated with the audio track; And Training the neural network based on the training data.

10. The computer-implemented method according to claim 9, wherein the neural network includes an audio separation network for performing separation of the input audio waveform into the plurality of audio tracks, and wherein the training of the audio separation network includes unsupervised mixture-invariant training.

11. The computer-implemented method according to claim 9, wherein the neural network includes a classifier configured to determine whether the input text description describes an audio track among the plurality of separated audio tracks, and wherein the training of the neural network is based on a matching loss of the classifier, and wherein the matching loss indicates the degree of matching between the text description and a given audio track in a given audio waveform.

12. The computer-implemented method according to claim 9, wherein the training of the neural network is performed at the computing device.

13. The computer-implemented method according to claim 1, wherein the receiving of the input audio waveform and the input text description includes receiving the input text description from a user via the interactive user interface.

14. The computer-implemented method according to claim 1, wherein the providing of the audio track includes providing user-selectable options via the interactive user interface to enhance or suppress the audio track corresponding to the input text description.

15. A computer-implemented method, comprising: Receiving, by a computing device, training data including a plurality of audio segments and text descriptions of the plurality of audios; Training a neural network based on the training data to perform the following operations: Receiving (i) an input audio waveform, and (ii) a text description of an audio track, Separating the input audio waveform into a plurality of audio tracks, and Determining whether the input text description describes an audio track among the plurality of separated audio tracks; And Providing, by the computing device, the trained neural network.

16. The computer-implemented method according to claim 15, wherein the training of the neural network is performed at the computing device.

17. The computer-implemented method according to claim 15, wherein the neural network includes: An audio separation network for performing separation of the input audio waveform into the plurality of audio tracks; And A classifier configured to determine whether the input text description describes an audio track among the plurality of separated audio tracks.

18. The computer-implemented method according to claim 17, wherein the training of the neural network includes unsupervised mixture-invariant training of the audio separation network.

19. The computer-implemented method according to claim 17, wherein the training of the neural network is based on a matching loss for training the classifier, wherein the matching loss indicates the degree of matching between the text description and a given audio track in a given audio waveform.

20. A computing device, comprising: One or more processors; And A data storage device having computer-executable instructions stored thereon, the computer-executable instructions, when executed by the one or more processors, cause the computing device to perform operations, the operations including: Receiving, by the computing device, an input audio waveform and an input text description; Separating, by a neural network, the input audio waveform into a plurality of audio tracks; Determining, by the neural network, whether the input text description describes an audio track among the plurality of separated audio tracks; and When determining that the input text description describes an audio track among the plurality of audio tracks, providing, by the computing device, the audio track corresponding to the input text description to an interactive user interface.