Method and system for target source separation
The acoustic processing system addresses the inflexibility of conventional systems by using a neural network with FiLM layer to extract target audio signals based on heterogeneous features, enhancing accuracy and efficiency in audio source separation.
Patent Information
- Application Number
- JP2025500628
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-10-09
- Filing Date
- 2023-03-31
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2043-03-31
AI Technical Summary
Conventional source separation systems lack flexibility in determining target audio sources and struggle to accurately extract specific audio signals based on heterogeneous semantic concepts, such as loudness, gender, and spatial location, due to mutually exclusive conditioning inputs.
An acoustic processing system that utilizes a neural network architecture with a feature-invariant linear modulation (FiLM) layer to process conditioning inputs, allowing for the extraction of target audio signals based on mutually inclusive and exclusive features, using a mixture of acoustic signals collected from multiple sources.
The system enables accurate and efficient extraction of target audio signals by mimicking human flexibility in selecting audio sources, improving performance over conventional systems by leveraging machine learning and neural networks to handle diverse acoustic scenarios.
Smart Images

Figure 0007760090000005 
Figure 0007760090000006 
Figure 0007760090000007
Abstract
Description
[Technical Field]
[0001] The present disclosure relates generally to target sound source separation, and more particularly to an acoustic processing system for extracting a target sound from a mixture of acoustic signals. [Background technology]
[0002] Conventional source separation systems for extracting target acoustic signals typically aim to isolate only specific types of sounds, such as for speech enhancement or instrument separation, where the target is determined by a training scheme and cannot be changed during testing. Conventional source separation approaches typically separate a sound mixture into only fixed sources (e.g., separating voices from background music) or all sources in the mixture without differentiating factors (e.g., separating individual speakers in a conference room), and then use post-processing to find the target signal. In recent years, conditioning-based approaches have emerged as a promising alternative, allowing auxiliary inputs such as class labels to indicate desired sources, but the available sets of conditions are typically mutually exclusive and lack flexibility.
[0003] For example, in the cocktail party problem, humans possess a preternatural ability to focus on sources of interest in complex acoustic scenes. They can adapt their focus depending on the situation, relying on attention mechanisms that regulate cortical responses to auditory stimuli. The field of audio source separation has made great strides toward replicating this ability, particularly with the advent of deep learning approaches. However, there remains a gap in the flexibility required to determine target audio sources. As mentioned above, early research developed "specialist" models designed to isolate only specific types of audio. More recent research, such as deep clustering and permutation invariant training (PIT), has focused on separating all audio sources in a mixture without any discriminatory factors. However, the problem of determining which of the extracted audio sources is the unresolved audio source of interest remains.
[0004] Therefore, there is a need for an advanced system that overcomes the above-mentioned drawbacks. To that end, there is a need for a technical solution to overcome the above-mentioned challenges. More specifically, there is a need for such a system that provides better performance than conventional acoustic processing systems for extracting target acoustic signals. Summary of the Invention
[0005] The present disclosure provides an improved sound processing system for identifying and extracting a target sound signal from a mixture of sounds. More specifically, the present disclosure provides a sound processing and training system configured to identify a target sound signal from a mixture of sounds based on mutually inclusive concepts such as loudness, gender, language, spatial location, etc.
[0006] To that end, some embodiments provide a conditional model configured to mimic human flexibility in selecting a target acoustic signal by focusing on extracting acoustics based on various, i.e., heterogeneous, semantic concepts and criteria, such as whether the speaker is close or far from the microphone, speaks softly or loudly, or speaks in a particular language. Some embodiments are based on the recognition that a mixture of acoustic signals is collected from multiple sources. In addition, a query is collected that identifies a target acoustic signal to be extracted from the mixture of acoustic signals. The query is associated with one or more identifiers that indicate mutually inclusive characteristics of the target acoustic signals.
[0007] To that end, a mixture of acoustic signals is collected with the aid of one or more microphones from multiple sound sources, the multiple sound sources corresponding to at least one of one or more speakers, people or individuals, industrial equipment, and vehicles.
[0008] Furthermore, each identifier present in the query having one or more identifiers belongs to a predetermined set of one or more identifiers and is extracted from the query, each extracted identifier defines at least one of mutually inclusive and mutually exclusive features of the target acoustic signal, and one or more logical operators are used to connect the extracted one or more identifiers.
[0009] Some embodiments are based on the recognition that the extracted one or more identifiers and one or more logical operators are converted into a digital representation, wherein the digital representation of the one or more identifiers is selected from a set of predetermined digital representations of a plurality of combinations of the one or more identifiers.
[0010] To that end, the digital representation corresponds to a conditioning input, which may be represented in any manner, such as by a text input, a speech input, by a one-hot or multi-hot conditional vector, etc. The conditioning input comprises one or more mutually inclusive features of the target acoustic signal.
[0011] Some embodiments are based on the recognition that a trained neural network is implemented to extract a target acoustic signal from a mixture of acoustic signals by mixing a digital representation with an intermediate output of an intermediate layer of the neural network. The neural network is trained for each set of predetermined digital representations of multiple combinations of one or more identifiers for extracting the target acoustic signal from the mixture of acoustic signals. To that end, during training, the extraction model is configured to generate one or more queries associated with one or more identifiers from the predetermined set of one or more identifiers.
[0012] To that end, in some embodiments, the neural network is based on an architecture comprising one or more interrelated blocks, where each block comprises at least a feature encoder, a conditioning network, a separation network, and a feature decoder. The conditioning network comprises a feature-invariant linear modulation (FiLM) layer that takes a mixture of acoustic signals as input and modulates the input into a conditioning input. The FiLM layer processes the conditioning input and sends the processed conditioning input to the separation network.
[0013] Accordingly, one embodiment discloses a computer-implemented method for extracting a target acoustic signal. The method includes collecting a mixture of acoustic signals from multiple sound sources. The method further includes selecting a query that identifies a target acoustic signal to be extracted from the mixture of acoustic signals. The method includes extracting each identifier present in a predetermined set of one or more identifiers from the query. The method includes determining one or more logical operators connecting the extracted one or more identifiers. The method further includes converting the extracted one or more identifiers and the one or more logical operators into a predetermined digital representation for querying the mixture of acoustic signals. The method includes implementing a neural network trained to extract the target acoustic signal identified by the digital representation from the mixture of acoustic signals by combining the digital representation with an intermediate output of a hidden layer of the neural network that processes the mixture of acoustic signals. The neural network is trained using machine learning to extract various acoustic signals identified in the set of predetermined digital representations. The method further includes outputting the extracted target acoustic signal.
[0014] Some embodiments provide an audio processing system configured to extract a target audio signal from a mixture of audio signals. The audio processing system includes at least one processor and a memory having instructions stored thereon that form executable modules of the audio processing system. The at least one processor is configured to collect the mixture of audio signals. Additionally, the at least one processor is configured to collect a query that identifies a target audio signal to be extracted from the mixture of audio signals. The query includes one or more identifiers. The at least one processor is further configured to extract each identifier of the one or more identifiers from the query, each identifier being present in a predetermined set of one or more identifiers. Each identifier defines at least one of mutually inclusive and mutually exclusive characteristics of the mixture of audio signals. The at least one processor is configured to determine one or more logical operators connecting the extracted one or more identifiers. Furthermore, the at least one processor is configured to convert the extracted one or more identifiers and the one or more logical operators into a predetermined digital representation for querying the mixture of audio signals. The at least one processor is further configured to execute a neural network trained to extract a target acoustic signal identified by the digital representation from the mixture of acoustic signals by combining the digital representation with an intermediate output of a hidden layer of the neural network, and to output the extracted target acoustic signal.
[0015] Various embodiments disclosed herein provide acoustic processing systems that can extract a target acoustic signal from a mixture of acoustic signals more accurately, efficiently, and in a shorter amount of time. Furthermore, various embodiments provide neural network-based acoustic processing systems that can be trained to extract a target acoustic signal based on mutually inclusive and / or mutually exclusive features of the target acoustic signal. The neural network can be trained using a combination of mutually inclusive and / or mutually exclusive feature datasets in the form of a predetermined set of one or more identifiers, providing advantages over existing neural networks.
[0016] Further features and advantages will become more readily apparent from the detailed description when taken in conjunction with the accompanying drawings. [Brief explanation of the drawings]
[0017] [Figure 1] FIG. 1 is a block diagram illustrating an environment for extraction of a target acoustic signal, according to some embodiments of the present disclosure. [Figure 2A] FIG. 1 is a block diagram illustrating an acoustic processing system for extracting a target acoustic signal, according to some embodiments of the present disclosure. [Figure 2B] FIG. 1 is a functional block diagram illustrating an acoustic processing system for extracting a target acoustic signal, according to some embodiments of the present disclosure. [Figure 2C] FIG. 2 is a block diagram illustrating a query interface of an audio processing system according to some embodiments of the present disclosure. [Figure 2D] FIG. 1 illustrates an example of a query interface for an audio processing system, according to some embodiments of the present disclosure. [Figure 3A] FIG. 1 is a block diagram illustrating a method for generating a digital representation according to some embodiments of the present disclosure. [Figure 3B] FIG. 2 is a block diagram illustrating multiple combinations of one or more identifiers, according to some embodiments of the present disclosure. [Figure 3C] 1A-1C are block diagrams illustrating various types of digital representations in accordance with some embodiments of the present disclosure. [Figure 4] FIG. 1 is a block diagram illustrating a neural network according to some embodiments of the present disclosure. [Figure 5A] FIG. 1 is a block diagram illustrating training of a neural network, according to some embodiments of the present disclosure. [Figure 5B] FIG. 1 is a block diagram illustrating training of a bridge-conditioned neural network, according to some embodiments of the present disclosure. [Figure 6] FIG. 1 is a block diagram illustrating an implementation of a neural network for extracting a target acoustic signal, according to some embodiments of the present disclosure. [Figure 7] FIG. 1 is a flow diagram illustrating training a neural network, according to some embodiments of the present disclosure. [Figure 8] FIG. 1 is a flow diagram illustrating a method performed by an audio processing system to perform signal processing, according to some embodiments of the present disclosure. [Figure 9] FIG. 1 is a block diagram illustrating an acoustic processing system for extracting a target acoustic signal, according to some embodiments of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0018] [Description of the embodiment] In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of the present disclosure. However, it will be apparent to those skilled in the art that the present disclosure may be practiced without these specific details. In other instances, devices and methods are shown only in block diagram form in order to avoid obscuring the disclosure. Various changes are contemplated that may be made in the function and arrangement of elements without departing from the spirit and scope of the disclosed subject matter as set forth in the appended claims.
[0019] As used in this specification and claims, the words "for example," "for example," and "such as," as well as "comprises," "has," "includes," and other verb forms thereof, when used in conjunction with a list of one or more components or other items, should be construed as open-ended. This means that the list should not be considered to exclude additional components or items. The term "based on" means based at least in part on. Furthermore, it should be understood that the terms and terminology used herein are for descriptive purposes and should not be considered limiting. Any headings used herein are for convenience only and do not have any legal or restrictive effect.
[0020] In the following description, specific details are provided to provide a thorough understanding of the embodiments. However, those skilled in the art will understand that the embodiments may be practiced without these specific details. For example, systems, processes, and other elements of the disclosed subject matter may be shown as components in block diagram form to avoid obscuring the embodiments in unnecessary detail. In other instances, well-known processes, structures, and techniques may be shown without unnecessary detail to avoid obscuring the embodiments. Furthermore, like reference numbers and names in the various drawings indicate like elements.
[0021] The present disclosure provides an audio processing system configured to identify a target audio signal from a mixture of audio signals based on concepts, including mutually inclusive concepts such as loudness, gender, language, and spatial location. That is, the same target audio signal can be identified using multiple different such concepts. The audio processing system collects a mixture of audio signals and selects a query that identifies a target audio signal to be extracted from the mixture of audio signals. The audio processing system then extracts one or more identifiers associated with the target audio signal from the query. The one or more identifiers indicate features of the target audio signal, including mutually inclusive and mutually exclusive features of the target audio signal. The one or more identifiers are used as conditioning inputs and converted into digital representations in the form of at least one of one-hot conditional vectors, multi-hot conditional vectors, text input, or audio input. The digital representations of the conditioning inputs are then used as inputs to a neural network to extract the target audio signal from the mixture of audio signals. The neural network is trained to extract the target audio signal identified by the digital representation from the mixture of audio signals by combining the digital representation with intermediate outputs from an intermediate layer of the neural network that processes the mixture of audio signals. The neural network is trained using machine learning to extract a target acoustic signal identified in a set of predetermined digital representations. Additionally, the neural network is trained based on an architecture having one or more interrelated blocks. The one or more interrelated blocks include at least one of a feature encoder, a conditioning network, a separation network, and a feature decoder. The conditioning network includes a feature-invariant linear modulation (FiLM) layer that takes an encoded feature representation of a mixture of acoustic signals as input and modulates the input based on the conditioning input, which is in the form of a digital representation. The FiLM layer processes the conditioning input and sends the processed conditioning input to the separation network, where the target acoustic signal is separated from the mixture of acoustic signals.Additionally, the FiLM layer repeats the process of sending conditioned inputs to the separation network to separate the target audio signal from the mixture of audio signals.
[0022] System Overview
[0023] 1 illustrates an environment 100 for extracting a target acoustic signal according to some embodiments of the present disclosure. The environment 100 includes multiple sound sources 102, one or more identifiers 104, one or more microphones 106, a mixture of acoustic signals 108, a network 110, and an acoustic processing system 112.
[0024] The multiple sound sources 102 may correspond to at least one of one or more speakers, such as a person or individual, industrial equipment, or a vehicle. A mixture of sound signals 108 is collected from the multiple sound sources 102 with the aid of one or more microphones 106. Each sound signal in the mixture of sound signals 108 is associated with one or more identifiers 104 or criteria that define some characteristic of that sound signal in the mixture of sound signals 108. For example, the one or more identifiers 104 may be used to mimic human flexibility in selecting which sound sources to process by focusing on extracting sounds from the mixture of sound signals 108 based on semantic concepts and criteria of various properties, i.e., heterogeneity. These heterogeneity criteria, in one example, include whether the speaker is close or far from one or more microphones 106, speaks softly or loudly, or speaks in a particular language. In this manner, one or more identifiers 104 are associated with the multiple sound sources 102. Other examples of the one or more identifiers 104 include at least one of the loudest source, the quietest source, the farthest source, the nearest source, a female speaker, a male speaker, and a language-specific source, etc.
[0025] A mixture 108 of acoustic signals associated with these one or more identifiers 104 may be transmitted over a network 110 to an audio processing system 112 .
[0026] In one embodiment of the present disclosure, the network 110 is the Internet. In another embodiment of the present disclosure, the network 110 is a wireless mobile network. The network 110 includes a set of channels, each of which supports a finite bandwidth. The finite bandwidth of each of the channels is based on the capacity of the network 110. Furthermore, the one or more microphones 106 are arranged in a pattern such that the acoustic signals of each of the multiple sound sources 102 are captured. The arrangement pattern of the one or more microphones 106 enables the sound processing system 112 to estimate localization information of the multiple sound sources 102 using the relative time differences between the microphones. The localization information may be provided in the form of a direction of arrival of the sound or a distance from the one or more microphones 106 to the sound source.
[0027] In operation, the sound processing system 112 is configured to collect a mixture of sound signals 108 from the multiple sound sources 102. Additionally, the sound processing system 112 is configured to collect a query that identifies a target sound signal to be extracted from the mixture of sound signals 108. The sound processing system 112 is further configured to extract, from the query, each identifier present in a predetermined set of one or more identifiers that define mutually inclusive and mutually exclusive features of the mixture of sound signals 108. The sound processing system 112 is further configured to determine one or more logical operators connecting the extracted one or more identifiers. The sound processing system 112 is further configured to convert the extracted one or more identifiers and the one or more logical operators into a predetermined digital representation for querying the mixture of sound signals 108. The sound processing system 112 is further configured to execute a trained neural network to extract, from the mixture of sound signals 108, the target sound signal identified by the digital representation by combining the digital representation with intermediate outputs of an intermediate layer of the neural network that processes the mixture of sound signals 108. The sound processing system 112 is described in further detail in Figures 2A and 2B.
[0028] 2A shows a block diagram of an acoustic processing system 112 for extracting a target acoustic signal 218 according to some embodiments of the present disclosure. The acoustic processing system 112 includes a memory 202, a processor 204, a database 206, a query interface 208, and an output interface 216. The memory 202 corresponds to at least one of RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassette, magnetic tape, magnetic disk storage, or any other storage medium that can be used to store desired information and that can be accessed by the acoustic processing system 112. The memory 204 includes non-transitory computer storage media in the form of volatile and / or non-volatile memory. The memory 204 can be removable, non-removable, or a combination thereof. Exemplary memory devices include solid-state memory, hard drives, optical disk drives, etc. The memory 202 stores instructions executed by the processor 204. The memory 202 includes a neural network 210 and a transformation module 212. The memory 202 is associated with a database 206 of the sound processing system 112. The sound processing system 112 collects sound signal mixtures 108 from multiple sound sources 102. The database 206 is configured to store the collected sound signal mixtures 108. The sound signal mixtures 108 correspond to sound signal mixtures having different characteristics, including, but not limited to, a sound source farthest from the one or more microphones 106, a sound source closest to the one or more microphones 106, a female speaker, a French-speaking speaker, etc. Additionally, the database 206 stores characteristics of each of the multiple sound sources 102. In one embodiment, the database 206 is queried to extract a target sound signal 218 using a query interface 208 of the sound processing system 112. Additionally, the database 206 stores a predetermined set of one or more identifiers and a set of predetermined digital representations associated with the target sound signals 218.
[0029] The sound processing system 112 is configured to use a query interface 208 to collect a query that identifies a target sound signal 218 to be extracted from the mixture of sound signals 108 .
[0030] Further, the sound processing system 112 is configured to extract, with the aid of an extraction model 214, from the collected queries, each identifier present within a predetermined set of one or more identifiers. In one example, an identifier corresponds to any characteristic of a sound source, such as a “loudest” speaker, a “female” speaker, etc. The identifiers are extracted from the collected queries using the extraction model 214. The collected queries are utilized by the extraction model 214 for analyzing the collected queries. The extraction model 214 is configured to identify each identifier from the collected queries based on the analysis of the collected queries. Each identifier present within a predetermined set of one or more identifiers. Furthermore, the predetermined set of one or more identifiers defines mutually inclusive and mutually exclusive characteristics of the mixture of acoustic signals 108. In one example, the predetermined set of one or more identifiers is stored in the database 206. The predetermined set of one or more identifiers is generated from a set of historical data associated with the mixture of acoustic signals 108. Additionally, a predetermined set of one or more identifiers may be generated from a set of historical data via one or more third party audio sources associated with sound processing system 112 .
[0031] The collected queries may include multiple combinations of one or more identifiers. In one example, the collected queries may be “female” and “French” speaker. Here, “female” speaker and “French” speaker are two identifiers. The multiple combinations of the one or more identifiers are selected using one or more logical operators. The one or more logical operators enable the sound processing system 112 to select the multiple combinations of the one or more identifiers. The extraction model 214 is configured to determine one or more logical operators connecting each of the one or more identifiers 104 extracted from a predetermined set of one or more identifiers. Additionally, the extraction model 214 is configured to generate one or more queries using the collected queries, the one or more identifiers 104, and the one or more determined logical operators. The one or more queries are further processed to generate conditioning inputs 222. In one example, the extraction model 214 generates queries such as “Which is the most distant acoustic signal?” and (&) “Which is the source of the English speech?”. Furthermore, the extraction model 214 utilizes all of the queries to generate conditioning inputs 222. Here, the conditioning input 222 is, for example, "Which is the most distant source of the English speech?" The conditioning input 222 is an input that includes multiple combinations of one or more identifiers of the target acoustic signal 218. The conditioning input 222 corresponds to a processed query that includes multiple combinations of one or more identifiers.
[0032] The conditioning input 222 is further utilized by the transformation module 212. The transformation module 212 is configured to transform the extracted one or more identifiers 220 into predetermined digital representations for querying the mixture of acoustic signals 108. In one example, the transformation module 212 selects the digital representations of the extracted one or more identifiers 220 from a set of predetermined digital representations of multiple combinations of one or more identifiers of the target acoustic signal 218 (as further described in FIGS. 2B and 3A). Furthermore, the one or more identifiers 104 may be used to train the neural network 210 to extract the target acoustic signal 218 from the mixture of acoustic signals 108 by generating various training combinations of the one or more identifiers during training. Furthermore, the one or more identifiers 104 are utilized by the extraction model 214 to generate one or more queries. The one or more queries are associated with mutually inclusive and mutually exclusive features of the target acoustic signal 218 used during training of the neural network 210. The extraction model 214 is configured to execute the trained neural network 210 to extract a target acoustic signal from the mixture of acoustic signals 108 by combining the digital representation of the one or more identifiers 104 with intermediate outputs of the hidden layers of the neural network 210. Further, the extracted target acoustic signal 218 is output from the output interface 216.
[0033] FIG. 2B shows a functional block diagram 200B of an acoustic processing system 112 for extracting a target acoustic signal 218, according to some embodiments of the present disclosure. The acoustic processing system 112 collects a mixture of acoustic signals 108. The mixture of acoustic signals 108 is collected from multiple sound sources 102 with the aid of one or more microphones 106 (as described in FIG. 1 ). The acoustic processing system 112 is configured to collect a query using a query interface 208 to identify a target acoustic signal 218 to be extracted from the mixture of acoustic signals 108. The query interface 208 is configured to collect a query to accept one or more identifiers 104 associated with the target acoustic signal 218 that indicate mutually inclusive and mutually exclusive characteristics of the target acoustic signal 218. The query is collected by the query interface 208. In one embodiment, the query is collected using a voice command with the aid of natural language processing techniques. The collected query is further analyzed (conditioning input 222) to identify one or more identifiers 104 for generating a processed query. For example, one identifier of the one or more identifiers 104 corresponds to the loudest speaker, and another identifier of the one or more identifiers 104 corresponds to a female speaker. Here, the mutually inclusive features of the target acoustic signal 218 correspond to "loudest" and "female," and the processed query corresponds to "loudest female speaker." The target acoustic signal 218 associated with the one or more identifiers 104 exhibiting the mutually inclusive and mutually exclusive features is the audio source having the loudest female voice from all of the multiple audio sources 102 (as described in more detail in FIG. 2C ).
[0034] The collected queries are then utilized by the extraction model 214 to extract each identifier present within a predetermined set of one or more identifiers that define mutually inclusive and mutually exclusive features of the mixture of acoustic signals 108. The extraction model 214 extracts each identifier to generate conditioning inputs 222 (as described above in FIG. 2A).
[0035] The conditioning input 222 is further utilized by the transformation module 212. The transformation module 212 converts the conditioning input 222 into a digital representation 224 (the transformation module is further described in FIG. 3A ). The digital representation 224 is then sent to the neural network 210 for training to extract a target acoustic signal 218 (the digital representation is further described in FIG. 3C ) and, during testing, to generate an output associated with the extracted target acoustic signal 218 from the mixture of acoustic signals 108. The neural network 210 is trained with machine learning with the help of one or more machine learning algorithms to extract various acoustic signals identified in a predetermined set of digital representations. The predetermined set of digital representations includes representations of various acoustic signals that may be extracted from a set of historical data or one or more third-party audio sources. The target acoustic signal 218 is extracted from these various acoustic signals present in the predetermined set of digital representations. In one embodiment, the one or more machine learning algorithms used to train the neural network 210 may include, but are not limited to, a voice activity detection (VAD) algorithm and a deep speech algorithm. Generally, a deep speech algorithm is used to automatically transcribe spoken speech. A deep speech algorithm takes digital speech as input and returns a "most likely" text transcript of the digital speech. Additionally, VAD is a technology that detects the presence or absence of human speech.
[0036]
number
[0037] In one embodiment, section 208A has drop-down lists that allow for the selection of appropriate identifiers, such as "French-speaking speaker" 208aa, "English-speaking speaker" 208bb, "male speaker" 208cc, "female speaker" 208dd, "loudest speaker" 208ee, and "softest speaker" 208ff. Section 208B has drop-down lists that allow for the selection of one or more logical operators. In one example, NOT operator 208c1 is selected and "male speaker" 208cc is selected from section 208A. In addition, AND operator 208b1 is selected from the drop-down list in section 208B. Furthermore, "loudest speaker" 208ee is selected from section 208A. The conditioning input 222 generated using the inputs selected from sections 208A, 208B, and NOT operator 208c1 corresponds to "should be loudest speaker 208ee, not male speaker 208cc." Additionally, the query interface 208 includes an "Add Identifier" section 208d for selecting multiple identifiers from the one or more identifiers 220. The "Add Identifier" section 208d may or may not be used.
[0038] Additionally, the query interface 208 has a voice interface 226 that allows a user to provide a voice command associated with the target acoustic signal 218. The voice command is analyzed using natural language processing techniques, and one or more identifiers 104 (which may be multiple combinations 304 of one or more identifiers) are extracted from the voice command to generate the conditioning input 222. The query interface 208 is utilized during training of the neural network 210 to accurately extract the target acoustic signal 218.
[0039] 2D shows an example of a query interface 208 of an audio processing system 112, according to some embodiments of the present disclosure. The query interface 208 has a section 228. The section 228 allows a user to type a logical expression 230 using one or more identifiers 220 and one or more logical operators. The one or more logical operators correspond to the logical operators described in FIG. 2C, such as AND operator 208b1, OR operator 208b2, and NOT operator 208C1. In one example, the user has typed the logical expression 230. The logical expression 230 is represented as follows:
[0040] (French-speaking speaker AND (NOT loudest speaker)) OR (male speaker)
[0041] The logic formula 230 indicates that the user is selecting a speaker who speaks French, who need not be the loudest speaker, and who may or may not be a male speaker.
[0042] Furthermore, the logical formula 230 is not limited to the formulas described above.
[0043] FIG. 3A shows a block diagram of a method for generating a digital representation 224 according to some embodiments of the present disclosure. In the illustrated example, the database 206 includes a set of predetermined digital representations 302. The set of predetermined digital representations 302 may be extracted from one or more third-party databases. The set of predetermined digital representations 302 may include a plurality of combinations 304 of one or more identifiers of the target acoustic signal 304. The plurality of combinations 304 of one or more identifiers corresponds to at least two or more characteristics of a particular sound source. In one example, the plurality of combinations 304 of one or more identifiers includes a “loudest” “female” speaker, a “softest” “male” “French-speaking” speaker, etc. (further described in FIG. 3B). The set of predetermined digital representations 302 is utilized by the conversion module 206 to convert the conditioning input 222 into the digital representation 224.
[0044] The transformation module 206 generates digital representations 224 of the conditioning inputs 222 of the extracted one or more identifiers 104 from the set of predetermined digital representations 302. For example, if the extracted one or more identifiers 104 correspond to the "loudest" "male" speaker, the transformation module 206 considers the "loudest" "male" speaker as the identifiers and transforms these identifiers into conditioning inputs for extracting the target acoustic signal 218 to generate the digital representations 224 of the conditioning inputs.
[0045] 3B shows an example block diagram 300B of a plurality of combinations of one or more identifiers 304 according to some embodiments of the present disclosure. The plurality of combinations of one or more identifiers 304 may include, but are not limited to, a French-speaking male speaker 304a, a furthest English-speaking speaker 304b, a loudest Spanish-speaking speaker 304c, and a nearest female speaker 304d. The plurality of combinations of one or more identifiers 304 is not limited to the above examples.
[0046] 3C shows a block diagram 300C of a digital representation 224 according to some embodiments of the present disclosure. The digital representation 224 corresponds to a transformed representation of the conditioning input 222. The digital representation 224 is represented by at least one of a one-hot conditional vector 306, a multi-hot conditional vector 308, a text description 310, etc.
[0047] In one example, digital representation 224 includes one-hot conditional vector 306. If conditioning input 222 is "extract the speaker farthest from the microphone," one-hot conditional vector 306 would contain a "1" in the position corresponding to the farthest sound source in the vector of features of the acoustic signal, and zeros for all other terms, such as closest speaker, male / female, loud / soft, etc. In another example, digital representation 224 includes multi-hot conditional vector 308. If conditioning input 222 is "extract the loudest female speaker," multi-hot conditional vector 308 would contain ones in the positions corresponding to loudest speaker and female speaker, and all other terms, such as male speaker, softer speaker, etc., would be set to zero.
[0048] In some examples, the conditioning input 222 is converted at runtime into a digital representation in the form of a one-hot vector or a multi-hot vector by one or more selections in a drop-down menu of possible options through rule-based parsing of the text input, or by first converting the utterance to text and then using rule-based parsing. Additionally, logical operators such as "and" and "or" may be combined to generate multi-hot vectors between conditions to generate multi-hot vector representations. In some embodiments, an additional one-hot dimension is added to indicate an "and" / "or" query for generating a digital representation for the conditioning input.
[0049] In yet another example, the digital representation 224 includes a text description 310, which is particularly important when the target acoustic signal 218 is not speech but a general sound source such as industrial equipment or a vehicle. In this situation, descriptions such as male / female and English / French cannot be used. In this case, the text description 310 is converted into an embedding vector, and the embedding vector is then input into the neural network 210 instead of the one-hot conditional vector 306. To generate the embedding vector from the text description 310, a model such as a word2vec model or a Bidirectional Representation For Transformer (BERT) model may be used. In general, word2vec is a technology for natural language processing. A word2vec model generally uses a neural network to learn word associations from a large corpus of text. In addition, a BERT model is designed to help computers or machines understand the meaning of ambiguous language in text by providing context using surrounding text. Regardless of the representation type of digital representation 224, neural network 210 is trained and guided by digital representation 224 to extract target acoustic signal 218. The training of neural network 210 is further explained in Figures 4 and 5A.
[0050] 4 is a block diagram illustrating an architecture 400 of a neural network 210 according to some embodiments of the present disclosure. The neural network 210 may be a neural network trained to extract a target acoustic signal 218, and in some instances, even extract localization information of the target acoustic signal 218. Furthermore, the training is based on the premise that the training data includes an unordered, heterogeneous set of training data components. For example, in the case of the neural network 210, a digital representation 224 is generated to train the neural network 210 to extract a target acoustic signal 410. The target acoustic signal 410 is the same as the target acoustic signal 218 of FIG. 2.
[0051] The neural network 210 is trained to extract a target acoustic signal 410 from the mixture of acoustic signals 108 by combining the conditioning inputs 222 with intermediate outputs of the hidden layers of the neural network 210. The neural network 210 includes one or more interrelated blocks, such as a feature encoder 404, a conditioning network 402, a separation network 406, and a feature decoder 408. In one example, the conditioning network 402 comprises a feature-invariant linear modulation (FiLM) layer (described in FIG. 6).
[0052] The conditioning network 402 takes as input the conditioning input 222 converted into a digital representation 224. The conditioning network processes the digital representation 224, which identifies the type of sound source to be extracted from the mixture of acoustic signals 108, into a form useful to the separation network 406. The feature encoder 404 receives the mixture of acoustic signals 108. The feature encoder 404 corresponds to a trained one-dimensional convolutional feature encoder (ConvID) (described below in FIG. 6). Furthermore, the feature encoder 404 is configured to convert the mixture of acoustic signals 108 into a matrix of features for further processing by the separation network 406. The separation network 406 corresponds to a convolutional block layer (described in detail in FIG. 6). The separation network 406 utilizes the conditioned input 222 and the matrix of features to separate the target acoustic signal 410 from the mixture of acoustic signals 108. The separation network 406 is configured to generate a latent representation of the target acoustic signal 410. The separation network 406 combines the conditioned inputs 222 with a matrix of features to generate a latent representation of a target acoustic signal 410 separated from the mixture of acoustic signals 108. The feature decoder 408 is typically the inverse process of the feature encoder 404, converting the latent representation of the target acoustic signal generated by the separation network 406 into an audio waveform in the form of the target speech signal 410.
[0053] Neural network 210 undergoes a training phase further illustrated in FIG. 5A.
[0054] 5A shows an example block diagram 500A of training a neural network 210 according to some embodiments of the present disclosure. To that end, during training, multiple combinations 304 of one or more identifiers are provided to the neural network 210 for training 504 of the neural network 210. The multiple combinations 304 of one or more identifiers are converted into a set of predetermined digital representations 302. In one embodiment, the neural network 210 is trained using the set of predetermined digital representations 302 of the multiple combinations of one or more identifiers. Additionally, the neural network 210 is trained with training data 502.
[0055] In one example, the training data 502 includes a first training data set 502a and a second training data set 502b. The first training data set 502a includes acoustic data recorded in reverberant conditions and includes spatial data of sound sources relative to one or more microphones 106, but does not include data associated with the language of the sound sources. The second training data set 502b has data for multiple languages but was recorded in non-reverberant conditions. Thus, the second training data set includes language-related data for the sound sources but does not include spatial data of the sound sources. The neural network 210 is trained using both the first training data set 502a and the second training data set 502b. Furthermore, the trained neural network 210 is configured to separate sound sources based on language in reverberant conditions by using conditioning inputs as described above, even if that combination was not present in the training data 502 during training 504. To enable this, the trained neural network 210 generates test mixtures 506 with all available combinations of feature conditions of the sound sources for a run 508 of the neural network 210. In one example, during run 508, if the conditions are language-specific but the recorded sound is reverberant, the trained neural network 210 extracts the target sound source based on the conditions (language) even though language-labeled reverberant data was not present in the training data 502 during training 504. This is particularly useful when bridge conditions exist between two different training data sets.
[0056] 5B shows an example block diagram 500B of training a neural network 210 using bridge conditions 502, according to some embodiments of the present disclosure. A plurality of mutually inclusive feature combinations 304 are provided to the neural network 210 for training 504 of the neural network 210. The plurality of mutually inclusive feature combinations 304 are in the form of digital representations 224. Additionally, the neural network 210 is trained with training data 502. In one example, the training data 502 includes a first training data set 502a and a second training data set 502b. The first training data set 502a includes gender data of sound sources but does not include energy data. The second training data set 502b includes energy data and gender data. The neural network 210 is trained with the first training data set 502a and the second training data set 502b.
[0057] In one example, the bridge condition 502c is the loudest speaker in the mixture of audio signals. Such energy conditioning is advantageous because it is easy to control the loudness of each audio source when generating the mixture of audio signals, and therefore training samples can often be easily introduced into the training data set 502. That is, when generating the mixture of audio signals during training 504 of the neural network 210, a simple gain can be applied to the isolated audio source examples, thus providing the ability to condition energy in any data set. The terms loudness and energy are used interchangeably to refer to some concept of the volume of an audio signal.
[0058] The trained neural network 210 generates test mixtures 506 with all feasible combinations of sound source feature conditions for execution 508 of the neural network 210. In one example, if only the first training data set 502a is accessed to extract gender-specific target sound sources, the neural network 210 will be able to accurately extract the gender-specific target sound sources thanks to the bridge conditions 502c used in the training data. Even if gender conditioning is not available in the first training data set 502a during training 504, the bridge conditions allow for gender conditioning in the first training data set 502a. Additionally, during execution 508, all feasible conditions are available for extracting the target sound sources. Execution 508 of the neural network 210 is further described in FIG. 6.
[0059] FIG. 6 shows a block diagram 600 for the implementation 508 of the neural network 210 to extract the target acoustic signal 218, according to some embodiments of the present disclosure. The neural network 210 inputs and outputs a time-domain signal and includes the following components: (1) a trained one-dimensional convolutional feature encoder (ConvID) 606 (hereinafter, feature encoder 606) configured to obtain an intermediate representation; (2) a feature-invariant linear modulation (FiLM) layer 602; (3) B intermediate blocks 604 for processing the intermediate representation; and (4) a trained one-dimensional transposed convolutional decoder 608 for converting back to the time-domain signal. The FiLM layer 602 corresponds to the B-FiLM layer. The trained one-dimensional convolutional feature encoder (ConvID) 606 corresponds to the feature encoder 404 in FIG. 4. The B intermediate blocks 604 correspond to the convolutional block layer 604. The convolutional block layer 604 corresponds to the separation network 406 in FIG. 4. The convolutional block layer 604 is a stack of U-net convolutional blocks. Each U-net block contains several convolutional blocks that learn high-level latent representations and several transposed convolutional blocks that translate from the high-level latent representations back to representations comparable to the U-net input. In one example, the combination of the FiLM layer 602 (B-FiLM layer) and B intermediate blocks 604 is repeated B times.
[0060] The mixture of acoustic signals 108 is sent to a feature encoder 606. The feature encoder 506 converts the mixture of acoustic signals 108 into a matrix of features for further processing by the FiLM layer 602 and the convolutional block layer 604. The FiLM layer 602 takes the matrix of features of the mixture of acoustic signals 108 as an input. In addition, the FiLM layer 602 takes the digital representation 224 (e.g., the one-hot conditional vector 306 shown in FIG. 3C) as an input. The FiLM layer 602 processes the input (the matrix of features and the one-hot conditional vector 306) and sends the processed input to the convolutional block layer 604. The convolutional layer 604 combines the matrix of features and the processed conditioning input to generate a latent representation of the target acoustic signal 218. The latent representation is sent to a trained one-dimensional transposed convolutional decoder 608 to separate the target acoustic signal 218 from other sound sources 610. The FiLM layer 602 and the convolutional block layer 604 are trained and executed to extract the target acoustic signal 218 and estimate localization information of the extracted target acoustic signal 218. The localization information of the target acoustic signal 218 indicates the location of the origin of the extracted target acoustic signal 218.
[0061]
number
[0062]
number
[0063]
number
[0064] 7 shows a flow diagram 700 illustrating the training of a neural network 210 to function as a heterogeneity separation model 712. The flow diagram 700 includes a database 206, an extraction model 214, conditioning inputs 222 converted into digital representations 224, a negative example selector 702, a positive example selector 704, a speech mixer 706a, a speech mixer 706b, the neural network 210 functioning as the heterogeneity separation model 712, and a loss function 714. The database 206 is a speech database containing a collection of disaggregated acoustic signals, such as various speech signals for human speech applications, and associated metadata for each disaggregated acoustic signal, such as speaker-to-microphone distance, signal level, language, etc.
[0065] The extraction model 214 is configured to generate one or more random queries associated with mutually inclusive features of the target acoustic signal 610. A one-hot conditional vector 306 or a multi-hot conditional vector 308 for the received one or more identifiers 220 is randomly selected based on the generated one or more random queries. The multi-hot conditional vector 308 can be a multi-hot "and" conditioning vector or a multi-hot "or" conditioning vector. For multi-hot "and" conditioning, all selected identifiers must be true for the acoustic signal to be the associated target acoustic signal. For multi-hot "or" conditioning, at least one of the selected identifiers must be true. For text description 310 conditioning, all acoustic signals in the database 206 are required to have one or more natural language descriptions of the corresponding acoustic signal.
[0066] In one example, an audio signal is randomly selected from the database 206 as a positive example, and the corresponding text description is used as the conditioning input 222. The conditioning input 222 converted to a digital representation 224 is sent to the heterogeneity separation model 712 for further processing. The conditioning input 222 converted to a digital representation 224 is sent to the negative example selector 702 and the positive example selector 704. The negative example selector 702 returns zero, one, or more acoustic signals from the database 206 that are not related to the given conditioning input used to train the heterogeneity separation model 712 on one or more random queries. In one embodiment, the negative example selector 702 may return zero or irrelevant acoustic signals so that the heterogeneity separation model 712 can be robust during inference.
[0067] The positive example selector 704 returns zero, one, or multiple acoustic signals from the database 206 that are relevant to a given conditioning input. In some cases, it is important to have the positive example selector return zero relevant audio signals so that the heterogeneous target acoustic extraction model can be robust to this case during inference.
[0068] Zero, one, or more acoustic signals from the positive example selector 704 are passed through an audio mixer 706a to obtain ground truth target acoustic signals for training. The acoustic signals returned from both the positive example selector 704 and the negative example selector 702 are also passed to an audio mixer 706b to generate an audio mixture signal 708 during training that is input to a heterogeneous separation model 712. The heterogeneous separation model 214 processes the digital representation 224 and the audio mixture signal 708 to extract a separated target acoustic signal 716.
[0069] The ground truth target audio signal is compared with the separated target acoustic signal 716 with the help of a loss function 714. In other words, the loss function 714 compares the ground truth target audio signal 710 with the separated target acoustic signal 716 returned by the heterogeneous separation model 712 using the loss function 714. In an example, the relevant loss function comparing the two acoustic signals (e.g., SNR, scale-invariant source-to-distortion ratio, mean squared error, etc.) may be calculated in the time domain, the frequency domain, or a weighted combination of time-domain and frequency-domain losses.
[0070] In some examples, there may be several sound sources, such as multiple people speaking at a business meeting or a party, and a machine listening device (e.g., a robot or a hearing aid-like device) that can focus on a specific person's speech may be needed. However, the machine listening device requires input from a user to identify which person to focus on, which is often context-dependent. For example, if two people are having a conversation, one male and one female, the user may provide input to the machine listening device to focus on the male speaker. In some examples, the machine listening device includes an audio processing system 112 that uses a neural network 210 to perform the task of identifying target acoustic signals using a heterogeneity separation model 712. If both speakers are male, the heterogeneity separation model 712 is utilized to describe the target person's speech, such as how far they are from the microphone or the volume of their speech relative to other competing speakers. The heterogeneous separation model 712 allows a control device to select the signal features for a given mixture of speakers that are best suited to separating the speaker of interest (particular sound source) given the particular circumstances.
[0071] The heterogeneity separation model 712 is trained to perform separation based on multiple criteria, as described above. Typically, the source separation model is trained using mixture / target pairs, where two or more separated source signals (e.g., speech waveforms) are combined to generate a mixture, and the separated signals are used as the target. This combination, also called a mixing process, takes each separated source signal, optionally applies some basic signal processing applications to the separated sources (e.g., applying gain, equalization, etc.), and then combines them to obtain a target audio mixture signal. The processed separated sources then serve as training targets for a given audio mixture signal. However, the heterogeneity separation model 712 uses triplets including (1) the audio mixture signal, (2) a digital representation represented, for example, by a one-hot conditional vector, and (3) a target signal corresponding to the description represented by the one-hot conditional vector.
[0072] Another example of the heterogeneous separation model 712 system may be combined with a system that uses multiple criteria to identify the signal features of all speakers present in a mixed signal without separating them. For example, it may be possible to detect the gender of the speaker and the language being spoken, even when the speech overlaps. The identified values of these criteria are used to conditionally extract the separated signals of the speakers present in the speech mixture. Furthermore, various criteria present in the speech mixture may be combined using a process similar to a "logical and" (i.e., a one-hot vector becomes a multi-hot vector with a "1" in the position of all relevant criteria) to separate the signals using all criteria. Each criterion may also be used independently to evaluate which conditioning criterion provides the best target signal separation performance for a given mixture.
[0073] 8 shows a flowchart 800 illustrating a method for identifying a target acoustic signal based on the above-described embodiments, according to some embodiments of the present disclosure. The method 800 is performed by the sound processing system 112. The flowchart begins at step 802. After step 802, in step 804, the method includes collecting a mixture of acoustic signals 108 from multiple sound sources 102 with the aid of one or more microphones 106. The multiple sound sources 102 correspond to at least one of a speaker, a person or individual, industrial equipment, and a vehicle. The mixture of acoustic signals 108 is collected from the multiple sound sources 102 with the aid of the one or more microphones 106 along with one or more identifiers 104.
[0074] In step 806, the method includes collecting, with the aid of a query interface 208 (as described in FIG. 2 ), a query that identifies a target acoustic signal 218 to be extracted from the mixture of acoustic signals 108. The query indicates mutually inclusive and mutually exclusive features of the target acoustic signal 218. The query is associated with one or more identifiers 104 of the plurality of audio sources 102. The one or more identifiers 104 include at least one of the loudest audio source, the quietest audio source, the farthest audio source, the nearest audio source, a female speaker, a male speaker, and a language-specific audio source.
[0075] In step 808, the method includes extracting from the query each identifier present in a predetermined set of one or more identifiers defining mutually inclusive and mutually exclusive features of the mixture of acoustic signals 108 with the help of an extraction model 214 (as illustrated in FIG. 2B). After step 808, in step 810, the method includes determining one or more logical operators connecting the extracted one or more identifiers 220 using a query interface 208 (as illustrated in FIG. 2C).
[0076] In step 812, the method includes converting the extracted one or more identifiers 220 into digital representations 224 with the aid of a transformation module 206 (as described in FIGS. 3A and 3C ). The transformation module 206 is configured to generate the digital representations 224 of the extracted one or more identifiers 220 from a set 302 of predetermined digital representations of a plurality of combinations 304 of mutually inclusive features of the target acoustic signal 218. The digital representations 224 are represented by one-hot conditional vectors 306 or multi-hot conditional vectors 308 (as described in FIG. 3C ) and a text description 310.
[0077] In step 814, the method includes executing a trained neural network 210 to extract a target acoustic signal 218 from the mixture of acoustic signals 108 with the aid of an extraction model 214. Additionally, the extraction model 214 is configured to generate one or more queries associated with mutually inclusive and mutually exclusive features of the target acoustic signals during training of the neural network 210. The neural network 210 is trained using a set 302 of predetermined digital representations of multiple combinations 304 of mutually inclusive features to extract the target acoustic signal 218. Furthermore, the neural network 210 is trained to generate localization information for the target acoustic signal 218 indicating a location of origin of a sound source among multiple sound sources 102 of the target acoustic signal 218.
[0078] In step 816, the method includes outputting the extracted target acoustic signal together with the localization information with the aid of the output interface 216. The method ends in step 814.
[0079] 9 shows a block diagram 900 of an audio processing system 112 for performing processing of a mixture of audio signals 108 according to some embodiments of the present disclosure. In some exemplary embodiments, the block diagram 900 includes one or more microphones 106 that collect data including the mixture of audio signals 108 of multiple audio sources 102 from an environment 902.
[0080] The sound processing system 112 includes a hardware processor 908. The hardware processor 908 is in communication with computer storage memory, such as memory 910. The memory 910 contains stored data, including algorithms, instructions, and other data, implemented by the hardware processor 908. It is contemplated that the hardware processor 908 may include two or more hardware processors depending on the requirements of a particular application. The two or more hardware processors may be internal or external. The sound processing system 112 may be incorporated into other components, including output interfaces and transceivers, among other devices.
[0081] In some alternative embodiments, the hardware processor 908 is connected to a network 904 that communicates with the mixture of acoustic signals 108. The network 904 includes, by way of non-limiting example and not limitation, one or more local area networks (LANs) and / or wide area networks (WANs). The network 904 also includes enterprise-wide computer networks, intranets, and the Internet. The acoustic processing system 112 includes one or more client devices, storage components, and data sources. Each of the one or more client devices, storage components, and data sources comprises one or more devices that cooperate in the distributed environment of the network 904.
[0082] In some other alternative embodiments, the hardware processor 908 is connected to a network-enabled server 914, which is connected to the client device 916. The network-enabled server 914 corresponds to a dedicated computer connected to a network that executes software intended to process client requests received from the client device 916 and provide appropriate responses on the client device 916. The hardware processor 908 is connected to an external memory device 918 that stores all necessary data used for target acoustic signal extraction, and to a transmitter 920. The transmitter 920 facilitates the transmission of data between the network-enabled server 914 and the client device 916. Additionally, an output 922 associated with the target acoustic signal and localization information for the target acoustic signal is generated.
[0083] The mixture of acoustic signals 108 is further processed by a neural network 210. The neural network 210 is trained with mutually inclusive feature combinations 906 of each acoustic signal. The mutually inclusive feature combinations 906 are provided to the neural network 210 for training thereof (as described in FIG. 7). The mutually inclusive feature combinations 906 are in the form of a digital representation 224.
[0084] Many variations and other embodiments of the disclosures described herein will come to mind to one skilled in the art to which these disclosures pertain having the benefit of the teachings presented in the foregoing description and the associated drawings. It is to be understood that the disclosure is not to be limited to the particular embodiments disclosed, and that variations and other embodiments are intended to be included within the scope of the appended claims. Moreover, while the foregoing description and the associated drawings describe exemplary embodiments in the context of certain illustrative combinations of elements and / or functions, it should be recognized that various combinations of elements and / or functions may be provided in alternative embodiments without departing from the scope of the appended claims. In this regard, combinations of elements and / or functions other than those expressly described above are also contemplated, for example, as may be set forth in some of the appended claims. Although specific terms are employed herein, they are used in an inclusive and descriptive sense only and not for purposes of limitation.
Claims
1. 1. An acoustic processing system for extracting a target acoustic signal, comprising: at least one processor; and a memory having instructions stored thereon, the instructions, when executed by the at least one processor, causing the sound processing system to: collecting a mixture of acoustic signals together with the target acoustic signal; collecting a query that identifies the target acoustic signal to be extracted from the mixture of acoustic signals, the query including one or more identifiers; and extracting each identifier of the one or more identifiers from the query using an extraction model configured to identify the one or more identifiers included in the query, wherein each identifier is in a predetermined set of one or more identifiers, and each identifier defines at least one of mutually inclusive and mutually exclusive features of the mixture of acoustic signals; and determining one or more logical operators to be applied to the extracted one or more identifiers; converting the extracted one or more identifiers and the one or more logical operators into a predetermined digital representation for querying the mixture of acoustic signals; implementing a neural network trained to extract the target acoustic signals identified by the digital representations from the mixture of acoustic signals by combining the digital representations with intermediate outputs of an intermediate layer of the neural network that processes the mixture of acoustic signals, the neural network being trained by machine learning to extract different acoustic signals identified in a predetermined set of the digital representations; and an acoustic processing system configured to output the extracted target acoustic signal;
2. 10. The sound processing system of claim 1, wherein the sound signals in the mixture of sound signals are collected from a plurality of sound sources with the aid of one or more microphones, each sound source of the plurality of sound sources corresponding to at least one of a speaker, a person or individual, industrial equipment, a vehicle, or a natural sound.
3. 2. The sound processing system of claim 1, wherein the predetermined set of one or more identifiers is associated with a plurality of sound sources, and each of the one or more identifiers in the predetermined set of one or more identifiers includes at least one of a loudest sound source identifier, a quietest sound source identifier, a farthest sound source identifier, a nearest sound source identifier, a female speaker identifier, a male speaker identifier, and a language-specific sound source identifier.
4. 2. The acoustic processing system of claim 1, wherein the one or more identifiers are combined using the one or more logical operators to extract the target acoustic signal having mutually inclusive and mutually exclusive characteristics, the one or more logical operators including at least one of a NOT operator, an AND operator, and an OR operator, and a NOT operator is used with any single identifier of the one or more identifiers.
5. The sound processing system of claim 1 , wherein the neural network is trained using a predetermined set of the digital representations of multiple combinations of multiple identifiers in the predetermined set.
6. The sound processing system of claim 1 , wherein the neural network is trained with a positive example selector and a negative example selector to extract the target sound signal.
7. The sound processing system of claim 1 , wherein the digital representation is represented by at least one of a one-hot conditional vector, a multi-hot conditional vector, and a text description.
8. The hidden layer of the neural network includes one or more associated blocks, each of which includes at least one of a feature encoder, a conditioning network, a separation network, and a feature decoder, the conditioning network including a feature-invariant linear modulation (FiLM) layer that takes as input the mixture of the acoustic signals and the digital representation and modulates the input into a conditioned input, the FiLM layer processing the conditioned input and generating the processed conditioned input.
10. The sound processing system of claim 1, wherein the isolation network is configured to provide a conditioned input.
9. 9. The sound processing system of claim 8, wherein the separation network includes a convolutional block layer that utilizes the conditioning input to separate the target sound signal from the mixture of sound signals, the separation network being configured to generate a latent representation of the target sound signal.
10. The sound processing system of claim 8 , wherein the feature decoder converts the latent representation of the target sound signal produced by the separation network into an audio waveform.
11. 1. A computer-implemented method for extracting a target acoustic signal, comprising: collecting a mixture of acoustic signals from a plurality of sound sources; selecting a query that identifies the target acoustic signal to be extracted from the mixture of acoustic signals, the query including one or more identifiers, the method further comprising: extracting each of the one or more identifiers from the query using an extraction model configured to identify the one or more identifiers included in the query, wherein each of the identifiers is in a predetermined set of one or more identifiers, and each identifier defines at least one of mutually inclusive and mutually exclusive features of the mixture of acoustic signals, the method further comprising: determining one or more logical operators to be applied to the extracted one or more identifiers; converting the extracted one or more identifiers and the one or more logical operators into a predetermined digital representation for querying the mixture of acoustic signals; and performing a neural network trained to extract the target acoustic signals identified by the digital representations from the mixture of acoustic signals by combining the digital representations with intermediate outputs of an intermediate layer of the neural network that processes the mixture of acoustic signals, the neural network being trained by machine learning to extract the target acoustic signals identified in a predetermined set of the digital representations, the method further comprising: a computer-implemented method comprising outputting the extracted target acoustic signal.
12. 12. The computer-implemented method of claim 11, wherein the mixture of acoustic signals is collected from multiple sound sources with the aid of one or more microphones, the multiple sound sources corresponding to at least one of a speaker, a person or individual, industrial equipment, and a vehicle.
13. The predetermined set of one or more identifiers is associated with a plurality of sound sources, and each of the one or more identifiers in the predetermined set of one or more identifiers is associated with a loudest 12. The computer-implemented method of claim 11, wherein the identifiers include at least one of a loudest sound source identifier, a quietest sound source identifier, a farthest sound source identifier, a nearest sound source identifier, a female speaker identifier, a male speaker identifier, and a language-specific sound source identifier.
14. 12. The computer-implemented method of claim 11, wherein the one or more identifiers are combined using the one or more logical operators to extract the target acoustic signal having mutually inclusive and mutually exclusive characteristics.
15. 15. The computer-implemented method of claim 14, wherein the neural network is trained with the set of predetermined digital representations of multiple combinations of multiple identifiers within the predetermined set.
16. 12. The computer-implemented method of claim 11, further comprising generating one or more queries associated with the mutually inclusive and mutually exclusive features of the target acoustic signal during training of the neural network.
17. 12. The computer-implemented method of claim 11, wherein the hidden layer of the neural network includes one or more associated blocks, each of the one or more associated blocks including at least one of a feature encoder, a conditioning network, a separation network, and a feature decoder, the conditioning network comprising a Feature Invariant Linear Modulation (FiLM) layer that takes the mixture of acoustic signals as an input and modulates the input into a conditioning input, the FiLM layer processes the conditioning input and sends the processed conditioning input to the separation network.
Citation Information
Patent Citations
Audio enhancement through supervised latent variable representation of target speech and noise
US20200349965A1
Signal processing device, signal processing method, signal processing program, learning device, learning method, and learning program
WO2022034675A1