Audio source separation using hyperbolic embedding
Hyperbolic embeddings in a neural network framework enhance audio source separation by projecting audio signals into a hyperbolic space, addressing the challenge of complex source separation in real-world scenarios and improving separation accuracy and visualization.
Patent Information
- Application Number
- CN202380082549.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-03-28
- Filing Date
- 2023-09-08
- Publication Date
- 2025-07-15
AI Technical Summary
Existing audio source separation technology is difficult to effectively separate multiple audio sources when dealing with complex audio scenarios, especially when the source definition is unclear, the lack of hierarchical representation of time-frequency units leads to poor separation results.
Hyperbolic embedding technology is used to map the time-frequency units into the hyperbolic space, and the audio source is separated using the hierarchical relationship of the hyperbolic space, and a specific sound source is selected through deterministic filtering and user interaction to realize the visualization and separation of the audio source.
Improves the accuracy and efficiency of audio source separation, can retain hierarchical information in low-dimensional embedded dimensions, reduces computing resource consumption, and provides user-friendly selection and visual interface.
Smart Images

Figure CN120322818A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to audio source separation, and more particularly, to a method and apparatus for audio source separation using hyperbolic embedding for time-frequency bins of spectrograms of audio mixed signals. Background Art
[0002] The diagnosis and analysis of sound signal mixtures requires separating these signals into sound sources of interest. With the introduction of deep learning techniques, significant performance improvements have been achieved in the field of audio source separation, most notably in the areas of speech enhancement, speech separation, and music separation. Specifically, deep learning techniques are used to separate human speech from background noise (or other speech), or to isolate specific instruments (e.g., vocals, drums, etc.). These techniques are successful when the concept of the source is well-defined, such as in the case of speech where the target is defined as the speech of a single speaker. However, in real-world scenarios, the source may not be well-defined and can be complex to process.
[0003] For example, in a scenario where a machine is operating in a large plant or factory while two people are talking nearby. In this scenario, the machine and the two people can each be an audio source. In fact, each part of the machine may also be an audio source. Therefore, segmenting this audio scene is difficult. Existing audio source separation systems typically use fixed and somewhat arbitrary source definitions. For example, music mixtures are usually separated into "vocals", "bass", "drums", and "other". In particular, existing models allow for separating sound scenes at multiple granularity levels (e.g., "drum kit" vs. "snare drum", "train" vs. "brake disc", etc.), but these models only consider the global hierarchy in terms of sound class labels and do not consider the hierarchy that may exist in each time-frequency bin of the sound mixture.
[0004] Hierarchical and tree-like structures are prevalent in many types of audio processing problems, such as instrument recognition and separation, speaker recognition, and sound event detection. However, all of these methods globally model the hierarchical information by computing a single embedding vector for the entire audio clip.
[0005] Audio source separation involves extracting the isolated sounds of individual sources from a sound signal mixture. For example, techniques that directly learn feature encoders and decoders based on waveform signals have achieved impressive performance. However, they lack interpretability compared to techniques based on time-frequency (T-F) representations. To overcome this, algorithms similar to deep clustering are used to assign embeddings to different sources by learning the embedding vectors of individual T-F units. Then the basic problem of these methods becomes how to best learn the discriminative embeddings of individual T-F units.
[0006] Therefore, an audio source separation technology for solving the above problems is needed. Summary of the Invention
[0007] The purpose of some embodiments is to achieve sound separation of audio sources in audio mixing. Additionally or alternatively, the purpose of some embodiments is to provide a system and method that allows a user to select a portion of a hyperbolic space corresponding to one or more sound sources and listen to the audio corresponding to the selected portion as a separated sound source. Another purpose of some embodiments is to use different hyperbolic hyperplanes learned for each sound category type to partition the hyperbolic space and allow the user to select each hyperbolic hyperplane. Additionally or alternatively, the purpose of some embodiments is to use different hyperbolic hyperplanes learned for each sound category type to partition the hyperbolic space and separate the audio mixing into output signals for each sound category. Additionally or alternatively, the purpose of some embodiments is to use deterministic filtering to separate the audio mixing. Additionally or alternatively, the purpose of some embodiments is to provide a neural network that is trained to map and classify embeddings for spectrogram units or time-frequency (T-F) units.
[0008] Generally, deep learning techniques are used to separate audio mixing, where the embeddings are represented in Euclidean space. However, Euclidean space lacks the ability to infer hierarchical representations from data without distortion.
[0009] Some embodiments are based on the understanding that hyperbolic space can be used instead of Euclidean space. For example, T-F units are mapped to hyperbolic embeddings and projected into hyperbolic space using a neural network. In this scenario, the embeddings on the hyperbolic manifold provide a framework for audio source separation, which compactly represent the hierarchical relationship between the sound sources and the time-frequency features. Using hyperbolic space, strong hierarchical embeddings can be obtained at a low embedding dimension compared to Euclidean embeddings, which can save computational resources during training and inference. Inspired by the recent successful modeling of hierarchical relationships in text and images using hyperbolic embeddings, the proposed framework obtains hyperbolic embeddings for each time-frequency unit of the spectrogram of the mixed signal and uses a hyperbolic softmax layer to estimate the T-F mask.
[0010] Furthermore, in hyperbolic space, the time-frequency regions including multiple overlapping audio sources are embedded towards the center of the hyperbolic space (i.e., the most uncertain region), and the regions of the hyperbolic space towards the edge correspond to the audio of a single sound source. This realizes a visual representation of the mixed audio signal including multiple audio sources. The visual representation of the mixed audio signal realizes the visualization of different audio signals existing in the mixed audio signal. From the visual representation of the mixed audio signal and the analysis of the hyperbolic space, a deterministic estimate of each sound can be inferred to efficiently balance between introducing artifacts and reducing interference when isolating each sound.
[0011] Audio Source Selection
[0012] For example, in some embodiments, given a hyperbolic embedding, an embedding classifier as part of a neural network can be used to classify these embeddings. However, the user cannot select the embeddings to be classified as required. Some embodiments are based on the recognition that the user should be able to select a specific sound source.
[0013] To this end, some embodiments of the present disclosure disclose an interface that accepts input from a user to select a region of a hyperbolic space, creates a T-F mask, and applies the created T-F mask to an original mixed audio signal to obtain a separated source corresponding to the selected region of the hyperbolic space. The interface enables visualization of individual sound sources in the mixed audio signal and selection of one or more audio sources for listening using a device such as (but not limited to) a speaker.
[0014] Deterministic filtering
[0015] In addition, in some embodiments, there is a lack of confidence in the sources of the separated hyperbolic embeddings. Some embodiments are based on the recognition that audio mixing can be separated using deterministic filtering. For example, some embodiments of the present disclosure disclose a hyperbolic separation network that performs two-stage source separation using deterministic filtering. Deterministic filtering uses T-F units near the boundary of the hyperbolic space of the learned hyperbolic embeddings because these units are likely to belong to a single source. The hyperbolic embeddings are filtered using an estimate of the distance of the hyperbolic embedding from the origin of the hyperbolic space. The greater the distance of the hyperbolic embedding from the origin, the more likely it belongs to a single audio source, and the smaller the distance of the hyperbolic embedding from the origin, the more likely it includes multiple overlapping audio sources. By setting an optimal threshold for the distance from the origin, deterministic filtering ensures no interference from overlapping audio sources.
[0016] Accordingly, one embodiment discloses a system that includes: an input interface that receives an input audio mix and transforms it into a time-frequency representation defined by values of time-frequency units; a processor that maps the values of the time-frequency units into a hyperbolic space by executing an embedding neural network trained to associate individual time-frequency units with high-dimensional embeddings and projecting the individual high-dimensional embeddings into the hyperbolic space; and an output interface that accepts a selection of at least a portion of the hyperbolic space and renders selected hyperbolic embeddings that fall within the selected portion of the hyperbolic space.
[0017] Accordingly, another embodiment discloses a method that includes the steps of: receiving an input audio mix and transforming it into a time-frequency representation defined by values of time-frequency units; mapping the values of the time-frequency units into a hyperbolic space by executing an embedding neural network trained to associate individual time-frequency units with high-dimensional embeddings and projecting the individual high-dimensional embeddings into the hyperbolic space; and accepting a selection of at least a portion of the hyperbolic space and rendering selected hyperbolic embeddings that fall within the selected portion of the hyperbolic space.
[0018] Accordingly, another embodiment discloses a non-transitory computer-readable storage medium having thereon a program that can be executed by a processor to perform a method, the method comprising the steps of: receiving an input audio mixture and transforming it into a time-frequency representation defined by values of time-frequency units; mapping the values of the time-frequency units into a hyperbolic space by executing an embedding neural network trained to associate each time-frequency unit with a high-dimensional embedding and projecting each high-dimensional embedding into the hyperbolic space; and receiving a selection of at least a portion of the hyperbolic space and rendering selected hyperbolic embeddings that fall within the selected portion of the hyperbolic space.
[0019] Embodiments of the present disclosure will be further described with reference to the accompanying drawings. The drawings shown are not necessarily to scale, and emphasis is generally placed on illustrating the principles of the embodiments of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 Figure 1 A block diagram showing an exemplary network environment for audio source separation using hyperbolic embeddings in accordance with some embodiments of the present disclosure.
[0021] Figure 2 Figure 2 A block diagram showing an exemplary audio processing system for audio source separation using hyperbolic embeddings in accordance with some embodiments of the present disclosure.
[0022] Figure 3 Figure 3 A flowchart showing an exemplary method of hyperbolic audio source separation in accordance with some embodiments of the present disclosure.
[0023] Figure 4 Figure 4 A flowchart showing the operation of an audio processing system for hyperbolic audio source separation in accordance with some embodiments of the present disclosure.
[0024] Figure 5A Figure 5A A spectrogram showing the time-frequency representation of an audio mixture signal in accordance with some embodiments of the present disclosure.
[0025] Figure 5B Figure 5B A Poincaré sphere of a hyperbolic space including classified hyperbolic embeddings in accordance with some embodiments of the present disclosure.
[0026] Figure 6 Figure 6 A Poincaré sphere having a decision boundary in accordance with some embodiments of the present disclosure.
[0027] Figure 7A Figure 7A An interface for selecting a hyperbolic embedding in hyperbolic space according to some embodiments of the present disclosure is shown.
[0028] Figure 7B Figure 7B An interface for displaying a masked spectrogram according to some embodiments of the present disclosure is shown.
[0029] Figure 8A Figure 8A An interface for selecting one or more hyperplanes in hyperbolic space according to some embodiments of the present disclosure is shown.
[0030] Figure 8B Figure 8B An interface for displaying a masked spectrogram according to some embodiments of the present disclosure is shown.
[0031] Figure 9 Figure 9 A flowchart depicting the operation of an audio processing system for hyperbolic audio source separation according to some embodiments of the present disclosure is shown.
[0032] Figure 10 Figure 10 A flowchart depicting the operation of an audio processing system for hyperbolic audio source separation using deterministic filtering according to some embodiments of the present disclosure is shown.
[0033] Figure 11 Figure 11 A flowchart depicting a hyperbolic separation network according to some embodiments of the present disclosure is shown.
[0034] Figure 12 Figure 12 A flowchart depicting an exemplary remote diagnosis application according to some embodiments of the present disclosure is shown.
[0035] Figure 13 Figure 13 An interface for a control knob for hyperbolic audio source separation according to some embodiments of the present disclosure is shown. Detailed Description
[0036] Some embodiments of the present disclosure provide a method and apparatus for audio source separation using hyperbolic embedding for time-frequency units of a spectrogram of an audio mixed signal. The audio mixed signal includes signals from different audio sources.
[0037] As used in this specification and the claims, the terms "such as" and "for example" and the verbs "comprising," "having," "including," and other verb forms when used in connection with a list of one or more components or other items shall each be construed as open-ended, meaning the list should not be considered as excluding other additional components or items. The term "based on" means at least partially based on. Additionally, it will be understood that the language and terminology used herein are for descriptive purposes and should not be regarded as limiting. Any headings utilized within this description are for convenience only and have no legal or limiting effect.
[0038] Figure 1 is a block diagram showing an exemplary network environment for audio source separation using hyperbolic embedding according to some embodiments of the present disclosure. Referring to Figure 1 , a network environment 100 is shown. The network environment 100 may include multiple sound sources that result in an audio mixed signal 102. The network environment 100 may also include an audio processing system 104, a communication network 106, and a database 108. In Figure 1 , the audio processing system 104, the communication network 106, and the database 108 are shown as separate devices. However, in some embodiments, without departing from the scope of the present disclosure, the entire functionality of the audio processing system 104, the communication network 106, and the database 108 may be incorporated into the audio processing system 104.
[0039] The audio mixed signal 102 may include audio signals from multiple audio sources. For example, the input audio mix may include the sounds of multiple instruments such as guitars, pianos, drums, and other similar instruments. In another example, the network environment 100 may correspond to an audio scene from an industrial site where multiple sounds originate. The sounds may include (but are not limited to) two people talking to each other, sounds from different parts of a machine. Specifically, the multiple audio sources cannot be identified in the time-domain or spectral-domain representation of the audio mixed signal 102.
[0040] The audio processing system 104 may include suitable logic, circuitry, code, and / or interfaces that may be configured to accept the audio mixed signal 102 as an input and output one or more audio signals as separated audio signals. One or more separated audio signals may include an audio signal 112 from audio source 1, an audio signal 114 from audio source 2, or an audio signal 116 from audio source n. The underlying technology for extracting one or more audio signals 112, 114, 116 from the audio mixed signal 102 is referred to Figure 4 is shown in detail. An exemplary audio processing system 104 is generated with reference to Figure 2 in detail.
[0041] The communication network 106 may include a communication medium through which the audio processing system 104 can communicate with the database 108 and other devices, which are omitted in this disclosure for the sake of brevity. In an embodiment, the user 110 may provide an input to the audio processing system 104 through the communication network 106. The communication network 106 may be one of a wired connection or a wireless connection. Examples of the communication network 106 may include (but are not limited to) the Internet, a cloud network, a Wi-Fi network, a personal area network (PAN), a local area network (LAN), or a metropolitan area network (MAN). Various devices in the network environment 100 may be configured to connect to the communication network 106 according to various wired and wireless communication protocols. Examples of these wired and wireless communication protocols may include (but are not limited to) at least one of the Transmission Control Protocol and Internet Protocol (TCP / IP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), File Transfer Protocol (FTP), ZigBee, EDGE, IEEE 802.11, Li-Fi, 802.16, IEEE 802.11s, IEEE 802.11g, multi-hop communication, a wireless access point (AP), device-to-device communication, a cellular communication protocol, and a Bluetooth (BT) communication protocol.
[0042] The database 108 may include suitable logic, circuitry, and / or interfaces, which may be configured to store a neural network model trained to separate audio sources from the audio mixed signal 102. In another embodiment, the database 108 may store program instructions to be executed by the audio processing system 104. Example implementations of the database 108 may include (but are not limited to) random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), a hard disk drive (HDD), a solid state drive (SSD), a CPU cache, and / or a Secure Digital (SD) card.
[0043] The user 110 may be a person capable of providing an input to the audio processing system 104 for audio source separation. The input modes may include (but are not limited to) touch input, gesture input, or input by pressing some buttons on a control device linked to the audio processing system 104.
[0044] Figure 2 A block diagram showing an exemplary audio processing system for audio source separation using hyperbolic embedding according to some embodiments of the present disclosure is shown. Figure 2 In combination with elements from Figure 1 will be described. Referring to Figure 2, showing a block diagram 200 of an audio processing system 104. The audio processing system 104 may include a processor 202, a memory 204, an output interface 206, and an input interface 208. The processor 202 may be communicatively coupled to the memory 204, the output interface 206, and the input interface 208. In some embodiments, the output interface 206 may include a display device.
[0045] The processor 202 may include suitable logic, circuitry, and / or interfaces that may be configured to execute program instructions associated with different operations to be performed by the audio processing system 104. The processor 202 may include one or more specialized processing units that may be implemented as an integrated processor or a processor cluster that jointly executes the functions of one or more specialized processing units. The processor 202 may be implemented based on a variety of processor technologies known in the art. Example implementations of the processor 202 may be an x86-based processor, a graphics processing unit (GPU), a reduced instruction set computing (RISC) processor, an application specific integrated circuit (ASIC) processor, a complex instruction set computing (CISC) processor, a microcontroller, a central processing unit (CPU), and / or other computing circuitry.
[0046] The memory 204 may include suitable logic, circuitry, and / or interfaces that may be configured to store program instructions to be executed by the processor 202. The memory 204 may also be configured to store a trained neural network, such as an embedded neural network or an embedded classifier. Without departing from the scope of the present disclosure, trained neural networks such as embedded neural networks and embedded classifiers may also be stored in the database 108. Example implementations of the memory 204 may include (but are not limited to) random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), hard disk drive (HDD), solid state drive (SSD), CPU cache, and / or secure digital (SD) card.
[0047] The output interface 206 may include suitable logic, circuitry, interfaces, and / or code that may be configured to receive an input and provide an output based on the received input. The output interface 206 may include various input and output devices that may be configured to communicate with the processor 202. For example, the audio processing system 104 may receive user input that selects a region on the hyperbolic space or a hyperplane on the hyperbolic space via the input interface 208. Examples of the input interface 208 may include (but are not limited to) a touch screen, a keyboard, a mouse, a joystick, or a microphone.
[0048] The input interface 208 may also include suitable logic, circuitry, interfaces, and / or code that may be configured to facilitate communication of the processor 202 with the database 108 and / or other communication devices via the communication network 106. The input interface 208 may be implemented using various known techniques to support wireless communication of the audio processing system 104 via the communication network 106. For example, the input interface 208 may include an antenna, a radio frequency (RF) transceiver, one or more amplifiers, a tuner, one or more oscillators, a digital signal processor, an encoder-decoder (CODEC) chipset, a subscriber identity module (SIM) card, local buffer circuitry, and the like.
[0049] The input interface 208 may be configured to communicate via wireless communication with a network such as the Internet, an intranet, a wireless network, a cellular telephone network, a wireless local area network (LAN), or a metropolitan area network (MAN). The wireless communication may be configured to use one or more of a variety of communication standards, protocols, and technologies such as the Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (W-CDMA), Long Term Evolution (LTE), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Wireless Fidelity (Wi-Fi) (e.g., IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, or IEEE 802.11n), Voice over Internet Protocol (VoIP), Light Fidelity (Li-Fi), or Worldwide Interoperability for Microwave Access (Wi-MAX).
[0050] As Figure 1 described, the functions or operations performed by the audio processing system 104 may be performed by the processor 202. The operations performed by the processor 202 are described in detail, for example, in Figure 3 and Figure 4 .
[0051] Figure 3 FIG. shows a flowchart depicting an exemplary method of hyperbolic audio source separation in accordance with some embodiments of the present disclosure. Figure 3 In conjunction with elements from Figure 1 and Figure 2 . Referring to Figure 3 , FIG. 300 is shown. The method shown in the flowchart 300 may be performed by any computing system, such as the audio processing system 104 that separates the audio mixed signal 102 into its multiple audio sources using hyperbolic embedding. The method may begin at 302 and proceed to 310.
[0052] At 302, the audio processing system 104 receives the audio mixed signal 102 directly or via the communication network 106. The audio mixed signal 102 may include audio signals from different sources. The source of the audio signals in the audio mixed signal 102 may be identified by at least one of, for example (but not limited to), pitch, amplitude, or timbre.
[0053] At 304, the audio processing system 104 transforms the received audio mixed signal into a spectrogram. A spectrogram is a time-frequency representation of the audio mixed signal 102. Alternatively, a spectrogram may be a visual representation of the spectrum of a signal changing over time. It represents the signal intensity or "loudness" of the various frequencies present in a particular waveform over time. A spectrogram also allows the visualization of the energy level of a sound over time.
[0054] At 306, the audio processing system 104 extracts the time-frequency (T-F) units of the audio mixed signal 102 from the generated spectrogram, maps the T-F units to high-dimensional embeddings using an embedded neural network, and projects the high-dimensional embeddings into hyperbolic space. Details of the embedded neural network are described in detail with reference to Figure 4 As an example, the embedded neural network may include (but not be limited to) multiple bidirectional long short-term memory (BLSTM) layers, convolutional layers, or transformer layers. Each T-F unit passes through the layers of the neural network, which transform its time-frequency data into D-dimensional embeddings. These embeddings are then mapped to high-dimensional embeddings or hyperbolic embeddings according to some embodiments of the present disclosure.
[0055] At 308, the audio processing system 104 accepts a user's selection of a region of the hyperbolic space. In an embodiment, the user may use a resizable selection tool to select a portion of the hyperbolic space. The shape and size of the selection tool depend on the user's selection. In another embodiment of the present disclosure, the user may select an entire hyperplane of the hyperbolic space. In an embodiment, the user may select one or more regions on the hyperbolic space. Selecting a region of the hyperbolic space results in the selection of the hyperbolic embeddings that fall within the selected portion of the hyperbolic space.
[0056] At 310, the audio processing system 104 renders the selected hyperbolic embeddings. Rendering the selected hyperbolic embeddings includes transforming the selected hyperbolic embeddings into separate output audio signals and sending the output audio signals to a speaker.
[0057] Figure 4 A flowchart showing the operation of an audio processing system for hyperbolic audio source separation according to some embodiments of the present disclosure. Figure 4 In combination with elements from Figure 1 , Figure 2 and Figure 3 are described. With reference to Figure 4, shows a flowchart 400. The flowchart 400 provides a description of the operations performed by the audio processing system 104 for audio source separation. Referring to Figure 1 and Figure 2 , the input interface 208 of the audio processing system 104 can perform functions related to the spectrogram extractor 404. The processor 202 can perform functions associated with the hyperbolic separation network 406 and the embedding classifier 412. The output interface 206 can perform functions related to the selection interface 414, the creation 416 of the T-F mask, the application 418 of the T-F mask, or the spectrogram inversion 420.
[0058] The audio processing system 104 obtains an audio mixed signal 402 from the environment. The environment can include (but is not limited to) an industrial site or factory, a smart AI assistant, a concert, etc., where multiple overlapping sound events coexist. Thus, the audio mixed signal 402 includes sounds from various sources that overlap in the time domain and the frequency domain. For example, machines operating in a large plant or factory generate such an audio mix while two people are talking nearby. The mix can contain audio from various machine sounds, background noise, and human voices. In an embodiment, the audio processing system 104 can include a set of sensors to detect audio signals from different sources in the environment. This set of sensors can include (but is not limited to) microphones, such as dynamic microphones, carbon microphones, ribbon microphones, or piezoelectric microphones. This set of sensors provides the audio signals detected from different sources as the audio mixed signal 402 to the spectrogram extractor 404 of the audio processing system 104 for further processing.
[0059] The spectrogram extractor 404 processes the audio mixed signal 402 to convert it into a spectrogram, which is a representation of the sound signal in a time-frequency format. A spectrogram is a visual way of representing the signal intensity or loudness of the different components of the audio mixed signal 402 over time at various frequencies present in the audio mixed signal 402. The details of the spectrogram are described with reference to Figure 5A the description of Figure 5A , shows a spectrogram 500A generated by the spectrogram extractor 404. The spectrogram 500A of the audio mixed signal 402 is sampled at regular intervals to generate time-frequency (T-F) units. In an example, a T-F unit represents the value of the spectrogram at different time instances. In Figure 5A , the T-F unit 502 represents the value 502 of the spectrogram. The values of the T-F units are extracted from the spectrogram and provided to the hyperbolic separation network 406 for transforming the values of the extracted T-F units into hyperbolic embeddings and projecting the hyperbolic embeddings into hyperbolic space.
[0060] In an embodiment, the hyperbolic separation network 406 includes an embedding neural network 408 and a hyperbolic projection 410. Without departing from the scope of the present disclosure, in another embodiment, the hyperbolic separation network 406 may additionally include an embedding classifier 412. For example, the embedding classifier 412 can be used to automatically segment the hyperbolic space using hyperbolic hyperplanes for different sound classes or audio sources in the audio mixed signal 402, as further described with reference to the description of FIG. 8. In another example, the embedding classifier 412 is used during the training and validation of the hyperbolic separation network 406. Details of the training and validation of the embedding neural network 408 are described with reference to Figure 11 the description of
[0061] The embedding neural network 408 takes the T-F units from the spectrogram as the input to its layers and outputs a D-dimensional hyperbolic embedding corresponding to the T-F units. In other words, the embedding neural network 408 performs a transformation of values from the time-frequency domain to the hyperbolic space. The embedding neural network 408 can be composed of multiple bidirectional long short-term memory (BLSTM) layers, which can be replaced by other types of deep network layers, such as convolutional layers, transformer layers, etc.
[0062] The hyperbolic projection 410 projects the D-dimensional hyperbolic embedding into the hyperbolic space. The hyperbolic space is represented as a Poincaré ball, as Figure 5B shown. The Poincaré ball is a hyperbolic geometry model, where lines are represented as circular arcs, and the endpoints of the circular arcs are perpendicular to the boundary of the disk (diameters are also allowed). Two non-intersecting arcs correspond to parallel rays, orthogonally intersecting arcs correspond to perpendicular lines, and arcs intersecting at the boundary are a pair of limiting rays.
[0063] Each hyperbolic embedding is visually indicated as a point at a specific position on the Poincaré ball. The position of the hyperbolic embedding on the Poincaré ball is calculated based on the audio mixed spectrogram and the values of the corresponding T-F units. The information associated with each hyperbolic embedding includes information about the original sound corresponding to the T-F units. Each value on the Poincaré ball includes some information about the original sound or the audio mixed signal 402. Advantageously, the hyperbolic embedding preserves the hierarchical information associated with the original sound, while the Euclidean embedding does not. The hierarchical representation in Euclidean space is not very compact and requires a large amount of spatial memory. On the other hand, representing the hierarchy on the Poincaré ball is advantageous because the hyperbolic representation requires less memory and has a smaller computational range as it compactly represents the hierarchy. Additionally, there is the advantage of obtaining a visual representation of the sound. Compared with non-hierarchical networks, low-dimensional embeddings can be used.
[0064] A Riemannian manifold is defined as a pair consisting of a manifold M and a Riemannian metric g, where g = (g x ) x∈M defines the local geometry g at each point x ∈ M in the tangent space T x M x(i.e., the inner product). Although g defines the local geometry on M, it also defines the global shortest path or geodesic (similar to a straight line in Euclidean space) between two given points on M. The exponential map exp x projects any vector v in the tangent space T x of M onto M such that exp x (v) ∈ M. Conversely, the logarithmic map projects any point in M back to the tangent space at x.
[0065] The L-dimensional hyperbolic space is an L-dimensional Riemannian manifold with a constant negative curvature -c. It can be described using multiple isometric models, mainly the Poincaré unit ball model defined in the space as assuming c > 0 such that corresponds to a ball with radius in Euclidean space. Its Riemannian metric is given by where is the so-called conformal factor and g E is the Euclidean metric. Given two points, the induced hyperbolic distance d c is obtained as
[0066]
[0067] where denotes the Möbius addition in
[0068]
[0069] One way to go back and forth between Euclidean space R L and hyperbolic space is to use the exponential and logarithmic maps at the origin 0, as
[0070] For v ∈ R L \{0} and it can be obtained as:
[0071]
[0072] By simply projecting the common Euclidean embedding in this way, the hyperbolic embedding at the output of a classical neural network such as f θ can be obtained as
[0073]
[0074] To define hyperbolic softmax based on these hyperbolic embeddings, Euclidean MLR can be generalized to the Poincaré ball. In Euclidean space, consider performing MLR on pairs obtained by computing the distance of the input embedding z ∈ R L (e.g., ) to each of the K-class hyperplanes, where the k-th class hyperplane is determined by the normal vector a k ∈ R L and a point p k ∈ R L on this hyperplane. Similarly, the Poincaré hyperplane can be defined as the union of all geodesics passing through the point p k and orthogonal to the normal vector a at p k in the tangent space k . Then, the distance from the hyperbolic embedding to each can be considered to define hyperbolic MLR, resulting in the following formula:
[0075]
[0076] The probability that the T-F unit (t, f) is dominated by the k-th source can be used to obtain the K source-specific mask values for each T-F unit in the input spectrum Figure X . Both pk and ak parameterize the k-th hyperbolic hyperplane and are trainable. The network f θ and all the parameters of the hyperplane can be optimized using classical source separation objective functions on the mask or the reconstructed signal.
[0077] The hyperbolic projection 410 is a formula that projects the embedding output of the embedding neural network 408 into the hyperbolic space or the Poincaré ball. Besides the benefits in hierarchical modeling, the hyperbolic layer can learn low-dimensional strong embeddings, representing uncertainty based on their distance from the origin in the hyperbolic space.
[0078] In an implementation, the audio processing system 104 accepts user input via the selection interface 414 to select a region of the hyperbolic space. In an example, the user can select a set of hyperbolic embeddings on the hyperbolic space. In another example, the user can select a hyperbolic hyperplane on the hyperbolic space. The hyperbolic hyperplane can correspond to a set of hyperbolic embeddings belonging to a specific sound class. Details of the selection interface 414 using the selection tool 702 are further described with reference to Figure 7A The selected region of the hyperbolic space can be used to create 416 T-F masks or select masks to apply to the audio mixed spectrogram. In another implementation, the selection interface 414 can also receive a weight selection of the energy of the hyperbolic embeddings included in the user-selected region.
[0079] If a hyperbolic hyperplane corresponding to a specific audio category is selected, the audio processing system 104 creates a 416 T-F mask using hyperbolic embeddings corresponding to the selected region in hyperbolic space based on a softmax operation, or if a region in hyperbolic space is selected, creates a 416 T-F mask using hyperbolic embeddings corresponding to the selected region in hyperbolic space based on a binary mask. The T-F mask is a real-valued matrix of the same size as the spectrogram. The T-F mask can include (but is not limited to) a binary mask or a soft mask. In a binary mask, the T-F units whose embeddings are within the selected part of hyperbolic space have a T-F mask value of 1, and the T-F units associated with embeddings located outside the selected part of hyperbolic space have a T-F mask value of 0. The soft mask takes non-binary continuous values. In the case of a soft mask, the mask value can include (but is not limited to) weights indicating the distance from the origin of hyperbolic space or the distance from the nearest classification hyperbolic plane. In another embodiment, the weight can depend on the position of the embedding relative to the hyperplane, i.e., whether it is on the same side as the center of hyperbolic space. In another embodiment, the weight can also depend on its position relative to one or more classification hyperbolic planes. Details of the T-F mask are further described with reference to Figure 7B the description of
[0080] Once the T-F mask is created, the audio processing system 104 applies the T-F mask to the spectrogram of the audio mixed signal 402 to generate a masked spectrogram of the audio mixed signal 402. In an example, the masked spectrogram of the audio mixed signal 402 includes only those T-F units corresponding to the hyperbolic embeddings selected by the user at the selection interface 414. Details of the masked spectrogram are further described with reference to Figure 7B the description of
[0081] The audio processing system 104 applies a spectrogram inversion 420 to the masked spectrogram to render an output signal 422 that is part of the original audio mixed signal 402 separated from the original audio mixed signal 402. In an example, the output interface 206 of the information processing system 104 can transform the selected hyperbolic embedding into a separated output audio signal (e.g., output signal 422). The output signal 422 can be an analog signal or a digital signal. In a particular embodiment, the audio processing system 104 can send the output signal 422 to a speaker. In another embodiment, the audio processing system 104 can include a speaker.
[0082] In a particular embodiment, the audio processing system 104 can include an embedding classifier 412 that is trained to classify hyperbolic embeddings in hyperbolic space into audio source categories, such as: drums, guitars, and trumpets. Details of the classification of hyperbolic embeddings into hyperbolic space are further described with reference to the description of FIG. 8.
[0083] A hierarchical representation is a method of representing data in a tree-like structure, where an increase in depth at each level represents an increase in features associated with a category. Considering the hyperbolic nature of audio signals, a hierarchical classifier is used, which consists of multiple polynomial logistic regression (MLR) layers, one MLR layer for each level in the audio sound category hierarchy (composed of parent and child categories). Two sets of masks are created, one set of N masks, containing the T-F masks for each of the N subcategories, and a second set of M masks, containing the masks for each of the M parent categories. For example, the parent categories include labels such as speech and music, while the subcategories include labels such as male speech, female speech, drums, and guitar.
[0084] After classifying the hyperbolic embeddings within the hyperbolic space, the user can use the selection interface 414 to select a specific region of the hyperbolic space. In some embodiments of the present disclosure, the selection interface can be configured to transform the selected hyperbolic embedding into a separated output audio signal and send the separated output audio signal to the memory 204, refer to Figure 2 . The memory 204 can store the separated output signal.
[0085] Once the user selects a region using the selection interface 414, the audio processing system 104 creates 416 T-F masks and applies 418 the T-F masks to the spectrogram of the audio mixed signal 402 to generate a masked spectrogram of the audio mixed signal 402. The details of the masked spectrogram are further described with reference to Figure 7B 's description.
[0086] The audio processing system applies a spectrogram inversion 916 to the masked spectrogram to render an output signal 918 that is part of the original audio mixed signal 902 separated from the original audio mixed signal 902.
[0087] In another embodiment, the audio processing system can render the hyperbolic embeddings contained in the user-selected region based on the weights of the energies of the hyperbolic embeddings.
[0088] The output signal 918 can be an analog signal or a digital signal. In a particular embodiment, the audio processing system can send the output signal 918 to a speaker. In another embodiment, the audio processing system can include a speaker.
[0089] Figure 5A A spectrogram showing the time-frequency representation of an audio mixed signal according to some embodiments of the present disclosure. Figure 5A Combined with elements from Figure 1 , Figure 2 , Figure 3 and Figure 4 for illustration. Refer to Figure 5A , a spectrogram 500A of the audio mixed signal 402 is shown. As Figure 4As shown, the spectrogram extractor 404 generates the spectrogram of the audio mixed signal 402 Figure X t,f , represented as spectrogram 500A. The spectrogram 500A shows the intensity of different spectral components of the audio mixed signal 402 over time. The spectrogram 500A includes T-F units as samples of the spectrogram 500A at different time instances. In the example, the T-F unit 502 represents the energy of the spectral component f of the audio mixed signal 402 at time t. For example, an industrial machine causes numerous varying energy waveforms during operation. Additionally, the waveform is affected by background noise from similar machines and human voices. The embedded neural network 408 takes the entire spectrogram 500A as input and transforms it into a set of high-dimensional embeddings associated with the individual T-F units in the spectrogram. For example, the T-F unit 502 is transformed into a hyperbolic embedding using Equation (6) represented by the hyperbolic embedding 504.
[0090]
[0091] Figure 5B Shows a Poincaré sphere of a hyperbolic space including a classified hyperbolic embedding according to some embodiments of the present disclosure. Figure 5B Combined with elements from Figure 1 , Figure 2 , Figure 3 and Figure 4 and Figure 5A are illustrated. Referring to Figure 5B , a Poincaré sphere 500B is shown, which includes the projection of the hyperbolic embedding in the hyperbolic space . The hyperbolic space is a Poincaré sphere 500B or a Poincaré disk classified according to hyperbolic geometry, and the hyperbolic geometry carries a hierarchical concept of audio source classification based on the position of the hyperbolic embedding relative to the origin of the Poincaré sphere 500B or the Poincaré disk. The Poincaré sphere 500B represents the hyperbolic projection including the hyperbolic embeddings corresponding to the individual T-F units from the spectrogram 500A. Each T-F unit from the spectrogram 500A is converted into a hyperbolic embedding and projected onto the hyperbolic space represented by the Poincaré sphere 500B For example, as Figure 5A shown, the T-F unit 502 is converted into the hyperbolic embedding 504 using Equation (6) and projected onto the Poincaré sphere 500B. In an embodiment, the hyperbolic projection 410 projects the hyperbolic embedding 504 onto the Poincaré sphere 500B. The hyperbolic embedding 504 projected onto the Poincaré sphere 500B is at a distance 506 from the hyperplane . The distance 506 of the hyperbolic embedding 504 from the hyperplane defines that the hyperbolic embedding 504 belongs to the hyperplane The probability of uncertainty of the corresponding sound category. Similarly, the distance of the hyperbolic embedding 504 from the origin of the Poincaré ball 500B defines the probability of uncertainty of the hyperbolic embedding 504 belonging to the sound category corresponding to the origin of the Poincaré ball 500B. In fact, embeddings located near or at the origin of the Poincaré ball 500B represent sounds that are inherently very likely to be uncertain or overlapping. In other words, time-frequency regions containing multiple overlapping sources are embedded towards the center of the Poincaré ball 500B (i.e., the most uncertain region). Therefore, the distance value of the hyperbolic embedding 504 from the origin of the Poincaré ball 500B defines the probability of uncertainty or certainty of the hyperbolic embedding 504 belonging to a specific audio category or audio source. In the example, the larger the distance value of the hyperbolic embedding 504 from the origin of the Poincaré ball 500B, the greater the probability that the hyperbolic embedding 504 belongs to a specific category. The smaller the distance value of the hyperbolic embedding 504 from the origin of the Poincaré ball 500B, the smaller the probability that the hyperbolic embedding 504 belongs to a specific category, or the higher the probability that the hyperbolic embedding 504 includes overlapping sound signals. Alternatively, hyperbolic embeddings on the edge of the Poincaré ball 500B are very likely to belong to a specific audio category.
[0092] The hyperbolic hyperplane includes parameters a k and p k , which are learned during training of the embedding network using audio data. These are similar to linear hyperplanes in Euclidean space. For example, a two-dimensional Euclidean hyperplane is a line. The hyperbolic embeddings are classified into corresponding audio source categories by hyperbolic hyperplanes, and the hyperbolic hyperplanes consider all geodesics passing through the point p k and orthogonal to the normal vector a k at p k in the tangent space. A geodesic is a curve that locally minimizes length. All embeddings within a specific hyperplane may have a higher probability of belonging to that specific audio source category. Embeddings located at the boundary of the hyperplane may have a mixed waveform. The distance of an embedding from each hyperplane can also be used to determine its audio source category. In one embodiment, hyperbolic embeddings located at the edge of the Poincaré ball 500B (rather than hyperbolic embeddings located at the origin of the Poincaré ball 500B) have a higher certainty of belonging to a category. For example, compared to an embedding of the category guitar near the hyperplane of the category guitar towards the origin, an embedding of the category guitar at the edge may have a higher certainty of belonging to the category guitar.
[0093] Figure 6 Shows a Poincaré ball with decision boundaries according to some embodiments of the present disclosure. Figure 6 In combination with elements from Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5A and Figure 5B are illustrated. Referring to Figure 6, showing the Poincaré sphere 602 and the masked spectrogram 604 of the audio mixed signal represented by the Poincaré sphere 602. Each region in the Poincaré sphere 602 is defined by a hyperbolic hyperplane and represents the dominant audio source category therein, such as guitar, drum, bass, male voice or female voice. The masked spectrogram 604 represents the masks for each individual subcategory (male voice, female voice, bass, drum or guitar), i.e., the positions from which the corresponding T-F units in each hyperbolic region are derived. Each hyperbolic hyperplane is marked by a decision boundary. These decision boundaries are learned during the training of the embedding classifier 412. The mixed embedding (scatterer) is represented by different symbols or patterns according to the maximum softmax layer output. The predicted T-F masks for each audio source are similarly plotted, resulting in the masked spectrogram 604.
[0094] Figure 7A Shows an interface for selecting a hyperbolic space according to some embodiments of the present disclosure. Figure 7A Combined with elements from Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5A , Figure 5B and Figure 6 are described. Referring to Figure 7A , an interface 700A is shown, which displays the hyperbolic embedding 710 projected into the hyperbolic space after classification on the 2D Poincaré sphere 712. The display may include (but is not limited to) a cathode ray tube (CRT), a color CRT monitor, a liquid crystal display (LCD), a light emitting diode (LED), a direct view storage tube (DVST), a plasma display, and a 3D display. In an embodiment, the interface 700A may be part of the selection interface 414, as Figure 4 shown. The interface 700A may include a selection tool 702 operated by the user. The shape of the selection may include (but is not limited to) an oval, an ellipse, a square, a rectangle, or any polygon. The user may change the size of the selection tool 702 as required. For example, if the user wants to select a larger area of the hyperbolic space, the user may provide an input to the interface 700A to increase the size of the selection tool 702. The selection tool 702 is resizable and movable, thus enabling the user to select a customized area from the hyperbolic space to select high-confidence T-F units only for a specific audio source. For example, the user may select only the embeddings in the area located on the left edge of the hyperbolic space. The selected embedding 704 may be the embedding selected by the user using the selection tool 702. Thus, the user is able to select to listen to the sources separated anywhere in the hyperbolic space. The hyperbolic embedding performs better than the Euclidean embedding at low embedding dimensions (e.g., 2D), so the concept of determinism is easy to interpret in this interface.
[0095] Figure 7BAn interface for displaying a masked spectrogram according to some embodiments of the present disclosure is shown. Figure 7B in combination with elements from Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5A , Figure 5B , Figure 6 and Figure 7A will be described. Referring to Figure 7B , a display 700B that displays a masked spectrogram 706 is shown. In some embodiments, the mask can be of two types: a binary mask, where all embeddings of T-F units within a selected portion of the hyperbolic space have a T-F mask value of 1, while all embeddings outside the selected portion of the hyperbolic space have a T-F mask value of 0. Another type of mask can include a soft mask, where continuous non-binary values in terms of weights represent the distances to the respective nearest classification hyperbolic planes and the distance to the origin of the hyperbolic space. After applying the mask, the T-F units corresponding to the selected embeddings are highlighted as 708 in the masked spectrogram 706. For example, if the user selects a hyperbolic region dominated by guitar category embeddings, the T-F units corresponding to the guitar category will be highlighted on the masked spectrogram 706. In the example, the interface 700A can provide the selected embeddings to an audio processing system 104, which uses the selected embeddings to generate a T-F mask 708. The audio processing system 104 can apply the generated T-F mask 708 on the spectrogram of the original audio mixed signal to generate a masked spectrogram 706. The audio processing system 104 applies a spectrogram inversion 420 to the masked spectrogram 708 to generate an output signal representing the audio signal corresponding to the selected embeddings 704 on the hyperbolic space.
[0096] Figure 8A An interface for selecting one or more hyperplanes in a hyperbolic space according to some embodiments of the present disclosure is shown. Figure 8A in combination with elements from Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5A , Figure 5B , Figure 6 , Figure 7A and Figure 7B will be described. Referring to Figure 8A, showing interface 800A, which displays the hyperbolic embedding projected into hyperbolic space on a 2D Poincaré sphere 802 after classification. The display may include (but is not limited to) a cathode ray tube (CRT), a color CRT monitor, a liquid crystal display (LCD), a light-emitting diode (LED), a direct view storage tube (DVST), a plasma display, and a 3D display. In this embodiment, the audio processing system 104 uses an embedding classifier 412 to classify each hyperbolic embedding in the Poincaré sphere 802 into one of different audio classes or audio sources. For example, the embedding classifier 412 classifies each hyperbolic embedding as belonging to one of the classes such as (but not limited to) guitar, bass, drum, male voice, or female voice. The learned hyperbolic hyperplane 808 is used to divide the 2D Poincaré sphere 802 for each class. In this embodiment, the user can provide an input to select one or more hyperplanes, as shown in the selection list 806. In another embodiment of the present disclosure, the user can move the position of the learned hyperbolic hyperplane to make different types of selections. In another embodiment of the present disclosure, the user can use the selection list 806 to select the desired hyperbolic audio source. The selection list may include (but is not limited to) a list of radio buttons, a check list, and a drop-down list. After selection, the selected embedding can be used to create a T-F mask to be applied to the audio mixing spectrogram 810. The separating hyperplane makes it easy to distinguish between audio sources, and the interface gives the user any number of audio sources as needed.
[0097] Figure 8B An interface for displaying a masked spectrogram according to some embodiments of the present disclosure is shown. Figure 8B Combined from Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5A , Figure 5B , Figure 6 , Figure 7A , Figure 7B and Figure 8A to illustrate. Referring to Figure 8B, a display 800B showing a masked spectrogram 810 is presented. After applying the mask, the T-F units corresponding to the selected embedding are highlighted as 812 in the masked spectrogram 810. For example, if the user selects the hyperbolic hyperplane corresponding to the guitar category embedding, the T-F units corresponding to the guitar category will be highlighted on the masked spectrogram 810. In the example, the interface 800A can provide the embeddings included in the selected hyperplane to the audio processing system 104, which uses the embeddings included in the selected hyperplane to generate the T-F mask 810. The audio processing system 104 can apply the generated T-F mask 812 on the spectrogram of the original audio mixed signal to generate the masked spectrogram 810. The audio processing system 104 applies a spectrogram inversion 420 to the masked spectrogram 810 to generate an output signal representing the audio signal corresponding to the embeddings included in the selected hyperplane of the hyperbolic space.
[0098] Figure 9 A flowchart illustrating the operation of an audio processing system for hyperbolic audio source separation according to some embodiments of the present disclosure is shown. Figure 9 Combined with elements from Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5A , Figure 5B , Figure 6 , Figure 7A , Figure 7B , Figure 8A and Figure 8B are described. Referring to Figure 9 , a flowchart 900 describing the operations performed by the audio processing system 104 is shown.
[0099] The audio processing system 104 obtains an audio mixed signal 902 from the environment. The environment can include (but is not limited to) industrial sites or factories, intelligent AI assistants, concerts, etc., where multiple overlapping sound events coexist. Thus, the audio mixed signal 902 includes sounds from various sources that overlap in the time domain and frequency domain. For example, machines operating in a large factory or plant generate such an audio mix while two people are talking nearby. The mix can contain audio from various machine sounds, background noise, and human voices. In an embodiment, the audio processing system can include a set of sensors to detect audio signals from different sources in the environment. This set of sensors can include (but is not limited to) microphones, such as dynamic microphones, carbon microphones, ribbon microphones, or piezoelectric microphones. This set of sensors provides the audio signals detected from different sources as the audio mixed signal 902 to the spectrogram extractor 904 of the audio processing system for further processing.
[0100] The spectrogram extractor 904 processes the audio mixed signal 902 to convert it into a spectrogram, which is a representation of the sound signal in a time-frequency format. The spectrogram is a visual way of representing the signal intensity or loudness of the different components of the audio mixed signal 902 over time at various frequencies present in the audio mixed signal 902. Details of the spectrogram are described with reference to Figure 5A the description of Figure 5A Figure 500A shows the spectrogram generated by the spectrogram extractor 904. The spectrogram 500A of the audio mixed signal 902 is sampled at regular intervals to generate time-frequency (T-F) units. In the example, the T-F units represent the values of the spectrogram at different time instances. In Figure 5A , the T-F unit 502 represents the value 502 of the spectrogram. The values of the T-F units are extracted from the spectrogram and provided to the separation network 906 for transforming the values of the extracted T-F units into hyperbolic embeddings and projecting the hyperbolic embeddings into hyperbolic space.
[0101] In an embodiment, the separation network 906 includes an embedding network 908 and a hyperbolic projection 910. Without departing from the scope of the present disclosure, in another embodiment, the separation network 906 may additionally include an embedding classifier 912. For example, the embedding classifier 912 can be used to automatically segment the hyperbolic space using hyperbolic hyperplanes for different sound classes or audio sources in the audio mixed signal 902, as further described with reference to the description of FIG. 8. In another example, the embedding classifier 912 is used during the training and validation of the separation network 906. Details of the training and validation of the embedding network 908 are described with reference to Figure 11 the description of
[0102] The embedding network 908 takes the T-F units from the spectrogram as the input to its layers and outputs a D-dimensional hyperbolic embedding corresponding to the T-F units. In other words, the embedding network 908 performs a transformation of values from the time-frequency domain to hyperbolic space.
[0103] The hyperbolic projection 910 projects the D-dimensional hyperbolic embedding into hyperbolic space.
[0104] Each hyperbolic embedding is visually indicated as a point at a specific position on the Poincaré sphere. The position of the hyperbolic embedding on the Poincaré sphere is calculated based on the entire spectrogram of the audio mixed signal and corresponds to a single T-F unit. The information associated with each hyperbolic embedding includes information about the original sound corresponding to the T-F unit. Each value on the Poincaré sphere includes some information about the original sound or the audio mixed signal 902.
[0105] The hyperbolic projection 910 is a formula that projects the embedding output of the embedding network 908 into a hyperbolic space or the Poincaré ball. Besides the benefits in hierarchical modeling, the hyperbolic layer can learn strong embeddings in low dimensions and represent uncertainty based on its distance to the origin in the hyperbolic space. The hyperbolic projection 910 is used to generate hyperbolic embeddings.
[0106] The audio processing system may include an embedding classifier 912, which is trained to classify hyperbolic embeddings in the hyperbolic space into audio source categories, such as drums, guitars, and trumpets. Details of the classification of hyperbolic embeddings in the hyperbolic space are described with reference to Figure 8A and Figure 8B illustrated. In the example, the embedding classifier 912 uses hyperbolic hyperplanes to partition the hyperbolic space according to the classification hierarchy.
[0107] The hierarchical representation is a method of representing data in a tree structure, where an increase in the depth of each level represents an increase in the features associated with the category. Considering the hyperbolic nature of the audio signal, a hierarchical classifier is used, which consists of multiple multinomial logistic regression (MLR) layers, with one MLR layer for each level in the audio sound category hierarchy (composed of parent categories and subcategories).
[0108] Masking is done by the audio processing system using the distances of hyperbolic embeddings in the hyperbolic space from the respective hyperbolic hyperplanes. For example, if the distance of a hyperbolic embedding from the hyperplane of the category guitar is the smallest, its masking will be corresponding to the audio source category guitar. Two sets of masks are created, one set of N masks, containing the T-F masks for each of the N subcategories, and a second set of M masks, containing the masks for each of the M parent categories. For example, the parent categories include labels such as speech and music, while the subcategories include labels such as male speech, female speech, drums, and guitars. In the example, T-F masks are created for each audio category based on the hyperbolic hyperplane.
[0109] Once the masks are created, the T-F masks 914 generated for each audio category by the audio processing system are applied to the spectrogram of the audio mixed signal 902 to generate a masked spectrogram of the audio mixed signal 902.
[0110] The audio processing system applies a spectrogram inversion 916 to the masked spectrogram to render the output signal 918 for each audio category based on the T-F masks created for the corresponding audio category. The output signal 918 can be an analog signal or a digital signal. In a particular embodiment, the audio processing system may send the output signal 918 to a speaker. In another embodiment, the audio processing system may include a speaker.
[0111] Figure 10 A flowchart showing the operation of an audio processing system for hyperbolic audio source separation using deterministic filtering according to some embodiments of the present disclosure is shown. Figure 10Combined from Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5A , Figure 5B , Figure 6 , Figure 7A , Figure 7B , Figure 8A and Figure 8B components. Referring to Figure 10 , a flowchart 1000 depicting the operations performed by the audio processing system 104 is shown.
[0112] The audio processing system 104 obtains an audio mixed signal 1002 from the environment. The environment can include (but is not limited to) an industrial site or factory, a smart AI assistant, a concert, etc., where multiple overlapping sound events coexist. Thus, the audio mixed signal 1002 includes sounds from various sources that overlap in the time domain and frequency domain. For example, machines operating in a large plant or factory generate such an audio mix while two people are talking nearby. The mix can contain audio from various machine sounds, background noise, and human voices. In an embodiment, the audio processing system can include a set of sensors to detect audio signals from different sources in the environment. This set of sensors can include (but is not limited to) microphones, such as dynamic microphones, carbon microphones, ribbon microphones, or piezoelectric microphones. This set of sensors provides the audio signals detected from different sources as the audio mixed signal 1002 to the spectrogram extractor 1004 of the audio processing system for further processing.
[0113] The spectrogram extractor 1004 processes the audio mixed signal 1002 to convert it into a spectrogram, which is a representation of the sound signal in a time-frequency format. The spectrogram 500A of the audio mixed signal 1002 is sampled at regular intervals to generate time-frequency (T-F) units. In an example, the T-F units represent the values of the spectrogram at different time instances. In Figure 5A , the T-F unit 502 represents the value 502 of the spectrogram. The values of the T-F units are extracted from the spectrogram and provided to the hyperbolic separation network 1006 for transforming the extracted values of the T-F units into hyperbolic embeddings and projecting the hyperbolic embeddings into hyperbolic space.
[0114] In an embodiment, the hyperbolic separation network 1006 includes an embedding network 1008 and a hyperbolic projection 1010. Without departing from the scope of the present disclosure, in another embodiment, the separation network 1006 can additionally include an embedding classifier 1012. For example, the embedding classifier 1012 can be used to automatically segment the hyperbolic space using hyperplanes for different sound classes or audio sources in the audio mixed signal 1002.
[0115] The embedding network 1008 takes the T-F units from the spectrogram as the input to its layers and outputs a D-dimensional hyperbolic embedding corresponding to the T-F units. In other words, the embedding network 1008 performs a transformation of values from the time-frequency domain to the hyperbolic space. The embedding network 1008 can be composed of multiple bidirectional long short-term memory (BLSTM) layers, which can be replaced by other types of deep network layers, such as convolutional layers, transformer layers, etc.
[0116] The hyperbolic projection 1010 projects the D-dimensional hyperbolic embedding into the hyperbolic space.
[0117] The hyperbolic projection 1010 is a formula that embeds the output of the embedding network 1008 into the hyperbolic space or the Poincaré ball.
[0118] The audio processing system may include an embedding classifier 1012, which is trained to classify the hyperbolic embeddings in the hyperbolic space into audio source categories, such as drums, guitars, and trumpets.
[0119] The audio processing system performs two-stage source separation using deterministic filtering 1014, such that the deterministic filtering 1014 is based on the deterministic information provided by the hyperbolic separation network 1006. In an example, the deterministic information is obtained from the information associated with the hyperbolic embeddings. The deterministic information is based on the number of hyperbolic embeddings near the edge of the Poincaré ball. In other words, the deterministic information indicates the confidence of the embeddings belonging to a specific audio category. In an example, if an embedding belongs to a specific category, it indicates that it is associated with a single audio source. The deterministic filtering 1014 is completed in two stages. According to some embodiments of the present disclosure, deterministic filters can be created using the T-F units near the edge of the hyperbolic space of the learned hyperbolic embeddings, because these units are highly likely to belong to a single audio source. The hyperbolic embeddings located near the edge of the Poincaré ball or the Poincaré disk correspond to audio categories that are more specific than the audio categories of the hyperbolic embeddings located closer to the origin of the Poincaré ball or the Poincaré disk in the classification hierarchy. The distance from the origin of the hyperbolic space to each hyperbolic embedding is used to derive a deterministic metric for processing, and a time-frequency mask is created based on the deterministic metric for processing. However, information will be missing in the separated sources thus obtained, because the low-confidence T-F units will not be included. Therefore, state-of-the-art generative models (such as those based on diffusion or generative adversarial networks) can be used to resynthesize the missing regions of the spectrogram while ensuring that there is no interference caused by the deterministic filtering provided by the hyperbolic separation model.
[0120] A generative adversarial network (or simply GAN) is a method of generative modeling using deep learning methods such as convolutional neural networks. Generative modeling is an unsupervised learning task in machine learning that involves automatically discovering and learning the patterns or regularities in the input data, such that the model can be used to generate or output new examples that may be drawn from the original data set.
[0121] The partial generation model 1018 can be used to resynthesize the missing regions of the spectrogram while ensuring that there is no interference caused by the deterministic filtering provided by the hyperbolic separation model.
[0122] After the deterministic filtering 1014 is completed and the missing regions of the spectrogram are predicted by the partial generation model 1018, the audio processing system uses the distances of the hyperbolic embeddings from the respective hyperbolic hyperplanes in the hyperbolic space for masking. For example, if the distance of the hyperbolic embedding from the hyperbolic hyperplane of the class guitar is the smallest, its mask corresponding to the audio source class guitar will have a larger value compared to the masks of other possible audio source classes. Two sets of masks are created, one set of N masks, including the T-F masks for each of the N subclasses, and a second set of M masks, including the masks for each of the M parent classes. For example, the parent classes include labels such as speech-like and music-like, while the subclasses include labels such as male speech-like, female speech-like, drums, and guitar.
[0123] Once the masks are created, the audio processing system applies the masks 1016 to the spectrogram of the audio mixed signal 1002 to generate a masked spectrogram of the audio mixed signal 1002.
[0124] The audio processing system applies a spectrogram inversion 1020 to the masked spectrogram to render an output signal 1022 that is part of the original audio mixed signal 1002 separated from the original audio mixed signal 1002. The output signal 1022 can be an analog signal or a digital signal. In a particular embodiment, the audio processing system can send the output signal 1022 to a speaker. In another embodiment, the audio processing system can include a speaker.
[0125] Figure 11 A flowchart depicting a hyperbolic separation network according to some embodiments of the present disclosure is shown. Figure 11 In combination with the elements from Figure 4 is illustrated. Referring to Figure 11 a flowchart 1100 is shown that depicts the operations performed by the hyperbolic separation network 406 as shown in Figure 4 .
[0126] The audio processing system obtains a mixed audio signal from the environment, which is converted by a spectrogram extractor into an audio amplitude spectrogram 1102. The audio amplitude spectrogram 1102 can be used as an input to an embedding network 1104. The embedding network 1104 can include a bidirectional long short-term memory (BLSTM) layer 1106 and a linear layer 1108. In other embodiments of the present disclosure, the embedding network 1104 can include multiple layers, which can include but are not limited to BLSTM layers, convolutional layers, transformer layers, and other deep network layers.
[0127] To memorize longer input data sequences, models based on Long Short-Term Memory (LSTM) contain "gates". However, BLSTM achieves additional training by traversing the input data twice (i.e., from left to right and from right to left). This improves the training of the neural network because the additional memory capacity addresses the temporal dependencies present in audio signals.
[0128] The training of the hyperbolic separation network can be done using a simple hierarchical source separation dataset that contains a mixture of two "parent" categories (music and speech) and five "leaf" categories (bass, drums, guitar, male speech, and female speech). As building blocks, the clean subsets of LibriSpeech and Slakh2100 are used, which is a dataset of 2100 synthetic music mixtures, each mixture containing bass, drums, and guitar tracks in addition to various other instruments. A dataset consisting of 1947 mixtures is constructed, each mixture being 60 seconds in length, totaling approximately 32 hours. For the training, validation, and test sets, the data split is 70%, 20%, and 10% respectively. The male speech source targets consist of male speech utterances randomly selected (without replacement) and concatenated sequentially (non-overlapping) until a 60-second track length is reached. Any signal with the last concatenated utterance exceeding this length is discarded. The same process is used for female speech. For Slakh2100, only the first 60 seconds of the bass, drums, and guitar tracks are selected for each track. Any track with a duration less than 60 seconds is discarded. All sources are summed without applying additional gain to form the total mixture as well as the speech and music sub-mixtures. This results in challenging input SDR values, with standard deviation values ranging from 2 - 13 dB depending on the category.
[0129] The model consists of four BLSTM layers with 600 units in each direction, followed by a dense layer to obtain the L-dimensional Euclidean embedding for each T-F unit. Dropout of 0.3 is applied to the output of each BLSTM layer except the last. For the hyperbolic model (c > 0), an exponential projection layer is placed after the dense layer to map the Euclidean embedding onto the Poincaré ball with curvature -c. Then an MLR layer (Euclidean or hyperbolic) with a softmax activation function is used to obtain the mask for each source category. Following the hierarchical softmax approach, there are two MLR layers: one MLR layer with K = 2 for the parent (speech / music) sources and a second MLR layer with K = 5 for the leaf categories. The mixing phase is used for resynthesis and comparison with multiple training targets. The ADAM optimizer is used for Euclidean parameters and the Riemannian ADAM implementation of geoopt is used for hyperbolic parameters. All models are trained with an initial learning rate of 10 - 3 within 300 epochs using 3.2-second chunks and a batch size of 10, halving the learning rate if the validation loss does not improve within 10 epochs. An STFT size of 32 ms is used, with 50% overlap and a square root Hann window.
[0130] The time-frequency (T-F) units are taken from the audio amplitude spectrogram 1102 and used as the input to the BLSTM layer 1106 of the embedding network 1104. The BLSTM layer 1106 memorizes the audio information corresponding to the T-F units and uses deep learning to compute an L-dimensional Euclidean embedding. The Euclidean space lacks the ability to infer hierarchical representations from data without distortion. Therefore, a hyperbolic space is used instead of the Euclidean space. Using the hyperbolic space, strong hierarchical embeddings are obtained at low embedding dimensions compared to Euclidean embeddings, which can save computational resources during training and inference.
[0131] Then, the linear layer 1108 converts the BLSTM output into D-dimensional embeddings for each T-F unit. These D-dimensional embeddings are used as the input to the hyperbolic projection 1110.
[0132] The embedding network 1104 is connected to the hyperbolic network layers: the hyperbolic projection 1110 and the embedding classifier 1112 to learn and classify the hyperbolic embeddings.
[0133] The hyperbolic projection 1110 takes the D-dimensional embeddings and projects them as hyperbolic embeddings onto the hyperbolic space in the Poincaré ball representation, as Figure 5B mentioned in.
[0134] The embedding classifier 1112 classifies the embeddings into audio source categories using hierarchical classification considering the hyperbolic nature of the audio signal. A hierarchical representation refers to a tree-like representation of the data, where the number of features associated with the data increases as the depth of the tree increases.
[0135] Polynomial logistic regression is a classification method that generalizes logistic regression to multi-class problems, i.e., problems with more than two possible discrete outcomes. It is a model for predicting the probabilities of different possible outcomes of a categorical distribution dependent variable given a set of independent variables (which can be real-valued, binary-valued, categorical-valued, etc.).
[0136] For classification, the embedding classifier 1112 includes multiple polynomial logistic regression (MLR) layers, one MLR layer for each level in the audio sound category hierarchy: the parent hyperbolic MLR layer 1114 and the child hyperbolic MLR layer 1116. The sub-layers can include categories such as male voice, female voice, guitar, and drum, and the parent categories can include categories such as voice and music.
[0137] The MLR layers in the embedding classifier 1112 classify the hyperbolic embeddings using the softmax function, which uses the distances of the hyperbolic embeddings from the respective hyperbolic hyperplanes in the Poincaré ball (as Figure 5B mentioned in the description of) to determine which category (parent category or sub-category) the T-F unit corresponding to the hyperbolic embedding belongs to.
[0138] After classification by the MLR layer, the embedding classifier 1112 creates two sets of masks: a set of N masks, including the T-F masks for each of the N subcategories, and a second set of M masks, including the masks for each of the M parent categories. The source-specific masks for the individual T-F units in the input spectrogram are obtained by using the probability that the embedding is dominated by the audio source through a clustering method. For example, if the probability that multiple embeddings belong to the guitar audio source is high, they can be clustered together to form the T-F mask for the class guitar.
[0139] Figure 12 A flowchart showing an exemplary remote diagnostic application in accordance with some embodiments of the present disclosure. Figure 12 Combined from Figure 1 、 Figure 2 、 Figure 3 、 Figure 4 、 Figure 5A 、 Figure 5B 、 Figure 6 、 Figure 7A 、 Figure 7B 、 Figure 8A 、 Figure 9 and Figure 11 to illustrate. Referring to Figure 12 , a flowchart 1200 depicting the working model of an exemplary remote diagnostic application is shown. In the example, the flowchart 1200 describes the anomaly detection applied by the audio processing system 104 to an audio signal.
[0140] For remote audio diagnostics, the machine 1202 can be considered. The machine 1202 can include (but is not limited to) a manufacturing unit, a vehicle, a tractor, and a large industrial unit. The audio signal from the machine 1202 can be recorded using the recorder 1204. The recorder 1204 can include an audio recording system having a microphone and a database storing the audio mixed signal 1206. The audio mixed signal 1206 can include the audio signals generated by the machine 1202 as well as other audio sources.
[0141] Then, the audio mixed signal 1206 is used as the input to the hyperbolic separation network 1208 of the audio processing system. The signal can be sent directly or through a communication network as mentioned in Figure 1 . The spectrogram extractor generates the spectrogram of the audio mixed signal 1206.
[0142] The hyperbolic separation network 1208 takes the individual T-F units from the spectrogram, embeds each T-F as a hyperbolic embedding using the underlying embedding neural network, and projects each hyperbolic embedding onto the hyperbolic space, as in Figure 11As mentioned. As an example, hyperbolic embeddings located near or at the origin of the hyperbolic space are very likely to represent sounds that are uncertain or essentially overlapping. In other words, time-frequency regions containing multiple overlapping sources are embedded towards the center of the hyperbolic space (i.e., the most uncertain region). Therefore, the distance value of the hyperbolic embedding from the origin of the hyperbolic space defines the probability of uncertainty or certainty that the hyperbolic embedding belongs to a specific audio category or audio source.
[0143] The hyperbolic interface 1210 can be provided to the user 1212 to use the hyperbolic projection of the hyperbolic embeddings generated by the hyperbolic separation network 1208 and select the embeddings corresponding to the T-F units preferably selected by the user, as mentioned in FIGS. 7 and 8. Once the selection is made, a T-F mask is applied to the original audio mixed spectrogram to generate a masked spectrogram. The masked spectrogram is inverted to obtain the separated source signal or the sound signal corresponding to the selection. The user 1212 provides the separated audio signal as an output to diagnose the noise 1214.
[0144] In another embodiment, the user can use the hyperbolic interface to detect anomalies in the audio mixed signal 1206. For example, the user 1212 can visually observe whether most of the hyperbolic embeddings are clustered near the origin of the hyperbolic space, and then can infer that the audio mixed signal 1206 includes mostly overlapping sounds. This can be regarded as an abnormal sound. In another example, if the user 1212 visually observes that most of the hyperbolic embeddings are clustered near the edge of the hyperbolic space, it can be inferred that the audio mixed signal 1206 includes mostly non-overlapping sounds. This can indicate that the audio mixed signal 1206 includes sounds from different sources, which can be listened to separately and thus are essentially non-overlapping.
[0145] In another embodiment, based on the number of hyperbolic embeddings located within a threshold distance from the origin of the Poincaré sphere or Poincaré disk being greater than or equal to a threshold, the hyperbolic interface 1210 can determine the input audio mix as an abnormal sound. The value of the threshold distance of the hyperbolic embedding from the origin, which is considered to belong to a specific audio category, is predetermined or can be determined during the training of the embedding classifier 1112, as shown in the description with reference to Figure 11 Moreover, the threshold for the number of hyperbolic embeddings for which it is inferred that the input audio mix is not an abnormal sound is predetermined or can be determined during the training of the embedding classifier 1112, as shown in the description with reference to Figure 11 the description.
[0146] In another embodiment, the inference regarding anomaly detection in the audio mixed signal 1206 can be automatically performed without user interference. For example, the audio processing system 104 can detect the positions of the respective hyperbolic embeddings in the hyperbolic space displayed using the hyperbolic interface 1210. Using the positions of the respective hyperbolic embeddings in the hyperbolic space, the audio processing system 104 can determine the distribution pattern of the hyperbolic embeddings in the hyperbolic space. Using the distribution pattern of the hyperbolic embeddings in the hyperbolic space, if most of the hyperbolic embeddings are clustered near the origin of the hyperbolic space, the audio processing system 104 can infer that the audio mixed signal 1206 includes mostly overlapping sounds. In another example, if most of the hyperbolic embeddings are clustered near the edge of the hyperbolic space, the audio processing system 104 infers that the audio mixed signal 1206 includes mostly non-overlapping sounds.
[0147] In another embodiment, the audio processing system 104 detects an anomaly in the machine 1202 based on the number of hyperbolic embeddings located within a threshold distance from the origin of the Poincaré sphere or Poincaré disk being greater than or equal to a threshold. The value of the threshold distance of a hyperbolic embedding, which is considered to belong to a specific audio class, is predetermined or can be determined during the training of the embedding classifier 1112, as shown in the description with reference to Figure 11 as described.
[0148] Accordingly, the audio processing system 104 is capable of detecting anomalies based on the distance of the embeddings from the origin of the hyperbolic space. In addition, the hyperbolic interface 1210 provides a visual representation of the sounds, which makes it easier to represent the hierarchy of the respective components present in the audio mixed signal 1206. In an embodiment, the processor 202 of the audio processing system 104 can control the machine 1202 based on the detected anomalies in the machine 1202.
[0149] Among many advantageous aspects, hyperbolic audio source separation provides a deterministic natural way of estimating classification values for the visualization of sounds and a more compact representation.
[0150] Figure 13 Shows a control knob for hyperbolic audio source separation according to some embodiments of the present disclosure. Figure 13 In conjunction with Figure 1 、 Figure 2 、 Figure 3 、 Figure 4 、 Figure 5A 、 Figure 5B 、 Figure 6 、 Figure 7A 、 Figure 7B 、 Figure 8A 、 Figure 9 、 Figure 11 and Figure 12 of the elements. With reference to Figure 13, a control knob 1302 is shown. In some embodiments, the control knob 1302 can also serve as a zoom control knob. According to an exemplary embodiment, the position of the zoom control knob 1302 is converted into a mixing weight for audio zoom. The zoom control knob 1302 is used to select a specific audio source or sound category from an audio mixing signal including audio from different sources. In this embodiment, the user can provide an input to select one or more audio sources separated by the embedded classifier 1112. As an example of user input, the user can move the position of the knob 1302 towards any one of points A - E. The position of the control knob 1302 can indicate the selection of an embedding included in a hyperplane corresponding to a specific audio category or audio source. In an example, when the user moves the control knob 1302 to position A, a T - F mask of the embedding included in the hyperplane corresponding to the guitar is created, a masked spectrogram is created based on the created T - F mask, and a separated sound signal corresponding to the category guitar is output after spectrogram inversion. In another example, when the user moves the control knob 1302 to position B, a T - F mask of the embedding included in the hyperplane corresponding to the bass is created, a masked spectrogram is created based on the created T - F mask, and a separated sound signal corresponding to the category bass is output after spectrogram inversion. In another example, when the user moves the control knob 1302 to position C, a T - F mask of the embedding included in the hyperplane corresponding to the drums is created, a masked spectrogram is created based on the created T - F mask, and a separated sound signal corresponding to the category drums is output after spectrogram inversion. In another example, when the user moves the control knob 1302 to position D, a T - F mask of the embedding included in the hyperplane corresponding to the female voice is created, a masked spectrogram is created based on the created T - F mask, and a separated sound signal corresponding to the category female voice is output after spectrogram inversion. In another example, when the user moves the control knob 1302 to position E, a T - F mask of the embedding included in the hyperplane corresponding to the male voice is created, a masked spectrogram is created based on the created T - F mask, and a separated sound signal corresponding to the category male voice is output after spectrogram inversion.
[0151] The following description provides only exemplary embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. On the contrary, the following description of the exemplary embodiments will provide those skilled in the art with a feasible description for implementing one or more exemplary embodiments. Various changes can be envisioned in the functions and arrangements of the elements without departing from the spirit and scope of the subject matter disclosed in the appended claims.
[0152] Specific details are given in the following description to provide a thorough understanding of the embodiments. However, one of ordinary skill in the art will understand that the embodiments may be practiced without these specific details. For example, systems, processes, and other elements in the disclosed subject matter may be shown as components in block diagrams to avoid obscuring the embodiments with unnecessary detail. In other instances, well-known processes, structures, and techniques may be shown without unnecessary detail to avoid obscuring the embodiments. Additionally, like reference numerals and designations in the various drawings indicate like elements.
[0153] Additionally, various embodiments may be described as a process, which is depicted as a flowchart, data flow diagram, structure diagram, or block diagram. Although a flowchart may describe operations as a sequential process, many operations may be performed in parallel or concurrently. Additionally, the order of operations may be rearranged. A process may terminate when its operations are completed, but may have additional steps not discussed or included in the figure. Furthermore, not all operations in any particular described process may appear in all embodiments. A process may correspond to a method, function, program, subroutine, subprogram, etc. When a process corresponds to a function, the termination of the function may correspond to the function returning to the calling function or the main function.
[0154] Furthermore, embodiments of the disclosed subject matter may be implemented, at least in part, manually or automatically. Execution or at least assistance with manual or automatic implementation may be accomplished by using a machine, hardware, software, firmware, middleware, microcode, a hardware description language, or any combination thereof. When implemented in software, firmware, middleware, or microcode, the program code or code segments for performing the required tasks may be stored in a machine-readable medium. A processor may execute the required tasks.
[0155] The various methods or processes outlined herein may be encoded as software that may be executed on one or more processors employing any of a variety of operating systems or platforms. Additionally, such software may be written using any of a number of suitable programming languages and / or programming or scripting tools, and also may be compiled into executable machine language code or intermediate code that may be executed on a frame or virtual machine. Generally, in various embodiments, the functionality of program modules may be combined or distributed as desired. The above-described embodiments of the present disclosure may be implemented in any of numerous ways. For example, the embodiments may be implemented using hardware, software, or a combination thereof. When implemented in software, the software code may be executed on any suitable processor or collection of processors, whether provided in a single computer or distributed among multiple computers. These processors may be implemented as integrated circuits having one or more processors therein. However, the processors may be implemented using circuitry in any suitable format.
[0156] Embodiments of the present disclosure may be embodied as a method, examples of which are provided. The acts performed as part of the method may be ordered in any suitable way. Accordingly, embodiments may be constructed to perform the acts in an order different than shown, which may include performing some acts simultaneously, although shown as sequential acts in the illustrative embodiments.
[0157] Although the present disclosure has been described with reference to particular preferred embodiments, it will be understood that various other adaptations and modifications may be made within the spirit and scope of the present disclosure. Accordingly, the aspects of the appended claims cover all such variations and modifications that fall within the true spirit and scope of the present disclosure.
Claims
1. An audio processing system, the audio processing system comprising: An input interface configured to receive an input audio mix and transform it into a time-frequency representation defined by values of time-frequency units; A processor configured to map the values of the time-frequency units into the hyperbolic space by executing an embedding neural network trained to associate each time-frequency unit with a high-dimensional embedding and projecting each high-dimensional embedding into the hyperbolic space; And An output interface configured to accept a selection of at least a portion of the hyperbolic space and render selected hyperbolic embeddings that fall within the selected portion of the hyperbolic space.
2. The audio processing system according to claim 1, wherein, For rendering the selected hyperbolic embeddings, the output interface is further configured to transform the selected hyperbolic embeddings into separate output audio signals and send the separate output audio signals to a memory, and The memory stores the separate output signals.
3. The audio processing system according to claim 2, wherein, The output interface is further configured to send the separate output signals to a speaker.
4. The audio processing system according to claim 1, wherein, The output interface is further configured to transform the selected hyperbolic embeddings into separate output audio signals by creating a time-frequency mask based on the selected hyperbolic embeddings and applying the time-frequency mask to the time-frequency representation of the input audio mix.
5. The audio processing system according to claim 4, wherein, The output interface is further configured to create the time-frequency mask based on a softmax operation.
6. The audio processing system according to claim 1, wherein, The hyperbolic space is a Poincaré ball or a Poincaré disk classified according to hyperbolic geometry, and the hyperbolic geometry carries the concept of a classification hierarchy of audio sources based on the position of the hyperbolic embeddings relative to the origin of the Poincaré ball or the Poincaré disk.
7. The audio processing system according to claim 6, wherein, The distance from the origin of the hyperbolic space to each hyperbolic embedding is used to derive a measure of certainty of the processing, and the creation of the time-frequency mask is based on the measure of certainty of the processing.
8. The audio processing system according to claim 6, wherein, The processor is further configured to determine a measure of certainty that the hyperbolic embedding belongs to only a single specific audio class in the classification hierarchy based on the distance of the hyperbolic embedding from the origin of the Poincaré ball or the Poincaré disk.
9. The audio processing system according to claim 6, wherein, The processor is further configured to determine the input audio mix as an abnormal sound based on the number of hyperbolic embeddings located within a threshold distance from the origin of the Poincaré ball or the Poincaré disk being greater than or equal to a threshold.
10. The audio processing system according to claim 6, wherein, The input audio mix is generated by components of a machine, The processor is further configured to detect an abnormality in the machine based on the number of hyperbolic embeddings located within a threshold distance from the origin of the Poincaré ball or the Poincaré disk being greater than or equal to a threshold.
11. The audio processing system according to claim 10, wherein, The processor is further configured to control the machine based on the detected abnormality in the machine.
12. The audio processing system according to claim 1, wherein, The hyperbolic space is classified according to a classification hierarchy of audio sources, and wherein the embedding neural network is trained end-to-end using a classifier trained to classify the hyperbolic embeddings according to the classification hierarchy.
13. The audio processing system according to claim 1, wherein, The output interface is operatively connected to a display device configured to display a visual representation of the hyperbolic embedding mapped to different locations of the hyperbolic space to allow selection of the portion of the hyperbolic space.
14. The audio processing system according to claim 1, wherein, the training data set for training the embedding neural network includes audio mixtures of at least two parent classes and at least five subclasses, the at least two parent classes include music and speech, and the at least five subclasses include bass, drums, guitar, male speech, and female speech.
15. The audio processing system according to claim 1, wherein, the processor is further configured to receive user input, and the size and shape of the selected portion of the hyperbolic space are based on the received user input.
16. The audio processing system according to claim 12, wherein, the classifier divides the hyperbolic space using hyperbolic hyperplanes according to the classification hierarchy, and the output interface is further configured to create a T-F mask for each audio class based on the hyperbolic hyperplane.
17. The audio processing system according to claim 16, wherein, The output interface is further configured to generate an output signal for each audio class based on the T-F mask created for the corresponding audio class.
18. The audio processing system according to claim 1, wherein, The processor is further configured to: accept a selection of weights of the energy of the selected hyperbolic embedding; and render the selected hyperbolic embedding based on the weights of the energy of the selected hyperbolic embedding.
19. An audio processing method, the audio processing method comprising the steps of: receiving an input audio mixture and transforming it into a time-frequency representation defined by values of time-frequency units; mapping the values of the time-frequency units into the hyperbolic space by performing an embedding neural network trained to associate each time-frequency unit with a high-dimensional embedding and projecting each high-dimensional embedding into the hyperbolic space; and accepting a selection of at least a portion of the hyperbolic space and rendering the selected hyperbolic embeddings that fall within the selected portion of the hyperbolic space.
20. A non-transitory computer-readable storage medium having embodied thereon a program executable by a processor for performing a method, the method comprising the steps of: receiving an input audio mixture and transforming it into a time-frequency representation defined by values of time-frequency units; mapping the values of the time-frequency units into the hyperbolic space by performing an embedding neural network trained to associate each time-frequency unit with a high-dimensional embedding and projecting each high-dimensional embedding into the hyperbolic space; and accepting a selection of at least a portion of the hyperbolic space and rendering the selected hyperbolic embeddings that fall within the selected portion of the hyperbolic space.
Citation Information
Cited By
Drosophila behavior recognition method based on hyperbolic spatial vision and language alignment
CN120822083A