Knowledge-based audio scene graph
By constructing an audio scene graph using audio segmentation and commonsense knowledge graphs, the problem of non-temporal correlation in audio analysis in existing technologies is solved, enabling more accurate understanding of audio event relationships and support for complex tasks.
Patent Information
- Application Number
- CN202480036551.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-06-10
- Filing Date
- 2024-06-11
- Publication Date
- 2026-01-30
AI Technical Summary
Existing technologies struggle to effectively utilize correlations outside of temporal order in audio signals, leading to limitations in audio analysis.
By using an audio segmentation model to segment audio clips into audio events and constructing an audio scene graph using a commonsense knowledge graph, edge weights are generated based on the temporal order and similarity measure of audio events to enrich the audio representation.
It realizes the generation of knowledge-based audio scene graphs, which can more accurately analyze and understand the relationships between audio events and support more complex downstream tasks and query responses.
Smart Images

Figure CN121444167A_ABST
Abstract
Description
[0001] Cross-references to related applications
[0002] This disclosure claims the benefit of priority to jointly owned Provisional Patent Application No. 63 / 508,199, filed June 14, 2023, and Non-Provisional Patent Application No. 18 / 738,243, filed June 10, 2024, the contents of which are expressly incorporated herein by reference in their entirety. Technical Field
[0003] This disclosure relates in general to knowledge-based audio scene graphs. Background Technology
[0004] Technological advancements have led to smaller and more powerful computing devices. For example, a wide variety of portable personal computing devices exist today, including small, lightweight, and easily portable cordless phones (such as mobile and smartphones, tablets, and laptops). These devices can transmit voice and data packets over wireless networks. Furthermore, many of these devices incorporate additional functionality, such as digital still cameras, digital camcorders, digital recorders, and audio file players. Moreover, such devices can process executable instructions, including software applications such as web browser applications that can be used to access the internet. Accordingly, these devices can include significant computing power.
[0005] Such computing devices often incorporate the functionality of receiving audio signals from one or more microphones. For example, an audio signal could represent user speech captured by a microphone, external sounds captured by a microphone, or a combination thereof. Typically, audio analysis determines the temporal order between sounds within an audio clip. However, sounds can be correlated in ways other than temporal order. Knowledge of such relationships can be useful in various types of audio analysis. Summary of the Invention
[0006] According to one embodiment of this disclosure, an apparatus includes a memory configured to store knowledge data. The apparatus also includes one or more processors coupled to the memory and configured to identify audio segments of audio data corresponding to audio events. The one or more processors are further configured to assign tags to audio segments. The tags for specific audio segments describe the corresponding audio events. The one or more processors are further configured to determine relationships between audio events based on the knowledge data. The one or more processors are also configured to construct an audio scene graph based on the temporal order of the audio events. The one or more processors are further configured to assign edge weights to the audio scene graph based on a similarity metric between audio events and relationships between audio events.
[0007] According to another specific embodiment of this disclosure, a method includes receiving audio data at a first device. The method further includes identifying audio segments of the audio data corresponding to audio events at the first device. The method also includes assigning labels to the audio segments at the first device. The labels of specific audio segments describe the corresponding audio events. The method further includes determining relationships between audio events based on knowledge data. The method also includes constructing an audio scene graph at the first device based on the temporal order of the audio events. The method further includes assigning edge weights to the audio scene graph at the first device based on a similarity measure between audio events and relationships between the audio events. The method also includes providing a representation of the audio scene graph to a second device.
[0008] According to another embodiment of this disclosure, a non-transitory computer-readable medium storage instruction, when executed by one or more processors, causes the one or more processors to identify audio segments of audio data corresponding to audio events. The instruction also causes the one or more processors to assign tags to audio segments. The tags of specific audio segments describe the corresponding audio events. The instruction further causes the one or more processors to determine relationships between audio events based on knowledge data. The instruction further causes the one or more processors to construct an audio scene graph based on the temporal order of the audio events. The instruction further causes the one or more processors to assign edge weights to the audio scene graph based on similarity measures between audio events and relationships between audio events.
[0009] According to another specific embodiment of this disclosure, an apparatus includes components for identifying audio segments of audio data corresponding to audio events. The apparatus also includes components for assigning labels to audio segments. The labels of specific audio segments describe the corresponding audio events. The apparatus further includes components for determining relationships between audio events based on knowledge data. The apparatus also includes components for constructing an audio scene graph based on the temporal order of the audio events. The apparatus further includes components for assigning edge weights to the audio scene graph based on a similarity measure between audio events and the relationships between the audio events.
[0010] Other aspects, advantages, and features of this disclosure will become apparent upon reading the entire application, which comprises the following sections: description of the drawings, detailed description, and claims. Attached Figure Description
[0011] FIG. 1 This is a block diagram illustrating specific exemplary aspects of a system operable to generate knowledge-based audio scene graphs, based on some examples of this disclosure.
[0012] FIG. 2 Based on some examples of this disclosure and FIG. 1 An illustrative diagram of the operation associated with the system's audio splitter.
[0013] FIG. 3 Based on some examples of this disclosure and FIG. 1 An illustrative diagram of the operations associated with the system's audio scene graph constructor.
[0014] FIG. 4 Based on some examples of this disclosure and FIG. 1 The system's event representation is an illustrative diagram of the operations associated with the generator.
[0015] FIG. 5A Based on some examples of this disclosure and FIG. 1 The diagram illustrates the exemplary aspects of the operations associated with the knowledge data analyzer of the system.
[0016] FIG. 5B Based on some examples of this disclosure and FIG. 1 An illustrative diagram of the operations associated with the system's audio scene graph updater.
[0017] FIG. 6A Based on some examples of this disclosure and FIG. 1 This is another illustrative aspect of the operations associated with the system's knowledge data analyzer.
[0018] FIG. 6B Based on some examples of this disclosure and FIG. 1 This is another illustrative aspect of the operation associated with the system's audio scene graph updater.
[0019] FIG. 7 Based on some examples of this disclosure and FIG. 1 The diagram illustrates the illustrative aspects of the operations associated with the system's encoder.
[0020] FIG. 8 Based on some examples of this disclosure and FIG. 1 A diagram illustrating an exemplary aspect of the operations associated with one or more graph transformer layers of a system.
[0021] FIG. 9 This is an illustration of an exemplary aspect of a system operable for updating a knowledge-based audio scene graph, based on some examples of this disclosure.
[0022] FIG. 10 Based on some examples of this disclosure FIG. 1 The system FIG. 9 An illustration of the graphical user interface (GUI) generated by the system or both.
[0023] FIG. 11This is another illustrative aspect of a system operable as an update of a knowledge-based audio scene graph, based on some examples of this disclosure.
[0024] FIG. 12 This is an illustration of an exemplary aspect of a system operable, based on some examples of this disclosure, for generating query results using a knowledge-based audio scene graph.
[0025] FIG. 13 This is a block diagram illustrating an exemplary aspect of a system operable to generate knowledge-based audio scene graphs, based on some examples of this disclosure.
[0026] FIG. 14 Examples of integrated circuits operable to generate knowledge-based audio scene graphs according to some examples of this disclosure are illustrated.
[0027] FIG. 15 This is an illustration of a mobile device operable to generate a knowledge-based audio scene graph, based on some examples of this disclosure.
[0028] FIG. 16 This is an illustration of a head-mounted device operable to generate a knowledge-based audio scene graph, based on some examples of this disclosure.
[0029] FIG. 17 This is an illustration of a wearable electronic device operable to generate a knowledge-based audio scene graph, based on some examples of this disclosure.
[0030] FIG. 18 This is a diagram illustrating, based on some examples of the present disclosure, operable as a voice-controlled speaker system for generating knowledge-based audio scene graphs.
[0031] FIG. 19 This is an illustration of a camera operable to generate a knowledge-based audio scene graph, based on some examples of this disclosure.
[0032] FIG. 20 This is an illustration of a head-mounted device (such as a virtual reality, mixed reality, or augmented reality head-mounted device) operable to generate a knowledge-based audio scene graph according to some examples of this disclosure.
[0033] FIG. 21 This is an illustration of a first example of a vehicle operable to generate a knowledge-based audio scene graph, based on some examples of this disclosure.
[0034] FIG. 22 This is a diagram illustrating a second example of a vehicle operable to generate a knowledge-based audio scene graph, based on some examples of this disclosure.
[0035] FIG. 23 Based on some examples of this disclosure, it can be derived byFIG. 1 This is a diagram illustrating a specific implementation of a method for generating knowledge-based audio scene graphs executed by the system.
[0036] FIG. 24 This is a block diagram illustrating specific exemplary examples of devices operable to generate knowledge-based audio scene graphs, based on some examples of this disclosure. Detailed Implementation
[0037] Audio analysis typically determines the temporal order of sounds within an audio clip. However, sounds can be associated in ways other than temporal order. For example, the sound of a door opening might be associated with the sound of a baby crying. To illustrate, if the door opening sound precedes the baby crying, the opening of the door might have startled the baby. Alternatively, if the door opening sound follows the baby crying, someone might have opened the door to enter the room where the baby was crying or to take the baby out of the room. Knowledge of such relationships can be useful in various types of audio analysis. For instance, an audio scene representation indicating that a baby crying sound might be associated with a preceding door opening sound could be used to answer the query "Why is the baby crying?" with the answer "The door is open." As another example, an audio scene representation indicating that a baby crying sound might be associated with a subsequent door opening sound could be used to answer the query "Why is the door open?" with the answer "The baby is crying."
[0038] Audio applications typically take audio clips as input and encode their representations using a convolutional neural network (CNN) architecture to derive a holistic encoded audio representation. This holistic encoded audio representation encodes all audio events of the audio clip into a single vector in a latent space. According to some examples described in this paper, the audio clips are encoded using an injected commonsense knowledge graph to enrich the encoded audio representation with information describing relationships between audio events captured within the audio clip. As a first step, the audio clip is segmented into audio events using an audio segmentation model, and audio segments are labeled using an audio tagger. Audio labels are fed as input to the commonsense knowledge graph to retrieve relationships between audio events. This relationship information enables the construction of an audio scene graph. According to some examples described in this paper, an audio graph transformer considers the multiplicity and directionality of the edges used to encode the audio representation. The audio scene graph is encoded using an encoder based on the audio graph transformer. Model performance can be tested on downstream tasks. In some specific implementations, the model (e.g., audio segmentation model, knowledge graph, audio graph transformer, or a combination thereof) can be updated based on the performance of the downstream tasks (e.g., a loss function associated with the downstream tasks).
[0039] A system and method for generating knowledge-based audio scene graphs are disclosed. For example, the audio scene graph generator identifies and labels audio segments corresponding to audio events. For illustration, a first audio event is detected in a first audio segment, a second audio event is detected in a second audio segment, and a third audio event is detected in a third audio segment. The first, second, and third audio segments are respectively assigned a first label associated with the first audio event, a second label associated with the second audio event, and a third label associated with the third audio event.
[0040] The audio scene graph generator constructs an audio scene graph based on the temporal order of audio events. For example, the audio scene graph includes a first node, a second node, and a third node corresponding to a first audio event, a second audio event, and a third audio event, respectively. During the initial audio scene graph construction phase, the audio scene graph generator adds edges between nodes that are temporally immediately following each other. For example, the audio scene graph generator adds a first edge connecting the first node to the second node based on the determination that the second audio event is temporally immediately following the first audio event. Similarly, the audio scene graph generator adds a second edge connecting the second node to the third node based on the determination that the third audio event is temporally immediately following the second audio event. The audio scene graph generator avoids adding edges between the first and third nodes based on the determination that the third audio event is not temporally immediately following the first audio event.
[0041] The audio scene graph generator generates event representations of audio events. It generates a first event representation for a first audio event, a second event representation for a second audio event, and a third event representation for a third audio event. In one example, the event representations of audio events are based on the tags and audio segments associated with the audio events.
[0042] During the second audio scene graph construction phase, the audio scene graph generator updates the audio scene graph based on knowledge data indicating relationships between audio events. In some examples, the knowledge data is based on human knowledge of relationships between various types of events. For illustration, the knowledge data indicates relationships between first and second audio events based on human input acquired during some prior knowledge data generation process, indicating that events like the first audio event can be related to events like the second audio event. In some examples, the knowledge data is generated by processing a large number of documents scraped from the internet.
[0043] In one example, the audio scene graph generator assigns edge weights to existing edges between nodes based on knowledge data. For illustration, the knowledge data indicates that a first audio event (e.g., the sound of a door opening) is related to a second audio event (e.g., the sound of a baby crying). In response to determining that the first and second audio events are related, the audio scene graph generator determines the edge weights based at least in part on a similarity metric associated with the representations of the first and second events. The audio scene graph generator assigns edge weights to the edges between the first and second nodes in the audio scene graph. In a particular aspect, the edge weights indicate the strength (e.g., probability) of the relationship between the first and second audio events. In one example, an edge weight closer to 1 indicates a strong correlation between the first and second audio events, while an edge weight closer to 0 indicates a weak correlation.
[0044] In one example, the audio scene graph generator adds edges between nodes based on knowledge data. For instance, in response to determining that the knowledge data indicates a first audio event and a third audio event are related and that the audio scene graph does not include any edges between the first and third nodes, the audio scene graph generator adds an edge between the first and third nodes. In response to determining that the first and third audio events are related, the audio scene graph generator determines the edge weights based at least in part on a similarity metric associated with the representations of the first and third events. The audio scene graph generator assigns edge weights to the edges between the first and third nodes in the audio scene graph. Thus, assigning edge weights adds knowledge-based information to the audio scene graph. The audio scene graph can be used to perform various downstream tasks, such as answering queries.
[0045] Specific aspects of this disclosure are described below with reference to the accompanying drawings. In this description, common features are designated by common reference numerals. As used herein, various terms are used only for the purpose of describing particular embodiments and are not intended to limit the scope of the embodiments. For example, the singular forms “a,” “an,” and “the” are intended to also include the plural forms unless the context clearly indicates otherwise. Furthermore, some features described herein are singular in some embodiments and plural in others. For example, FIG. 13 It describes a system that includes one or more processors ( FIG. 13 The device 1302 (of which the “processor” 1390) indicates that in some embodiments, device 1302 includes a single processor 1390, while in other embodiments, device 1302 includes multiple processors 1390. For ease of reference herein, such features are generally introduced as “one or more” features and are subsequently referred to in the singular or optional plural form (as indicated by “(multiple)”), unless the aspect described relates to multiples of features.
[0046] In some figures, multiple instances of a particular type of feature are used. Although these features are physically and / or logically different, the same reference numerals are used for each feature, and these different instances are distinguished by adding letters to the reference numerals. Reference numerals are used without distinguishing letters when a feature is referenced herein as a group or a type of feature (e.g., when a specific feature among these features is not referenced). However, reference numerals are used with distinguishing letters when a specific feature among multiple features of the same type is mentioned herein. For example, see reference... FIG. 2 The figure illustrates several audio segments, which are associated with reference numerals 112A, 112B, 112C, 112D, and 112E. When referring to a specific audio segment (such as audio segment 112A), the distinguishing letter "A" is used. However, when referring to any individual audio segment among these audio segments, or when referring to these audio segments as a group, reference numeral 112 is used, without the distinguishing letter.
[0047] As used herein, the term “comprise” may be used interchangeably with “include”. Additionally, the term “wherein” may be used interchangeably with “where”. As used herein, “exemplary” indicates an example, specific implementation, and / or aspect, and should not be construed as restrictive or indicating a preference or preferred implementation. As used herein, ordinal terms used to modify elements (such as structures, components, operations, etc.) (e.g., “first,” “second,” “third,” etc.) do not themselves indicate any priority or order of that element relative to another element, but merely distinguish that element from another element with the same name (but using ordinal terms). As used herein, the term “set” refers to one or more specific elements within a set of specific elements, while the term “multiple” refers to multiple (e.g., two or more) specific elements.
[0048] As used herein, “coupling” can include “communicationally coupled,” “electrically coupled,” or “physically coupled,” and may also (or alternatively) include any combination thereof. Two devices (or components) may be coupled directly or indirectly (e.g., communicationally coupled, electrically coupled, or physically coupled) via one or more other devices, components, wires, buses, networks (e.g., wired networks, wireless networks, or combinations thereof). As an illustrative, non-limiting example, two electrically coupled devices (or components) may be included in the same device or in different devices and may be connected via electronics, one or more connectors, or inductive coupling. In some specific implementations, two communicationally coupled (e.g., electrically connected) devices (or components) may transmit and receive signals (e.g., digital or analog signals) directly or indirectly via one or more wires, buses, networks, etc. As used herein, “direct coupling” can include two devices coupled (e.g., communicationally coupled, electrically coupled, or physically coupled) without intermediate components.
[0049] In this disclosure, terms such as “determine,” “calculate,” “estimate,” “shift,” and “adjust” can be used to describe how one or more operations are performed. It should be noted that such terms should not be construed as restrictive, and similar operations can be performed using other techniques. Additionally, as mentioned herein, “generate,” “calculate,” “estimate,” “use,” “select,” “access,” and “determine” can be used interchangeably. For example, “generating,” “calculating,” “estimate,” or “determining” a parameter (or signal) can refer to actively generating, estimating, calculating, or determining that parameter (or signal), or it can refer to using, selecting, or accessing a parameter (or signal) that has already been generated (e.g., by another component or device).
[0050] refer to FIG. 1 A specific exemplary aspect of a system configured to generate knowledge-based audio scene graphs is disclosed and is generally designated as 100. System 100 includes an audio scene graph generator 140 configured to process audio data 110 based on knowledge data 122 to generate an audio scene graph 162. According to some specific embodiments, the audio scene graph generator 140 is coupled to a graph encoder 120 configured to encode the audio scene graph 162 to generate an encoded graph 172.
[0051] The audio scene graph generator 140 includes an audio scene segmenter 102 configured to determine audio segments 112 of audio data 110 corresponding to audio events. In a particular implementation, the audio scene segmenter 102 includes an audio segmentation model (e.g., a machine learning model). The audio scene segmenter 102 is configured to assign event labels 114 to the audio segments 112 describing the corresponding audio events (e.g., including an audio tagger configured to perform this operation). The audio scene segmenter 102 is coupled to an audio scene graph updater 118 via an audio scene graph builder 104, an event representation generator 106, and a knowledge data analyzer 108. The audio scene graph builder 104 is configured to generate an audio scene graph 162 based on the temporal order of audio events detected by the audio scene segmenter 102. The event representation generator 106 is configured to generate event representations 146 of the detected audio events based on the corresponding audio segments 112 and the corresponding event labels 114. The knowledge data analyzer 108 is configured to generate event pair relation data 152, indicating any relationships between audio event pairs, based on the knowledge data 122. The audio scene graph updater 118 is configured to assign edge weights to edges between nodes of the audio scene graph 162 based on the event representation 146 and the event pair relation data 152.
[0052] In some specific implementations, the audio scene graph generator 140 corresponds to one of various types of devices, or is included in one of various types of devices. In an exemplary example, the audio scene graph generator 140 is integrated into a head-mounted device, such as a reference device. FIG. 16 Further described. In other examples, the audio scene graph generator 140 is integrated into at least one of the following: as referenced FIG. 15 The described mobile phone or tablet computer device, as shown in the reference FIG. 17 The wearable electronic devices described, as in the reference FIG. 18 The described voice control loudspeaker system, as in the reference FIG. 19 The described camera device or as in the reference FIG. 20 The described virtual reality, mixed reality, or augmented reality head-mounted device. In another illustrative example, an audio scene graph generator 140 is integrated into a vehicle, such as a reference... FIG. 21 and FIG. 22 Further description.
[0053] During operation, the audio scene graph generator 140 acquires audio data 110. In some examples, the audio data 110 corresponds to an audio stream received from a network device. In some examples, the audio data 110 corresponds to an audio signal received from one or more microphones. In some examples, the audio data 110 is retrieved from a storage device. In some examples, the audio data 110 is obtained from an audio generation application. In some examples, the audio scene graph generator 140 processes the audio data 110 while receiving portions of it (e.g., real-time processing). In some examples, the audio scene graph generator 140 accesses all portions of the audio data 110 before initiating processing of the audio data 110 (e.g., offline processing).
[0054] Audio scene segmenter 102 identifies the audio segment 112 of audio data 110 corresponding to an audio event and assigns the event label 114 to the audio segment 112, as shown in the reference. FIG. 2 Further description: The event label 114 of a specific audio segment 112 describes the corresponding audio event. For example, the audio scene segmenter 102 identifies the audio segment 112 of the audio data 110 as corresponding to an audio event. The audio scene segmenter 102 assigns an event label 114 to the audio segment 112 that describes (e.g., identifies) the audio event.
[0055] In certain implementations, knowledge data 122 indicates relationships between pairs of event tags 114 to indicate the existence of relationships between corresponding pairs of audio events. In some implementations, audio scene segmenter 102 is configured to identify audio segments 112 corresponding to audio events associated with a set of event tags 114 included in knowledge data 122. In response to identifying audio segment 112 as corresponding to an audio event associated with a specific event tag 114 in the set of event tags 114, audio scene segmenter 102 assigns the specific event tag 114 to audio segment 112.
[0056] Audio scene segmenter 102 generates data indicating the audio segment time sequence 164 of audio segment 112, as shown in the reference. FIG. 2 Further described. For example, the audio segment timing sequence 164 indicates that the first audio segment 112 corresponds to the first playback time associated with the first audio frame to the second playback time associated with the second audio frame, the second audio segment 112 corresponds to the third playback time associated with the third audio frame to the fourth playback time associated with the fourth audio frame, and so on.
[0057] The audio scene graph constructor 104 performs the initial audio scene graph construction phase. For example, the audio scene graph constructor 104 constructs an audio scene graph 162 based on the audio segment time sequence 164, as shown in the reference. FIG. 3Further described. For illustration, the audio scene graph constructor 104 adds nodes to the audio scene graph 162 corresponding to the audio events, and adds edges between pairs of nodes that are temporally immediately following each other, as indicated by the audio segment time sequence 164. The audio scene graph constructor 104 provides the audio scene graph 162 to the audio scene graph updater 118.
[0058] The audio scene graph constructor 104 generates an event representation 146 for audio events based on audio segments 112 and event labels 114, as shown in the reference. FIG. 4 Further described. In one example, audio segment 112 is identified as associated with an audio event described by event label 114. Audio scene graph constructor 104 generates an event representation 146 of the audio event based on audio segment 112 and event label 114. Audio scene graph constructor 104 provides event representation 146 to audio scene graph updater 118.
[0059] Knowledge data analyzer 108 determines the relationships between audio events based on knowledge data 122, as shown in the reference. FIG. 5A and FIG. 6A Further described. In one example, the knowledge data analyzer 108 generates event pair relation data 152 based on knowledge data 122, indicating relationships between audio events corresponding to event labels 114. For illustration, the knowledge data analyzer 108 determines for each specific event label 114 whether the knowledge data 122 indicates one or more relationships between that specific event label 114 and the remaining event labels in the event labels 114. In response to determining that the knowledge data 122 indicates one or more relationships between a first event label 114 (corresponding to a first audio event) and a second event label 114 (corresponding to a second audio event), the knowledge data analyzer 108 generates event pair relation data 152 indicating one or more relationships between the first event label 114 (e.g., the first audio event) and the second event label 114 (e.g., the second audio event). The knowledge data analyzer 108 provides the event pair relation data 152 to the audio scene graph updater 118.
[0060] The audio scene graph updater 118 performs a second audio scene graph construction phase. For example, the audio scene graph updater 118 obtains data generated during the initial audio scene graph construction phase and uses that data to perform the second audio scene graph construction phase. In some implementations, the initial audio scene graph construction phase may be performed at a first device, which provides the data to a second device, and the second device performs the second audio scene graph construction phase.
[0061] During the second audio scene graph construction phase, the audio scene graph updater 118 can selectively add one or more edges to the audio scene graph 162 based on the relationships indicated by the event pair relation data 152, as shown in the reference. FIG. 5B and FIG. 6B Further described. For example, in response to determining that event pair relationship data 152 indicates at least one relationship between a first audio event and a second audio event, and that audio scene graph 162 does not include any edge between a first node corresponding to the first audio event and a second node corresponding to the second audio event, audio scene graph updater 118 adds an edge between the first node and the second node.
[0062] During the second audio scene graph construction phase, the audio scene graph updater 118 also assigns edge weights to the audio scene graph 162 based on a similarity metric associated with the event representation 146 and the relationships indicated by the event pair relation data 152, as referenced. FIG. 5B and FIG. 6B Further described. In the first example, event-to-relationship data 152 indicates a single relationship between a first audio event (e.g., first event label 114) and a second audio event (e.g., second event label 114), as described in reference. FIG. 5A As described. In the first example, the audio scene graph updater 118 determines the first edge weight based on an event similarity metric associated with the first event representation 146 of the first audio event and the second event representation 146 of the second audio event, as referenced. FIG. 5B Further described. The audio scene graph updater 118 assigns the first edge to the edge between the first node and the second node of the audio scene graph 162.
[0063] In the second example, event-to-relationship data 152 indicates multiple relationships between a first audio event and a second audio event, and each of these relationships has an associated relationship label, as shown in the reference. FIG. 6A Further described. In the second example, the audio scene graph updater 118 determines edge weights based on an event similarity metric and a relation similarity metric associated with a relation (e.g., relation label). The audio scene graph updater 118 assigns edge weights to the edge between the first node and the second node. Each of these edges corresponds to a corresponding relation among these relations. Assigning edge weights to the audio scene graph 162 adds information about the relation strengths, which are determined based on the relations indicated by knowledge data 122.
[0064] According to some specific implementations, the audio scene graph 162 is provided to the graph encoder 120. The graph encoder 120 encodes the audio scene graph 162 to generate an encoded graph 172, as shown in the reference. FIG. 7 to FIG. 8Further description. In a particular aspect, the encoded graph 172 preserves the directional information of the edges of the audio scene graph 162.
[0065] Depending on the specific implementation, the graph updater is configured to update the audio scene graph 162 based on various inputs. In one example, the graph updater is based on user feedback (such as reference...) FIG. 9 to FIG. 10 Further description), analysis of visual data (as referenced) FIG. 11 (Further description), the performance of one or more downstream tasks, or a combination thereof, to update the audio scene (Figure 162).
[0066] Depending on the specific implementation, audio scene diagram 162 or encoding diagram 172 is used to perform one or more downstream tasks. For example, audio scene diagram 162 or encoding diagram 172 can be used to generate a response to a query, as shown in the reference... FIG. 12 Further described. As another example, audio scene graph 162 (or coded graph 172) can be used to initiate one or more actions. For illustration, in response to determining an edge weight greater than a threshold indicating a relationship between a detected baby crying sound and a detected door opening sound (e.g., someone entering the room to change a diaper) in audio scene graph 162 (or coded graph 172), a baby care application can activate a baby wipe warmer. In a particular implementation, the graph updater updates audio scene graph 162 based on the performance of one or more downstream tasks (e.g., a loss function associated with one or more downstream tasks).
[0067] The technical advantages of the audio scene graph generator 140 include the generation of a knowledge-based audio scene graph 162. The audio scene graph 162 can be used to perform various types of analysis on the audio scene represented by the audio scene graph 162. For example, the audio scene graph 162 can be used to generate responses to queries, initiate one or more actions, or combinations thereof.
[0068] Although the audio scene segmenter 102, audio scene graph constructor 104, event representation generator 106, knowledge data analyzer 108, audio scene graph updater 118, and graph encoder 120 are described as separate components, in some examples, two or more of the audio scene segmenter 102, audio scene graph constructor 104, event representation generator 106, knowledge data analyzer 108, audio scene graph updater 118, and graph encoder 120 can be combined into a single component.
[0069] In some implementations, the audio scene graph generator 140 and the graph encoder 120 can be integrated into a single device. In other implementations, the audio scene graph generator 140 can be integrated into a first device, and the graph encoder 120 can be integrated into a second device.
[0070] FIG. 2 Figure 200 is an illustrative aspect of the operation associated with the audio scene splitter 102 according to some examples of this disclosure. The audio scene splitter 102 acquires audio data 110, as shown in the reference... FIG. 1 As described.
[0071] Audio scene segmenter 102 performs audio event detection on audio data 110 to identify audio segments 112 corresponding to audio events and assigns corresponding labels to audio segments 112. In example 202, audio scene segmenter 102 identifies audio segment 112A (e.g., white noise sound) extending from a first playback time (e.g., 0 seconds) to a second playback time (e.g., 2 seconds) as associated with a first audio event (e.g., white noise). Audio scene segmenter 102 assigns an event label 114A (e.g., "white noise") describing the first audio event to audio segment 112A. Similarly, the audio scene segmenter 102 assigns event labels 114B (e.g., "doorbell"), 114C (e.g., "music"), 114D (e.g., "baby crying"), and 114E (e.g., "door opening") to audio segments 112B (e.g., the sound of a doorbell), 112C (e.g., the sound of music), 112D (e.g., the sound of a baby crying), and 112E (e.g., the sound of a door opening), respectively. It should be understood that audio segment 112, comprising five audio segments, is provided as an illustrative example; in other examples, audio segment 112 may include fewer or more than five audio segments.
[0072] Audio scene segmenter 102 generates data indicating the audio segment timing sequence 164 of audio segment 112. For example, audio segment timing sequence 164 indicates that audio segment 112A (e.g., white noise sound) is identified as continuing from a first playback time (e.g., 0 seconds) to a second playback time (e.g., 2 seconds). Similarly, audio segment timing sequence 164 indicates that audio segment 112B (e.g., a doorbell sound) is identified as continuing from a second playback time (e.g., 2 seconds) to a third playback time (e.g., 5 seconds).
[0073] In some examples, gaps may exist between consecutively identified audio segments 112. For illustration, audio segment timing sequence 164 indicates that audio segment 112C (e.g., music) is identified as continuing from a fourth playback time (e.g., 7 seconds) to a fifth playback time (e.g., 11 seconds). The gap between the third playback time (e.g., 5 seconds) and the fourth playback time (e.g., 7 seconds) may correspond to a silence or unidentifiable sound between audio segment 112B (e.g., a doorbell sound) and audio segment 112C (e.g., music sound).
[0074] In some examples, audio segment 112 may overlap with one or more other audio segments 112. For example, audio segment timing sequence 164 indicates that audio segment 112D is identified as continuing from a sixth playback time (e.g., 9 seconds) to a seventh playback time (e.g., 13 seconds). The sixth playback time is between the fourth and fifth playback times, and the seventh playback time is after the fourth playback time, indicating that audio segment 112D (e.g., the sound of a baby crying) at least partially overlaps with audio segment 112C (e.g., the sound of music).
[0075] FIG. 3 Figure 300 is an illustrative aspect of the operation associated with the audio scene graph constructor 104 according to some examples of this disclosure. The audio scene graph constructor 104 is configured to construct an audio scene graph 162 based on the audio segment time sequence 164 of the audio segment 112 and the event label 114 assigned to the audio segment 112.
[0076] The audio scene graph constructor 104 adds node 322 to the audio scene graph 162. Node 322 corresponds to the audio event associated with event label 114. For example, the audio scene graph constructor 104 adds node 322A to the audio scene graph 162 corresponding to the audio event associated with event label 114A. Similarly, the audio scene graph constructor 104 adds nodes 322B, 322C, 322D, and 322E to the audio scene graph 162, respectively, corresponding to event labels 114B, 114C, 114D, and 114E.
[0077] Node 322A is associated with audio segment 112A of the assigned event label 114A. Similarly, nodes 322B, 322C, 322D, and 322E are associated with audio segments 112B, 112C, 112D, and 112E, respectively.
[0078] The audio scene graph constructor 104 adds edges 324 between pairs of nodes 322 that are temporally immediately following each other in the audio segment temporal sequence 164 and associated with event labels 114. For example, in response to determining that node 322A is associated with audio segment 112A that extends from a first playback time (e.g., 0 seconds) to a second playback time (e.g., 2 seconds), the audio scene graph constructor 104 identifies a temporally immediately following audio segment that overlaps with audio segment 112A or has a start playback time that is closest to the second playback time among audio segments whose start playback time is greater than or equal to the second playback time. For illustration, the audio scene graph constructor 104 identifies audio segment 112B that extends from the second playback time (e.g., 2 seconds) to a third playback time (e.g., 5 seconds) as an audio segment that temporally follows audio segment 112A. In response to determining that audio segment 112B immediately follows audio segment 112A in time, the audio scene graph constructor 104 adds an edge 324A from node 322A associated with audio segment 112A to node 322B associated with audio segment 112B.
[0079] Similarly, in response to determining that audio segment 112C is associated with the start playback time (e.g., 7 seconds) of the audio segment start playback time that is closest to the third playback time (e.g., 5 seconds) among audio segment start playback times greater than or equal to the third playback time, the audio scene graph constructor 104 identifies audio segment 112C as the audio segment that immediately follows audio segment 112B in time. In response to determining that audio segment 112C immediately follows audio segment 112B in time, the audio scene graph constructor 104 adds an edge 324B from node 322B associated with audio segment 112B to node 322C associated with audio segment 112C.
[0080] In response to determining that audio segment 112D at least partially overlaps with audio segment 112C, the audio scene graph constructor 104 determines that audio segment 112D immediately follows audio segment 112C in time. In response to determining that audio segment 112D at least partially overlaps with audio segment 112C, the audio scene graph constructor 104 adds an edge 324C from node 322C (associated with audio segment 112C) to node 322D (associated with audio segment 112D), and adds an edge 324D from node 322D to node 322C.
[0081] The audio scene graph constructor 104 continues to add edges 324 to the audio scene graph 162 in this manner until an end node is reached. For example, in response to determining that audio segment 112E is associated with the start playback time (e.g., 14 seconds) of the audio segment start playback time that is closest to the end playback time (e.g., 13 seconds) of audio segment 112D whose start playback time is greater than or equal to the end playback time, the audio scene graph constructor 104 determines that audio segment 112E is temporally immediately following audio segment 112D. In response to determining that audio segment 112E is temporally immediately following audio segment 112D, the audio scene graph constructor 104 adds an edge 324E from node 322D associated with audio segment 112D to node 322E associated with audio segment 112E.
[0082] The audio scene graph constructor 104 determines that the construction of the audio scene graph 162 is complete based on determining that node 322E corresponds to the last audio segment 112 in the audio segment time sequence 164. In a particular aspect, in response to determining that audio segment 112E has the maximum start playback time in audio segment 112, the audio scene graph constructor 104 determines that audio segment 112E corresponds to the last audio segment 112.
[0083] FIG. 4 Figure 400 is an illustrative aspect of the operation associated with event representation generator 106 according to some examples of this disclosure. Event representation generator 106 is configured to generate an event representation 146 of an audio event detected in an audio segment 112 of an assigned event tag 114. Event representation generator 106 includes a combiner 426 coupled to event audio representation generator 422 and event tag representation generator 424.
[0084] Event audio representation generator 422 is configured to process audio segment 112 to generate an audio embedding 432 representing audio segment 112. Audio embedding 432 may correspond to a lower-dimensional representation of audio segment 112. In one example, audio embedding 432 includes an audio feature vector comprising feature values of audio features. Audio features may include spectral information, such as frequency content varying over time, and statistical properties, such as Mel-frequency cepstral coefficients (MFCCs). In some implementations, event audio representation generator 422 includes a machine learning model (e.g., a deep neural network) trained on labeled audio data to generate audio embeddings. According to some implementations, event audio representation generator 422 preprocesses audio segment 112 before generating audio embedding 432. Preprocessing may include resampling, normalization, filtering, or a combination thereof.
[0085] Event label representation generator 424 is configured to process event label 114 to generate a text embedding 434 representing event label 114. Text embedding 434 may correspond to a numerical representation that captures the semantic meaning and contextual information of event label 114. In one example, text embedding 434 includes a text feature vector comprising feature values of text features. In some implementations, event label representation generator 424 includes a machine learning model (e.g., a deep neural network) trained on labeled text to generate text embeddings. According to some implementations, event label representation generator 424 preprocesses event label 114 before generating text embedding 434. Preprocessing may include converting text to lowercase, removing punctuation, handling special characters, tokenizing event label 114 into individual words or sub-word units, or a combination thereof.
[0086] Combiner 426 is configured to combine (e.g., splice) audio embedding 432 and text embedding 434 to generate an event representation 146 of an audio event detected in audio segment 112 and described by event label 114. In one example, event representation generator 106 thus generates a first event representation 146 corresponding to audio segment 112A and event label 114A, a second event representation 146 corresponding to audio segment 112B and event label 114B, a third event representation 146 corresponding to audio segment 112C and event label 114C, a fourth event representation 146 corresponding to audio segment 112D and event label 114D, a fifth event representation 146 corresponding to audio segment 112E and event label 114E, and so on.
[0087] FIG. 5A Figure 500 is an illustrative aspect of the operation associated with knowledge data analyzer 108 according to some examples of this disclosure. Knowledge data analyzer 108 can access knowledge data 122. In a particular implementation, knowledge data 122 is based on human knowledge of relationships between various types of events. In some examples, knowledge data analyzer 108 obtains knowledge data 122 from storage devices, network devices, websites, databases, users, or combinations thereof.
[0088] Knowledge data 122 indicates relationships between audio events. In one example, knowledge data 122 includes a knowledge graph comprising nodes 522 corresponding to audio events and edges 524 corresponding to relationships. For example, knowledge data 122 includes node 522A representing a first audio event (e.g., the sound of a baby crying) described by event label 114D and node 522B representing a second audio event (e.g., the sound of a door opening) described by event label 114E. Knowledge data 122 includes an edge 524A between nodes 522A and 522B, indicating that the first audio event is related to the second audio event. It should be understood that knowledge data 122 indicating a relationship between two audio events is provided as an illustrative example; in other examples, knowledge data 122 may indicate relationships between additional audio events. It should be understood that knowledge data 122, which includes a graph representation of relationships between audio events, is provided as an illustrative example; in other examples, other types of representations may be used to indicate relationships between audio events.
[0089] In a specific implementation, in response to receiving event tag 114, knowledge data analyzer 108 generates event pairs for each specific event tag and for each other event tag. In one example, the count of event pairs is given by: (n*(n-1)) / 2, where n = the count of event tag 114. For example, knowledge data analyzer 108 generates 10 event pairs for 5 events (e.g., (5*4) / 2=10).
[0090] The knowledge data analyzer 108 determines for each event pair whether the knowledge data 122 indicates that the corresponding event is relevant. For example, the knowledge data analyzer 108 generates an event pair that includes a first audio event described by event label 114D and a second audio event described by event label 114E.
[0091] The knowledge data analyzer 108 determines that node 522A is associated with the first audio event (described by event label 114D) based on a comparison between event label 114D and the node event label associated with node 522A. Knowledge data 122, including nodes associated with the same event label 114D generated by the audio scene segmenter 102, is provided as an illustrative example. In this example, the knowledge data analyzer 108 determines that node 522A is associated with the first audio event based on an exact match between event label 114D and the node event label associated with node 522A.
[0092] In some examples, knowledge data 122 may include node event labels that are different from event labels 114 generated by audio scene segmenter 102. In these examples, knowledge data analyzer 108 determines that node 522A is associated with the first audio event based on determining that a similarity metric between event label 114D and the node event label associated with node 522A meets a similarity criterion. For example, knowledge data analyzer 108 determines that node 522A is associated with the first audio event based on determining that event label 114D has the maximum similarity to that node event label compared to other node event labels and that the similarity between event label 114D and that node event label is greater than a similarity threshold. In a particular implementation, knowledge data analyzer 108 determines the similarity between event label 114 and a specific node event label based on a comparison of the text embedding 434 of event label 114 with the text embedding of a specific node event label (e.g., a node event label embedding). For example, the similarity between event label 114 and a specific node event label may be based on the Euclidean distance between the text embedding 434 and the node event text embedding in the embedding space. In another example, the similarity between event label 114 and a specific node event label can be based on the cosine similarity between text embedding 434 and the node event text embedding.
[0093] Similarly, the knowledge data analyzer 108 determines that node 522B is associated with the second audio event (described by event label 114E) based on a comparison of event label 114E and the node event label associated with node 522B. In response to determining that knowledge data 122 indicates that node 522A is connected to node 522B via edge 524A, the knowledge data analyzer 108 determines that the first audio event is associated with the second audio event and generates event pair relation data 152 indicating that the first audio event described by event label 114D is associated with the second audio event described by event label 114E. Alternatively, in response to determining that there is no direct edge connecting node 522A and node 522B, the knowledge data analyzer 108 determines that the first audio event is not associated with the second audio event and generates event pair relation data 152 indicating that the first audio event described by event label 114D is not associated with the second audio event described by event label 114E. Similarly, the knowledge data analyzer 108 generates event pair relation data 152 that indicates whether the remaining event pairs (e.g., the remaining 9 event pairs) are related.
[0094] It should be understood that, as an illustrative example, knowledge data 122 is described as indicating a relationship without directional information; in another example, knowledge data 122 may indicate directional information of the relationship. For illustration, knowledge data 122 may include a directed edge 24 from node 522B to node 522A to indicate that a correspondence applies when the audio event indicated by event label 114E (e.g., opening a door) is earlier than the audio event indicated by event label 114D (e.g., a baby crying). In this example, in response to determining that event label 114D is associated with an earlier audio segment (e.g., audio segment 112D) than the audio segment 112E associated with event label 114E, and that knowledge data 122 includes an edge 524 from node 522A to node 522B, knowledge data analyzer 108 generates event pair relationship data 152 indicating the relationship between event label pairs 114D-E. Alternatively, in this example, in response to determining that event label 114D (e.g., baby crying) is associated with an earlier audio segment (e.g., audio segment 112D) than the audio segment 112E associated with event label 114E (e.g., door opening) and that knowledge data 122 does not include any edge 524 from node 522A to node 522B, knowledge data analyzer 108 generates event pair relation data 152 indicating that event label pair 114D-E is not related, regardless of whether the edge in the other direction from node 522B to node 522A is included in knowledge data 122.
[0095] FIG. 5B This is an illustrative diagram 550 illustrating an aspect of the operation associated with the audio scene graph updater 118 according to some examples of this disclosure. The audio scene graph updater 118 is configured to assign edge weights to edges 324 of the audio scene graph 162 based on event pair relation data 152 and event representations 146. The audio scene graph updater 118 includes a total edge weight (OW) generator 510, which is configured to generate total edge weights 528 based on a similarity metric for a pair of event representations 146.
[0096] In response to receiving event pair relationship data 152 indicating that the event pairs are related, the audio scene graph updater 118 generates total edge weights 528 corresponding to the event pairs. For example, in response to determining that the event pair relationship data 152 indicates that a first audio event described by event label 114D is related to a second audio event described by event label 114E, the audio scene graph updater 118 uses the total edge weight generator 510 to determine the total edge weights 528 associated with the first audio event and the second audio event.
[0097] The audio scene graph updater 118 obtains the event representation 146D of the first audio event and the event representation 146E of the second audio event. Event representation 146D is based on audio segment 112D and event label 114D, and event representation 146E is based on audio segment 112E and event label 114E, as shown in the reference. FIG. 4 As described.
[0098] The total edge weight generator 510 determines the total edge weight 528 (e.g., 0.7) corresponding to the similarity metric associated with event representation 146D and event representation 146E. In one example, the similarity metric is based on the cosine similarity between event representation 146D and event representation 146E.
[0099] In response to determining that the event pair relationship data 152 indicates that the knowledge data 122 indicates a single relationship between the first audio event (described by event label 114D) and the second audio event (described by event label 114E), the audio scene graph updater 118 assigns the total edge weight 528 as an edge weight 526A (e.g., 0.7) to the edge 324E between node 322D (associated with event label 114D) and node 322E (associated with event label 114E).
[0100] In a specific implementation, the audio scene graph updater 118 indicates a single relationship between the first audio event and the second audio event based on the deterministic knowledge data 122 and the audio scene graph 162 includes a single edge (e.g., a one-way edge) between node 322D and node 322E to assign the total edge weight 528 as edge weight 526A to edge 324E.
[0101] If the audio scene graph 162 includes multiple edges (e.g., bidirectional edges), the audio scene graph updater 118 can split the total edge weight among the multiple edges. For example, the audio scene graph updater 118 determines a total edge weight (e.g., 1.2) corresponding to a first audio event (e.g., the sound of music) associated with node 322C and a second audio event (e.g., the sound of a baby crying) associated with node 322D. In response to determining that knowledge data 122 indicates a single relationship between the first audio event (e.g., the sound of music) and the second audio event (e.g., the sound of a baby crying), and that the audio scene graph 162 includes two edges (e.g., edge 324C and edge 324D) between nodes 322C and 322D, the audio scene graph updater 118 splits the total edge weight (e.g., 1.2) into edge weight 526B (e.g., 0.6) and edge weight 526C (e.g., 0.6). The audio scene graph updater 118 assigns edge weight 526B to edge 324C and edge weight 526C to edge 324D.
[0102] In a specific implementation, in response to the determination of event pair relationship data 152 indicating a relationship between a pair of audio events that are not directly connected in the audio scene graph 162, the audio scene graph updater 118 adds an edge between the pair of audio events and assigns an edge weight to the edge. For example, in response to the determination of event pair relationship data 152 indicating that a first audio event (e.g., the sound of a doorbell) is related to a second audio event (e.g., the sound of a door opening), and the audio scene graph 162 indicating that there is no edge between node 322B associated with the first audio event and node 322C associated with the second audio event, the audio scene graph updater 118 adds an edge 324F between node 322B and node 322E. The direction of edge 324F is based on the temporal order of the first audio event relative to the second audio event. For example, the audio scene graph updater 118 adds an edge 324F from node 322B to node 322E based on the determination of audio segment temporal order 164 indicating that the first audio event (e.g., the sound of a doorbell) precedes the second audio event (e.g., the sound of a door opening). The total edge weight generator 510 determines the total edge weight (e.g., 0.9) corresponding to the first audio event (e.g., the sound of a doorbell) and the second audio event (e.g., the sound of a door opening) and assigns the total edge weight to edge 324F.
[0103] Therefore, the audio scene graph updater 118 assigns edge weights to edges corresponding to audio event pairs based on the similarity between the event representations of the audio event pairs. Audio event pairs with similar audio embeddings and similar text embeddings are more likely to be related.
[0104] In a specific example where knowledge data 122 includes directional information about relations, if the temporal order of audio events associated with the direction of edge 324E matches the temporal order of the relations of audio events indicated by knowledge data 122, then the audio scene graph updater 118 assigns the total edge weight 528 as edge weight 526A.
[0105] FIG. 6A The diagram 600 is an illustrative aspect of the operation associated with the knowledge data analyzer 108, based on some examples of this disclosure.
[0106] Knowledge data 122 indicates multiple relationships between at least some audio events. In one example, knowledge data 122 includes node 522A representing a first audio event (e.g., the sound of a baby crying) described by event label 114D and node 522B representing a second audio event (e.g., the sound of a door opening) described by event label 114E. Knowledge data 122 includes edge 524A between nodes 522A and 522B, which indicates a first relationship between the first and second audio events. Knowledge data 122 also includes edge 524B between nodes 522A and 522B, which indicates a second relationship between the first and second audio events. Edge 524A is associated with relationship label 624A describing the first relationship (e.g., being woken up by…). Edge 524B is associated with relationship label 624B describing the second relationship (e.g., sudden noise).
[0107] In response to determining that knowledge data 122 indicates that node 522A is connected to node 522B via multiple edges (e.g., edge 524A and edge 524B), knowledge data analyzer 108 determines that a first audio event is related to a second audio event and generates event pair relation data 152 indicating multiple relationships between the first audio event described by event label 114D and the second audio event described by event label 114E. For example, event pair relation data 152 indicates that the audio event pair corresponding to event labels 114D and 114E has multiple relationships indicated by relation labels 624A and 624B.
[0108] It should be understood that, as an illustrative example, knowledge data 122 is described as indicating relations without directional information; in another example, knowledge data 122 may indicate directional information of the relations. For illustration, knowledge data 122 may include a directed edge 524 from node 522B to node 522A to indicate that the corresponding relation indicated by relation label 624A (e.g., being awakened by…) applies when the audio event indicated by event label 114E (e.g., opening a door) is earlier than the audio event indicated by event label 114D (e.g., a baby crying). In this example, in response to determining that event label 114D is associated with an earlier audio segment (e.g., audio segment 112D) than the audio segment 112E associated with event label 114E, and that knowledge data 122 includes an edge 524 from node 522A to node 522B, knowledge data analyzer 108 generates event pair relation data 152 indicating the relationship between event label pairs 114D-E. Alternatively, in this example, in response to determining that event label 114D is associated with an earlier audio segment (e.g., audio segment 112D) than audio segment 112E associated with event label 114E and that knowledge data 122 does not include any edge 524 from node 522A to node 522B, knowledge data analyzer 108 generates event pair relation data 152 indicating that event label pair 114D-E is not related, regardless of whether the edge in the other direction from node 522B to node 522A is included in knowledge data 122.
[0109] FIG. 6B The diagram 650 is an illustrative aspect of the operation associated with the audio scene graph updater 118 according to some examples of this disclosure. The audio scene graph updater 118 is configured to assign edge weights to edges 324 between nodes 322 corresponding to pairs of audio events with multiple relationships.
[0110] The audio scene graph updater 118 includes a total edge weight generator 510 coupled to an edge weight generator 616. The audio scene graph updater 118 also includes a relation similarity measure generator 614, which is coupled to an event pair text representation generator 610, a relation text embedding generator 612, and an edge weight generator 616.
[0111] Event-to-text representation generator 610 is configured to generate event-to-text embeddings 634 based on text embeddings 434 of audio event pairs. For example, event-to-text representation generator 610 generates event-to-text embeddings 634 for a first audio event (e.g., the sound of a baby crying) and a second audio event (e.g., the sound of a door opening). Event-to-text embeddings 634 are based on text embeddings 434D of event label 114D describing the first audio event and text embeddings 434E of event label 114E describing the second audio event. In one example, text embedding 434D includes a first feature value of the feature set, and text embedding 434E includes a second feature value of the feature set. In this example, event-to-text embedding 634 includes a third feature value of the feature set. The third feature value is based on the first and second feature values. For example, the first feature value includes the first feature value of the first feature, the second feature value includes the second feature value of the first feature, and the third feature value includes the third feature value of the first feature. The third feature value is based on the first and second feature values (e.g., the average of the first and second feature values). In a particular implementation, in response to determining that knowledge data 122 indicates that audio event pairs include multiple relations, event pair text representation generator 610 generates event pair text embedding 634.
[0112] The relational text embedding generator 612 generates relational text embeddings 644 for multiple relations of audio event pairs. For example, in response to determining that event pair relation data 152 indicates multiple relation tags for audio event pairs, the relational text embedding generator 612 generates a relational text embedding 644 for each of the multiple relation tags. For example, the relational text embedding generator 612 generates relational text embeddings 644A and 644B, respectively, corresponding to relation tags 624A and 624B. In a particular implementation, the relational text embedding generator 612 performs reference... FIG. 4 The event label indicates a similar operation described by generator 424.
[0113] The relational text embedding 644 may correspond to a numerical representation that captures the semantic meaning and contextual information of the relational label 624. In one example, the relational text embedding 644 includes a text feature vector comprising feature values of text features. In some implementations, the relational text embedding generator 612 includes a machine learning model (e.g., a deep neural network) trained on labeled text to generate text embeddings. According to some implementations, the relational text embedding generator 612 preprocesses the relational label 624 before generating the relational text embedding 644. Preprocessing may include converting the text to lowercase, removing punctuation, handling special characters, tokenizing the relational label 624 into individual words or sub-word units, or a combination thereof.
[0114] The relation similarity measure generator 614 generates a relation similarity measure 654 based on the event-to-text embedding 634 and the relation text embedding 644. For example, the relation similarity measure generator 614 determines a relation similarity measure 654A (e.g., cosine similarity) between the relation text embedding 644A and the event-to-text embedding 634. Similarly, the relation similarity measure generator 614 determines a relation similarity measure 654B (e.g., cosine similarity) between the relation text embedding 644B and the event-to-text embedding 634.
[0115] Edge weight generator 616 is configured to determine edge weights 526 for multiple relations based on relation similarity metric 654 and total edge weight 528. For example, edge weight generator 616 determines edge weight 526A based on the total edge weight 528 and the ratio of relation similarity metric 654A to the sum of relation similarity metrics 654 (e.g., edge weight 526A = total edge weight 528 * (relation similarity metric 654A / sum of relation similarity metrics 654)). Similarly, edge weight generator 616 generates edge weight 526B based on the total edge weight 528 and the ratio of relation similarity metric 654B to the sum of relation similarity metrics 654 (e.g., edge weight 526B = total edge weight 528 * (relation similarity metric 654B / sum of relation similarity metrics 654)).
[0116] The audio scene graph updater 118 assigns an edge weight 526A (e.g., 0.3) and a relation label 624A (e.g., "woke up by...") to edge 324E. The audio scene graph updater 118 adds one or more edges between nodes 322D and 322E for the remaining relation labels of multiple relations, and assigns the relation label and edge weight to each of the added edges. For example, the audio scene graph updater 118 adds an edge 324G between nodes 322D and 322E. Edge 324G has the same direction as edge 324E. The audio scene graph updater 118 assigns an edge weight 526B (e.g., 0.4) and a relation label 624B (e.g., "burst noise") to edge 324G.
[0117] If the audio scene graph 162 includes multiple edges (e.g., bidirectional edges), the audio scene graph updater 118 can split the edge weights 526 of a specific relationship among the multiple edges. For example, the audio scene graph updater 118 assigns a first part (e.g., half) of the edge weight 526A (e.g., 0.3) and the relationship label 624A to the edge 324E from node 322D to node 322E, and assigns the remaining part (e.g., half) of the edge weight 526A (e.g., 0.3) and the relationship label 624A to the edge from node 322E to node 322D.
[0118] Therefore, the audio scene graph updater 118 assigns a portion of the total edge weight 528 as edge weights to edges corresponding to the relations based on the similarity between the event-to-text embedding 634 and the corresponding relation text embedding 644. Relations whose relation labels have a relation text embedding that is more similar to the event-to-text embedding 634 are more likely to be accurate (e.g., have greater strength). For example, a first audio event (e.g., a baby crying) and a second audio event (e.g., music) have a first relation with a first relation label (e.g., restless because of…) and a second relation with a second relation label (e.g., eavesdropping). The first relation embedding with the first relation label (e.g., restless because of…) is more similar to the event-to-text embedding 634 than the second relation embedding with the second relation label (e.g., eavesdropping). The first relation may be stronger than the second relation.
[0119] In a specific example where knowledge data 122 includes directional information about relations, if the temporal order of audio events associated with the direction of edge 324E matches the temporal order of the corresponding relations of audio events indicated by knowledge data 122, then the audio scene graph updater 118 assigns edge weights 526A.
[0120] FIG. 7 This is an illustrative diagram of some examples of operation associated with the graph encoder 120 according to the present disclosure. The graph encoder 120 includes a position encoding generator 750 coupled to the graph transformer 770.
[0121] Graph encoder 120 is configured to encode audio scene graph 162 to generate coded graph 172. Position encoder generator 750 is configured to generate position codes 756 for nodes 322 of audio scene graph 162. Graph transformer 770 is configured to encode audio scene graph 162 based on position codes 756 to generate coded graph 172.
[0122] According to some specific implementations, the position encoding generator 750 is configured to determine the time position 754 of node 322. For example, the position encoding generator 750 determines the time position 754 based on the audio segment time sequence 164 of the audio segment 112 corresponding to node 322. For illustration, the position encoding generator 750 assigns a first time position 754 (e.g., 1) to node 322A associated with audio segment 112A, which has the earliest playback start time (e.g., 0 seconds) as indicated by the audio segment time sequence 164. Similarly, the position encoding generator 750 assigns a second time position 754 (e.g., 2) to node 322B associated with audio segment 112A, which has a second earliest playback time (e.g., 2 seconds) as indicated by the audio segment time sequence 164, and so on. The position encoding generator 750 assigns time position 754D (e.g., 4) to node 322D corresponding to the playback start time of audio segment 112D, and assigns time position 754E (e.g., 5) to node 322E corresponding to the playback start time of audio segment 112E.
[0123] According to some specific implementations, the position encoding generator 750 determines the Laplace position code 752 of node 322 of audio scene graph 162. For example, the position encoding generator 750 generates a Laplace position code 752D, which indicates the position of node 322D relative to other nodes in audio scene graph 162. As another example, the position encoding generator 750 generates a Laplace position code 752E, which indicates the position of node 322E relative to other nodes in audio scene graph 162.
[0124] Position code generator 750 generates position code 756 based on time position 754, Laplace position code 752, or a combination thereof. For example, position code generator 750 generates position code 756D based on time position 754D, Laplace position code 752D, or both. For illustration, position code 756D can be a combination (e.g., concatenation) of time position 754D and Laplace position code 752D. In a particular implementation, position code 756D corresponds to a weighted sum of the code of time position 754D and Laplace position code 752D according to the following formula: Position code 756D = w1 * code of time position 754D + w2 * Laplace position code 752D, where w1 and w2 are weights. Similarly, position code generator 750 generates time position 754E based on time position 754E, Laplace position code 752E, or both. The position encoding generator 750 provides the position encoding 756 to the graph transformer 770.
[0125] Graph transformer 770 includes an input generator 772 coupled to one or more graph transformer layers 774. The input generator 772 is configured to generate node embeddings 782 of nodes 322 of the audio scene graph 162. For example, the input generator 772 generates a node embedding 782D of node 322D. In a particular aspect, the node embedding 782D is based on audio segment 112D, event label 114D, audio embedding 432 of audio segment 112D, text embedding 434 of event label 114D, event representation 146D, or a combination thereof. Similarly, the input generator 772 generates a node embedding 782E of node 322E.
[0126] The input generator 772 is also configured to generate edge embeddings 784 of edge 324 of the audio scene graph 162. For example, the input generator 772 generates an edge embedding 784DE of edge 324E from node 322D to node 322E. In a particular aspect, the edge embedding 784DE is based on any relation label 624 associated with edge 324E, the edge weight 526A associated with edge 324E, or both. In the example where the audio scene graph 162 includes edge 324 from node 322E to node 322D, the input generator 772 generates an edge embedding 784ED of edge 324.
[0127] In the example where audio scene graph 162 includes multiple edges from node 322D to node 322E corresponding to multiple relationships, edge embedding 784 includes multiple edge embeddings corresponding to the multiple edges.
[0128] Input generator 772 provides node embeddings 782 and edge embeddings 784 to one or more graph transformer layers 774. The one or more graph transformer layers 774 process the node embeddings 782 and edge embeddings 784 based on position encoding 756 to generate an encoded graph 172, as referenced. FIG. 8 Further description
[0129] FIG. 8This is an illustrative diagram of an aspect of the operation associated with one or more graph transformer layers 774 according to some examples of this disclosure. Each graph transformer layer 774 includes one or more heads 804 (e.g., one or more attention heads). Each of the one or more heads 804 includes a product and scaling layer 810 coupled to a softmax layer 814 via a dot product layer 812. The softmax layer 814 is coupled to a dot product layer 816. The one or more heads 804 of the graph transformer layer are coupled to a stitching layer 818 and a stitching layer 820 of the graph transformer layer. For example, the dot product layer 816 of each of the one or more heads 804 is coupled to the stitching layer 818, and the dot product layer 812 of each of the one or more heads 804 is coupled to the stitching layer 820.
[0130] The graph transformer layer includes a stitching layer 818, which is coupled to an addition and normalization layer 834 via an addition and normalization layer 822 and a feedforward network 828. The graph transformer layer includes a stitching layer 820, which is coupled to an addition and normalization layer 836 via an addition and normalization layer 824 and a feedforward network 830. The graph transformer includes a stitching layer 820, which is coupled to an addition and normalization layer 838 via an addition and normalization layer 826 and a feedforward network 832.
[0131] Node embedding 782, edge embedding 784, and position encoding 756 are provided as input to the initial graph transformer layer in one or more graph transformer layers 774. The output of the previous graph transformer layer is provided as input to the subsequent graph transformer layer. The output of the last graph transformer layer corresponds to the encoded graph 172.
[0132] The combination of the position code 756D and the node embedding 782D of node 322D is provided to the head 804 as query vector 809. The combination of the node embedding 782E and the position code 756E of node 322E is provided to the head 804 as key vector 811 and value vector 813. If the audio scene graph 162 includes an edge from node 322D to node 322E, then the edge embedding 784DE is provided to the head 804 as edge vector 815. If the audio scene graph 162 includes an edge from node 322E to node 322D, then the edge embedding 784ED is provided to the head 804 as edge vector 845.
[0133] The product and scaling layer 810 of layer 804 generates the product of query vector 809 and key vector 811, and performs scaling on the product. The dot product layer 812 generates the dot product of the output of the product and scaling layer 810 with a combination (e.g., concatenation) of edge vectors 815 and 845. The output of the dot product layer 812 is provided to each of the softmax layer 814 and the concatenation layer 820. The softmax layer 814 performs normalization on the output of the dot product layer 812. The dot product layer 816 generates the dot product of the output of the softmax layer 814 with the value vector 813. The sum 817 of the outputs of the dot product layer 816 is provided to the concatenation layer 818.
[0134] The stitching layer 818 sums the dot product layers 816 of each head in one or more heads 804 of the stitching graph transformer layer to generate output 819. The stitching layer 820 sums the outputs of the dot product layers 812 of each head in one or more heads 804 of the stitching graph transformer layer to generate output 821. The summation and normalization layer 822 performs summation and normalization of the query vector 809 and the output 819, and the generated output is provided to each of the feedforward network 828 and the summation and normalization layers 834.
[0135] Addition and normalization layer 824 performs addition and normalization of edge embedding 784DE and output 821, and the resulting output is provided to each of feedforward network 830 and addition and normalization layer 836. Addition and normalization layer 826 performs addition and normalization of edge embedding 784ED and output 821, and the resulting output is provided to each of feedforward network 832 and addition and normalization layer 838.
[0136] Addition and normalization layer 834 performs addition and normalization on the output of addition and normalization layer 822 and the output of feedforward network 828 to generate node embedding 882D corresponding to node 322D. A similar operation can be performed to generate node embedding 882 corresponding to node 322E. Addition and normalization layer 836 performs addition and normalization on the output of addition and normalization layer 824 and the output of feedforward network 830 to generate edge embedding 884DE. Addition and normalization layer 838 performs addition and normalization on the output of addition and normalization layer 826 and the output of feedforward network 832 to generate edge embedding 884ED.
[0137] According to some specific implementations, the layer update equation of the graph transformer layer (l) is given by the following equation:
[0138] Equation 1
[0139] Equation 2
[0140] Equation 3
[0141] Equation 4
[0142] Where i represents a node (e.g., node 322D). The output 819 of the splicing layer 818 of the graph transformer layer (l) is represented by ||, where || represents splicing, k=1 to H represents the number of attention heads, and j represents the set of neighboring elements (N) included in node i. i (a node directly connected to that node, e.g., node 322E). This represents a value vector (e.g., value vector 813), and This represents the node embedding of node j (e.g., node embedding 782E). This indicates that the output of the splicing layer 818 is 819. This represents the output of the product and scaling layer 810. This represents the output of the dot product layer 812. This represents the output of the softmax layer 814. This indicates that the query vector is 809. This represents the node embedding of node i (e.g., node embedding 782D). Represents the key vector 811. This represents the dimension of the key vector 811. Let the edge vector (e.g., edge vector 815) represent the first edge embedding (e.g., edge embedding 784DE), and The edge vector (e.g., edge vector 845) represents the second edge embedding (e.g., edge embedding 784ED). This represents the output of the addition performed by the addition and normalization layer 822, and This represents the output of the addition performed by each of the addition and normalization layers 824 and 826. Output and It is passed to a separate feedforward network, which is preceded and followed by residual connections and normalization layers, as given by the following equation:
[0143] Equation 5
[0144] Equation 6
[0145] Equation 7
[0146] Equation 8
[0147] Equation 9
[0148] Equation 10
[0149] Equation 11
[0150] Equation 12
[0151] Equation 13
[0152] in This represents the output of the addition and normalization layer 822. Corresponding to the intermediate representation, (For example, node embedding 882D) represents the output of the addition and normalization layer 834. This represents the output of the addition and normalization layer 824. This indicates the middle part. (For example, edge embedding 884DE) represents the output of the summation and normalization layer 836. This represents the output of the addition and normalization layer 826. This indicates the intermediate representation, and (For example, edge embedding 884EE) represents the output of the summation and normalization layer 838. , and ReLU represents the intermediate representation, and ReLU represents the rectified linear unit activation function.
[0153] If one or more graph transformer layers 774 include subsequent graph transformer layers, then node embedding 882D, node embedding 882 corresponding to node 322E, edge embedding 884DE, and edge embedding 884ED are provided as input to the subsequent graph transformer layers. For example, node embedding 882D is provided as query vector 809 to the subsequent graph transformer layer, and node embedding 882 corresponding to node 322E is provided as key vector 811 and value vector 813 to the subsequent graph transformer layer. In some aspects, a combination of edge embedding 884DE and edge embedding 884ED is provided as input to the dot product layer 812 of the head 804 of the subsequent graph transformer layer. Edge embedding 884DE is provided as input to the summation and normalization layer 824 of the subsequent graph transformer layer. Edge embedding 884ED is provided as input to the summation and normalization layer 826 of the subsequent graph transformer layer.
[0154] The node embedding 882D, edge embedding 884DE, and edge embedding 884ED of the last graph transformer layer in one or more graph transformer layers 774 are included in the encoded graph 172. Similar operations are performed for the other nodes 322 of the audio scene graph 162.
[0155] One or more graph transformer layers 774 that process two edge embeddings (e.g., edge embedding 784DE and edge embedding 784ED) of a pair of nodes (e.g., nodes 322D and 322E) are provided as an illustrative example. In other examples, the audio scene graph 162 may include fewer or more than two edges between a pair of nodes, and one or more graph transformer layers 774 process the corresponding edge embeddings of that pair of nodes. For illustration, one or more graph transformer layers 774 may include one or more additional edge layers, wherein each edge layer includes a first summation and normalization layer coupled to a second summation and normalization layer of the feedforward network. A splicing layer 820 of the graph transformer layers is coupled to the first summation and normalization layer of each of these edge layers.
[0156] refer to FIG. 9 A specific exemplary aspect of a system configured to update a knowledge-based audio scene graph is disclosed, and is generally specified as 900. In this specific aspect, FIG. 1 System 100 includes one or more components of system 900.
[0157] System 900 includes a graph updater 962 coupled to audio scene graph generator 140. Graph updater 962 is configured to update audio scene graph 162 based on user feedback 960. In a particular implementation, user feedback 960 is based on video data 910 associated with audio data 110. For example, audio data 110 and video data 910 represent scene environment 902. In a particular aspect, scene environment 902 corresponds to a physical environment, a virtual environment, or a combination thereof, wherein video data 910 corresponds to an image of scene environment 902, and audio data 110 corresponds to audio of scene environment 902.
[0158] Audio scene graph generator 140 generates audio scene graph 162 based on audio data 110, as shown in the reference. FIG. 1 As described. During forward pass 920, audio scene graph generator 140 provides audio scene graph 162 to graph updater 962, and graph updater 962 provides audio scene graph 162 to user interface 916. For example, user interface 916 includes user device, display device, graphical user interface (GUI), or a combination thereof. For illustration, graph updater 962 generates a GUI including a representation of audio scene graph 162 and provides the GUI to display device.
[0159] User 912 provides user input 914 to indicate graph update 917 of audio scene graph 162. In a particular implementation, user 912 provides user input 914 in response to viewing an image represented by video data 910. Graph updater 962 is configured to update audio scene graph 162 based on user input 914, video data 910, or both. In a first example, based on determining that video data 910 indicates a strong correlation between a second audio event (e.g., the sound of a door opening) and a first audio event (e.g., the sound of a doorbell), user 912 provides user input 914 indicating an edge weight 526A (e.g., 0.9) for the edge 324F from node 322B corresponding to the first audio event to node 322C corresponding to the second audio event. In the second example, based on the determination of video data 910 indicating that a second audio event (e.g., a baby crying) has a relationship with a first audio event (e.g., music) that is not indicated in the audio scene diagram 162, user 912 provides user input 914 indicating the relationship of a new edge from node 322C corresponding to the first audio event to node 322D corresponding to the second audio event, edge weight 526B (e.g., 0.8), relationship label, or a combination thereof. In the third example, based on the determination of video data 910 indicating that an audio event (e.g., the sound of a car passing by) is associated with a corresponding audio segment, user 912 provides user input 914 indicating the association between the audio segment and the audio event.
[0160] In response to receiving a graph update 917 (e.g., corresponding to user input 914), the graph updater 962 updates the audio scene graph 162 based on the graph update 917. In a first example, the graph updater 962 assigns edge weight 526A to edge 324F. In a second example, the graph updater 962 adds an edge 324H from node 322C to node 322D and assigns edge weight 526B, relation label, or both to edge 324H.
[0161] In some specific implementations, the audio scene graph generator 140 performs backpropagation 922 based on graph update 917. For example, graph updater 962 provides graph update 917 to audio scene graph generator 140. In a particular aspect, audio scene graph generator 140 updates knowledge data 122 based on graph update 917. In a first example, audio scene graph generator 140 updates knowledge data 122 to indicate the relevance of a first audio event (e.g., described by event label 114B) associated with node 322B and a second audio event (e.g., described by event label 114E) associated with node 322E. In a particular aspect, audio scene graph generator 140 updates the similarity measure associated with the first audio event (e.g., described by event label 114B) and the second audio event (e.g., described by event label 114E) to correspond to edge weights 526A. In the second example, the audio scene graph generator 140 updates the knowledge data 122 to add a relationship from a first audio event (e.g., described by event label 114C) associated with node 322C to a second audio event (e.g., described by event label 114D). If indicated by graph update 917, the audio scene graph generator 140 assigns relationship labels to the relationships in the knowledge data 122. In a specific aspect, the audio scene graph generator 140 updates the similarity metric associated with the relationship between the first audio event (e.g., described by event label 114C) and the second audio event (e.g., described by event label 114D) to correspond to the edge weight 526B. In the third example, the audio scene graph generator 140 updates the audio scene segmenter 102 based on the graph update 917 indicating that an audio event has been detected in an audio segment.
[0162] The audio scene graph generator 140 uses an updated audio scene segmenter 102, updated knowledge data 122, updated similarity metrics, or a combination thereof, in subsequent processing of the audio data 110. The technical advantages of backpropagation 922 include dynamic adjustments to the audio scene graph 162 based on graph updates 917.
[0163] FIG. 10 This is an illustrative illustration of a graphical user interface (GUI) 1000 according to some examples of this disclosure. In a particular aspect, the GUI 1000 is generated by a GUI generator, which is coupled to... FIG. 1 System 100 FIG. 9 The system 900 or both audio scene graph generator 140. In a certain respect, FIG. 9 The graph updater 962 or user interface 916 includes a GUI generator.
[0164] GUI 1000 includes audio input 1002 and submission input 1004. User 912 uses audio input 1002 to select audio data 110 and activates submission input 1004 to provide audio data 110 to audio scene graph generator 140. In response to the activation of submission input 1004, audio scene graph generator 140 generates audio scene graph 162 based on audio data 110, as shown in the reference. FIG. 1 As described.
[0165] The GUI generator updates GUI 1000 to include a representation of the audio scene graph 162. Depending on the specific implementation, GUI 1000 includes update input 1006. In one example, user 912 uses GUI 1000 to update the representation of the audio scene graph 162, such as by adding or updating edge weights, adding or removing edges, adding or updating relation labels, etc. User 912 activates update input 1006 to generate user input 914 corresponding to the update to the representation of the audio scene graph 162. Graph updater 962 updates the audio scene graph 162 based on user input 914, as shown in the reference... FIG. 9 The technical advantages of GUI 1000, as described in Figure 162, include user authentication, user updates, or both.
[0166] refer to FIG. 11 A specific exemplary aspect of a system configured to update a knowledge-based audio scene graph is disclosed, and is generally designated as 1100. In this specific aspect, FIG. 1 System 100 includes one or more components of system 1100.
[0167] System 1100 includes a visual analyzer 1160 coupled to a graph updater 962. The visual analyzer 1160 is configured to detect visual relationships in video data 910 and generate a graph update 917 based on the visual relationships to update the audio scene graph 162.
[0168] The visual analyzer 1160 includes a spatial analyzer 1114 coupled to a fully connected layer 1120 and an object detector 1116 coupled to the fully connected layer 1120. The fully connected layer 1120 is coupled to an audio scene graph analyzer 1124 via a visual relation encoder 1122.
[0169] Video data 910 represents video frame 1112. In a specific aspect, spatial analyzer 1114 uses multiple convolutional layers (C) to perform spatial mapping across video frame 1112. Object detector 1116 performs object detection and recognition on video frame 1112 to generate feature vectors 1118 corresponding to the detected objects. In a specific aspect, the output of spatial analyzer 1114 and feature vectors 1118 are concatenated to generate the input of fully connected layer 1120. The output of fully connected layer 1120 is provided to visual relation encoder 1122. In a specific aspect, visual relation encoder 1122 includes multiple transformer encoder layers. Visual relation encoder 1122 processes the output of fully connected layer 1120 to generate visual relation encoding 1123 representing the visual relations detected in video data 910. Audio scene graph analyzer 1124 generates graph update 917 based on visual relation encoding 1123 and audio scene graph 162 (or encoded graph 172).
[0170] In a particular aspect, the audio scene graph analyzer 1124 includes one or more graph transformer layers. In a particular implementation, the audio scene graph analyzer 1124 generates visual node embeddings and visual edge embeddings based on visual relation encoding 1123, and processes visual node embeddings, visual edge embeddings, node embeddings of encoded graph 172, edge embeddings of encoded graph 172, or combinations thereof, to generate a graph update 917. In a particular example, the audio scene graph analyzer 1124 determines, based on video data 910, that an audio event has been detected in a corresponding audio segment, and generates a graph update 917 to indicate that an audio event has been detected in the audio segment. Graph updater 962 updates the audio scene graph 162 based on the graph update 917, as referenced. FIG. 9 As described. In a particular aspect, graph updater 962 performs backpropagation 922 based on graph update 917, as referenced. FIG. 9 The technical advantages of the visual analyzer 1160, as described, include automatic updating of the audio scene graph 162 based on video data 910.
[0171] FIG. 12 This is an illustrative diagram illustrating an exemplary aspect of a system operable, based on some examples of this disclosure, for generating query result 1226 using audio scene diagram 162. In a particular aspect, FIG. 1 System 100 includes one or more components of system 1200.
[0172] System 1200 includes a decoder 1224 coupled to a query encoder 1220 and a graph encoder 120. The query encoder 1220 is configured to encode a query 1210 to generate an encoded query 1222. The decoder 1224 is configured to generate a query result 1226 based on the encoded query 1222 and the encoded graph 172. In a particular aspect, a combination (e.g., concatenation) of the encoded query 1222 and the encoded graph 172 is provided as input to the decoder 1224, and the decoder 1224 generates the query result 1226.
[0173] It should be understood that the use of audio scene graph 162 to generate query result 1226 is provided as an illustrative example, and in other examples, audio scene graph 162 can be used to perform one or more downstream tasks of various types. Technical advantages of using audio scene graph 162 include the ability to generate query result 1226 corresponding to a more complex query 1210 based on information from knowledge data 122 injected into audio scene graph 162.
[0174] FIG. 13 This is a block diagram illustrating an exemplary aspect of a system 1300 operable to generate a knowledge-based audio scene graph, according to some examples of this disclosure. System 1300 includes a device 1302, wherein one or more processors 1390 include an always-on power domain 1303 and a second power domain 1305, such as an on-demand power domain. In some specific implementations, a first stage 1340 and a buffer 1360 of a multi-stage system 1320 are configured to operate in an always-on mode, and a second stage 1350 of the multi-stage system 1320 is configured to operate in an on-demand mode.
[0175] The always-on power domain 1303 includes a buffer 1360 and a first stage 1340, which includes a keyword detector 1342. The buffer 1360 is configured to store audio data 110, video data 910, or both, for access by components of the multi-stage system 1320. In a particular aspect, device 1302 is coupled to (e.g., includes) a camera 1310, a microphone 1312, or both. In a particular embodiment, microphone 1312 is configured to generate audio data 110. In a particular embodiment, camera 1310 is configured to generate video data 910.
[0176] The second power domain 1305 includes a second stage 1350 of the multi-stage system 1320 and also includes activation circuitry 1330. The second stage 1350 includes an audio scene graph system 1356, which includes an audio scene graph generator 140. In some embodiments, the audio scene graph system 1356 may also include one or more of a graph encoder 120, a graph updater 962, a user interface 916, a visual analyzer 1160, or a query encoder 1220.
[0177] The first stage 1340 of the multi-stage system 1320 is configured to generate at least one of a wake-up signal 1322 or an interrupt 1324 to initiate one or more operations at the second stage 1350. In a particular embodiment, the first stage 1340 generates at least one of a wake-up signal 1322 or an interrupt 1324 in response to a keyword detector 1342 detecting a phrase in the audio data 110 corresponding to a command to activate the audio scene graph system 1356. In a particular embodiment, the first stage 1340 generates at least one of a wake-up signal 1322 or an interrupt 1324 in response to receiving user input or a command from another device indicating that the audio scene graph system 1356 will be activated. In one example, the wake-up signal 1322 is configured to switch the second power domain 1305 from a low-power mode 1332 to an active mode 1334 to activate one or more components of the second stage 1350.
[0178] For example, activation circuit 1330 may include or be coupled to power management circuitry, clock circuitry, head switch or foot switch circuitry, buffer control circuitry, or any combination thereof. Activation circuit 1330 may be configured to initiate power-on of the second stage 1350, such as by selectively applying or increasing the supply voltage of the second stage 1350, the supply voltage of the second power domain 1305, or both. As another example, activation circuit 1330 may be configured to selectively enable or disable the clock signal to the second stage 1350, such as preventing or enabling circuit operation without removing power.
[0179] Output 1352 generated by the second stage 1350 of the multi-stage system 1320 is provided to one or more applications 1354. In certain aspects, output 1352 includes at least one of an audio scene graph 162, an encoded graph 172, a graph update 917, a GUI 1000, an encoded query 1222, or a combination of encoded query 1222 and encoded graph 172. One or more applications 1354 may be configured to perform one or more downstream tasks based on output 1352. For illustration, one or more applications 1354 may include decoder 1224, a voice interface application, an integrated assistance application, a vehicle navigation and entertainment application, or a home automation system, as illustrative and non-limiting examples.
[0180] By selectively activating the second level 1350 based on user input, commands, or the results of processing audio data 110 at the first level 1340 of the multi-level system 1320, the total power consumption associated with generating a knowledge-based audio scene graph can be reduced.
[0181] FIG. 14A specific embodiment 1400 of an integrated circuit 1402 including one or more processors 1490 is depicted. The one or more processors 1490 include an audio scene graph system 1356. In some embodiments, the one or more processors 1490 also include a keyword detector 1342.
[0182] Integrated circuit 1402 includes an audio input 1404, such as one or more bus interfaces, to enable receiving audio data 110 for processing. Integrated circuit 1402 also includes a video input 1408, such as one or more bus interfaces, to enable receiving video data 910 for processing. Integrated circuit 1402 also includes a signal output 1406, such as a bus interface, to enable sending output signals 1452, such as an audio scene graph 162, an encoded graph 172, a graph update 917, an encoded query 1222, a combination of encoded graph 172 and encoded query 1222, a query result 1226, or a combination thereof.
[0183] Integrated circuit 1402 enables the audio scene graph system 1356 to be implemented as a component in a system that includes a microphone, such as... FIG. 15 The mobile phone or tablet computer described, such as FIG. 16 The described head-mounted device, such as FIG. 17 The described wearable electronic devices, such as FIG. 18 The described voice control loudspeaker system, such as FIG. 19 The camera equipment described, such as FIG. 20 The virtual reality, mixed reality, or augmented reality headsets described, or such as FIG. 21 or FIG. 22 The vehicles depicted.
[0184] As an illustrative and non-restrictive example, FIG. 15A specific implementation 1500 of a mobile device 1502 (such as a phone or tablet) is depicted. The mobile device 1502 includes a camera 1510, a microphone 1520, and a display screen 1504. Components of one or more processors 1490 (including an audio scene graph system 1356, a keyword detector 1342, or both) are integrated into the mobile device 1502 and illustrated using dashed lines to indicate internal components of the mobile device 1502 that are not typically visible to the user. In a particular example, the keyword detector 1342 operates to detect user voice activity and then processes the user voice activity to perform one or more operations at the mobile device 1502, such as launching a graphical user interface or otherwise displaying additional information associated with the user's voice at the display screen 1504 (e.g., via an integrated "smart assistant" application). In an illustrative example, in response to the keyword detector 1342 detecting a specific phrase, an audio scene graph generator 140 is activated to generate an audio scene graph 162. In a particular aspect, the audio scene graph system 1356 uses a decoder 1224 to generate a query result 1226 indicating which application might be useful to the user, and activates the application indicated in the query result 1226.
[0185] FIG. 16 A specific implementation 1600 of a head-mounted device 1602 is depicted. The head-mounted device 1602 includes a microphone 1620. Components of one or more processors 1490 (including an audio scene graph system 1356) are integrated into the head-mounted device 1602. In a particular example, the audio scene graph system 1356 operates to detect user voice activity and then processes the user voice activity to perform one or more operations at the head-mounted device 1602, such as generating an audio scene graph 162, performing one or more downstream tasks based on the audio scene graph 162, sending the audio scene graph 162 to a second device (not shown) for further processing, or a combination thereof.
[0186] FIG. 17A specific implementation 1700 of a wearable electronic device 1702 (illustrated as a "smartwatch") is depicted. An audio scene graph system 1356, a keyword detector 1342, a camera 1710, a microphone 1720, or a combination thereof are integrated into the wearable electronic device 1702. In a particular example, the keyword detector 1342 operates to detect user voice activity and then processes the user voice activity to perform one or more operations at the wearable electronic device 1702, such as launching a graphical user interface or otherwise displaying other information associated with the user's voice on a display screen 1704 of the wearable electronic device 1702. For illustration, the display screen 1704 may be configured to display notifications based on user voice detected by the wearable electronic device 1702. In a particular example, the wearable electronic device 1702 includes a haptic device that provides haptic notifications (e.g., vibration) in response to the detection of user voice activity. For example, a haptic notification may cause a user to look at the wearable electronic device 1702 to see a displayed notification indicating that a keyword spoken by the user has been detected. Therefore, the wearable electronic device 1702 can alert users with hearing impairments or users wearing head-mounted devices to the detection of user voice activity. In a specific aspect, in response to the keyword detector 1342 detecting a specific phrase in the user's voice activity, the audio scene graph system 1356 generates an audio scene graph 162 and uses the audio scene graph 162 to perform one or more downstream tasks.
[0187] FIG. 18 This is a specific implementation 1800 of a wireless speaker and voice activation device 1802. The wireless speaker and voice activation device 1802 may have wireless network connectivity and be configured to perform auxiliary operations. A camera 1810, a microphone 1820, and one or more processors 1890, including an audio scene mapping system 1356 and a keyword detector 1342, are included in the wireless speaker and voice activation device 1802. The wireless speaker and voice activation device 1802 also includes a speaker 1804. During operation, in response to a verbal command identified as user voice received via operation of the keyword detector 1342, the wireless speaker and voice activation device 1802 may perform an assistant operation, such as via execution of a voice activation system (e.g., an integrated assistant application). This assistant operation may include adjusting the temperature, playing music, turning on lights, etc. For example, the assistant operation may be performed in response to receiving a command following a keyword or key phrase (e.g., “Hello, assistant”). In a particular aspect, in response to the keyword detector 1342 detecting a specific phrase in the user's voice activity, the audio scene graph system 1356 generates an audio scene graph 162 and uses the audio scene graph 162 to perform one or more downstream tasks, such as generating query results 1226.
[0188] FIG. 19A specific implementation 1900 of a portable electronic device corresponding to a camera device 1902 is depicted. An audio scene graph system 1356, a keyword detector 1342, an image sensor 1910, a microphone 1920, or a combination thereof, are included in the camera device 1902. During operation, in response to receiving a spoken command identified as user speech via operation of the keyword detector 1342, the camera device 1902 can perform operations in response to the spoken user command, such as adjusting image or video capture settings, image or video playback settings, or image or video capture instructions, as illustrative examples. In a particular aspect, in response to the keyword detector 1342 detecting a specific phrase in the user's voice activity, the audio scene graph system 1356 generates an audio scene graph 162 and uses the audio scene graph 162 to perform one or more downstream tasks, such as adjusting camera settings based on the detected audio scene.
[0189] FIG. 20 A specific implementation 2000 of a portable electronic device corresponding to a virtual reality, mixed reality, or augmented reality head-mounted device 2002 is depicted. For example, the head-mounted device 2002 corresponds to an extended reality head-mounted device. An audio scene graph system 1356, a keyword detector 1342, a camera 2010, a microphone 2020, or a combination thereof are integrated into the head-mounted device 2002. In a particular aspect, the head-mounted device 2002 includes a microphone 2020 to capture a user's speech, ambient sounds, or a combination thereof. The keyword detector 1342 may perform user voice activity detection based on audio signals received from the microphone 2020 of the head-mounted device 2002. A visual interface device is positioned in front of the user's eyes to enable the display of augmented reality images or scenes, mixed reality images or scenes, or virtual reality images or scenes to the user when wearing the head-mounted device 2002. In a particular example, the visual interface device is configured to display a notification indicating user voice detected in the audio signal. In a specific aspect, in response to the keyword detector 1342 detecting a specific phrase in the user's voice activity, the audio scene graph system 1356 generates an audio scene graph 162 and uses the audio scene graph 162 to perform one or more downstream tasks.
[0190] FIG. 21A specific implementation 2100 of vehicle 2102 is depicted, exemplified as a manned or unmanned aerial device (e.g., a package delivery drone). A keyword detector 1342, an audio scene graph system 1356, a camera 2110, a microphone 2120, or a combination thereof, are integrated into vehicle 2102. Keyword detector 1342 can perform user voice activity detection based on audio signals received from microphone 2120 of vehicle 2102, such as delivery instructions for an authorized user of vehicle 2102. In a particular aspect, in response to keyword detector 1342 detecting a specific phrase in user voice activity, audio scene graph system 1356 generates an audio scene graph 162 and uses audio scene graph 162 to perform one or more downstream tasks, such as generating query results 1226.
[0191] FIG. 22 Another specific embodiment 2200 of a vehicle 2202, exemplified as an automobile, is depicted. The vehicle 2202 includes one or more processors 1490, which include an audio scene graph system 1356, a keyword detector 1342, or both. The vehicle 2202 also includes a camera 2210, a microphone 2220, or both. The microphone 2220 is positioned to capture the speech of the operator of the vehicle 2202. The keyword detector 1342 can perform user voice activity detection based on audio signals received from the microphone 2220 of the vehicle 2202.
[0192] In some implementations, user voice activity detection can be performed based on audio signals received from an internal microphone (e.g., microphone 2220), such as voice commands from authorized passengers. For example, user voice activity detection can be used to detect voice commands from the operator of vehicle 2202 (e.g., a parent setting the volume to 5 or setting the destination of the autonomous vehicle) and ignore the voices of other passengers (e.g., a child setting the volume to 10 or other passengers discussing another location). In some implementations, user voice activity detection can be performed based on audio signals received from an external microphone (e.g., microphone 2220) (such as an authorized user of the vehicle).
[0193] In a particular implementation, in response to receiving a spoken command identified as user speech via the operation of keyword detector 1342, the voice activation system initiates one or more operations on vehicle 2202 based on one or more keywords detected in the microphone signal (e.g., “unlock,” “start engine,” “play music,” “display weather forecast,” or another voice command), such as by providing feedback or information via display 2222 or one or more speakers. In a particular aspect, in response to keyword detector 1342 detecting a specific phrase in user voice activity, audio scene graph system 1356 generates audio scene graph 162 and uses audio scene graph 162 to perform one or more downstream tasks, such as generating query results 1226.
[0194] refer to FIG. 23 This illustrates a specific implementation of method 2300 for generating knowledge-based audio scene graphs. In a particular aspect, one or more operations of method 2300 are performed by at least one of the following: FIG. 1 The audio scene segmenter 102, audio scene graph constructor 104, knowledge data analyzer 108, event representation generator 106, audio scene graph updater 118, audio scene graph generator 140, and system 100 are included. FIG. 4 The event audio representation generator 422, event label representation generator 424, combiner 426, total edge weight generator 510 (Figure 5), event pair text representation generator 610 (Figure 6), relation text embedding generator 612, relation similarity measure generator 614, and edge weight generator 616 are all included. FIG. 7 Position encoding generator 750, graph transformer 770, input generator 772, one or more graph transformer layers 774 FIG. 13 The audio scene diagram includes system 1356, second stage 1350, second power domain 1305, one or more processors 1390, device 1302, and system 1300. FIG. 14 One or more processors 1490, integrated circuits 1402, FIG. 15 Mobile devices 1502 FIG. 16 Head-mounted device 1602, FIG. 17 Wearable electronic devices 1702 FIG. 18 Wireless speakers and voice activation devices 1802 FIG. 19 Camera equipment 1902, FIG. 20 Head-mounted devices 2002 FIG. 21 Transportation 2102 FIG. 22 The means of transport 2202 or a combination thereof.
[0195] Method 2300 includes identifying the audio data segment corresponding to the audio event at 2302. For example,FIG. 1 The audio scene segmenter 102 identifies the audio segment 112 of the audio data 110 corresponding to the audio event, as shown in the reference. FIG. 1 and FIG. 2 As described.
[0196] Method 2300 also includes assigning labels to these segments at 2304. For example, FIG. 1 The audio scene segmenter 102 assigns event label 114 to audio segment 112, as shown in the reference. FIG. 1 and FIG. 2 As described. The event label 114 for a specific audio segment 112 describes the corresponding audio event.
[0197] Method 2300 also includes determining relationships between audio events at 2306 based on knowledge data. For example, knowledge data analyzer 108 generates event pair relationship data 152 indicating relationships between audio events based on knowledge data 122, as referenced. FIG. 1 , FIG. 5A and FIG. 6A As described.
[0198] Method 2300 also includes constructing an audio scene graph at 2308 based on the temporal order of audio events. For example, FIG. 1 The audio scene graph constructor 104 constructs the audio scene graph 162 based on the audio segment time sequence 164 of the audio segments 112 corresponding to the audio events, as shown in the reference. FIG. 1 and FIG. 3 As described.
[0199] Method 2300 also includes assigning edge weights to the audio scene graph at 2310 based on a similarity metric between audio events and the relationships between audio events. For example, audio scene graph updater 118 assigns edge weights 526 to audio scene graph 162 based on the total edge weights 528 corresponding to the similarity metric between audio events and the relationships between audio events indicated by event pair relation data 152, as referenced. FIG. 5B and FIG. 6B As described.
[0200] The technical advantages of method 2300 include the generation of a knowledge-based audio scene graph 162. The audio scene graph 162 can be used to perform various types of analysis on the audio scene represented by the audio scene graph 162. For example, the audio scene graph 162 can be used to generate responses to queries, initiate one or more actions, or combinations thereof.
[0201] FIG. 23Method 2300 can be implemented by a field-programmable gate array (FPGA) device, an application-specific integrated circuit (ASIC), a processing unit (such as a central processing unit (CPU)), a digital signal processor (DSP), a controller, another hardware device, a firmware device, or any combination thereof. As an example, FIG. 23 Method 2300 can be executed by a processor that executes instructions, such as reference... FIG. 24 As described.
[0202] refer to FIG. 24 A block diagram depicting a specific, exemplary embodiment of the device is provided, and is generally designated as 2400. In various embodiments, device 2400 may have the same... FIG. 24 The illustrated example has more or fewer components. In the illustrative embodiment, device 2400 may correspond to... FIG. 13 Equipment 1302, including FIG. 14 The equipment of integrated circuit 1402, FIG. 15 Mobile devices 1502 FIG. 16 Head-mounted device 1602, FIG. 17 Wearable electronic devices 1702 FIG. 18 Wireless speakers and voice activation devices 1802 FIG. 19 Camera equipment 1902, FIG. 20 Head-mounted devices 2002 FIG. 21 Transportation 2102 FIG. 22 The means of transport 2202 or combinations thereof. In an exemplary embodiment, device 2400 may perform the reference FIG. 15 One or more operations as described.
[0203] In a particular implementation, device 2400 includes a processor 2406 (e.g., a CPU). Device 2400 may include one or more additional processors 2410 (e.g., one or more DSPs). In a particular aspect, FIG. 16 One or more processors 1390, FIG. 17 One or more processors 1490, FIG. 18 One or more processors 1890 or combinations thereof correspond to processor 2406, processor 2410 or combinations thereof. Processor 2410 may include a speech and music decoder-decoder (codec) 2408, which includes a speech decoder (“phonecoder”) encoder 2436, a phonecoder decoder 2438 or both. Processor 2410 includes an audio scene graph system 1356, a keyword detector 1342, one or more applications 1354 or combinations thereof.
[0204] Device 2400 may include memory 2486 and codec 2434. Memory 2486 may include instructions 2456 executable by one or more additional processors 2410 (or processor 2406) to implement the functionality described by reference audio scene graph system 1356, keyword detector 1342, one or more applications 1354, or a combination thereof. In a particular aspect, memory 2486 is configured to store data used or generated by audio scene graph system 1356, keyword detector 1342, one or more applications 1354, or a combination thereof. In one example, memory 2486 is configured to store... FIG. 19 The data includes: audio data 110, knowledge data 122, audio segments 112, event tags 114, audio segment time sequence 164, audio scene graph 162, event representation 146, event pair relation data 152, and encoding graph 172. FIG. 20 Audio embedding 432, text embedding 434, FIG. 21 Total edge weight 528, edge weight 526 FIG. 22 Relationship tag 624 FIG. 24 Relational text embedding 644, event-to-text embedding 634, relational similarity measurement 654. FIG. 1 Laplacian position encoding 752, time position 754, position encoding 756, node embedding 782, edge embedding 784. FIG. 13 Input and output, FIG. 14 User input 914, image update 917, video data 910 FIG. 15 GUI 1000 FIG. 16 Video frame 1112, feature vector 1118, visual relation encoding 1123. FIG. 17 Query 1210, Code Query 1222, Query Result 1226 FIG. 18 The output 1352 or a combination thereof. Device 2400 may include a modem 2470 coupled to antenna 2452 via transceiver 2450.
[0205] Device 2400 may include a display 2428 coupled to display controller 2426. One or more speakers 2492, one or more microphones 2420, one or more cameras 2418, or combinations thereof may be coupled to codec 2434. Codec 2434 may include digital-to-analog converter (DAC) 2402, analog-to-digital converter (ADC) 2404, or both. In a particular embodiment, codec 2434 may receive analog signals from one or more microphones 2420, use ADC 2404 to convert the analog signals into digital signals, and provide the digital signals to speech and music codec 2408. Speech and music codec 2408 may process digital signals, and these digital signals may be further processed by audio scene graph system 1356, keyword detector 1342, one or more applications 1354, or combinations thereof. In a particular embodiment, speech and music codec 2408 may provide digital signals to codec 2434. The codec 2434 can use the digital-to-analog converter 2402 to convert digital signals into analog signals, and can provide analog signals to one or more speakers 2492.
[0206] In a particular aspect, one or more microphones 2420 are configured to generate audio data 110. In a particular aspect, one or more cameras 2418 are configured to generate... FIG. 19 Video data 910. In certain aspects, one or more microphones 2420 include FIG. 20 Microphone 1312 FIG. 21 Microphone 1520 FIG. 22 Microphone 1620 FIG. 24 Microphone 1720 FIG. 1 Microphone 1820 FIG. 13 The microphone 1920 FIG. 14 Microphone 2020 FIG. 15 Microphone 2120 FIG. 16 Microphone 2220 or a combination thereof. In a particular aspect, one or more cameras 2418 include FIG. 17 Camera 1310 FIG. 18 Camera 1510 FIG. 19 Camera 1710, FIG. 20 Camera 1810, FIG. 21 Image sensor 1910, FIG. 22 Camera 2010 FIG. 24 Camera 2110 FIG. 1 Camera 2210 or a combination thereof.
[0207] In a particular embodiment, device 2400 may be included in a system-in-package (SoC) or system-on-a-chip (SoC) 2422. In a particular embodiment, memory 2486, processor 2406, processor 2410, display controller 2426, codec 2434, and modem 2470 are included in the SoC or SoC 2422. In a particular embodiment, input device 2430 and power supply 2444 are coupled to the SoC or SoC 2422. Furthermore, in a particular embodiment, such as... FIG. 13 As illustrated, the display 2428, input device 2430, one or more speakers 2492, one or more cameras 2418, one or more microphones 2420, antenna 2452, and power supply 2444 are external to the system-in-package or system-on-chip device 2422. In a particular implementation, each of the display 2428, input device 2430, one or more speakers 2492, one or more cameras 2418, one or more microphones 2420, antenna 2452, and power supply 2444 may be coupled to components of the system-in-package or system-on-chip device 2422, such as an interface or controller.
[0208] Device 2400 may include smart speakers, speaker bars, mobile communication devices, smartphones, cellular phones, laptops, computers, tablets, personal digital assistants, display devices, televisions, game consoles, music players, radios, digital video players, digital video disc (DVD) players, tuners, cameras, navigation devices, vehicles, head-mounted devices, augmented reality head-mounted devices, mixed reality head-mounted devices, virtual reality head-mounted devices, extended reality (XR) head-mounted devices, XR devices, mobile phones, air vehicles, home automation systems, voice-activated devices, wireless speakers and voice-activated devices, portable electronic devices, automobiles, computing devices, communication devices, Internet of Things (IoT) devices, virtual reality (VR) devices, base stations, mobile devices, or any combination thereof.
[0209] In conjunction with the described specific implementation, an apparatus includes a component for identifying an audio segment of audio data corresponding to an audio event. For example, the component for identifying the audio segment may correspond to... FIG. 14 Audio scene segmenter 102, audio scene graph generator 140, system 100, FIG. 15 The audio scene diagram includes system 1356, second stage 1350, second power domain 1305, one or more processors 1390, device 1302, and system 1300. FIG. 16 Integrated circuit 1402, one or more processors 1490, FIG. 17 Mobile devices 1502 FIG. 18 Head-mounted device 1602, FIG. 19Wearable electronic devices 1702 FIG. 20 One or more processors 1890, wireless speakers and voice activation devices 1802, FIG. 21 Camera equipment 1902, FIG. 22 Head-mounted devices 2002 FIG. 24 Transportation 2102 FIG. 1 Transportation vehicle 2202 FIG. 13 The processor 2406, processor 2410, device 2400, one or more other circuits or components, or any combination thereof, configured to identify audio segments of audio data corresponding to an audio event.
[0210] The device also includes components for assigning tags to audio segments, where the tag description for a specific audio segment corresponds to an audio event. For example, the component for assigning tags could correspond to... FIG. 14 Audio scene segmenter 102, audio scene graph generator 140, system 100, FIG. 15 The audio scene diagram includes system 1356, second stage 1350, second power domain 1305, one or more processors 1390, device 1302, and system 1300. FIG. 16 Integrated circuit 1402, one or more processors 1490, FIG. 17 Mobile devices 1502 FIG. 18 Head-mounted device 1602, FIG. 19 Wearable electronic devices 1702 FIG. 20 One or more processors 1890, wireless speakers and voice activation devices 1802, FIG. 21 Camera equipment 1902, FIG. 22 Head-mounted devices 2002 FIG. 24 Transportation 2102 FIG. 1 Transportation vehicle 2202 FIG. 13 The processor 2406, processor 2410, device 2400, are configured to assign tags to one or more other circuits or components, or any combination thereof, of audio segments.
[0211] The device also includes components for determining relationships between audio events based on knowledge data. For example, the component for determining relationships could correspond to... FIG. 14 Knowledge Data Analyzer 108, Audio Scene Graph Generator 140, System 100 FIG. 15 The audio scene diagram includes system 1356, second stage 1350, second power domain 1305, one or more processors 1390, device 1302, and system 1300. FIG. 16 Integrated circuit 1402, one or more processors 1490, FIG. 17 Mobile devices 1502FIG. 18 Head-mounted device 1602, FIG. 19 Wearable electronic devices 1702 FIG. 20 One or more processors 1890, wireless speakers and voice activation devices 1802, FIG. 21 Camera equipment 1902, FIG. 22 Head-mounted devices 2002 FIG. 24 Transportation 2102 FIG. 1 Transportation vehicle 2202 FIG. 13 The processor 2406, processor 2410, device 2400, one or more other circuits or components, or any combination thereof, configured to determine the relationship between audio events.
[0212] The device also includes components for constructing an audio scene graph based on the temporal sequence of audio events. For example, a component for identifying audio segments could correspond to... FIG. 14 Audio scene graph constructor 104, audio scene graph generator 140, system 100, FIG. 15 The audio scene diagram includes system 1356, second stage 1350, second power domain 1305, one or more processors 1390, device 1302, and system 1300. FIG. 16 Integrated circuit 1402, one or more processors 1490, FIG. 17 Mobile devices 1502 FIG. 18 Head-mounted device 1602, FIG. 19 Wearable electronic devices 1702 FIG. 20 One or more processors 1890, wireless speakers and voice activation devices 1802, FIG. 21 Camera equipment 1902, FIG. 22 Head-mounted devices 2002 FIG. 24 Transportation 2102 FIG. 1 Transportation vehicle 2202 FIG. 13 The processor 2406, processor 2410, device 2400, one or more other circuits or components, or any combination thereof, configured to construct an audio scene graph based on the temporal sequence of audio events.
[0213] The device also includes components for assigning edge weights to the audio scene graph based on a similarity metric between audio events and the relationships between audio events. For example, the component for assigning edge weights could correspond to... FIG. 14 Audio scene graph updater 118, audio scene graph generator 140, system 100, FIG. 15 The audio scene diagram includes system 1356, second stage 1350, second power domain 1305, one or more processors 1390, device 1302, and system 1300. FIG. 16Integrated circuit 1402, one or more processors 1490, FIG. 17 Mobile devices 1502 FIG. 18 Head-mounted device 1602, FIG. 19 Wearable electronic devices 1702 FIG. 20 One or more processors 1890, wireless speakers and voice activation devices 1802, FIG. 21 Camera equipment 1902, FIG. 22 Head-mounted devices 2002 FIG. 24 Transportation 2102 FIG. 1 Transportation vehicle 2202 FIG. 13 FIG. 14 FIG. 15 FIG. 16 FIG. 17 FIG. 18 FIG. 19 FIG. 20 FIG. 21 FIG. 22 FIG. 24 FIG. 1 FIG. 13 FIG. 14 FIG. 15 FIG. 16 FIG. 17 FIG. 18 FIG. 19 FIG. 20 FIG. 21 FIG. 22 FIG. 24 FIG. 1 FIG. 13 FIG. 14 FIG. 15 FIG. 16 FIG. 17 FIG. 18 FIG. 19 FIG. 20 FIG. 21 FIG. 22 FIG. 24 FIG. 1 FIG. 13 FIG. 14 FIG. 15 FIG. 16 FIG. 17 FIG. 18 FIG. 19 FIG. 20 FIG. 21 FIG. 22 FIG. 24 FIG. 1 FIG. 13 FIG. 14 FIG. 15 FIG. 16 FIG. 17 FIG. 18 FIG. 19 FIG. 20 FIG. 21 FIG. 22 FIG. 24 FIG. 1 FIG. 13 FIG. 14 FIG. 15 FIG. 16 FIG. 17 FIG. 18 FIG. 19 FIG. 20 FIG. 21 FIG. 22 FIG. 24 FIG. 1 FIG. 13 FIG. 14 FIG. 15 FIG. 16 FIG. 17 FIG. 18 FIG. 19 FIG. 20 FIG. 21 FIG. 22 FIG. 24 FIG. 1 FIG. 13 FIG. 14 FIG. 15 FIG. 16 FIG. 17 FIG. 18 FIG. 19 FIG. 20 FIG. 21 FIG. 22 FIG. 24 FIG. 1 FIG. 13 FIG. 14 FIG. 15 FIG. 16 FIG. 17 FIG. 18 FIG. 19 FIG. 20 FIG. 21 FIG. 22 FIG. 24 FIG. 1 FIG. 13 FIG. 14 FIG. 15 FIG. 16 FIG. 17 FIG. 18 FIG. 19 FIG. 20 FIG. 21 FIG. 22 FIG. 24 FIG. 1 FIG. 13 FIG. 14 FIG. 15 FIG. 16 FIG. 17 FIG. 18 FIG. 19 FIG. 20 FIG. 21 FIG. 22 FIG. 24 FIG. 1 FIG. 13 FIG. 14 FIG. 15 FIG. 16 FIG. 17 FIG. 18 FIG. 19 FIG. 20 FIG. 21 FIG. 22 FIG. 24 FIG. 1 FIG. 13 FIG. 14 FIG. 15 FIG. 16 FIG. 17 FIG. 18 FIG. 19 FIG. 20 FIG. 21 FIG. The processors 2406, 2410, and 2400 are configured to assign edge weights to one or more other circuits or components, or any combination thereof, of the audio scene graph based on similarity measures between audio events and relationships between audio events.
[0214] In some implementations, a non-transitory computer-readable medium (e.g., a computer-readable storage device, such as memory 2486) includes instructions (e.g., instruction 2456) that, when executed by one or more processors (e.g., one or more processors 2410 or 2406), cause the one or more processors to identify audio segments (e.g., audio segment 112) of audio data (e.g., audio data 110) corresponding to audio events. The instructions also cause the one or more processors to assign tags (e.g., event tag 114) to the audio segments. The tags for specific audio segments describe the corresponding audio events. The instructions further cause the one or more processors to determine relationships between audio events based on knowledge data (e.g., knowledge data 122) (e.g., indicated by event-to-relationship data 152). The instructions also cause the one or more processors to construct an audio scene graph (e.g., audio scene graph 162) based on the temporal order of the audio events (e.g., audio segment temporal order 164). The instruction also causes the one or more processors to assign edge weights (e.g., edge weights 526) to the audio scene graph based on a similarity metric between audio events (e.g., total edge weight 528) and the relationship between the audio events.
[0215] Specific aspects of this disclosure are described below in a collection of related embodiments:
[0216] According to Embodiment 1, an apparatus includes: a memory configured to store knowledge data; and one or more processors coupled to the memory and configured to: identify audio segments of audio data corresponding to audio events; assign tags to the audio segments, wherein the tags of specific audio segments describe corresponding audio events; determine relationships between the audio events based on the knowledge data; construct an audio scene graph based on the temporal order of the audio events; and assign edge weights to the audio scene graph based on a similarity metric between the audio events and the relationships between the audio events.
[0217] Example 2 includes the device according to Example 1, wherein the one or more processors are further configured to: generate a first event representation of a first audio event in the audio events, wherein the audio scene graph is constructed to include a first node corresponding to the first audio event; generate a second event representation of a second audio event in the audio events, wherein the audio scene graph is constructed to include a second node corresponding to the second audio event; and assign a first edge weight to a first edge between the first node and the second node based on determining that the knowledge data at least indicates a first relationship between the first audio event and the second audio event, wherein the first edge weight is based on a first similarity measure associated with the first event representation and the second event representation.
[0218] Example 3 includes the device according to Example 1 or Example 2, wherein one or more processors are further configured to: determine a first audio embedding of a first audio segment in the audio segments, the first audio segment corresponding to the first audio event; and determine a first text embedding of a first tag in the tags, the first tag being assigned to the first audio segment, wherein the first event represents an event based on the first audio embedding and the first text embedding.
[0219] Example 4 includes the device according to Example 3, wherein one or more processors are configured to generate the first event representation based on the concatenation of the first audio embedding and the first text embedding.
[0220] Example 5 includes a device according to any one of Examples 2 to 4, wherein one or more processors are configured to determine the first similarity metric based on the cosine similarity between the first event representation and the second event representation.
[0221] Example 6 includes a device according to any one of Examples 2 to 5, wherein the one or more processors are further configured to determine the first edge weight based on determining the knowledge data indicating multiple relationships between the first audio event and the second audio event, and further based on a relationship similarity metric of the multiple relationships.
[0222] Example 7 includes a device according to any one of Examples 2 to 6, wherein one or more processors are further configured to: generate event-pair text embeddings of the first audio event and the second audio event based on determining the knowledge data indicating multiple relationships between the first audio event and the second audio event; generate event-pair text embeddings of the first audio event and the second audio event, wherein the event-pair text embeddings are based on a first text embedding of a first tag and a second text embedding of a second tag, wherein the first tag is assigned to a first audio segment corresponding to the first audio event, and wherein the second tag is assigned to a second audio segment corresponding to the second audio event; generate relation text embeddings of the multiple relationships; generate a relation similarity measure based on the event-pair text embeddings and the relation text embeddings; and further determine the first edge weight based on the relation similarity measure.
[0223] Example 8 includes the device according to Example 7, wherein one or more processors are configured to determine a first relation similarity metric for the first relation based on the event-pair text embedding and a first relation text embedding for the first relation, wherein the first edge weight is based on the ratio of the first relation similarity metric to the sum of the relation similarity metrics.
[0224] Example 9 includes the device according to Example 8, wherein one or more processors are configured to determine a first relation similarity measure based on the cosine similarity between the event-pair text embedding and the first relation text embedding.
[0225] Example 10 includes a device according to any one of Examples 1 to 9, wherein one or more processors are further configured to encode the audio scene graph to generate an encoded graph, and to use the encoded graph to perform one or more downstream tasks.
[0226] Example 11 includes a device according to any one of Examples 1 to 10, wherein one or more processors are configured to update the audio scene graph based on user input, video data, or both.
[0227] Example 12 includes a device according to any one of Examples 1 to 11, wherein one or more processors are configured to generate a graphical user interface (GUI) including a representation of the audio scene graph; provide the GUI to a display device; receive user input; and update the audio scene graph based on the user input.
[0228] Example 13 includes a device according to any one of Examples 1 to 12, wherein one or more processors are configured to detect visual relationships in video data associated with the audio data; and to update the audio scene map based on the visual relationships.
[0229] Example 14 includes the device according to Example 13 and also includes a camera configured to generate the video data.
[0230] Example 15 includes a device according to any one of Examples 1 to 14, wherein one or more processors are further configured to update the knowledge data in response to an update of the audio scene graph.
[0231] Example 16 includes a device according to any one of Examples 1 to 15, wherein one or more processors are further configured to update the similarity metric in response to an update of the audio scene graph.
[0232] Example 17 includes the device according to any one of Examples 1 to 16 and further includes a microphone configured to generate the audio data.
[0233] According to Embodiment 18, a method includes: receiving audio data at a first device; identifying audio segments of the audio data corresponding to audio events at the first device; assigning tags to the audio segments at the first device, wherein the tags of specific audio segments describe corresponding audio events; determining relationships between the audio events based on knowledge data; constructing an audio scene graph at the first device based on the temporal order of the audio events; assigning edge weights to the audio scene graph at the first device based on a similarity metric between the audio events and the relationships between the audio events; and providing a representation of the audio scene graph to a second device.
[0234] Example 19 includes the method according to Example 18, and further includes: generating a first event representation of a first audio event in the audio events, wherein the audio scene graph is constructed to include a first node corresponding to the first audio event; generating a second event representation of a second audio event in the audio events, wherein the audio scene graph is constructed to include a second node corresponding to the second audio event; and determining a first edge weight based on a first similarity metric associated with the first event representation and the second event representation, based on determining that the knowledge data at least indicates a first relationship between the first audio event and the second audio event, wherein the first edge weight is assigned to a first edge between the first node and the second node.
[0235] Example 20 includes the method according to Example 18 or Example 19, and further includes: determining a first audio embedding of a first audio segment in the audio segments, the first audio segment corresponding to the first audio event; and determining a first text embedding of a first tag in the tags, the first tag being assigned to the first audio segment, wherein the first event is represented based on the first audio embedding and the first text embedding.
[0236] Example 21 includes the method according to Example 20, wherein the first event represents a concatenation based on the first audio embedding and the first text embedding.
[0237] Example 22 includes the method according to any one of Examples 19 to 21, wherein the first similarity measure is based on the cosine similarity between the first event representation and the second event representation.
[0238] Example 23 includes the method according to any one of Examples 19 to 22, and further includes: determining the first edge weight based on a relation similarity metric of the plurality of relations between the first audio event and the second audio event, based on determining the knowledge data indicating a plurality of relations.
[0239] Example 24 includes the method according to any one of Examples 19 to 23, and further includes determining multiple relationships between the first audio event and the second audio event based on the knowledge data: generating event-pair text embeddings for the first audio event and the second audio event, wherein the event-pair text embeddings are based on a first text embedding of a first tag and a second text embedding of a second tag, wherein the first tag is assigned to a first audio segment corresponding to the first audio event, and wherein the second tag is assigned to a second audio segment corresponding to the second audio event; generating relation text embeddings for the multiple relationships; generating a relation similarity measure based on the event-pair text embeddings and the relation text embeddings; and further determining the first edge weight based on the relation similarity measure.
[0240] Example 25 includes the method according to Example 24, and further includes determining a first relation similarity measure of the first relation based on the event pair text embedding and the first relation text embedding of the first relation, wherein the first edge weight is based on the ratio of the first relation similarity measure to the sum of the relation similarity measures.
[0241] Example 26 includes the method according to Example 25, wherein the first relation similarity measure is based on the cosine similarity between the event-pair text embedding and the first relation text embedding.
[0242] Example 27 includes the method according to any one of Examples 18 to 26, and further includes: encoding the audio scene graph to generate an encoded graph, and using the encoded graph to perform one or more downstream tasks.
[0243] Example 28 includes the method according to any one of Examples 18 to 27, and further includes updating the audio scene graph based on user input, video data, or both.
[0244] Example 29 includes the method according to any one of Examples 18 to 28, and further includes: generating a graphical user interface (GUI) including a representation of the audio scene graph; providing the GUI to a display device; receiving user input; and updating the audio scene graph based on the user input.
[0245] Example 30 includes the method according to any one of Examples 18 to 29, and further includes: detecting visual relationships in video data associated with the audio data; and updating the audio scene graph based on the visual relationships.
[0246] Example 31 includes the method according to Example 30, and further includes receiving the video data from a camera.
[0247] Example 32 includes the method according to any one of Examples 18 to 31, and further includes updating the knowledge data in response to an update of the audio scene graph.
[0248] Example 33 includes the method according to any one of Examples 18 to 32, and further includes updating the similarity metric in response to an update of the audio scene graph.
[0249] Example 34 includes the method according to any one of Examples 18 to 33, and further includes receiving the audio data from a microphone.
[0250] According to embodiment 35, an apparatus includes: a memory configured to store instructions; and a processor configured to execute the instructions to perform a method according to any one of embodiments 18 to 34.
[0251] According to Embodiment 36, a non-transitory computer-readable medium storage instruction, when executed by a processor, causes the processor to perform the method according to any one of Embodiments 18 to 34.
[0252] According to embodiment 37, an apparatus includes components for performing the method according to any one of embodiments 18 to 34.
[0253] According to embodiment 38, a non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to: identify audio segments of audio data corresponding to audio events; assign labels to the audio segments, wherein the labels of specific audio segments describe corresponding audio events; determine relationships between the audio events based on knowledge data; construct an audio scene graph based on the temporal order of the audio events; and assign edge weights to the audio scene graph based on a similarity metric between the audio events and the relationships between the audio events.
[0254] Example 39 includes a non-transitory computer-readable medium according to Example 38, wherein the instructions, when executed by the one or more processors, further cause the one or more processors to: encode the audio scene graph to generate an encoded graph, and use the encoded graph to perform one or more downstream tasks.
[0255] According to embodiment 40, an apparatus includes: components for identifying audio segments of audio data corresponding to audio events; components for assigning labels to the audio segments, wherein the labels of specific audio segments describe corresponding audio events; components for determining relationships between the audio events based on knowledge data; components for constructing an audio scene graph based on the temporal order of the audio events; and components for assigning edge weights to the audio scene graph based on a similarity metric between the audio events and the relationships between the audio events.
[0256] Example 41 includes the apparatus according to Example 40, wherein at least one of the components for identifying the audio segment, the components for assigning the tag, the components for determining the relationship, the components for constructing the audio scene graph, and the components for assigning the edge weights are integrated into at least one of a computer, mobile phone, communication device, vehicle, head-mounted device, or extended reality (XR) device.
[0257] Those skilled in the art will also understand that the various exemplary logic blocks, configurations, modules, circuits, and algorithm steps described in connection with the specific embodiments disclosed herein can be implemented as electronic hardware, computer software executed by a processor, or a combination of both. The various exemplary components, blocks, configurations, modules, circuits, and steps have been generally described above in terms of their functionality. Whether such functionality is implemented as hardware or processor-executable instructions depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each specific application, and such implementation decisions shall not be construed as departing from the scope of this disclosure.
[0258] The steps of the methods or algorithms described in conjunction with the specific embodiments disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of both. The software module may reside in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, removable disks, compressed optical disc read-only memory (CD-ROM), or any other form of non-transitory storage medium known in the art. An exemplary storage medium is coupled to a processor such that the processor can read information from and write information to the storage medium. Alternatively, the storage medium may be integral with the processor. The processor and storage medium may reside in an application-specific integrated circuit (ASIC). The ASIC may reside in a computing device or a user terminal. Alternatively, the processor and storage medium may reside as discrete components in a computing device or a user terminal.
[0259] The prior description of the disclosed aspects is provided to enable those skilled in the art to make or use the disclosed aspects. Various modifications to these aspects will be apparent to those skilled in the art, and the principles defined herein can be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but should be granted the broadest scope that may be consistent with the principles and novel features as defined by the following claims.
Claims
1. A device comprising: a memory configured to store knowledge data; and one or more processors coupled to the memory and configured to: obtain a first audio embedding for a first audio segment of audio data, the first audio segment corresponding to a first one of audio events; obtain a first text embedding for a first label assigned to the first audio segment; obtain a first event representation for the first audio event, the first event representation based on a combination of the first audio embedding and the first text embedding; obtain a second event representation for a second one of the audio events; determine relationships between the audio events based on the knowledge data; and construct an audio scene graph based on a temporal order of the audio events, the audio scene graph constructed to include a first node corresponding to the first audio event and a second node corresponding to the second audio event.
2. The device of claim 1, wherein the one or more processors are configured to: obtain audio segments of audio data identified as corresponding to the audio events, the audio segments including the first audio segment and a second audio segment, wherein the second audio segment corresponds to the second audio event; and obtain labels assigned to the audio segments, a label for a particular audio segment describing a corresponding audio event, wherein the labels include the first label.
3. The device of claim 1, wherein the one or more processors are configured to obtain edge weights assigned to the audio scene graph based on similarity measures between the audio events and the relationships between the audio events.
4. The device of claim 1, wherein the one or more processors are configured to assign a first edge weight to a first edge between the first node and the second node based on determining that the knowledge data indicates at least a first relationship between the first audio event and the second audio event, wherein the first edge weight is based on a first similarity measure associated with the first event representation and the second event representation.
5. The device of claim 4, wherein the one or more processors are configured to determine the first similarity measure based on a cosine similarity between the first event representation and the second event representation.
6. The device of claim 4, wherein the one or more processors are configured to determine the first edge weight further based on relationship similarity measures for a plurality of relationships between the first audio event and the second audio event based on determining that the knowledge data indicates the plurality of relationships.
7. The device of claim 4, wherein the one or more processors are configured to determine the first edge weight based on determining that the knowledge data indicates a plurality of relationships between the first audio event and the second audio event: generating an event pair text embedding of the first audio event and the second audio event, wherein the event pair text embedding is based on the first text embedding and a second text embedding of a second label, wherein the second label is assigned to a second audio segment corresponding to the second audio event; generating a relationship text embedding of the plurality of relationships; generating a relationship similarity measure based on the event pair text embedding and the relationship text embedding; and determining the first edge weight further based on the relationship similarity measure.
8. The device of claim 7, wherein the one or more processors are configured to determine a first relationship similarity measure of the first relationship based on the event pair text embedding and a first relationship text embedding of the first relationship, wherein the first edge weight is based on a ratio of the first relationship similarity measure to a sum of the relationship similarity measures.
9. The device of claim 8, wherein the one or more processors are configured to determine the first relationship similarity measure based on a cosine similarity between the event pair text embedding and the first relationship text embedding.
10. The device of claim 4, wherein the one or more processors are configured to update the first similarity measure in response to an update of the audio scene graph.
11. The device of claim 1, wherein the one or more processors are configured to encode the audio scene graph to generate an encoded graph, and use the encoded graph to perform one or more downstream tasks.
12. The device of claim 1, wherein the one or more processors are configured to update the audio scene graph based on user input, video data, or both.
13. The device of claim 1, wherein the one or more processors are configured to: generate a graphical user interface (GUI) that includes a representation of the audio scene graph; provide the GUI to a display device; receive user input; and update the audio scene graph based on the user input.
14. The device of claim 1, wherein the one or more processors are configured to: detect a visual relationship in video data, the video data being associated with the audio data; and update the audio scene graph based on the visual relationship.
15. The device of claim 14, further comprising a camera configured to generate the video data.
16. The device of claim 1, wherein the one or more processors are configured to update the knowledge data in response to an update of the audio scene graph.
17. The device of claim 1, further comprising a microphone configured to generate the audio data.
18. A method comprising: obtaining, at a first device, a first audio embedding of a first audio segment of audio data, the first audio segment corresponding to a first audio event of audio events; obtaining, at the first device, a first text embedding of a first label assigned to the first audio segment; obtaining, at the first device, a first event representation of the first audio event, the first event representation based on a combination of the first audio embedding and the first text embedding; obtaining, at the first device, a second event representation of a second audio event of the audio events; determining relationships between the audio events based on knowledge data; constructing, at the first device, an audio scene graph based on a temporal order of the audio events, the audio scene graph constructed to include a first node corresponding to the first audio event and a second node corresponding to the second audio event; and providing, to a second device, a representation of the audio scene graph.
19. The method of claim 18, further comprising determining, based on a determination that the knowledge data indicates at least a first relationship between the first audio event and the second audio event, a first edge weight based on a first similarity measure associated with the first event representation and the second event representation, wherein the first edge weight is assigned to a first edge between the first node and the second node.
20. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to: obtain a first audio embedding of a first audio segment of audio data, the first audio segment corresponding to a first audio event of audio events; obtain a first text embedding of a first label assigned to the first audio segment; obtain a first event representation of the first audio event, the first event representation based on a combination of the first audio embedding and the first text embedding; obtain a second event representation of a second audio event of the audio events; determine relationships between the audio events based on knowledge data; and construct an audio scene graph based on a temporal order of the audio events, the audio scene graph constructed to include a first node corresponding to the first audio event and a second node corresponding to the second audio event.