Processing network-graph embeddings for audio-processing operations
The AGrapE technique transforms audio analysis into a graph-representation problem using GNNs and GraphSAGE, providing an efficient solution for edge devices by reducing resource consumption and enabling complex audio applications.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-11-18
- Publication Date
- 2026-03-12
AI Technical Summary
Existing RNN-based audio-processing models require significant computational resources and memory, posing a bottleneck for deploying sophisticated audio applications on edge devices with limited hardware resources.
Utilizing a Graph Neural Network (GNN) approach, AGrapE, to model audio frames as nodes in a network graph, leveraging graph-processing models like GraphSAGE to generate refined feature embeddings, which are then processed by lightweight audio-operation models for efficient audio processing.
This method reduces computational and memory demands, enabling advanced audio applications like keyword detection and noise cancellation on edge devices without performance compromise.
Smart Images

Figure US2024056419_12032026_PF_FP_ABST
Abstract
Description
PATENT Qualcomm Ref. No. 2407274WOPROCESSING NETWORK-GRAPH EMBEDDINGS FOR AUDIO-PROCESSING OPERATIONSFIELD
[0001] The present disclosure generally relates to generating network graphs that represent sequences of audio frames. For example, aspects of the present disclosure relate to systems and techniques for processing audio-frame network graphs using a graph-processing machine-learning model to generate refined feature embeddings, in which the refined feature embeddings can be used for certain audio-processing operations.BACKGROUND
[0002] Audio processing can involve analyzing, modifying, and synthesizing audio signals to perform various types of operations. The audio-processing operations can encompass various techniques, including digital signal processing (DSP) to improve clarity or remove unwanted noise from the audio data. In some cases, feature-extraction techniques can be used to transform raw audio signals into representations that can be processed for speech recognition, music classification, or audio-event detection. In recent years, machine learning has significantly advanced audio processing, enabling more sophisticated applications such as automatic speech recognition, emotion detection, and real-time audio enhancement. These audio applications rely on machine-learning models to identify patterns in the audio data and make predictions based on the audio data, thereby expanding the possibilities in audio processing across various industries, from consumer electronics to healthcare.
[0003] An artificial neural network is an example type of a machine-learning model that attempts to replicate, using computer technology, logical reasoning performed by the biological neural networks that constitute animal brains. Deep neural networks, such as convolutional neural networks, are widely used for numerous applications, such as object detection, object classification, object tracking, big data analysis, among others.1 Polsinelli Ref. No. 094922-825835PATENTQualcomm Ref. No. 2407274WOSUMMARY
[0004] The following presents a simplified summary relating to one or more aspects disclosed herein. Thus, the following summary should not be considered an extensive overview relating to all contemplated aspects, nor should the following summary be considered to identify key or critical elements relating to all contemplated aspects or to delineate the scope associated with any particular aspect. Accordingly, the following summary has the sole purpose to present certain concepts relating to one or more aspects relating to the mechanisms disclosed herein in a simplified form to precede the detailed description presented below.
[0005] Systems and techniques are described herein for processing network-graph embeddings for audio-processing operations. According to some examples, an apparatus for processing network-graph embeddings for audio-processing operations is provided. The apparatus includes one or more memories configured to store audio data, the audio data including a sequence of audio frames; and one or more processors coupled to the one or more memories and configured to: process the audio data to determine a plurality of features for the sequence of audio frames; process the plurality of features for the sequence of audio frames using a graph-processing machinelearning model to encode a plurality of nodes of a network graph representing the sequence of audio frames into a set of refined feature embeddings; process the set of refined feature embeddings using an audio-operation machine-learning model to generate a set of output values associated with one or more audio-processing operations, wherein an output value of the set of output values is associated with one or more audio frames of the sequence of audio frames; and output the set of output values for the one or more audio-processing operations.
[0006] In another illustrative example, a method for processing network-graph embeddings for audio-processing operations is provided. The method includes: accessing audio data including a sequence of audio frames; processing the audio data to determine a plurality of features for the sequence of audio frames; processing the plurality of features for the sequence of audio frames using a graph-processing machine-learning model to encode a plurality of nodes of a network graph representing the sequence of audio frames into a set of refined feature embeddings; processing the set of refined feature embeddings using an audio-operation machine-learning model to generate a set of output values associated with one or more audio-processing operations, wherein2 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO an output value of the set of output values is associated with one or more audio frames of the sequence of audio frames; and outputting the set of output values for the one or more audioprocessing operations.
[0007] In another illustrative example, a non-transitory computer-readable medium is provided that includes instructions that, when executed by at least one processor, cause the at least one processor to: access audio data including a sequence of audio frames; process the audio data to determine a plurality of features for the sequence of audio frames; process the plurality of features for the sequence of audio frames using a graph-processing machine-learning model to encode a plurality of nodes of a network graph representing the sequence of audio frames into a set of refined feature embeddings; process the set of refined feature embeddings using an audio-operation machine-learning model to generate a set of output values associated with one or more audioprocessing operations, wherein an output value of the set of output values is associated with one or more audio frames of the sequence of audio frames; and output the set of output values for the one or more audio-processing operations.
[0008] In another illustrative example, an apparatus for processing network-graph embeddings for audio-processing operations is provided. The apparatus includes: means for accessing audio data including a sequence of audio frames; means for processing the audio data to determine a plurality of features for the sequence of audio frames; means for processing the plurality of features for the sequence of audio frames using a graph-processing machine-learning model to encode a plurality of nodes of a network graph representing the sequence of audio frames into a set of refined feature embeddings; means for processing the set of refined feature embeddings using an audiooperation machine-learning model to generate a set of output values associated with one or more audio-processing operations, wherein an output value of the set of output values is associated with one or more audio frames of the sequence of audio frames; and means for outputting the set of output values for the one or more audio-processing operations.
[0009] In some aspects, one or more of apparatuses described herein include a mobile device (e.g., a mobile telephone or so-called “smart phone” or other mobile device), a wireless communication device, a vehicle or a computing device, system, or component of the vehicle, an extended reality device (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a3 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO mixed reality (MR) device), a wearable device, a personal computer, a laptop computer, a server computer, a camera, or other device. In some aspects, the one or more processors include an image signal processor (ISP). In some aspects, the apparatus includes a camera or multiple cameras for capturing one or more images. In some aspects, the apparatus includes an image sensor that captures the image data. In some aspects, the apparatus further includes a display for displaying the image, one or more notifications (e.g., associated with processing of the image), and / or other displayable data.
[0010] Aspects generally include a method, apparatus, system, computer program product, non- transitory computer-readable medium, user device, user equipment, wireless communication device, and / or processing system as substantially described with reference to and as illustrated by the drawings and specification.
[0011] Some aspects include a device having a processor configured to perform one or more operations of any of the methods summarized above. Further aspects include processing devices for use in a device configured with processor-executable instructions to perform operations of any of the methods summarized above. Further aspects include a non-transitory processor-readable storage medium having stored thereon processor-executable instructions configured to cause a processor of a device to perform operations of any of the methods summarized above. Further aspects include a device having means for performing functions of any of the methods summarized above.
[0012] The foregoing has outlined rather broadly the features and technical advantages of examples according to the disclosure in order that the detailed description that follows may be better understood. Additional features and advantages will be described hereinafter. The conception and specific examples disclosed may be readily utilized as a basis for modifying or designing other structures for carrying out the same purposes of the present disclosure. Such equivalent constructions do not depart from the scope of the appended claims. Characteristics of the concepts disclosed herein, both their organization and method of operation, together with associated advantages will be better understood from the following description when considered in connection with the accompanying figures. Each of the figures is provided for the purposes of illustration and description, and not as a definition of the limits of the claims. The foregoing,4 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO together with other features and aspects, will become more apparent upon referring to the following specification, claims, and accompanying drawings.
[0013] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in isolation to determine the scope of the claimed subject matter. The subject matter should be understood by reference to appropriate portions of the entire specification of this patent, any or all drawings, and each claim.BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The accompanying drawings are presented to aid in the description of various aspects of the disclosure and are provided solely for illustration of the aspects and not limitation thereof. So that the above-recited features of the present disclosure can be understood in detail, a more particular description, briefly summarized above, may be had by reference to aspects, some of which are illustrated in the appended drawings. It is to be noted, however, that the appended drawings illustrate only certain typical aspects of this disclosure and are therefore not to be considered limiting of its scope, for the description may admit to other equally effective aspects. The same reference numbers in different drawings may identify the same or similar elements.
[0015] FIG. 1 illustrates an example implementation of a system-on-a-chip (SoC), in accordance with some examples;
[0016] FIG. 2 illustrates an example schematic diagram fortraining a graph-processing machinelearning model and an audio-processing model for performing audio-processing operations, in accordance with some examples;
[0017] FIG. 3 illustrates an example process for generating network graphs from the plurality of features, in accordance with some examples;
[0018] FIG. 4 illustrates an example representation of the GraphSAGE algorithm configured to process network graphs, in accordance with some examples;
[0019] FIG. 5 is a flowchart illustrating an example process for processing network-graph embeddings for audio-processing operations, in accordance with aspects of the present disclosure;5 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO
[0020] FIG. 6 illustrates an example schematic diagram for processing network-graph embeddings to detect keywords from audio data, in accordance with some examples;
[0021] FIG. 7 illustrates an example schematic diagram for processing network-graph embeddings to perform in-ear detection, in accordance with some examples;
[0022] FIG. 8 illustrates an example schematic diagram for processing network-graph embeddings to identify metadata associated with audio data, in accordance with some examples;
[0023] FIG. 9 illustrates an example schematic diagram for processing network-graph embeddings to perform active-noise cancellation, in accordance with some examples;
[0024] FIG. 10 illustrates another example schematic diagram for training a supervised graphprocessing machine-learning model performing audio-processing operations, in accordance with some examples;
[0025] FIG. 11 is a block diagram illustrating an example of a deep learning network, in accordance with some examples;
[0026] FIG. 12 is a block diagram illustrating an example of a convolutional neural network, in accordance with some examples; and
[0027] FIG. 13 is a diagram illustrating an example system architecture for implementing certain aspects described herein.DETAILED DESCRIPTION
[0028] Certain aspects and examples of this disclosure are provided below. Some of these aspects and examples may be applied independently and some of them may be applied in combination as would be apparent to those of skill in the art. In the following description, for the purposes of explanation, specific details are set forth in order to provide a thorough understanding of aspects and examples of the disclosure. However, it will be apparent that various aspects and examples may be practiced without these specific details. The figures and description are not intended to be restrictive.6 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO
[0029] The ensuing description provides exemplary aspects and examples only, and is not intended to limit the scope, applicability, or configuration of the disclosure. Rather, the ensuing description of the exemplary aspects and examples will provide those skilled in the art with an enabling description for implementing aspects and examples of the disclosure. It should be understood that various changes may be made in the function and arrangement of elements without departing from the scope of the application as set forth in the appended claims.
[0030] Applying artificial intelligence and machine-learning techniques to audio applications can often involve processing sequences of audio frames. The sequential data structure of audio data can lend itself to the use of recurrent neural networks (RNNs), such as Gated Recurrent Units (GRUs) and Long Short-Term Memory networks (LSTMs). These machine-learning architectures can be used in existing audio-processing systems because they can incorporate temporal dependencies by using hidden states to retain information from previous inputs. However, RNN- based models typically require a significant demand for memory and computational resources to maintain hidden states and handle the recurrent computations. This requirement can pose various challenges, particularly for implementing audio processing on edge devices (e g., True Wireless Stereo (TWS) earbuds), in which hardware resources are often limited.
[0031] Using RNN-based models can thus become a bottleneck for deploying sophisticated machine-learning solutions on edge devices. Since many audio-operation machine-learning models rely heavily on the recurrent architectures, the problem is pervasive across the industry. To effectively utilize audio-processing machine-learning frameworks on various computing and edge devices, it is desirable to use computationally efficient solutions. Reducing consumption of computing resources in audio processing can pave the way for deploying more complex and efficient ML-DSP (Digital Signal Processing) hybrid solutions on edge devices. As a result, deploying complex audio processing in edge devices can enable advanced audio applications without compromising on performance or battery life of the edge devices. Accordingly, identifying optimizations that can reduce the computational burden of these machine-learning models without sacrificing accuracy is challenging in the field of on-edge audio machine-learning space.
[0032] Systems, apparatuses, processes (also referred to as methods), and computer-readable media (collectively referred to as “systems and techniques”) are described herein for generating7 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO network-graph embeddings that represent sequences of audio frames. These refined feature embeddings can be processed using machine-learning models (e.g., an audio-operation machinelearning model) to generate outputs that can be utilized for audio-processing operation.
[0033] The systems and techniques can utilize an Audio frames relational Graph neural network for on-Edge Audio Applications (AGrapE) to identify various audio-processing operations based on analyzing sequences of audio frames. In contrast to existing RNN-based methods that heavily depend on maintaining hidden states and recurrent connections, AGrapE can use graph neural networks (GNNs) to model relationships between audio frames. By modeling audio frames as graph nodes and their relationships through edges, AGrapE can circumvent the need for computationally expensive recurrent connections, making the present techniques more suitable for edge devices with limited resources.
[0034] In some cases, each graph node of the network graph can represent an audio frame of a sequence of audio frames. The node can include initial node embeddings derived from a plurality of features associated with the corresponding audio frame. The features can include digital signal processing (DSP) features such as RMS energy, spectral bandwidth, zero-crossing rate, and other extracted features during audio pre-processing. The feature-based representation in the network graph can provide a robust foundation for the graph-based learning process, capturing various audio characteristics in a compact format. The edges between nodes can be generated based on distance metrics, in which the distance metric is indicative of a degree of similarity between two nodes. For instance, the distance metric can include cosine similarity, Euclidean distance, or any other suitable metrics to quantify the relationship between audio frames.
[0035] In some cases, when an edge is formed between the nodes, an edge weight can be assigned to the edge. The edge weight can be processed by the graph-processing machine-learning model to encode the plurality of nodes of a network graph. For example, the edge weight can be the same distance metric calculated between the nodes. In another example, the edge weight can be another distance metric calculated based on one or more other distance metrics associated with other edges of the network graph. In yet another example, the edge weight can be recalculated to be another type of metric that is different from the distance metric. Alternatively, no edge weight is assigned to the edge once it is formed. The edges can thus encode similarity or dissimilarity8 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO between audio frames, effectively transforming the sequential audio data into a relational graph structure.
[0036] Once the graph is constructed, AGrapE can process the network graph using a graphprocessing machine-learning model, such as an unsupervised GraphSAGE (Graph Sample and Aggregated) algorithm. The parameters of the graph-processing machine-learning model can be learned to generate refined feature embeddings for each node (audio frame) in the network graph. The graph-processing machine-learning model can be trained to aggregate features from a given node's local neighborhood, thereby enabling efficient learning on large network graphs by sampling a fixed-size neighborhood for each node. This makes it highly suitable for audio applications on edge devices, where memory and computational resources are constrained. By leveraging the graph-processing machine-learning models, AGrapE can generate refined feature embeddings that capture relational information between audio frames without the need for large memory footprints or high computational overheads typically associated with recurrent neural networks.
[0037] AGrapE can then process the refined feature embeddings using audio-operation machinelearning models. The audio-operation machine-learning models can be a Logistic Regression or Linear Regression model that can process the refined feature embeddings to perform downstream tasks like classification or regression. The audio-operation machine-learning models can be configured to be lightweight and can run efficiently on edge devices, further reducing computational requirements while maintaining high performance. The choice of the audiooperation machine-learning models can align with the goal of optimizing for resource-limited environments, ensuring that the present techniques are deployable on a wide range of devices, from TWS earbuds to other low-power, on-edge hardware devices.
[0038] By reframing audio analysis as a graph-representation problem, the present techniques can provide a scalable and efficient solution for on-edge audio applications. The present techniques can leverage the power of network graphs to handle sequential data efficiently while avoiding the pitfalls of existing RNN-based techniques. As a result, the present techniques can offer a path forward for deploying more sophisticated machine-learning audio operations on edge devices (e.g.,9 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO keyword detection, in-ear detection, noise cancellation), empowering advanced audio applications without the constraints of memory and processing power limitations.
[0039] Various aspects of the present disclosure will be described with respect to the figures.
[0040] FIG. 1 illustrates an example implementation of a system-on-a-chip (SOC) 100, which may include a central processing unit (CPU) 102 or a multi-core CPU, configured to perform one or more of the functions described herein. Parameters or variables (e.g., neural signals and synaptic weights), system parameters associated with a computational device (e.g., neural network with weights), delays, frequency bin information, task information, among other information may be stored in a memory block associated with a neural processing unit (NPU) 108, in a memory block associated with a CPU 102, in a memory block associated with a graphics processing unit (GPU) 104, in a memory block associated with a digital signal processor (DSP) 106, in a memory block 118, and / or may be distributed across multiple blocks. Instructions executed at the CPU 102 may be loaded from a program memory associated with the CPU 102 or may be loaded from a memory block 118.
[0041] The SOC 100 may also include additional processing blocks tailored to specific functions, such as a GPU 104, a DSP 106, a connectivity block 110, which may include fifth generation (5G) connectivity, fourth generation long term evolution (4G LTE) connectivity, WiFi connectivity, USB connectivity, Bluetooth connectivity, and the like, and a multimedia processor 112 that may, for example, detect and recognize gestures. In one implementation, the NPU is implemented in the CPU 102, DSP 106, and / or GPU 104. The SOC 100 may also include a sensor processor 114, image signal processors (ISPs) 116, and / or storage 120.
[0042] The SOC 100 may be based on an ARM instruction set. In an aspect of the present disclosure, the instructions loaded into the CPU 102 may comprise code to search for a stored multiplication result in a lookup table (LUT) corresponding to a multiplication product of an input value and a filter weight. The instructions loaded into the CPU 102 may also comprise code to disable a multiplier during a multiplication operation of the multiplication product when a lookup table hit of the multiplication product is detected. In addition, the instructions loaded into the CPU10 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO102 may comprise code to store a computed multiplication product of the input value and the filter weight when a lookup table miss of the multiplication product is detected.
[0043] SOC 100 and / or components thereof may be configured to perform image processing using machine learning techniques according to aspects of the present disclosure discussed herein. For example, SOC 100 and / or components thereof may be configured to perform depth completion according to aspects of the present disclosure. In some cases, by using a graph-based neural network with a segmentation input and a depth input each associated with a same image, aspects of the present disclosure can increase the accuracy and efficiency of generating dense depth maps from an image input and a sparse depth input.
[0044] SOC 100 can be part of a computing device or multiple computing devices. In some examples, SOC 100 can be part of an electronic device (or devices) such as a camera system (e.g., a digital camera, an IP camera, a video camera, a security camera, etc.), a telephone system (e.g., a smartphone, a cellular telephone, a conferencing system, etc.), a desktop computer, an XR device (e.g., a head-mounted display, etc.), a smart wearable device (e.g., a smart watch, smart glasses, etc.), a laptop or notebook computer, a tablet computer, a set-top box, a television, a display device, a system-on-chip (SoC), a digital media player, a gaming console, a video streaming device, a server, a drone, a computer in a car, an Internet-of-Things (loT) device, or any other suitable electronic device(s).
[0045] In some implementations, the CPU 102, the GPU 104, the DSP 106, the NPU 108, the connectivity block 110, the multimedia processor 112, the one or more sensors 114, the ISPs 116, the memory block 118 and / or the storage 120 can be part of the same computing device. For example, in some cases, the CPU 102, the GPU 104, the DSP 106, the NPU 108, the connectivity block 110, the multimedia processor 112, the one or more sensors 114, the ISPs 116, the memory block 118 and / or the storage 120 can be integrated into a smartphone, laptop, tablet computer, smart wearable device, video gaming system, server, and / or any other computing device. In other implementations, the CPU 102, the GPU 104, the DSP 106, the NPU 108, the connectivity block 110, the multimedia processor 112, the one or more sensors 114, the ISPs 116, the memory block 118 and / or the storage 120 can be part of two or more separate computing devices.11 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO
[0046] Machine learning (ML) can be considered a subset of artificial intelligence (Al). ML systems can include algorithms and statistical models that computer systems can use to perform various tasks by relying on patterns and inference, without the use of explicit instructions. One example of a ML system is a neural network (also referred to as an artificial neural network), which may include an interconnected group of artificial neurons (e.g., neuron models). Neural networks may be used for various applications and / or devices, such as image and / or video coding, image analysis and / or computer vision applications, Internet Protocol (IP) cameras, Internet of Things (loT) devices, vehicles, service robots, among others.
[0047] Individual nodes in a neural network may emulate biological neurons by taking input data and performing simple operations on the data. The results of the simple operations performed on the input data are selectively passed on to other neurons. Weight values are associated with each vector and node in the network, and these values constrain how input data is related to output data. For example, the input data of each node may be multiplied by a corresponding weight value, and the products may be summed. The sum of the products may be adjusted by an optional bias, and an activation function may be applied to the result, yielding the node’ s output signal or “output activation” (sometimes referred to as a feature map or an activation map). The weight values may initially be determined by an iterative flow of training data through the network (e.g., weight values are established during a training phase in which the network learns how to identify particular classes by their typical input data characteristics).
[0048] Different types of neural networks exist, such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), multilayer perceptron (MLP) neural networks, transformer neural networks, among others. For instance, convolutional neural networks (CNNs) are a type of feed-forward artificial neural network. Convolutional neural networks may include collections of artificial neurons that each have a receptive field (e.g., a spatially localized region of an input space) and that collectively tile an input space. RNNs work on the principle of saving the output of a layer and feeding this output back to the input to help in predicting an outcome of the layer. A GAN is a form of generative neural network that can learn patterns in input data so that the neural network model can generate new synthetic outputs that reasonably could have been from the original dataset. A GAN can include two neural networks12 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO that operate together, including a generative neural network that generates a synthesized output and a discriminative neural network that evaluates the output for authenticity. In MLP neural networks, data may be fed into an input layer, and one or more hidden layers provide levels of abstraction to the data. Predictions may then be made on an output layer based on the abstracted data.
[0049] Deep learning (DL) is one example of a machine learning technique and can be considered a subset of ML. Many DL approaches are based on a neural network, such as an RNN or a CNN, and utilize multiple layers. The use of multiple layers in deep neural networks can permit progressively higher-level features to be extracted from a given input of raw data. For example, the output of a first layer of artificial neurons becomes an input to a second layer of artificial neurons, the output of a second layer of artificial neurons becomes an input to a third layer of artificial neurons, and so on. Layers that are located between the input and output of the overall deep neural network are often referred to as hidden layers. The hidden layers learn (e.g., are trained) to transform an intermediate input from a preceding layer into a slightly more abstract and composite representation that can be provided to a subsequent layer, until a final or desired representation is obtained as the final output of the deep neural network.
[0050] As noted above, a neural network is an example of a machine learning system, and can include an input layer, one or more hidden layers, and an output layer. Data is provided from input nodes of the input layer, processing is performed by hidden nodes of the one or more hidden layers, and an output is produced through output nodes of the output layer. Deep learning networks typically include multiple hidden layers. Each layer of the neural network can include feature maps or activation maps that can include artificial neurons (or nodes). A feature map can include a filter, a kernel, or the like. The nodes can include one or more weights used to indicate an importance of the nodes of one or more of the layers. In some cases, a deep learning network can have a series of many hidden layers, with early layers being used to determine simple and low-level characteristics of an input, and later layers building up a hierarchy of more complex and abstract characteristics.
[0051] A deep learning architecture may learn a hierarchy of features. If presented with visual data, for example, the first layer may learn to recognize relatively simple features, such as edges,13 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO in the input stream. In another example, if presented with auditory data, the first layer may learn to recognize spectral power in specific frequencies. The second layer, taking the output of the first layer as input, may learn to recognize combinations of features, such as simple shapes for visual data or combinations of sounds for auditory data. For instance, higher layers may learn to represent complex shapes in visual data or words in auditory data. Still higher layers may learn to recognize common visual objects or spoken phrases. Deep learning architectures may perform especially well when applied to problems that have a natural hierarchical structure. For example, the classification of motorized vehicles may benefit from first learning to recognize wheels, windshields, and other features. These features may be combined at higher layers in different ways to recognize cars, trucks, and airplanes.
[0052] Neural networks may be designed with a variety of connectivity patterns. In feed-forward networks, information is passed from lower to higher layers, with each neuron in a given layer communicating to neurons in higher layers. A hierarchical representation may be built up in successive layers of a feed-forward network, as described above. Neural networks may also have recurrent or feedback (also called top-down) connections. In a recurrent connection, the output from a neuron in a given layer may be communicated to another neuron in the same layer. A recurrent architecture may be helpful in recognizing patterns that span more than one of the input data chunks that are delivered to the neural network in a sequence. A connection from a neuron in a given layer to a neuron in a lower layer is called a feedback (or top-down) connection. A network with many feedback connections may be helpful when the recognition of a high-level concept may aid in discriminating the particular low-level features of an input.
[0053] As previously noted, systems and techniques are described herein for generating network-graph embeddings that represent sequences of audio frames. FIG. 2 illustrates an example schematic diagram 200 for training a graph-processing machine-learning model and an audioprocessing model for performing audio-processing operations, in accordance with some examples.
[0054] To initiate training of the graph-processing machine-learning model and the audioprocessing model, a training system can access audio data 202. The audio data 202 can be accessed using various techniques. For example, at least one microphone (e g., a single microphone or multiple microphones) of a computing device can be used to capture the audio data 202. In addition14 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO to microphones, specialized sensors such as hydrophones for underwater audio capture or contact microphones for picking up vibrations from solid surfaces can be used to capture the audio data 202. In some cases, the audio data 202 can be extracted from pre-existing audio recordings, such as those stored in digital files (e.g., MP3, WAV). Additionally or alternatively, an edge device can capture the audio data 202 from one or more streaming sources, in which the audio data can be captured by intercepting and storing portions of streaming data being transmitted in real-time over the Internet.
[0055] The audio data 202 can include a sequence of audio frames. For example, the sequence can include 1, 2, 3, 4, 5, 10, 12, 15, 30, or more than 30 audio frames, in which each audio frame can include time or sequence data that can indicate a position of the audio frame relative to other audio frames of the sequence of audio frames. In some cases, the audio data 202 can include a plurality of sequences of audio frames, in which each sequence can include a predetermined number of audio frames. An audio frame can include a set of samples that represent an audio signal at a particular time point. Each sample within the audio frame can be a digital representation of an analog audio signal, in which a resolution of the sample can be determined by its corresponding bit depth (e.g., 16-bit depth).
[0056] In some cases, the audio frame can be associated with a particular time duration. The time duration of the audio frame can be determined by a sample rate of the audio data, which is the number of samples taken per second, measured in Hertz (Hz). The duration of the audio frame can be calculated using the formula: Frame Duration = Number of samples per frame / Sample Rate. As an illustrative example, with a sample rate of 44.1 kHz (44,100 samples per second) and 1024 samples in an audio frame, the duration of the audio frame of the sequence is approximately 0.0232 seconds, or 23.2 milliseconds (ms). Example duration of an audio frame can be 1 ms, 2 ms, 3 ms, 4 ms, 5 ms, 10 ms, 15 ms, 20 ms, 22 ms, 23 ms, 24 ms, 25 ms, 50 ms, 100 ms, 500 ms, 1 second, 5 second, or more than 5 seconds.
[0057] At block 204, the training system can process the audio data to determine a plurality of features for the sequence of audio frames. The plurality of features can identify various aspects associated with the audio data. For example, the plurality of features can include root mean square (RMS) energy of the sequence of audio frames, in which the RMS energy can be a quantitative15 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO value that represents loudness or intensity of the sequence audio frames. In another example, the plurality of features can include a spectral bandwidth of the sequence of audio frames, in which the spectral bandwidth quantifies a range of frequencies present in an audio signal within each audio frame. Other example features can include spectral centroid, zero-crossing rate for identifying pitch and noisiness of the audio frames, Mel-frequency cepstral coefficients (MFCCs) used for speech and music recognition, Spectral Centroid, PCEN, Spectrogram-based features, and LPCC used for ML-based speech processing in natural -language processing.
[0058] Various signal-processing techniques can be used to determine the plurality of features from the sequence of audio frames. For example, Fourier Transform (FT) or Short-Time Fourier Transform (STFT) can convert each audio frame from a time domain into a frequency domain, which can then be used to determine frequency-based features associated with the sequence of audio frames. Other signal-processing techniques can include a wavelet transform and autocorrelation.
[0059] At block 206, the training system can generate a training dataset based on the plurality of features associated with each audio frame of the sequence of audio frames. For example, each frame is represented by a set of features (e.g., Fl, F2, F3) and a corresponding ground-truth label (L), resulting in a dataset with dimensions such as (12, 4) — 12 audio frames and four columns (three features and one ground-truth label).
[0060] At block 208, the training system can generate a network graph 210 representing the audio frames of the training dataset. In some examples, the network graph 210 can include a plurality of nodes and one or more edges between nodes of the plurality of nodes. In some cases, a node of the plurality of nodes includes an initial embedding associated with a set of features of the plurality of features determined for an audio frame of the sequence of audio frames. An edge of the one or more edges can identify a degree of similarity between two audio frames represented by two nodes of the plurality of nodes.
[0061] In some cases, to generate the edge(s) of the graph, the training system can include determining a distance metric between respective sets of features associated with the two audio frames represented by the two nodes. For example, the distance metric can include a Euclidean16 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WODistance or a Manhattan Distance. The distance metric can be compared to a distance threshold. The edge can be generated based on a determination that the distance metric exceeds the distance threshold. The distance metric exceeding the distance threshold can be indicative of the degree of similarity of the two audio frames represented by the two nodes. Additionally or alternatively, other types of metrics can be used to determine the degree of similarity between two nodes and generate a corresponding edge. Examples of these other metrics can include edge weights, Cosine similarity, and Jaccard index.
[0062] In some cases, when an edge is formed between the nodes, an edge weight can be assigned to the edge. The edge weight can be processed by the graph-processing machine-learning model to encode the plurality of nodes of a network graph. For example, the edge weight can be the same distance metric calculated between the nodes. In another example, the edge weight can be another distance metric calculated based on one or more other distance metrics associated with other edges of the network graph. In yet another example, the edge weight can be recalculated to be another type of metric that is different from the distance metric. Alternatively, no edge weight is assigned to the edge once it is formed.
[0063] FIG. 3 illustrates an example process 300 for generating network graphs from the plurality of features, in accordance with some examples. At step 302, a network graph can be initialized. To initialize the network graph, a graph-generating application can import one or more libraries associated with network graph operations. After the import of the libraries, the graphgenerating application can initialize an empty graph using a function provided by the chosen library. The initial network graph can be a new graph object with no nodes or edges. After the network graph is initialized, the graph-generating application can define one or more properties associated with the network graph. For example, the graph-generating application can define the network graph to include undirected edges. The network graph can be configured to receive new nodes and edges to be added. For example, network-graph operations such as add_node() and add_edge() can be used to incorporate nodes and define relationships between the nodes.
[0064] At step 304, a plurality of nodes for the network graph can be generated. In some cases, a node of the plurality of nodes can represent an audio frame of the training dataset. For example, a network graph representing a sequence of 12 audio frames can include 12 nodes, in which each17 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO node can represent a corresponding audio frame of the sequence. In some implementations, the graph-generating application can perform an add_node() function, in which the node can be defined to represent the audio frame.
[0065] At step 306, the plurality of features can be appended to each node of the plurality of nodes. From the training dataset, the ground-truth labels can be removed, such that only the plurality of features are used for graph construction. Removing the ground-truth labels leads to a refined training dataset that includes only the features (Fl, F2, F3) for each audio frame. As a result, each node can represent an individual audio frame and corresponding features.
[0066] At step 308, edges between nodes can be formed based on a distance metric determined using the features associated with the plurality of nodes. Various distance metrics can be determined, such as cosine similarity, Euclidean distance, or Manhattan distance. For example, cosine similarity can be used as an example distance metric to measure the distance between the two nodes. If the distance metric exceeds a similarity threshold, an edge can be created between those nodes. The similarity threshold can include any numerical value determined to identify a degree of similarity between two nodes. The numerical value of the similarity threshold can be 0.1, 0.2, 0.5, 0.75, 0.9, 1, 5, 10, 15, 50, 75, 90, 100, and so on.
[0067] After forming edges between nodes based on the selected distance metric, edge weights are assigned. The edge weight can be processed by the graph-processing machine-learning model to encode the plurality of nodes of a network graph. The weight can either be the value of the same distance metric (e.g., cosine similarity score) or a different metric value as needed for the subsequent operations (e.g., processing by the graph-processing machine-learning model). In another example, the edge weight can be another distance metric calculated based on one or more other distance metrics associated with other edges of the network graph. The flexibility of selecting various edge weights can facilitate fine-tuning the network graph and enhance the graphprocessing machine-learning model's ability to capture meaningful relationships between audio frames.
[0068] The resulting network graph 210 can effectively represent the audio data as a network of audio frames (nodes) connected by edges that indicate their similarity or relationship. The network18 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO graph can then be processed using graph-processing machine-learning models, such as a GraphSAGE algorithm. The outputs from the graph-processing machine-learning models can be used for downstream processing such as like classification or regression. By leveraging graphbased methods rather than existing recurrent neural networks, the present techniques offer a more efficient and scalable solution for audio processing, especially in edge environments with limited computational and memory resources.
[0069] Returning to FIG. 2, the training system can train a graph-processing machine-learning model 214 to encode the plurality of nodes of the network graph 210 into a set of refined feature embeddings 216 (block 212). For example, an artificial neural network trained using unsupervised training can be used to generate the set of refined feature embeddings. In some cases, to normalize the quantitative features of the sequence of audio frames, the graph-processing machine-learning model can utilize Z-score normalization and min-max scaling to normalize statistical properties (mean, standard deviation, min, max) from the initial embeddings of the nodes of the network graph. The normalization can transform the features of the audio frames to be comparable across different scales.
[0070] In some cases, the graph-processing machine-learning model can be a GraphSAGE (Graph Sample and Aggregated) algorithm. FIG. 4 illustrates an example representation 400 of the GraphSAGE algorithm configured to process network graphs, in accordance with some examples. The GraphSAGE can be an inductive framework for generating low-dimensional node embeddings for large-scale graph data. GraphSAGE can generate node embeddings in higher-dimensional spaces, which can be particularly useful in instances where the initially derived features are deemed inadequate. In some cases, GraphSAGE can generate embeddings for unseen nodes, making it highly suitable for dynamic graphs where nodes and edges may change or grow over time. An example pseudocode of the GraphSAGE can be found below:Algorithm 1: GraphSAGE embedding generation (i.e., forward propagation) algorithmInput : Graph <J(V, £); input features {x^ Vv £ V}; depth K; weight matrices Wk, V / < e {1, .... if}; non-linearity o; differentiable aggregator functions AGGREGATEfc, Vfc £ {1, . neighborhood function TV: v -> 2VOutput: Vector representations z„ for all v E V1 h' «- xv, Vu £ V2 for k = 1 ... K do19 Polsinelli Ref. No. 094922-825835PATENTQualcomm Ref. No. 2407274WOend zu«- h„ , Wv e V
[0071] The GraphSAGE can be trained to learn a function that generates refined feature embeddings by sampling and aggregating features from a node's local neighborhood (e.g., other nodes that are connected to the given node). The neighborhood-based approach can facilitate scalable learning and avoids the computational limitations of full-batch training methods, which often become infeasible for very large graphs.
[0072] To train GraphSAGE, the model can process input data (e.g., the network graph 210 of FIG. 2) by iteratively aggregating features from a node's neighbors to compute its embedding. At each layer of the model, a node's initial embedding is updated by combining its own features (e.g., spectral bandwidth, RMS) with a sampled subset of its corresponding neighbors' features. For example, for each node, GraphSAGE samples a fixed size set of neighboring nodes, which helps to control the memory footprint and computational complexity. This sampling strategy ensures that the model remains efficient even when the graph has high-degree nodes with many neighbors, as it limits the number of nodes involved in the aggregation process.
[0073] Once the neighbors are sampled, GraphSAGE applies an aggregation function to combine the initial embeddings of the neighbor nodes. Several aggregation functions can be used, such as mean, LSTM, and pooling. The mean aggregator can compute the element-wise mean of the sampled neighbors' initial embeddings. The LSTM aggregator can use a recurrent neural network-based approach to combine the neighbors' initial embeddings, capturing more complex dependencies among them. The pooling aggregator can apply a non-linear transformation, such as a fully connected neural network layer, to each neighbor's initial embeddings and then taking an element-wise max or mean of the transformed features. These different types of aggregators can allow GraphSAGE to capture varying levels of complexity and relationships of audio frames associated with the audio data.20 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO
[0074] After the aggregation step, the aggregated neighborhood embeddings are concatenated with the target node's own initial embeddings. This combined embedding is then passed through a fully connected layer with a non-linear activation function, such as ReLU (Rectified Linear Unit), to produce the refined feature embeddings for the node. The above process can be repeated for multiple layers, allowing the refined feature embeddings to incorporate information from nodes that are several hops away in the graph. By stacking these layers, GraphSAGE can capture more complex and deeper relational structures in the network graph 210.
[0075] The final output of GraphSAGE is a set of refined feature embeddings, one for each node in the graph. These embeddings can be used for various downstream tasks, such as node classification, link prediction, and clustering. Because GraphSAGE generates embeddings in an inductive manner, it can be applied to unseen data, allowing it to generalize well to new nodes and edges that were not present during training. This inductive capability is achieved by learning how to aggregate local neighborhood information rather than memorizing the entire graph structure of the network graph 210.
[0076] For training GraphSAGE using unsupervised techniques, the weights of the neural network layers can be learned during training using backpropagation, optimizing the model to minimize a loss function, such as the margin loss or negative sampling loss, depending on the task at hand. The unsupervised training process for GraphSAGE can include sampling a batch of nodes from the graph and generates embeddings for each node by aggregating features from its sampled neighbors. For each node, the training system can select a set of positive samples, which are nodes that are directly connected to it via edges. The objective is to ensure that the embeddings of these neighboring nodes are close to the embedding of the target node in the learned vector space. Simultaneously, the model samples a set of negative samples, which are nodes that are not directly connected to the target node. The goal is to push the negative samples farther away from the target node's embedding, effectively learning to distinguish between nearby and distant nodes.
[0077] The model uses a loss function to optimize these embedding relationships. A commonly used loss function for unsupervised GraphSAGE is the negative sampling loss or contrastive loss. The negative sampling loss is designed to maximize the similarity between embeddings of nodes that share an edge (positive samples) and minimize the similarity between embeddings of nodes21 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO that do not share an edge (negative samples). For a pair of nodes u and v, where v is a positive sample and v' is a negative sample, the loss function could be formulated as: - - log log ( ff (zM- zv'))>where zuand zvare the embeddings of nodes u and v, and G is the sigmoid function.
[0078] The first term ensures that the dot product between embeddings of positive samples is maximized, while the second term ensures that the dot product between embeddings of negative samples is minimized.
[0079] During training, GraphSAGE iteratively updates the parameters of the neural network that governs the aggregation functions, using the gradients computed from the unsupervised loss function. These gradients guide the model to adjust the node embeddings such that they capture the local structure and features of the graph, even without labels. In this manner, the embeddings learned by GraphSAGE reflect the network graph's underlying structure and can be used for downstream tasks like clustering, visualization, or node ranking.
[0080] Returning to FIG. 2, the training system can train an audio-operation machine-learning model 220 to generate a set of output values associated with one or more audio-processing operations (block 218). In particular, the audio-operation machine-learning model 220 can be trained to process the set of refined feature embeddings 216 to generate the set of output values. In some instances, an output value of the set of output values is associated with one or more audio frames of the sequence of audio frames. For example, each output value of the set can correspond to a classification associated with a corresponding audio frame of the sequence of audio frames. In another example, an output value of the set can correspond to a classification associated with two or more frames of the sequence. In some cases, the set of output values can be aggregated to generate an aggregate output value that identifies one or more characteristics associated with the audio data.
[0081] Training the audio-operation machine-learning model can include adjusting parameters to learn various characteristics associated with the refined feature embeddings. For example, a22 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO parameter can be a value that the machine-learning model learns from the training data. Parameters correspond to components of the machine-learning model that are optimized during training to minimize the loss function, which measures how well the machine-learning model's predictions match the actual output. In the context of neural networks, parameters typically include weights and biases that are adjusted iteratively through gradient descent (for example) to improve the machine-learning model's accuracy in making predictions.
[0082] A weight is a type of parameter in machine learning models that represents the strength or importance of a connection between units, such as neurons in adjacent layers. Each weight can be a scalar value that, when multiplied with the input signal, influences the contribution of that input to the neuron's output. During training, weights can be adjusted to reduce prediction errors by increasing or decreasing the influence of certain inputs. For example, in a neural network, weights can be updated based on an error of the output relative to the expected result, allowing the parameters to learn which features of the input data are more significant for accurate predictions.
[0083] A bias is another type of parameter in machine learning models that allows the machinelearning model to make more flexible decisions by shifting the activation function, which determines the output of a neuron. Unlike weights, which scale the input data, a bias can be added to the weighted sum of inputs before applying the activation function. Biases can facilitate fitting the data better by providing a baseline that can move the output up or down, regardless of the input values. Bias can be particularly useful for handling tasks where the optimal decision boundary does not pass through the origin of the input space, thereby allowing the machine-learning model to capture patterns and relationships more effectively. Biases can be trained alongside weights, which can allow neural networks to approximate complex functions.
[0084] The audio-operation machine-learning model 220 can include any type of machinelearning model configured to generate output that can be subsequently used for various audioprocessing operations. The audio-operation machine-learning model can include, but not limited to, a classifier (e.g., single-variate or multivariate that is based on k-nearest neighbors, Naive Bayes, Logistic regression, support vector machine, decision trees, an ensemble network of classifiers, and / or the like), regression model (e.g., such as, but not limited to, linear regressions, logarithmic regressions, Lasso regression, Ridge regression, and / or the like), clustering model23 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO(e.g., such as, but not limited to, models based on k-means, hierarchical clustering, DBSCAN, biclustering, expectation-maximization, random forest, and / or the like), deep learning model (e.g., such as, but not limited to, neural networks, convolutional neural networks, recurrent neural networks, long short-term memory (LSTM), multilayer perceptions, etc.), combinations thereof (e.g., disparate-type ensemble networks, etc.), or the like.
[0085] As an illustrative example, the audio-operation machine-learning model 220 can be an artificial neural network selected from a models database. The neural network can be defined by an example neural network description for machine learning in a neural controller, which can be the same as a processing unit inside a mobile device. Neural network description can include a full specification of the neural network, including the neural architecture. For example, the neural network description can include a description or specification of architecture of the neural network (e.g., the layers, layer interconnections, number of nodes in each layer, etc.); an input and output description which indicates how the input and output are formed or processed; an indication of the activation functions in the neural network, the operations or filters in the neural network, etc.; neural network parameters such as weights, biases, etc. and so forth.
[0086] The neural network can reflect the architecture defined in neural network description. In this non-limiting example, the neural network includes an input layer, which includes input data, which can be any type of data such as media content (images, videos, etc.), numbers, text, etc., associated with the corresponding input data (e.g., the refined feature embeddings with reference to FIGS. 1-3). The neural network can include one or more hidden layers. The hidden layers can include n number of hidden layers, where n is an integer greater than or equal to one. The number of hidden layers can include as many layers as needed for a desired processing outcome and / or rendering intent. The neural network further includes an output layer that provides an output resulting from the processing performed by the hidden layers.
[0087] The neural network, in this example, is a multi-layer neural network of interconnected nodes. Each node can represent a piece of information. Information associated with the nodes is shared among the different layers and each layer retains information as information is processed. In some cases, the neural network can include a feed-forward neural network, in which case there are no feedback connections where outputs of the neural network are fed back into itself.24 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO
[0088] Information can be exchanged between nodes through node-to-node interconnections between the various layers. Nodes of the input layer can activate a set of nodes in the first hidden layer. For example, as shown, each input node of the input layer is connected to each node of the first hidden layer. Nodes of the first hidden layer can transform the information of each input node by applying activation functions to the information. The information derived from the transformation can then be passed to and can activate the nodes of the next hidden layer, which can perform their own designated functions. Example functions include convolutional, up- sampling, data transformation, pooling, and / or any other suitable functions. The output of the hidden layer (e.g.,) can then activate nodes of the next hidden layer, and so on. The output of the last hidden layer can activate one or more nodes of the output layer, at which point an output is provided. In some cases, while nodes in the neural network are shown as having multiple output lines, a node has a single output and all lines shown as being output from a node represent the same output value.
[0089] In some cases, each node or interconnection between nodes can have a weight that is a set of parameters derived from training the neural network. For example, an interconnection between nodes can represent a piece of information learned about the interconnected nodes. The interconnection can have a numeric weight that can be tuned (e.g., based on a training dataset), allowing the neural network to be adaptive to inputs and able to learn as more data is processed.
[0090] In some instances, the neural network is pre-trained to process the features from the data in the input layer using different hidden layers in order to provide the output through the output layer. The neural network can include any suitable neural or deep learning type of network. One example includes a convolutional neural network (CNN), which includes an input layer and an output layer, with multiple hidden layers between the input and out layers. The hidden layers of a CNN include a series of convolutional, nonlinear, pooling (for downsampling), and fully connected layers. In other examples, the neural network can represent any other neural or deep learning network, such as an autoencoder, deep belief nets (DBNs), recurrent neural networks (RNNs), etc.
[0091] The training system can train the audio-operation machine-learning model using another training dataset the includes, for each audio frame, the refined feature embeddings 216 and a25 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO corresponding ground-truth label. In some cases, the training dataset can be split into two classes of data called training data set and test data set. For example, 70% of the accessed data from the pool may be used as part of the training data set while the remaining 30% of the accessed data from the pool may be used as part of the test data set. The percentages according to which the pool of the data are split into training data set and test data set is not limited to 70 / 30 and may be set according to a configurable accuracy requirement and / or error tolerance (e.g., the split can be 50 / 50, 60 / 40, 70 / 30, 80 / 20, 90 / 10, etc. between the two data sets).
[0092] The training system can then use the other training dataset to train the audio-operation machine-learning model by calculating a loss based on a comparison between an output generated from the machine-learning model and a corresponding label of the training data. With each output generated by the audio-operation machine-learning model, the label can thus be used to correct the output of the audio-operation machine-learning model. In some instances, manual feedback is further utilized to adjust the corresponding parameters of the machine-learning model. As noted, weights of different nodes of the audio-operation machine-learning model may be adjusted / tuned during the training process to improve resulting output.
[0093] During training, weights of nodes associated with the audio-operation machine-learning model can be adjusted using a training process called backpropagation. B ackpropagation can include a forward pass, a loss function, a backward pass, and a weight update. The forward pass, loss function, backward pass, and parameter update can be performed for one training iteration. The process can be repeated for a certain number of iterations for each set of training media data until the weights of the layers are accurately tuned. In particular, the training of the audio-operation machine-learning model (e.g., adjustment of the weights) can be performed until a corresponding loss (e.g., a mean square error) reaches a minimum threshold.
[0094] Once trained, the training system can test the audio-operation machine-learning model using the test data set. Examples of testing methods can include regression testing, unit testing, beta testing, and alpha testing. Once the result of testing the audio-operation machine-learning model is satisfactory (e.g., when outputs of the testing stage is greater than or equal to a threshold or incorrect detections are less than a threshold), the training system can deploy the trained audiooperation machine-learning model (which may also be referred to as a trained machine learning26 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO model or machine trained neural network) to a computing device (e.g., an edge device), which can use the trained audio-operation machine-learning model to perform one or more audio-processing operations.
[0095] In some cases, the graph-processing machine-learning model 214 and the audiooperation machine-learning model 220 can be trained individually at different time points. For example, the graph-processing machine-learning model 214 can be trained at a first time point and the audio-operation machine-learning model 220 can be trained at a second time point. Additionally or alternatively, the graph-processing machine-learning model 214 and the audiooperation machine-learning model 220 can be trained together at the same time point, in which the loss determined for the audio-operation machine-learning model 220 can be backpropagated for adjusting parameters of the graph-processing machine-learning model 214.
[0096] In some aspects, training of one or more of the machine learning systems or neural networks described herein (e.g., the graph-processing machine-learning model 214 of FIG. 2, the audio-operation machine-learning model 220 of FIG. 2) can be performed using online training (e.g., in some case on-device training), offline training, and / or various combinations of online and offline training. In some cases, online may refer to time periods during which the input data (e.g., the network graph 210 of FIG. 2) is processed, for instance for generating refined feature embeddings by the systems and techniques described herein. In some examples, offline may refer to idle time periods or time periods during which input data is not being processed. Additionally, offline may be based on one or more time conditions (e.g., after a particular amount of time has expired, such as a day, a week, a month, etc.) and / or may be based on various other conditions such as network and / or server availability, etc., among various others. In some aspects, offline training of a machine learning model (e.g., a neural network model) can be performed by a first device (e.g., a server device) to generate a pre-trained model, and a second device can receive the trained model from the first device. In some cases, the second device (e.g., a mobile device, an XR device, a vehicle or system / component of the vehicle, or other device) can perform online (or on- device) training of the pre-trained model to further adapt or tune the parameters of the model.
[0097] FIG. 5 is a flowchart illustrating an example process 500 for processing network-graph embeddings for audio-processing operations, using one or more of the techniques described herein.27 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WOThe process 500 can be performed by a computing device (or apparatus), or a component of the computing device (e.g., a chipset, a processor such as a neural processing unit (NPU), a digital signal processor (DSP), a graphics processing unit (GPU), a central processing unit (CPU), etc.), utilizing or implementing the neural network model (e.g., the graph-processing machine-learning model 214 of FIG. 2, the audio-operation machine-learning model 220 of FIG. 2).
[0098] At block 502, the process 500 can include accessing audio data. The audio data can be accessed using various techniques. For example, at least one microphone of a computing device can be used to capture the audio data. In addition to microphones, specialized sensors such as hydrophones for underwater audio capture or contact microphones for picking up vibrations from solid surfaces can be used to capture the audio data. In some cases, the audio data can be extracted from pre-existing audio recordings, such as those stored in digital files (e.g., MP3, WAV). Additionally or alternatively, the audio data can be captured from one or more streaming sources, in which the audio data can be captured by intercepting and storing portions of streaming data being transmitted in real-time over the Internet.
[0099] The audio data can include a sequence of audio frames. In some cases, each audio frame can include time or sequence data that can indicate a position of the audio frame relative to other audio frames of the sequence of audio frames. In some cases, the audio data can include a plurality of sequences of audio frames, in which each sequence can include a predetermined number of audio frames. An audio frame can include a set of samples that represent an audio signal at a particular time point. Each sample within the audio frame can be a digital representation of an analog audio signal, in which a resolution of the sample can be determined by its corresponding bit depth (e.g., 16-bit depth). In some cases, the audio frame can be associated with a particular time duration (e.g., 2 ms, 23 ms). As noted above, example duration of an audio frame can be 1 ms, 2 ms, 3 ms, 4 ms, 5 ms, 10 ms, 15 ms, 20 ms, 22 ms, 23 ms, 24 ms, 25 ms, 50 ms, 100 ms, 500 ms, 1 second, 5 second, or more than 5 seconds.
[0100] At block 504, the process 500 can include processing the audio data to determine a plurality of features for the sequence of audio frames. The plurality of features can identify various aspects associated with the audio data. For example, the plurality of features can include root mean square (RMS) energy of the sequence of audio frames, in which the RMS energy can be a28 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO quantitative value that represents loudness or intensity of the sequence audio frames. In another example, the plurality of features can include a spectral bandwidth of the sequence of audio frames, in which the spectral bandwidth quantifies a range of frequencies present in an audio signal within each audio frame. Other example features can include spectral centroid, zero-crossing rate for identifying pitch and noisiness of the audio frames, MFCCs used for speech and music recognition, PCEN, Spectrogram-based features, and LPCC used for ML -based speech processing in naturallanguage processing.
[0101] At block 506, the process 500 can include processing the plurality of features using a graph-processing machine-learning model (e.g., the trained graph-processing machine learning model 214 of FIG. 2) to encode a plurality of nodes of a network graph into a set of refined feature embeddings. The network graph can represent the sequence of audio frames. In some examples, the network graph can include the plurality of nodes and one or more edges between nodes of the plurality of nodes. In some cases, a node of the plurality of nodes includes an initial embedding associated with a set of features of the plurality of features determined for an audio frame of the sequence of audio frames. An edge of the one or more edges can identify a degree of similarity between two audio frames represented by two nodes of the plurality of nodes. For example, a network graph representing a sequence of 12 audio frames can include 12 nodes, in which each node can include an initial embedding of a corresponding audio frame of the sequence and edges can connect a subset of nodes that are similar to each other.
[0102] In some cases, to generate the edge of the one or more edges, the process 500 can include determining a distance metric between respective sets of features associated with the two audio frames represented by the two nodes. For example, the distance metric can include a Euclidean Distance or a Manhattan Distance. The distance metric can be compared to a distance threshold. The edge can be generated based on a determination that the distance metric exceeds the distance threshold. The distance metric exceeding the distance threshold can be indicative of the degree of similarity of the two audio frames represented by the two nodes. Additionally or alternatively, other types of metrics can be used to determine the degree of similarity between two nodes and generate a corresponding edge. Examples of these other metrics can include Cosine similarity, and Jaccard index.29 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO
[0103] In some cases, when an edge is formed between the nodes, an edge weight can be assigned to the edge. The edge weight can be processed by the graph-processing machine-learning model to encode the plurality of nodes of a network graph. For example, the edge weight can be the same distance metric calculated between the nodes. In another example, the edge weight can be another distance metric calculated based on one or more other distance metrics associated with other edges of the network graph. In yet another example, the edge weight can be recalculated to be another type of metric that is different from the distance metric. Alternatively, no edge weight is assigned to the edge once it is formed. The edges can thus encode similarity or dissimilarity between audio frames, effectively transforming the sequential audio data into a relational graph structure.
[0104] The network graph representing the sequence of audio frames can be obtained using different approaches. In some cases, the network graph can be downloadable from another device. For example, a smartphone can process the plurality of features for the sequence of audio frames to generate the network graph, which can then be downloaded by an Extended Reality (XR) headset. In another example, a server communicating with an edge device (e g., XR headset, a car, earbuds) can process the plurality of features for the sequence of audio frames to generate the network graph, at which the edge device can download the network directly from the server. In some cases, the edge device can generate the network graph in real-time, such that the uses of the machine-learning models and subsequent audio-processing operations can be performed locally within the edge device in real-time.
[0105] In some cases, an artificial neural network trained using unsupervised training can be used to generate the set of refined feature embeddings. In some cases, to normalize the quantitative features of the sequence of audio frames, the graph-processing machine-learning model can utilize Z-score normalization and min-max scaling to normalize statistical properties (mean, standard deviation, min, max) from the initial embeddings of the nodes of the network graph. The normalization can transform the features of the audio frames to be comparable across different scales.
[0106] At block 508, the process 500 can include processing the set of refined feature embeddings using an audio-operation machine-learning model (e.g., the trained audio-operation30 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO machine learning model 220 of FIG. 2) to generate a set of output values associated with one or more audio-processing operations. In some instances, an output value of the set of output values is associated with one or more audio frames of the sequence of audio frames. For example, each output value of the set can correspond to a classification associated with a corresponding audio frame of the sequence of audio frames. In another example, an output value of the set can correspond to a classification associated with two or more frames of the sequence. In some cases, the set of output values can be aggregated to generate an aggregate output value that identifies one or more characteristics associated with the audio data.
[0107] The audio-operation machine-learning model can include any type of machine-learning model configured to generate output that can be subsequently used for various audio-processing operations. The audio-operation machine-learning model can include, but not limited to, a classifier (e.g., single-variate or multivariate that is based on k-nearest neighbors, Naive Bayes, Logistic regression, support vector machine, decision trees, an ensemble network of classifiers, and / or the like), regression model (e.g., such as, but not limited to, linear regressions, logarithmic regressions, Lasso regression, Ridge regression, and / or the like), clustering model (e.g., such as, but not limited to, models based on k-means, hierarchical clustering, DBSCAN, biclustering, expectation-maximization, random forest, and / or the like), deep learning model (e.g., such as, but not limited to, neural networks, convolutional neural networks, recurrent neural networks, long short-term memory (LSTM), multilayer perceptions, etc.), combinations thereof (e.g., disparatetype ensemble networks, etc.), or the like.
[0108] In some cases, the audio-operation machine-learning model can be trained using supervised training (e.g., a binary classification model) or unsupervised training (e.g., / / -means clustering algorithm). As an example supervised training process, the audio-operation machinelearning model can be trained based on: (i) accessing a training dataset including a sequence of training audio frames, wherein a training audio frame of the sequence of training audio frames is associated with a ground-truth label; (ii) iteratively transforming, applying the graph-processing machine-learning model, and applying the audio-operation machine-learning model to generate a second set of output values corresponding to the sequence of training audio frames; (iii) determining a loss between the ground-truth label and a corresponding output value of the second31 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO set of output values; and (iv) adjusting parameters of the audio-operation machine-learning model based on the loss.
[0109] In some cases, the graph-processing machine-learning model and the audio-operation machine-learning model can be trained individually at different time points. For example, the graph-processing machine-learning model can be trained at a first time point and the audiooperation machine-learning model is trained at a second time point. Additionally or alternatively, the graph-processing machine-learning model and the audio-operation machine-learning model can be trained together at the same time point, in which the loss determined for the audio-operation machine-learning model can be backpropagated for adjusting parameters of the graph-processing machine-learning model.
[0110] At block 510, the process 500 can include outputting the set of output values for the one or more audio-processing operations. In some cases, the process 500 can include performing the one or more audio-processing operations based on the set of output values. The one or more audioprocessing operations can be associated with various types of operations that can be performed by a computing device. For example, the one or more audio-processing operations can include detecting one or more keywords from the audio data. In another example, the one or more audioprocessing operations can include detecting whether an audio-accessory device is engaged with a subject (e.g., in-ear detection). In yet another example, the one or more audio-processing operations can include identifying metadata (e.g., a song title, artists, an album name, speakers of a podcast) associated with the audio data. In yet another example, the one or more audio-processing operations can include using the set of output values to perform active-noise cancellation on portions of the audio data.[0U1] For example, FIG. 6 illustrates an example schematic diagram 600 for processing network-graph embeddings to detect keywords from audio data, in accordance with some examples. In particular, the graph-processing machine-learning models (e.g., graph neural networks) can be used to process audio data and detect keywords from the audio data in real-time. As an illustrative example, an audio-processing operation can be associated with detecting keywords “Call Johnny”, at which a smart-speaker device can perform subsequent computing32 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO operations such as accessing contacts of a user to identify a phone number associated with the name “Johnny” and initiating a phone call using the phone number.
[0112] With references to the process 500 of FIG. 5, a process for detecting keywords can include accessing audio data including a sequence of audio frames (block 602). In some cases, at least one microphone or other specialized sensors can be used to capture the audio data. The process can also include determining a plurality of features for the sequence of audio frames (block 604). For example, the plurality of features can include RMS energy, spectral bandwidth, spectral centroid, and / or zero-crossing rate.
[0113] The process can also include processing the plurality of features using a graph-processing machine-learning model to generate refined feature embeddings (block 606). Processing the plurality of features can include: (i) obtaining a network graph (e.g., downloading a previously generated network graph, generating the network graph in real-time) representing the sequence of audio frames; and (ii) processing the network graph using the graph-processing machine-learning model (e.g., a GraphSAGE model) to generate refined feature embeddings. The process can also include processing the refined feature embeddings using an audio-operation machine-learning model (e.g., a binary classification model) to detect the presence of keywords from the audio data (block 608). Continuing with the example, the audio-operation machine-learning model can generate a classification that is indicative of whether the keywords “Call Johnny” are included in the audio data. In some cases, the audio-operation machine-learning model can generate a classification for each audio frame, at which the classifications are aggregated to determine a final output (e.g., whether the keywords were detected).
[0114] In another example, FIG. 7 illustrates an example schematic diagram 700 for processing network-graph embeddings to perform in-ear detection, in accordance with some examples. In particular, the graph-processing machine-learning models (e.g., graph neural networks) can be used to process audio data and detect, in real-time, whether one or more edge devices (e.g., earbuds, headphones) are engaged with a user. As an illustrative example, an audio-processing operation can be associated with the earbuds detecting that they have been engaged with the user, at which the earbuds can initiate transmission of audio content from a paired computing device.33 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO
[0115] With references to the process 500 of FIG. 5, a process for in-ear detection can include accessing internal audio data including a first sequence of audio frames and external audio data including a second sequence of audio frames (block 702). In some cases, an internal microphone (or multiple internal microphones) of the earbuds can be used to capture the internal audio data, and an external microphone (or multiple external microphones) of the earbuds can be used to capture the external audio data. The process can also include determining a plurality of features for the first and second sequences of audio frames (block 704). In some instances, the plurality of features can be determined by aggregating the features associated with the first and second sequences of audio frames. The plurality of features can include RMS energy, spectral bandwidth, spectral centroid, and / or zero-crossing rate.
[0116] The process can also include processing the plurality of features using a graph-processing machine-learning model to generate refined feature embeddings (block 706). Processing the plurality of features can include: (i) obtaining a network graph (e.g., downloading a previously generated network graph, generating the network graph in real-time) representing the first and second sequences of audio frames; and (ii) processing the network graph using the graphprocessing machine-learning model (e.g., a GraphSAGE model) to generate refined feature embeddings. The process can also include processing the refined feature embeddings using an audio-operation machine-learning model (e.g., a binary classification model) to perform in-ear detection (block 708). Continuing with the example, the audio-operation machine-learning model can generate a classification that is indicative of whether the earbuds are engaged with the user. In some cases, the audio-operation machine-learning model can generate a classification for each audio frame, at which the classifications are aggregated to determine a final output (e.g., whether the earbuds are engaged with the user).
[0117] In yet another example, FIG. 8 illustrates an example schematic diagram 800 for processing network-graph embeddings to identify metadata associated with audio data, in accordance with some examples. The graph-processing machine-learning models (e.g., graph neural networks) can be used to process audio data and identify metadata from the audio data in real-time. As an illustrative example, an audio-processing operation can be configured to identify a song title, an artist name, and a genre associated with the audio data.34 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO
[0118] With references to the process 500 of FIG. 5, a process for identifying metadata can include accessing audio data including a sequence of audio frames (block 802). In some cases, at least one microphone or other specialized sensors can be used to capture the audio data. The process can also include determining a plurality of features for the sequence of audio frames (block 804). For example, the plurality of features can include RMS energy, spectral bandwidth, spectral centroid, and / or zero-crossing rate.
[0119] The process can also include processing the plurality of features using a graph-processing machine-learning model to generate refined feature embeddings (block 806). Processing the plurality of features can include: (i) obtaining a network graph (e.g., downloading a previously generated network graph, generating the network graph in real-time) representing the sequence of audio frames; and (ii) processing the network graph using the graph-processing machine-learning model (e.g., a GraphSAGE model) to generate refined feature embeddings. The process can also include processing the refined feature embeddings using an audio-operation machine-learning model (e.g., a clustering algorithm) to identify metadata from the audio data (block 808).
[0120] Continuing with the example, the audio-operation machine-learning model can generate a feature vector that represents the refined feature embeddings of the sequence of audio frames, project the feature vector into an / / -dimensional space, identify the closest cluster to the feature vector by measuring the distance between the feature vector and centroids (mean positions) of existing clusters in the / / -dimensional space. Once the cluster is identified, the clustering algorithm can compare the plurality of features to those of songs within the cluster. The song in the cluster with the highest similarity to the plurality of features is likely the match, and metadata corresponding to the song can be provided. Continuing with the example, the metadata of the audio data can include “Song: Let It Be”, “Artist: The Beatles”, and “Genre: Rock.”
[0121] In another example, FIG. 9 illustrates an example schematic diagram 900 for processing network-graph embeddings to perform active-noise cancellation, in accordance with some examples. In particular, the graph-processing machine-learning models (e.g., graph neural networks) can be used to process audio data and perform active-noise cancellation in real-time by finding out filter coefficients. As an illustrative example, an audio-processing operation can be associated with headphones performing active-noise cancellation on external audio data captured35 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO from at least one external microphone and / or internal audio data captured from at least one internal microphone.
[0122] With references to the process 500 of FIG. 5, a process for active-noise cancellation can include accessing external audio data including a first sequence of audio frames (block 902). In some cases, at least one external microphone of the headphones can be used to capture the external audio data. The process can also include determining a plurality of features for the first sequence of audio frames (block 904). In some instances, the plurality of features can be determined by aggregating the features associated with the first and second sequences of audio frames. The plurality of features can include RMS energy, spectral bandwidth, spectral centroid, and / or zerocrossing rate.
[0123] The process can also include processing the plurality of features using a graph-processing machine-learning model to generate a first set of refined feature embeddings (block 906). Processing the plurality of features can include: (i) obtaining a network graph (e.g., downloading a previously generated network graph, generating the network graph in real-time) representing the first sequence of audio frames; and (ii) processing the network graph using the graph-processing machine-learning model (e.g., a GraphSAGE model) to generate the first set of refined feature embeddings. The process can also include processing the first set of refined feature embeddings an audio-operation machine-learning model (e g., a linear / polynomial regression model) to generate a first set of noise signals to be further processed by an adaptive filter (block 908). The adaptive filter can then process first set of noise signals to generate external anti-noise signals for the external audio data (block 910), at which the external anti -noise signals can be played simultaneously with the audio data to perform a feedforward active noise-cancellation (block 912).
[0124] The process can further include accessing internal audio data including a second sequence of audio frames (block 914). In some cases, at least one internal microphone of the headphones can be used to capture the internal audio data. The process can also include determining another plurality of features for the second sequence of audio frames (block 916). The process can return to block 906 and process the other plurality of features using the graphprocessing model and the audio-operation machine-learning model to generate a second set of noise signals (block 908). The adaptive filter can then process second set of noise signals to36 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO generate internal anti-noise signals for the internal audio data (block 910), at which the internal anti-noise signals can be played simultaneously with the audio data to perform a feedback active noise-cancellation (block 912). In some cases, the internal anti-noise signals and the external antinoise signals can be simultaneously generated and played with the audio data to perform a hybrid active noise-cancellation.
[0125] In some cases, an output device can be used to generate an output based on the one or more audio-processing operations. As an illustrative example, the output device can display the keywords from the audio data on a screen associated with the output device. In another example, the output device can be a headphone or earbuds that can process the output to remove background noise when streaming the audio data. In yet another example, the output device can additionally pause the audio data in response to detecting that the audio-accessory device is disengaged with a subject. Process 500 can terminate thereafter.
[0126] As noted above, the processes described herein (e.g., process 500 and / or any other process described herein) may be performed by a computing device or apparatus utilizing or implementing the neural network model (e.g., the graph-processing machine-learning model 214 of FIG. 2, the audio-operation machine-learning model 220 of FIG. 2). In one example, the process 500 can be performed by the electronic device of FIG. 1. In another example, the process 500 can be performed by the computing system having the computing device architecture of the computing system 1300 shown in FIG. 13 utilizing or implementing the neural network model (e.g., the graphprocessing machine-learning model 214 of FIG. 2, the audio-operation machine-learning model 220 of FIG. 2). For instance, a computing device with the computing device architecture of the computing system 1300 shown in FIG. 13 can implement the operations of FIG. 5 and / or the components and / or operations described herein with respect to any of FIGS. 2 through 5.
[0127] The computing device can include any suitable device, such as a mobile device (e.g., a mobile phone), a desktop computing device, a tablet computing device, an XR device (e.g., a VR headset, an AR headset, AR glasses, etc.), a wearable device (e.g., a network-connected watch or smartwatch, or other wearable device), a server computer, a vehicle or computing device of the vehicle, a robotic device, a laptop computer, a smart television, a camera, and / or any other computing device with the resource capabilities to perform the processes described herein,37 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO including the process 500 and / or any other process described herein. In some cases, the computing device or apparatus may include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other component(s) that are configured to carry out the steps of processes described herein. In some examples, the computing device may include a display, a network interface configured to communicate and / or receive the data, any combination thereof, and / or other component(s). The network interface may be configured to communicate and / or receive Internet Protocol (IP) based data or other type of data.
[0128] The components of the computing device can be implemented in circuitry. For example, the components can include and / or can be implemented using electronic circuits or other electronic hardware, which can include one or more programmable electronic circuits (e.g., microprocessors, graphics processing units (GPUs), digital signal processors (DSPs), central processing units (CPUs), and / or other suitable electronic circuits), and / or can include and / or be implemented using computer software, firmware, or any combination thereof, to perform the various operations described herein.
[0129] The process 500 is illustrated as a logical flow diagram, the operation of which represents a sequence of operations that can be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computerexecutable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the recited operations. Generally, computerexecutable instructions include routines, programs, objects, components, data structures, and the like that perform particular functions or implement particular data types. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations can be combined in any order and / or in parallel to implement the processes.
[0130] Additionally, the process 500 and / or any other process described herein may be performed under the control of one or more computer systems configured with executable instructions and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) executing collectively on one or more processors, by hardware, or combinations thereof. As noted above, the code may be stored on a computer-38 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO readable or machine-readable storage medium, for example, in the form of a computer program comprising a plurality of instructions executable by one or more processors. The computer- readable or machine-readable storage medium may be non -transitory.
[0131] In some implementations, a single supervised graph-processing machine-learning model can be used to perform certain types of audio-processing operations (e.g., a binary classification). FIG. 10 illustrates another example schematic diagram 1000 for training a supervised graphprocessing machine-learning model performing audio-processing operations, in accordance with some examples.
[0132] To initiate training of the supervised graph-processing machine-learning model, a training system can access audio data 1002. The audio data 1002 can be accessed using various techniques. For example, at least one microphone (e.g., a single microphone or multiple microphones) of a computing device can be used to capture the audio data 1002. In addition to microphones, specialized sensors such as hydrophones for underwater audio capture or contact microphones for picking up vibrations from solid surfaces can be used to capture the audio data 1002. In some cases, the audio data 1002 can be extracted from pre-existing audio recordings, such as those stored in digital files (e.g., MP3, WAV). Additionally or alternatively, an edge device can capture the audio data 1002 from one or more streaming sources, in which the audio data can be captured by intercepting and storing portions of streaming data being transmitted in real-time over the Internet.
[0133] The audio data 1002 can include a sequence of audio frames. For example, the sequence can include 1, 2, 3, 4, 5, 10, 12, 15, 30, or more than 30 audio frames, in which each audio frame can include time or sequence data that can indicate a position of the audio frame relative to other audio frames of the sequence of audio frames. In some cases, the audio data 1002 can include a plurality of sequences of audio frames, in which each sequence can include a predetermined number of audio frames. An audio frame can include a set of samples that represent an audio signal at a particular time point. Each sample within the audio frame can be a digital representation of an analog audio signal, in which a resolution of the sample can be determined by its corresponding bit depth (e.g., 16-bit depth).39 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO
[0134] In some cases, the audio frame can be associated with a particular time duration. The time duration of the audio frame can be determined by a sample rate of the audio data, which is the number of samples taken per second, measured in Hertz (Hz). The duration of the audio frame can be calculated using the formula: Frame Duration = Number of samples per frame / Sample Rate. As an illustrative example, with a sample rate of 44.1 kHz (44,100 samples per second) and 1024 samples in an audio frame, the duration of the audio frame of the sequence is approximately 0.0232 seconds, or 23.2 milliseconds (ms). Example duration of an audio frame can be 1 ms, 2 ms, 3 ms, 4 ms, 5 ms, 10 ms, 15 ms, 20 ms, 22 ms, 23 ms, 24 ms, 25 ms, 50 ms, 100 ms, 500 ms, 1 second, 5 second, or more than 5 seconds.
[0135] At block 1004, the training system can process the audio data to determine a plurality of features for the sequence of audio frames. The plurality of features can identify various aspects associated with the audio data. Additional details associated with determining the plurality of features are further described in block 204 of FIG. 2.
[0136] At block 1006, the training system can generate a training dataset based on the plurality of features associated with each audio frame of the sequence of audio frames. For example, each frame is represented by a set of features (e.g., Fl, F2, F3) and a corresponding ground-truth label (L), resulting in a dataset with dimensions such as (12, 4) — 12 audio frames and four columns (three features and one ground-truth label).
[0137] At block 1008, the training system can generate a network graph 1010 representing the audio frames of the training dataset. In some examples, the network graph 1010 can include a plurality of nodes and one or more edges between nodes of the plurality of nodes. In some cases, a node of the plurality of nodes includes an initial embedding associated with a set of features of the plurality of features determined for an audio frame of the sequence of audio frames. An edge of the one or more edges can identify a degree of similarity between two audio frames represented by two nodes of the plurality of nodes.
[0138] In some cases, to generate the edges of the graph, the training system can include determining a distance metric between respective sets of features associated with the two audio frames represented by the two nodes. For example, the distance metric can include a Euclidean40 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WODistance or a Manhattan Distance. The distance metric can be compared to a distance threshold. The edge can be generated based on a determination that the distance metric exceeds the distance threshold. The distance metric exceeding the distance threshold can be indicative of the degree of similarity of the two audio frames represented by the two nodes. Additionally or alternatively, other types of metrics can be used to determine the degree of similarity between two nodes and generate a corresponding edge. Examples of these other metrics can include edge weights, Cosine similarity, and Jaccard index.
[0139] In some cases, when an edge is formed between the nodes, an edge weight can be assigned to the edge. The edge weight can be processed by the graph-processing machine-learning model to encode the plurality of nodes of a network graph. For example, the edge weight can be the same distance metric calculated between the nodes. In another example, the edge weight can be another distance metric calculated based on one or more other distance metrics associated with other edges of the network graph. In yet another example, the edge weight can be recalculated to be another type of metric that is different from the distance metric. Alternatively, no edge weight is assigned to the edge once it is formed. Additional details associated with generating the network graph 1010 are further described in process 300 of FIG. 3.
[0140] After the network graph 1010 is generated, each node of the network graph 1010 can be assigned with a corresponding ground-truth label. For example, the node representing a particular audio frame can be assigned with a ground-truth label that classifies whether the audio frame includes one or more keywords (e.g., the keyword detection process described with respect to FIG. 6). The iterative labeling of the nodes of the network graph 1010 can thus facilitate a supervised graph-processing machine-learning model 1014 to be trained to directly generate a classification for each audio frame.
[0141] The training system can then train a supervised graph-processing machine-learning model 1014 to process the plurality of labeled nodes of the network graph 1010 into a set of output values 1016 (block 1012). In particular, the supervised graph-processing machine-learning model 1014 can be trained based on a loss determined between an output value of the set of output values and a corresponding ground-truth label. In some instances, an output value of the set of output values is associated with one or more audio frames of the sequence of audio frames. For example,41 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO each output value of the set can correspond to a classification associated with a corresponding audio frame of the sequence of audio frames. In another example, an output value of the set can correspond to a classification associated with two or more frames of the sequence. In some cases, the set of output values can be aggregated to generate an aggregate output value that identifies one or more characteristics associated with the audio data, which can then be utilized for audioprocessing operations. Using the supervised graph-processing machine-learning model 1014 for audio-processing operations can be advantageous in some situations, as it can reduce the time and amount of training two different models (e g., the graph-processing machine-learning model and the audio-operation machine-learning model of FIG. 2). In some implementations, the supervised graph-processing machine-learning model 1014 can be a Graphs AGE (Graph Sample and Aggregated) algorithm. Additional details associated with implementing the Graphs AGE algorithm are further described in FIG. 4.
[0142] FIG. 11 is an illustrative example of a deep learning neural network 1100 that can be used for processing network-graph embeddings for audio-processing operations as described in the process 300 of FIG. 3. An input layer 1120 includes input data. In one illustrative example, the input layer 1120 can include data representing the pixels of an input video frame. The neural network 1100 includes multiple hidden layers 1122a, 1122b, through 1122n. The hidden layers 1122a, 1122b, through 1122n include “n” number of hidden layers, where “n” is an integer greater than or equal to one. The number of hidden layers can be made to include as many layers as needed for the given application. The neural network 1100 further includes an output layer 1124 that provides an output resulting from the processing performed by the hidden layers 1122a, 1122b, through 1122n. In one illustrative example, the output layer 1124 can provide a classification for an object in an input video frame. The classification can include a class identifying the type of object (e.g., a person, a dog, a cat, or other object).
[0143] The neural network 1100 is a multi-layer neural network of interconnected nodes. Each node can represent a piece of information. Information associated with the nodes is shared among the different layers and each layer retains information as information is processed. In some cases, the neural network 1100 can include a feed-forward network, in which case there are no feedback connections where outputs of the network are fed back into itself. In some cases, the neural network42 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO1100 can include a recurrent neural network, which can have loops that allow information to be carried across nodes while reading in input.
[0144] Information can be exchanged between nodes through node-to-node interconnections between the various layers. Nodes of the input layer 1120 can activate a set of nodes in the first hidden layer 1122a. For example, as shown, each of the input nodes of the input layer 1120 is connected to each of the nodes of the first hidden layer 1122a. The nodes of the hidden layers 1122a, 1122b, through 1122n can transform the information of each input node by applying activation functions to the information. The information derived from the transformation can then be passed to and can activate the nodes of the next hidden layer 1122b, which can perform their own designated functions. Example functions include convolutional, up-sampling, data transformation, and / or any other suitable functions. The output of the hidden layer 1122b can then activate nodes of the next hidden layer, and so on. The output of the last hidden layer 1122n can activate one or more nodes of the output layer 1124, at which an output is provided. In some cases, while nodes (e.g., node 1126) in the neural network 1100 are shown as having multiple output lines, a node has a single output and all lines shown as being output from a node represent the same output value.
[0145] In some cases, each node or interconnection between nodes can have a weight that is a set of parameters derived from the training of the neural network 1100. Once the neural network 1100 is trained, it can be referred to as a trained neural network, which can be used to classify one or more objects. For example, an interconnection between nodes can represent a piece of information learned about the interconnected nodes. The interconnection can have a tunable numeric weight that can be tuned (e.g., based on a training dataset), allowing the neural network 1100 to be adaptive to inputs and able to learn as more and more data is processed.
[0146] The neural network 1100 is pre-trained to process the features from the data in the input layer 1120 using the different hidden layers 1122a, 1122b, through 1122n in order to provide the output through the output layer 1124. In an example in which the neural network 1100 is used to identify objects in images, the neural network 1100 can be trained using training data that includes both images and labels. For instance, training images can be input into the network, with each training image having a label indicating the classes of the one or more objects in each image43 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO(basically, indicating to the network what the objects are and what features they have). In one illustrative example, a training image can include an image of a number 2, in which case the label for the image can be [0 0 1 0 0 0 0 0 0 0],
[0147] In some cases, the neural network 1100 can adjust the weights of the nodes using a training process called backpropagation. Backpropagation can include a forward pass, a loss function, a backward pass, and a weight update. The forward pass, loss function, backward pass, and parameter update is performed for one training iteration. The process can be repeated for a certain number of iterations for each set of training images until the neural network 1100 is trained well enough so that the weights of the layers are accurately tuned.
[0148] For the example of identifying objects in images, the forward pass can include passing a training image through the neural network 1100. The weights are initially randomized before the neural network 1100 is trained. The image can include, for example, an array of numbers representing the pixels of the image. Each number in the array can include a value from 0 to 255 describing the pixel intensity at that position in the array. In one example, the array can include a 28 x 28 x 3 array of numbers with 28 rows and 28 columns of pixels and 3 color components (such as red, green, and blue, or luma and two chroma components, or the like).
[0149] For a first training iteration for the neural network 1100, the output will likely include values that do not give preference to any particular class due to the weights being randomly selected at initialization. For example, if the output is a vector with probabilities that the object includes different classes, the probability value for each of the different classes may be equal or at least very similar (e.g., for ten possible classes, each class may have a probability value of 0.1). With the initial weights, the neural network 1100 is unable to determine low level features and thus cannot make an accurate determination of what the classification of the object might be. A loss function can be used to analyze error in the output. Any suitable loss function definition can be used. One example of a loss function includes a mean squared error (MSE). The MSE is defined as EtotaL= (target — output)2, which calculates the sum of one-half times a ground truth output (e.g., the actual answer) minus the predicted output (e.g., the predicted answer) squared. The loss can be set to be equal to the value of Etotai.44 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO
[0150] The loss (or error) will be high for the first training images since the actual values will be much different than the predicted output. The goal of training is to minimize the amount of loss so that the predicted output is the same as the training label. The neural network 1100 can perform a backward pass by determining which inputs (weights) most contributed to the loss of the network, and can adjust the weights so that the loss decreases and is eventually minimized.
[0151] A derivative of the loss with respect to the weights (denoted as dIJdW, where W are the weights at a particular layer) can be computed to determine the weights that contributed most to the loss of the network. After the derivative is computed, a weight update can be performed by updating all the weights of the filters. For example, the weights can be updated so that they change in the opposite direction of the gradient. The weight update can be denoted as w = wtwhere w denotes a weight, wi denotes the initial weight, and r| denotes a learning rate. The learning rate can be set to any suitable value, with a high learning rate including larger weight updates and a lower value indicating smaller weight updates.
[0152] The neural network 1100 can include any suitable deep network. One example includes a convolutional neural network (CNN), which includes an input layer and an output layer, with multiple hidden layers between the input and out layers. An example of a CNN is described below with respect to FIG. 11. The hidden layers of a CNN include a series of convolutional, nonlinear, pooling (for downsampling), and fully connected layers. The neural network 1100 can include any other deep network other than a CNN, such as an autoencoder, a deep belief nets (DBNs), a Recurrent Neural Networks (RNNs), among others.
[0153] FIG. 12 is an illustrative example of a convolutional neural network 1200 (CNN 1200). The input layer 1220 of the CNN 1200 includes data representing an image. For example, the data can include an array of numbers representing the pixels of the image, with each number in the array including a value from 0 to 255 describing the pixel intensity at that position in the array. Using the previous example from above, the array can include a 28 x 28 x 3 array of numbers with 28 rows and 28 columns of pixels and 3 color components (e.g., red, green, and blue, or luma and two chroma components, or the like). The image can be passed through a convolutional hidden layer 1222a, an optional non-linear activation layer, a pooling hidden layer 1222b, and fully45 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO connected hidden layers 1222c to get an output at the output layer 1224. While only one of each hidden layer is shown in FIG. 12, one of ordinary skill will appreciate that multiple convolutional hidden layers, non-linear layers, pooling hidden layers, and / or fully connected layers can be included in the CNN 1200. As previously described, the output can indicate a single class of an object or can include a probability of classes that best describe the object in the image.
[0154] The first layer of the CNN 1200 is the convolutional hidden layer 1222a. The convolutional hidden layer 1222a analyzes the image data of the input layer 1220. Each node of the convolutional hidden layer 1222a is connected to a region of nodes (pixels) of the input image called a receptive field. The convolutional hidden layer 1222a can be considered as one or more filters (each filter corresponding to a different activation or feature map), with each convolutional iteration of a filter being a node or neuron of the convolutional hidden layer 1222a. For example, the region of the input image that a filter covers at each convolutional iteration would be the receptive field for the filter. In one illustrative example, if the input image includes a 28*28 array, and each filter (and corresponding receptive field) is a 5*5 array, then there will be 24*24 nodes in the convolutional hidden layer 1222a. Each connection between a node and a receptive field for that node learns a weight and, in some cases, an overall bias such that each node learns to analyze its particular local receptive field in the input image. Each node of the hidden layer 1222a will have the same weights and bias (called a shared weight and a shared bias). For example, the filter has an array of weights (numbers) and the same depth as the input. A filter will have a depth of 3 for the video frame example (according to three color components of the input image). An illustrative example size of the filter array is 5 x 5 x 3, corresponding to a size of the receptive field of a node.
[0155] The convolutional nature of the convolutional hidden layer 1222a is due to each node of the convolutional layer being applied to its corresponding receptive field. For example, a filter of the convolutional hidden layer 1222a can begin in the top-left corner of the input image array and can convolve around the input image. As noted above, each convolutional iteration of the filter can be considered a node or neuron of the convolutional hidden layer 1222a. At each convolutional iteration, the values of the filter are multiplied with a corresponding number of the original pixel values of the image (e.g., the 5x5 filter array is multiplied by a 5x5 array of input pixel values at46 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO the top-left corner of the input image array). The multiplications from each convolutional iteration can be summed together to obtain a total sum for that iteration or node. The process is next continued at a next location in the input image according to the receptive field of a next node in the convolutional hidden layer 1222a.
[0156] For example, a filter can be moved by a step amount to the next receptive field. The step amount can be set to 1 or other suitable amount. For example, if the step amount is set to 1, the filter will be moved to the right by 1 pixel at each convolutional iteration. Processing the filter at each unique location of the input volume produces a number representing the filter results for that location, resulting in a total sum value being determined for each node of the convolutional hidden layer 1222a.
[0157] The mapping from the input layer to the convolutional hidden layer 1222a is referred to as an activation map (or feature map). The activation map includes a value for each node representing the filter results at each locations of the input volume. The activation map can include an array that includes the various total sum values resulting from each iteration of the filter on the input volume. For example, the activation map will include a 24 x 24 array if a 5 x 5 filter is applied to each pixel (a step amount of 1) of a 28 x 28 input image. The convolutional hidden layer 1222a can include several activation maps in order to identify multiple features in an image. The example shown in FIG. 12 includes three activation maps. Using three activation maps, the convolutional hidden layer 1222a can detect three different kinds of features, with each feature being detectable across the entire image.
[0158] In some examples, a non-linear hidden layer can be applied after the convolutional hidden layer 1222a. The non-linear layer can be used to introduce non-linearity to a system that has been computing linear operations. One illustrative example of a non-linear layer is a rectified linear unit (ReLU) layer. AReLU layer can apply the function f(x) = max(0, x) to all of the values in the input volume, which changes all the negative activations to 0. The ReLU can thus increase the nonlinear properties of the CNN 1200 without affecting the receptive fields of the convolutional hidden layer 1222a.47 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO
[0159] The pooling hidden layer 1222b can be applied after the convolutional hidden layer 1222a (and after the non-linear hidden layer when used). The pooling hidden layer 1222b is used to simplify the information in the output from the convolutional hidden layer 1222a. For example, the pooling hidden layer 1222b can take each activation map output from the convolutional hidden layer 1222a and generates a condensed activation map (or feature map) using a pooling function. Max-pooling is one example of a function performed by a pooling hidden layer. Other forms of pooling functions be used by the pooling hidden layer 1222a, such as average pooling, L2-norm pooling, or other suitable pooling functions. A pooling function (e.g., a max-pooling filter, an L2- norm filter, or other suitable pooling filter) is applied to each activation map included in the convolutional hidden layer 1222a. In the example shown in FIG. 12, three pooling filters are used for the three activation maps in the convolutional hidden layer 1222a.
[0160] In some examples, max-pooling can be used by applying a max-pooling filter (e.g., having a size of 2x2) with a step amount (e.g., equal to a dimension of the filter, such as a step amount of 2) to an activation map output from the convolutional hidden layer 1222a. The output from a max-pooling filter includes the maximum number in every sub-region that the filter convolves around. Using a 2x2 filter as an example, each unit in the pooling layer can summarize a region of 2^2 nodes in the previous layer (with each node being a value in the activation map). For example, four values (nodes) in an activation map will be analyzed by a 2x2 max-pooling filter at each iteration of the filter, with the maximum value from the four values being output as the “max” value. If such a max-pooling filter is applied to an activation filter from the convolutional hidden layer 1222a having a dimension of 24x24 nodes, the output from the pooling hidden layer 1222b will be an array of 12x12 nodes.
[0161] In some examples, an L2-norm pooling filter could also be used. The L2-norm pooling filter includes computing the square root of the sum of the squares of the values in the 2x2 region (or other suitable region) of an activation map (instead of computing the maximum values as is done in max-pooling), and using the computed values as an output.
[0162] Intuitively, the pooling function (e.g., max-pooling, L2-norm pooling, or other pooling function) determines whether a given feature is found anywhere in a region of the image, and discards the exact positional information. This can be done without affecting results of the feature48 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO detection because, once a feature has been found, the exact location of the feature is not as important as its approximate location relative to other features. Max-pooling (as well as other pooling methods) offer the benefit that there are many fewer pooled features, thus reducing the number of parameters needed in later layers of the CNN 1200.
[0163] The final layer of connections in the network is a fully-connected layer that connects every node from the pooling hidden layer 1222b to every one of the output nodes in the output layer 1224. Using the example above, the input layer includes 28 x 28 nodes encoding the pixel intensities of the input image, the convolutional hidden layer 1222a includes 3*24*24 hidden feature nodes based on application of a 5*5 local receptive field (for the filters) to three activation maps, and the pooling layer 1222b includes a layer of 3* 12* 12 hidden feature nodes based on application of max-pooling filter to 2*2 regions across each of the three feature maps. Extending this example, the output layer 1224 can include ten output nodes. In such an example, every node of the 3x12x12 pooling hidden layer 1222b is connected to every node of the output layer 1224.
[0164] The fully connected layer 1222c can obtain the output of the previous pooling layer 1222b (which should represent the activation maps of high-level features) and determines the features that most correlate to a particular class. For example, the fully connected layer 1222c layer can determine the high-level features that most strongly correlate to a particular class, and can include weights (nodes) for the high-level features. A product can be computed between the weights of the fully connected layer 1222c and the pooling hidden layer 1222b to obtain probabilities for the different classes. For example, if the CNN 1200 is being used to predict that an object in a video frame is a person, high values will be present in the activation maps that represent high-level features of people (e.g., two legs are present, a face is present at the top of the object, two eyes are present at the top left and top right of the face, a nose is present in the middle of the face, a mouth is present at the bottom of the face, and / or other features common for a person).
[0165] In some examples, the output from the output layer 1224 can include an M-dimensional vector (in the prior example, M=10), where M can include the number of classes that the program has to choose from when classifying the object in the image. Other example outputs can also be provided. Each number in the N-dimensional vector can represent the probability the object is of a certain class. In one illustrative example, if a 10-dimensional output vector represents ten49 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO different classes of objects is [0 0 0.05 0.8 0 0.15 0 0 0 0], the vector indicates that there is a 5% probability that the image is the third class of object (e.g., a dog), an 120% probability that the image is the fourth class of object (e.g., a human), and a 15% probability that the image is the sixth class of object (e.g., a kangaroo). The probability for a class can be considered a confidence level that the object is part of that class.
[0166] FIG. 13 is a diagram illustrating an example of a system for implementing certain aspects of the present disclosure. In particular, FIG. 13 illustrates an example of computing system 1300, which can be for example any computing device making up a computing system, a camera system, or any component thereof in which the components of the system are in communication with each other using connection 1305. Connection 1305 can be a physical connection using a bus, or a direct connection into processor 1310, such as in a chipset architecture. Connection 1305 can also be a virtual connection, networked connection, or logical connection.
[0167] In some examples, computing system 1300 is a distributed system in which the functions described in this disclosure can be distributed within a datacenter, multiple data centers, a peer network, etc. In some examples, one or more of the described system components represents many such components each performing some or all of the function for which the component is described. In some examples, the components can be physical or virtual devices.
[0168] Example system 1300 includes at least one processing unit (CPU or processor) 1310 and connection 1305 that couples various system components including system memory 1315, such as read-only memory (ROM) 1320 and random access memory (RAM) 1325 to processor 1310. Computing system 1300 can include a cache 1312 of high-speed memory connected directly with, in close proximity to, or integrated as part of processor 1310.
[0169] Processor 1310 can include any general purpose processor and a hardware service or software service, such as services 1332, 1334, and 1336 stored in storage device 1330, configured to control processor 1310 as well as a special-purpose processor where software instructions are incorporated into the actual processor design. Processor 1310 may essentially be a completely self- contained computing system, containing multiple cores or processors, a bus, memory controller, cache, etc. A multi-core processor may be symmetric or asymmetric.50 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO
[0170] To enable user interaction, computing system 1300 includes an input device 1345, which can represent any number of input mechanisms, such as a microphone for speech, a touch-sensitive screen for gesture or graphical input, keyboard, mouse, motion input, speech, etc. Computing system 1300 can also include output device 1335, which can be one or more of a number of output mechanisms. In some instances, multimodal systems can enable a user to provide multiple types of input / output to communicate with computing system 1300. Computing system 1300 can include communications interface 1340, which can generally govern and manage the user input and system output.
[0171] The communication interface may perform or facilitate receipt and / or transmission wired or wireless communications using wired and / or wireless transceivers, including those making use of an audio jack / plug, a microphone jack / plug, a universal serial bus (USB) port / plug, an Apple® Lightning® port / plug, an Ethernet port / plug, a fiber optic port / plug, a proprietary wired port / plug, a BLUETOOTH® wireless signal transfer, a BLUETOOTH® low energy (BLE) wireless signal transfer, an IBEACON® wireless signal transfer, a radio-frequency identification (RFID) wireless signal transfer, near-field communications (NFC) wireless signal transfer, dedicated short range communication (DSRC) wireless signal transfer, 1302.11 Wi-Fi wireless signal transfer, wireless local area network (WLAN) signal transfer, Visible Light Communication (VLC), Worldwide Interoperability for Microwave Access (WiMAX), Infrared (IR) communication wireless signal transfer, Public Switched Telephone Network (PSTN) signal transfer, Integrated Services Digital Network (ISDN) signal transfer, 3G / 4G / 5G / LTE cellular data network wireless signal transfer, ad- hoc network signal transfer, radio wave signal transfer, microwave signal transfer, infrared signal transfer, visible light signal transfer, ultraviolet light signal transfer, wireless signal transfer along the electromagnetic spectrum, or some combination thereof.
[0172] The communications interface 1340 may also include one or more Global Navigation Satellite System (GNSS) receivers or transceivers that are used to determine a location of the computing system 1300 based on receipt of one or more signals from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the USbased Global Positioning System (GPS), the Russia-based Global Navigation Satellite System (GLONASS), the China-based BeiDou Navigation Satellite System (BDS), and the Europe-based51 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WOGalileo GNSS. There is no restriction on operating on any particular hardware arrangement, and therefore the basic features here may easily be substituted for improved hardware or firmware arrangements as they are developed.
[0173] Storage device 1330 can be a non-volatile and / or non-transitory and / or computer- readable memory device and can be a hard disk or other types of computer readable media which can store data that are accessible by a computer, such as magnetic cassettes, flash memory cards, solid state memory devices, digital versatile disks, cartridges, a floppy disk, a flexible disk, a hard disk, magnetic tape, a magnetic strip / stripe, any other magnetic storage medium, flash memory, memristor memory, any other solid-state memory, a compact disc read only memory (CD-ROM) optical disc, a rewritable compact disc (CD) optical disc, digital video disk (DVD) optical disc, a blu-ray disc (BDD) optical disc, a holographic optical disk, another optical medium, a secure digital (SD) card, a micro secure digital (microSD) card, a Memory Stick® card, a smartcard chip, a EMV chip, a subscriber identity module (SIM) card, a mini / micro / nano / pico SIM card, another integrated circuit (IC) chip / card, random access memory (RAM), static RAM (SRAM), dynamic RAM (DRAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash EPROM (FLASHEPROM), cache memory (L1 / L2 / L3 / L4 / L5 / L#), resistive random-access memory (RRAM / ReRAM), phase change memory (PCM), spin transfer torque RAM (STT-RAM), another memory chip or cartridge, and / or a combination thereof.
[0174] The storage device 1330 can include software services, servers, services, etc., that when the code that defines such software is executed by the processor 1310, it causes the system to perform a function. In some examples, a hardware service that performs a particular function can include the software component stored in a computer-readable medium in connection with the necessary hardware components, such as processor 1310, connection 1305, output device 1335, etc., to carry out the function. The term “computer-readable medium” includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other mediums capable of storing, containing, or carrying instruction(s) and / or data. A computer-readable medium may include a non-transitory medium in which data can be stored and that does not include carrier waves and / or transitory electronic signals propagating wirelessly or over wired connections.52 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WOExamples of a non-transitory medium may include, but are not limited to, a magnetic disk or tape, optical storage media such as compact disk (CD) or digital versatile disk (DVD), flash memory, memory or memory devices. A computer-readable medium may have stored thereon code and / or machine-executable instructions that may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, or the like.
[0175] In some examples the computer-readable storage devices, mediums, and memories can include a cable or wireless signal containing a bit stream and the like. However, when mentioned, non-transitory computer-readable storage media expressly exclude media such as energy, carrier signals, electromagnetic waves, and signals per se.
[0176] Specific details are provided in the description above to provide a thorough understanding of the aspects and examples provided herein. However, it will be understood by one of ordinary skill in the art that the aspects and examples may be practiced without these specific details. For clarity of explanation, in some instances the present technology may be presented as including individual functional blocks comprising devices, device components, steps or routines in a method embodied in software, or combinations of hardware and software. Additional components may be used other than those shown in the figures and / or described herein. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form in order not to obscure the aspects and examples in unnecessary detail. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary detail in order to avoid obscuring the aspects and examples.
[0177] Individual aspects and examples may be described above as a process or method which is depicted as a flowchart, a flow diagram, a data flow diagram, a structure diagram, or a block diagram. Although a flowchart may describe the operations as a sequential process, many of the operations can be performed in parallel or concurrently. In addition, the order of the operations53 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO may be re-arranged. A process is terminated when its operations are completed, but could have additional steps not included in a figure. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, its termination can correspond to a return of the function to the calling function or the main function.
[0178] Processes and methods according to the above-described examples can be implemented using computer-executable instructions that are stored or otherwise available from computer- readable media. Such instructions can include, for example, instructions and data which cause or otherwise configure a general purpose computer, special purpose computer, or a processing device to perform a certain function or group of functions. Portions of computer resources used can be accessible over a network. The computer executable instructions may be, for example, binaries, intermediate format instructions such as assembly language, firmware, source code. Examples of computer-readable media that may be used to store instructions, information used, and / or information created during methods according to described examples include magnetic or optical disks, flash memory, USB devices provided with non-volatile memory, networked storage devices, and so on.
[0179] Devices implementing processes and methods according to these disclosures can include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and can take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, the program code or code segments to perform the necessary tasks (e.g., a computer-program product) may be stored in a computer-readable or machine- readable medium. A processor(s) may perform the necessary tasks. Typical examples of form factors include laptops, smart phones, mobile phones, tablet devices or other small form factor personal computers, personal digital assistants, rackmount devices, standalone devices, and so on. Functionality described herein also can be embodied in peripherals or add-in cards. Such functionality can also be implemented on a circuit board among different chips or different processes executing in a single device, by way of further example.
[0180] The instructions, media for conveying such instructions, computing resources for executing them, and other structures for supporting such computing resources are example means for providing the functions described in the disclosure.54 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO
[0181] In the foregoing description, aspects of the application are described with reference to specific examples thereof, but those skilled in the art will recognize that the application is not limited thereto. Thus, while illustrative aspects and examples of the application have been described in detail herein, it is to be understood that the inventive concepts may be otherwise variously embodied and employed, and that the appended claims are intended to be construed to include such variations, except as limited by the prior art. Various features and aspects of the above-described application may be used individually or jointly. Further, aspects and examples can be utilized in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of the specification. The specification and drawings are, accordingly, to be regarded as illustrative rather than restrictive. For the purposes of illustration, methods were described in a particular order. It should be appreciated that in alternate aspects and examples, the methods may be performed in a different order than that described.
[0182] One of ordinary skill will appreciate that the less than (“<”) and greater than (“>”) symbols or terminology used herein can be replaced with less than or equal to (“<”) and greater than or equal to (“>”) symbols, respectively, without departing from the scope of this description.
[0183] Where components are described as being “configured to” perform certain operations, such configuration can be accomplished, for example, by designing electronic circuits or other hardware to perform the operation, by programming programmable electronic circuits (e.g., microprocessors, or other suitable electronic circuits) to perform the operation, or any combination thereof.
[0184] The phrase “coupled to” refers to any component that is physically connected to another component either directly or indirectly, and / or any component that is in communication with another component (e.g., connected to the other component over a wired or wireless connection, and / or other suitable communication interface) either directly or indirectly.
[0185] Claim language or other language in the disclosure reciting “at least one of’ a set and / or “one or more” of a set indicates that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, claim language reciting “at least one of A and B” or “at least one of A or B” means A, B, or A and B. In another example, claim language reciting “at55 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO least one of A, B, and C” or “at least one of A, B, or C” means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language “at least one of’ a set and / or “one or more” of a set does not limit the set to the items listed in the set. For example, claim language reciting “at least one of A and B” or “at least one of A or B” can mean A, B, or A and B, and can additionally include items not listed in the set of A and B.
[0186] The various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the examples disclosed herein may be implemented as electronic hardware, computer software, firmware, or combinations thereof. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.
[0187] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices such as general purposes computers, wireless communication device handsets, or integrated circuit devices having multiple uses including application in wireless communication device handsets and other devices. Any features described as modules or components may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, then the techniques may be realized at least in part by a computer-readable data storage medium comprising program code including instructions that, when executed, performs one or more of the methods, algorithms, and / or operations described above. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may comprise memory or data storage media, such as random access memory (RAM) such as synchronous dynamic random access memory (SDRAM), read-only memory (ROM), nonvolatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), FLASH memory, magnetic or optical data storage media, and the like. The techniques56 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO additionally, or alternatively, may be realized at least in part by a computer-readable communication medium that carries or communicates program code in the form of instructions or data structures and that can be accessed, read, and / or executed by a computer, such as propagated signals or waves.
[0188] The program code may be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, an application specific integrated circuits (ASICs), field programmable logic arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. Such a processor may be configured to perform any of the techniques described in this disclosure. A general purpose processor may be a microprocessor; but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. Accordingly, the term “processor,” as used herein may refer to any of the foregoing structure, any combination of the foregoing structure, or any other structure or apparatus suitable for implementation of the techniques described herein.
[0189] Illustrative aspects of the present disclosure include:
[0190] Aspect 1. An apparatus for processing network-graph embeddings for audio-processing operations, the apparatus comprising: one or more memories configured to store audio data, the audio data including a sequence of audio frames; and one or more processors coupled to the one or more memories and configured to: process the audio data to determine a plurality of features for the sequence of audio frames; process the plurality of features for the sequence of audio frames using a graph-processing machine-learning model to encode a plurality of nodes of a network graph representing the sequence of audio frames into a set of refined feature embeddings; process the set of refined feature embeddings using an audio-operation machine-learning model to generate a set of output values associated with one or more audio-processing operations, wherein an output value of the set of output values is associated with one or more audio frames of the sequence of audio frames; and output the set of output values for the one or more audio-processing operations.57 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO
[0191] Aspect 2. The apparatus of Aspect 1, wherein the network graph includes the plurality of nodes and one or more edges between nodes of the plurality of nodes, wherein a node of the plurality of nodes includes an initial embedding associated with a set of features of the plurality of features determined for an audio frame of the sequence of audio frames, and wherein an edge of the one or more edges identifies a degree of similarity between two audio frames represented by two nodes of the plurality of nodes.
[0192] Aspect 3. The apparatus of Aspect 2, wherein, to generate the edge of the one or more edges, the one or more processors are configured to: determine an edge weight between respective sets of features associated with the two audio frames represented by the two nodes; compare the edge weight to a similarity threshold; and generate the edge based on a determination that the edge weight exceeds the similarity threshold, wherein the edge weight exceeding the similarity threshold is indicative of the degree of similarity of the two audio frames represented by the two nodes.
[0193] Aspect 4. The apparatus of any one of Aspects 1 to 3, wherein the graph-processing machine-learning model is an artificial neural network trained using unsupervised training.
[0194] Aspect 5. The apparatus of any one of Aspects 1 to 4, wherein the audio-operation machine-learning model is trained using supervised training, and wherein the audio-operation machine-learning model was trained based on: access of a training dataset including a sequence of training audio frames, wherein a training audio frame of the sequence of training audio frames was associated with a ground-truth label; iterative applications of the graph-processing machinelearning model and the audio-operation machine-learning model to generate a second set of output values corresponding to the sequence of training audio frames; determination of a loss between the ground-truth label and a corresponding output value of the second set of output values; and adjustment of parameters of the audio-operation machine-learning model based on the loss.
[0195] Aspect 6. The apparatus of any one of Aspects 1 to 5, wherein the one or more audioprocessing operations include detecting one or more keywords from the audio data.58 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO
[0196] Aspect 7. The apparatus of any one of Aspects 1 to 6, wherein the one or more audioprocessing operations include detecting whether an audio-accessory device is engaged with a subject.
[0197] Aspect 8. The apparatus of any one of Aspects 1 to 7, wherein the one or more audioprocessing operations include identifying metadata associated with the audio data.
[0198] Aspect 9. The apparatus of any one of Aspects 1 to 8, wherein the one or more audioprocessing operations include using the set of output values to perform active-noise cancellation on portions of the audio data.
[0199] Aspect 10. The apparatus of any one of Aspects 1 to 9, wherein the graph-processing machine-learning model is trained at a first time point and the audio-operation machine-learning model is trained at a second time point, and wherein the first time point is different from the second time point.
[0200] Aspect 11. The apparatus of any one of Aspects 1 to 10, wherein the plurality of features include at least one of root mean square (RMS) energy or a spectral bandwidth of the sequence of audio frames.
[0201] Aspect 12. The apparatus of any one of Aspects 1 to 11, wherein the one or more processors are configured to perform the one or more audio-processing operations based on the set of output values.
[0202] Aspect 13. The apparatus of any one of Aspects 1 to 12, further comprising an output device configured to generate an output based on the one or more audio-processing operations.
[0203] Aspect 14. The apparatus of any one of Aspects 1 to 13, further comprising at least one microphone configured to capture the audio data.
[0204] Aspect 15. A method of processing network-graph embeddings for audio-processing operations, the method comprising: accessing audio data including a sequence of audio frames; processing the audio data to determine a plurality of features for the sequence of audio frames; processing the plurality of features for the sequence of audio frames using a graph-processing59 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO machine-learning model to encode a plurality of nodes of a network graph representing the sequence of audio frames into a set of refined feature embeddings; processing the set of refined feature embeddings using an audio-operation machine-learning model to generate a set of output values associated with one or more audio-processing operations, wherein an output value of the set of output values is associated with one or more audio frames of the sequence of audio frames; and outputting the set of output values for the one or more audio-processing operations.
[0205] Aspect 16. The method of Aspect 15, wherein the network graph represents the sequence of audio frames, the network graph including the plurality of nodes and one or more edges between nodes of the plurality of nodes, wherein a node of the plurality of nodes includes an initial embedding associated with a set of features of the plurality of features determined for an audio frame of the sequence of audio frames, and wherein an edge of the one or more edges identifies a degree of similarity between two audio frames represented by two nodes of the plurality of nodes.
[0206] Aspect 17. The method of Aspect 16, wherein, to generate the edge of the one or more edges, the method further comprises: determining an edge weight between respective sets of features associated with the two audio frames represented by the two nodes; comparing the edge weight to a similarity threshold; and generating the edge based on a determination that the edge weight exceeds the similarity threshold, wherein the edge weight exceeding the similarity threshold is indicative of the degree of similarity of the two audio frames represented by the two nodes.
[0207] Aspect 18. The method of any one of Aspects 15 to 17, wherein the graph-processing machine-learning model is an artificial neural network trained using unsupervised training.
[0208] Aspect 19. The method of any one of Aspects 15 to 18, wherein the audio-operation machine-learning model is trained using supervised training, and wherein the audio-operation machine-learning model was trained based on: accessing a training dataset including a sequence of training audio frames, wherein a training audio frame of the sequence of training audio frames is associated with a ground-truth label; iteratively applying the graph-processing machine-learning model and the audio-operation machine-learning model to generate a second set of output values corresponding to the sequence of training audio frames; determining a loss between the ground-60 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO truth label and a corresponding output value of the second set of output values; and adjusting parameters of the audio-operation machine-learning model based on the loss.
[0209] Aspect 20. The method of any one of Aspects 15 to 19, wherein the one or more audioprocessing operations include detecting one or more keywords from the audio data.
[0210] Aspect 21. The method of any one of Aspects 15 to 20, wherein the one or more audioprocessing operations include detecting whether an audio-accessory device is engaged with a subject.
[0211] Aspect 22. The method of any one of Aspects 15 to 21, wherein the one or more audioprocessing operations include identifying metadata associated with the audio data.
[0212] Aspect 23. The method of any one of Aspects 15 to 22, wherein the one or more audioprocessing operations include using the set of output values to perform active-noise cancellation on portions of the audio data.
[0213] Aspect 24. The method of any one of Aspects 15 to 23, wherein the graph-processing machine-learning model is trained at a first time point and the audio-operation machine-learning model is trained at a second time point, and wherein the first time point is different from the second time point.
[0214] Aspect 25. The method of any one of Aspects 15 to 24, wherein the plurality of features include at least one of root mean square (RMS) energy or a spectral bandwidth of the sequence of audio frames.
[0215] Aspect 26. The method of any one of Aspects 15 to 25, further comprising performing the one or more audio-processing operations based on the set of output values.
[0216] Aspect 27. The method of any one of Aspects 15 to 26, further comprising using an output device to generate an output based on the one or more audio-processing operations.
[0217] Aspect 28. The method of any one of Aspects 15 to 27, further comprising using at least one microphone to capture the audio data.61 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO
[0218] Aspect 29. A non-transitory computer-readable medium storing executable instructions that, when executed by at least one processor, cause one or more processors to: access audio data including a sequence of audio frames; process the audio data to determine a plurality of features for the sequence of audio frames; process the plurality of features for the sequence of audio frames using a graph-processing machine-learning model to encode a plurality of nodes of a network graph representing the sequence of audio frames into a set of refined feature embeddings; process the set of refined feature embeddings using an audio-operation machine-learning model to generate a set of output values associated with one or more audio-processing operations, wherein an output value of the set of output values is associated with one or more audio frames of the sequence of audio frames; and output the set of output values for the one or more audio-processing operations.
[0219] Aspect 30. The non-transitory computer-readable medium of Aspect 29, wherein the network graph represents the sequence of audio frames, the network graph including the plurality of nodes and one or more edges between nodes of the plurality of nodes, wherein a node of the plurality of nodes includes an initial embedding associated with a set of features of the plurality of features determined for an audio frame of the sequence of audio frames, and wherein an edge of the one or more edges identifies a degree of similarity between two audio frames represented by two nodes of the plurality of nodes.
[0220] Aspect 31. The non-transitory computer-readable medium of Aspect 30, wherein, to generate the edge of the one or more edges, the one or more processors are configured to: determine an edge weight between respective sets of features associated with the two audio frames represented by the two nodes; compare the edge weight to a similarity threshold; and generate the edge based on a determination that the edge weight exceeds the similarity threshold, wherein the edge weight exceeding the similarity threshold is indicative of the degree of similarity of the two audio frames represented by the two nodes.
[0221] Aspect 32. The non-transitory computer-readable medium of any one of Aspects 29 to 31, wherein the graph-processing machine-learning model is an artificial neural network trained using unsupervised training.62 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO
[0222] Aspect 33. The non-transitory computer-readable medium of any one of Aspects 29 to32, wherein the audio-operation machine-learning model is trained using supervised training, and wherein the audio-operation machine-learning model was trained based on: access of a training dataset including a sequence of training audio frames, wherein a training audio frame of the sequence of training audio frames is associated with a ground-truth label; iterative applications of the graph-processing machine-learning model and the audio-operation machine-learning model to generate a second set of output values corresponding to the sequence of training audio frames; determination of a loss between the ground-truth label and a corresponding output value of the second set of output values; and adjustment of parameters of the audio-operation machine-learning model based on the loss.
[0223] Aspect 34. The non-transitory computer-readable medium of any one of Aspects 29 to33, wherein the one or more audio-processing operations include detecting one or more keywords from the audio data.
[0224] Aspect 35. The non-transitory computer-readable medium of any one of Aspects 29 to34, wherein the one or more audio-processing operations include detecting whether an audioaccessory device is engaged with a subject.
[0225] Aspect 36. The non-transitory computer-readable medium of any one of Aspects 29 to35, wherein the one or more audio-processing operations include identifying metadata associated with the audio data.
[0226] Aspect 37. The non-transitory computer-readable medium of any one of Aspects 29 to36, wherein the one or more audio-processing operations include using the set of output values to perform active-noise cancellation on portions of the audio data.
[0227] Aspect 38. The non-transitory computer-readable medium of any one of Aspects 29 to37, wherein the graph-processing machine-learning model is trained at a first time point and the audio-operation machine-learning model is trained at a second time point, and wherein the first time point is different from the second time point.63 Polsinelli Ref. No. 094922-825835PATENTQualcomm Ref. No. 2407274WO
[0228] Aspect 39. The non-transitory computer-readable medium of any one of Aspects 29 to 38, wherein the plurality of features include at least one of root mean square (RMS) energy or a spectral bandwidth of the sequence of audio frames.
[0229] Aspect 40. The non-transitory computer-readable medium of any one of Aspects 29 to 39, wherein the one or more processors are configured to perform the one or more audio-processing operations based on the set of output values.
[0230] Aspect 41. An apparatus including one or more means for performing operations according to any of Aspects 15 to 28.64 Polsinelli Ref. No. 094922-825835
Claims
PATENT Qualcomm Ref. No. 2407274WOCLAIMSWHAT IS CLAIMED IS:
1. An apparatus for processing network-graph embeddings for audio-processing operations, the apparatus comprising: one or more memories configured to store audio data, the audio data including a sequence of audio frames; and one or more processors coupled to the one or more memories and configured to: process the audio data to determine a plurality of features for the sequence of audio frames; process the plurality of features for the sequence of audio frames using a graphprocessing machine-learning model to encode a plurality of nodes of a network graph representing the sequence of audio frames into a set of refined feature embeddings; process the set of refined feature embeddings using an audio-operation machinelearning model to generate a set of output values associated with one or more audioprocessing operations, wherein an output value of the set of output values is associated with one or more audio frames of the sequence of audio frames; and output the set of output values for the one or more audio-processing operations.
2. The apparatus of claim 1, wherein the network graph includes the plurality of nodes and one or more edges between nodes of the plurality of nodes, wherein a node of the plurality of nodes includes an initial embedding associated with a set of features of the plurality of features determined for an audio frame of the sequence of audio frames, and wherein an edge of the one or more edges identifies a degree of similarity between two audio frames represented by two nodes of the plurality of nodes.
3. The apparatus of claim 2, wherein, to generate the edge of the one or more edges, the one or more processors are configured to: determine a distance metric between respective sets of features associated with the two audio frames represented by the two nodes; compare the distance metric to a distance threshold; and65 Polsinelli Ref. No. 094922-825835PATENTQualcomm Ref. No. 2407274WO generate the edge based on a determination that the distance metric exceeds the distance threshold, wherein the distance metric exceeding the distance threshold is indicative of the degree of similarity of the two audio frames represented by the two nodes.
4. The apparatus of claim 1, wherein the graph-processing machine-learning model is an artificial neural network trained using unsupervised training.
5. The apparatus of claim 1, wherein the audio-operation machine-learning model is trained using supervised training, and wherein the audio-operation machine-learning model was trained based on: access of a training dataset including a sequence of training audio frames, wherein a training audio frame of the sequence of training audio frames was associated with a ground-truth label; iterative applications of the graph-processing machine-learning model and the audiooperation machine-learning model to generate a second set of output values corresponding to the sequence of training audio frames; determination of a loss between the ground-truth label and a corresponding output value of the second set of output values; and adjustment of parameters of the audio-operation machine-learning model based on the loss.
6. The apparatus of claim 1, wherein the one or more audio-processing operations include detecting one or more keywords from the audio data.
7. The apparatus of claim 1, wherein the one or more audio-processing operations include detecting whether an audio-accessory device is engaged with a subject.
8. The apparatus of claim 1, wherein the one or more audio-processing operations include identifying metadata associated with the audio data.
9. The apparatus of claim 1, wherein the one or more audio-processing operations include using the set of output values to perform active-noise cancellation on portions of the audio data.66 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO10. The apparatus of claim 1, wherein the graph-processing machine-learning model is trained at a first time point and the audio-operation machine-learning model is trained at a second time point, and wherein the first time point is different from the second time point.
11. The apparatus of claim 1, wherein the plurality of features include at least one of root mean square (RMS) energy or a spectral bandwidth of the sequence of audio frames.
12. The apparatus of claim 1, wherein the one or more processors are configured to perform the one or more audio-processing operations based on the set of output values.
13. The apparatus of claim 1, further comprising an output device configured to generate an output based on the one or more audio-processing operations.
14. The apparatus of claim 1, further comprising at least one microphone configured to capture the audio data.
15. A method of processing network-graph embeddings for audio-processing operations, the method comprising: accessing audio data including a sequence of audio frames; processing the audio data to determine a plurality of features for the sequence of audio frames; processing the plurality of features for the sequence of audio frames using a graphprocessing machine-learning model to encode a plurality of nodes of a network graph representing the sequence of audio frames into a set of refined feature embeddings; processing the set of refined feature embeddings using an audio-operation machinelearning model to generate a set of output values associated with one or more audio-processing operations, wherein an output value of the set of output values is associated with one or more audio frames of the sequence of audio frames; and outputting the set of output values for the one or more audio-processing operations.
16. The method of claim 15, wherein the network graph represents the sequence of audio frames, the network graph including the plurality of nodes and one or more edges between nodes of the plurality of nodes, wherein a node of the plurality of nodes includes an initial67 Polsinelli Ref. No. 094922-825835PATENT Qualcomm Ref. No. 2407274WO embedding associated with a set of features of the plurality of features determined for an audio frame of the sequence of audio frames, and wherein an edge of the one or more edges identifies a degree of similarity between two audio frames represented by two nodes of the plurality of nodes.
17. The method of claim 15, wherein the audio-operation machine-learning model is trained using supervised training, and wherein the audio-operation machine-learning model was trained based on: accessing a training dataset including a sequence of training audio frames, wherein a training audio frame of the sequence of training audio frames is associated with a ground-truth label; iteratively applying the graph-processing machine-learning model and the audiooperation machine-learning model to generate a second set of output values corresponding to the sequence of training audio frames; determining a loss between the ground-truth label and a corresponding output value of the second set of output values; and adjusting parameters of the audio-operation machine-learning model based on the loss.
18. The method of claim 15, wherein the one or more audio-processing operations include using the set of output values to perform active-noise cancellation on portions of the audio data.
19. A non-transitory computer-readable medium storing executable instructions that, when executed by at least one processor, cause one or more processors to: access audio data including a sequence of audio frames; process the audio data to determine a plurality of features for the sequence of audio frames; process the plurality of features for the sequence of audio frames using a graphprocessing machine-learning model to encode a plurality of nodes of a network graph representing the sequence of audio frames into a set of refined feature embeddings; process the set of refined feature embeddings using an audio-operation machine-learning model to generate a set of output values associated with one or more audio-processing68 Polsinelli Ref. No. 094922-825835PATENTQualcomm Ref. No. 2407274WO operations, wherein an output value of the set of output values is associated with one or more audio frames of the sequence of audio frames; and output the set of output values for the one or more audio-processing operations.
20. The non-transitory computer-readable medium of claim 19, wherein the network graph represents the sequence of audio frames, the network graph including the plurality of nodes and one or more edges between nodes of the plurality of nodes, wherein a node of the plurality of nodes includes an initial embedding associated with a set of features of the plurality of features determined for an audio frame of the sequence of audio frames, and wherein an edge of the one or more edges identifies a degree of similarity between two audio frames represented by two nodes of the plurality of nodes.69 Polsinelli Ref. No. 094922-825835