Matching audio using machine learning-based audio representations
Through machine learning-based audio representation technology, detecting and matching input audio with stored audio, the problem of low audio matching efficiency in the prior art is solved, and the application performance of the audio processing system is improved.
Patent Information
- Application Number
- CN202380072292.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-18
- Filing Date
- 2023-10-04
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to effectively match the input audio with the stored audio, resulting in poor performance of audio processing systems in applications such as keyword detection, speech recognition, and speech quality evaluation.
Using machine learning-based audio representations, by detecting input audio clips, generating their representations, and comparing them with multiple stored audio clip representations, the representation of the target audio clip is determined, and then the indexes associated with them are obtained, grouped and sent.
It realizes efficient matching of input audio with stored audio, and improves the performance of the audio processing system in applications such as keyword detection, speech recognition and speech quality evaluation.
Smart Images

Figure CN120077435A_ABST
Abstract
Description
Technical Field
[0001] This application relates to processing audio data. For example, systems and techniques are described for using machine learning-based audio representations (e.g., embedding vectors) to match input audio to stored audio and perform one or more functions based on the result of the match. Background Art
[0002] Electronic devices such as smart phones, tablet computers, wearable electronic devices, smart TVs, etc. have become increasingly popular among consumers. These devices can provide audio (e.g., voice or speech, music, etc.) and / or data communication functionality via wireless or wired networks. In addition, such electronic devices may include other features that provide a variety of functions designed to enhance user convenience. Digital audio includes a large amount of data to meet the needs of consumers and audio providers.
[0003] Speech is an example of audio. Speech applications may rely on being able to effectively model speech using speech models. Speech models can be used by applications such as speech decoding, voice conversion, keyword localization, speech quality assessment, etc. The speech quality, low bit rate, and detection capabilities of these systems depend on the quality of the underlying model. Summary of the Invention
[0004] Systems and techniques for processing audio data are described herein. In some aspects, the systems and techniques described herein relate to an apparatus for encoding audio information, the apparatus including: at least one memory; and at least one processor coupled to the at least one memory and configured to: detect an input audio segment; process the input audio segment to generate a representation of the input audio segment; compare the representation of the input audio segment with a plurality of representations stored in the at least one memory, the plurality of representations representing a plurality of audio segments; based on comparing the representation with the plurality of representations, determine one or more target representations of one or more target audio segments from the plurality of representations stored in the at least one memory; determine one or more indexes associated with the one or more target audio segments; group the one or more indexes; and transmit the grouped one or more indexes.
[0005] In some aspects, the systems and techniques described herein relate to a method for encoding audio information, the method comprising: detecting an input audio segment; processing the input audio segment to generate a representation of the input audio segment; comparing the representation of the input audio segment with a plurality of representations stored in the at least one memory, the plurality of representations representing a plurality of audio segments; based on comparing the representation with the plurality of representations, determining one or more target representations of one or more target audio segments from the plurality of representations stored in the at least one memory; determining one or more indexes associated with the one or more target audio segments; grouping the one or more indexes; and transmitting the grouped one or more indexes.
[0006] In some aspects, the systems and techniques described herein relate to a non-transitory computer-readable medium having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to: detect an input audio segment; process the input audio segment to generate a representation of the input audio segment; compare the representation of the input audio segment with a plurality of representations stored in the at least one memory, the plurality of representations representing a plurality of audio segments; based on comparing the representation with the plurality of representations, determine one or more target representations of one or more target audio segments from the plurality of representations stored in the at least one memory; determine one or more indexes associated with the one or more target audio segments; group the one or more indexes; and transmit the grouped one or more indexes.
[0007] In some aspects, the systems and techniques described herein relate to an apparatus for encoding audio information. The apparatus includes: means for detecting an input audio segment; means for processing the input audio segment to generate a representation of the input audio segment; means for comparing the representation of the input audio segment with a plurality of representations stored in the at least one memory, the plurality of representations representing a plurality of audio segments; means for, based on comparing the representation with the plurality of representations, determining one or more target representations of one or more target audio segments from the plurality of representations stored in the at least one memory; means for determining one or more indexes associated with the one or more target audio segments; means for grouping the one or more indexes; and means for transmitting the grouped one or more indexes.
[0008] In some aspects, the systems and techniques described herein relate to an apparatus for decoding audio information, the apparatus including: at least one memory; and at least one processor coupled to the at least one memory and configured to: receive one or more packetized indexes associated with one or more target audio segments; depacketize the one or more packetized indexes to generate one or more indexes associated with the one or more target audio segments; retrieve the one or more target audio segments from the at least one memory based on the one or more indexes; and combine the one or more target audio segments to generate decoded audio.
[0009] In some aspects, the systems and techniques described herein relate to a method for decoding audio information, the method including: receiving one or more packetized indexes associated with one or more target audio segments; depacketizing the one or more packetized indexes to generate one or more indexes associated with the one or more target audio segments; retrieving the one or more target audio segments from at least one memory based on the one or more indexes; and combining the one or more target audio segments to generate decoded audio.
[0010] In some aspects, the systems and techniques described herein relate to a non-transitory computer-readable medium having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to: receive one or more packetized indexes associated with one or more target audio segments; depacketize the one or more packetized indexes to generate one or more indexes associated with the one or more target audio segments; retrieve the one or more target audio segments from at least one memory based on the one or more indexes; and combine the one or more target audio segments to generate decoded audio.
[0011] In some aspects, the systems and techniques described herein relate to an apparatus for decoding audio information. The apparatus includes: means for receiving one or more packetized indexes associated with one or more target audio segments; means for depacketizing the one or more packetized indexes to generate one or more indexes associated with the one or more target audio segments; means for retrieving the one or more target audio segments from at least one memory based on the one or more indexes; and means for combining the one or more target audio segments to generate decoded audio.
[0012] In some aspects, one or more of the devices described herein are the following, part of the following, and / or include the following: a mobile device or wireless communication device (e.g., a mobile phone or other mobile device), an extended reality (XR) device or system (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a vehicle or a computing device or component of a vehicle, a wearable device (e.g., a network-connected watch or other wearable device), a camera, a personal computer, a laptop computer, a server computer or server device (e.g., an edge or cloud-based server, a personal computer acting as a server device, a mobile device such as a mobile phone acting as a server device, an XR device acting as a server device, a vehicle acting as a server device, a network router, or other device acting as a server device), another device, or a combination thereof. In some aspects, the device includes one camera or multiple cameras for capturing one or more images. In some aspects, the device further includes a display for displaying one or more images, notifications, and / or other displayable data. In some aspects, the device may include one or more sensors (e.g., one or more inertial measurement units (IMUs), such as one or more gyroscopes, one or more gyro testers, one or more accelerometers, any combination thereof, and / or other sensors). In some aspects, the device may include a receiver configured to receive information or data, a transmitter configured to send information or data, and / or a transceiver configured to receive and send information or data.
[0013] The above aspects related to any one of the method, device, and computer-readable medium can be used alone or in any suitable combination.
[0014] This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used alone to determine the scope of the claimed subject matter. The subject matter should be understood with reference to the appropriate portions of the entire specification of this patent, any or all of the drawings, and each claim.
[0015] The foregoing and other features and embodiments will become more apparent when referring to the following specification, claims, and drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Examples of various specific implementations are described in detail below with reference to the following drawings:
[0017] Figure 1 is a block diagram illustrating an example system-on-chip (SoC) that may include an audio processing system according to some examples;
[0018] Figure 2 is a diagram illustrating an example of an audio processing system according to aspects of the present disclosure;
[0019] Figure 3 is a block diagram illustrating an example operation of an audio representation search and comparison engine of an audio processing engine according to aspects of the present disclosure; Figure 2 of the audio processing engine;
[0020] Figure 4A is a diagram illustrating an example of a speech encoder and a speech decoder according to aspects of the present disclosure;
[0021] Figure 4B is a diagram illustrating an example of the encoding of an index according to aspects of the present disclosure;
[0022] Figure 5 is a diagram illustrating an example of audio framing using fixed-length segments according to aspects of the present disclosure;
[0023] Figure 6 is a diagram illustrating an example of audio framing using variable-length segments according to aspects of the present disclosure;
[0024] Figures 7A to 7C is a diagram illustrating an example of a neural network according to some examples;
[0025] Figure 8 is a block diagram illustrating an example of a deep convolutional network (DCN) according to aspects of the present disclosure;
[0026] Figure 9 is a flowchart illustrating an example of a process for encoding audio information according to aspects of the present disclosure;
[0027] Figure 10 is a flowchart illustrating an example of a process for decoding audio information according to aspects of the present disclosure;
[0028] Figure 11 is an example computing device architecture of an example computing device that can implement the various technologies described herein. Detailed Description
[0029] Certain aspects and implementations of the present disclosure are provided below. Some of these aspects and implementations can be applied independently, and some of them can be applied in combination, which will be apparent to those skilled in the art. In the following description, specific details are set forth for purposes of explanation in order to provide a thorough understanding of the implementations of the present application. However, it will be apparent that the various implementations can be practiced without these specific details. The accompanying drawings and description are not intended to be restrictive.
[0030] The following description merely provides example embodiments and is not intended to limit the scope, applicability, or configuration of the present disclosure. On the contrary, the following description of the example embodiments will provide those skilled in the art with an enabling description for implementing the example embodiments. It should be understood that various changes may be made to the functions and arrangements of the elements without departing from the spirit and scope of the present application as set forth in the appended claims.
[0031] Systems and devices can benefit from the ability to compare received audio information (e.g., speech or voice information, music, recorded sounds, etc.) with stored audio information to determine the similarity between the received audio information and the stored audio information. For example, a system can use such comparisons to perform keyword detection, speech recognition, speech quality assessment, and other tasks. However, audio information can include a large amount of data, and storing the data in such cases can result in high storage costs. There is a need for more efficient systems and techniques for matching input audio with stored audio.
[0032] Described herein are systems, apparatuses, electronic devices, methods (also referred to as processes), and computer-readable media (collectively referred to herein as "systems and techniques") for using machine learning-based audio representations (e.g., embedding vectors) to match input audio with stored audio and perform one or more functions based on the results of the match. For example, a machine learning system can process stored audio (e.g., pulse code modulation (PCM) speech samples) to generate a representation of the stored audio (e.g., a feature vector or an embedding vector) that represents the characteristics of the audio data. The representation of the stored speech can be stored in an audio representation storage device (e.g., an audio representation database). Such representations can be referred to as deep audio representations (e.g., deep speech representations). The same machine learning system can process input audio to generate a representation of the input audio (e.g., a feature vector or an embedding vector) that represents the characteristics of the input audio data. The systems and techniques can then compare the representation generated for the input audio with the representations generated for the stored audio to determine one or more closest matching representations in the audio representation storage device.
[0033] Audio can include speech or voice, music, recorded sounds, any combination thereof, and / or other types of audio. For illustrative purposes, speech will be used to describe the examples herein. However, those of ordinary skill in the art will understand that the systems and techniques described herein are applicable to any type of audio information.
[0034] Systems and techniques can perform one or more functions based on matching representations from an audio representation storage device. In one illustrative example, the systems and techniques can be used to implement a speech encoder, a speech decoder, or a combined speech encoder-decoder (codec). In such examples, the representations stored in the audio representation storage device can represent speech from multiple people (or “speakers”). For example, the speech encoder can include an audio memory storing speech from multiple people. The speech decoder can also include an audio storage device storing speech from multiple people. A machine learning model on the encoder can process the speech stored at the encoder to generate a representation (e.g., a feature or an embedding vector) of the stored speech. For example, the stored speech can be divided into segments. The segments can be fixed-length segments or can be variable-length segments. In the case where the audio is divided into variable-length segments, the encoder can resample the segments to convert the segments with variable lengths into audio segments with fixed lengths (e.g., such that the speech segments have the same length or duration). The machine learning model can generate one or more corresponding representations for each segment (e.g., generate 12 feature vectors with 512 values for each segment with a duration or length of 120 milliseconds).
[0035] The encoder can receive speech input from a person speaking. A machine learning model (the same model used to generate the representations stored in the audio representation storage device) can generate one or more representations (e.g., one or more features or embedding vectors) of the speech input. For example, similar to the stored speech, the encoder can divide the input speech into segments (e.g., fixed-length or variable-length segments) and can generate one or more corresponding representations for each segment. Then, the encoder can determine the representations of the speech stored in the audio representation storage device that match the representation of the speech input. For example, a distance metric (e.g., (e.g., mean squared error (MSE), root mean squared error (RMSE), sum of squared errors (SSE), or other distance metrics) can be used to compare the representations of the stored speech with the representation of the speech input. The encoder can determine an index identifying the location in the audio storage device of the speech segment corresponding to the stored representation that is determined to match the segment of the speech input.
[0036] Once the index of the matching utterance is determined, the encoder can encode the index into a bitstream (e.g., by quantizing and / or entropy coding the index value). The encoder can then store the bitstream and / or send the bitstream to the speech decoder. The speech decoder can decode the bitstream (e.g., by performing inverse quantization and / or entropy decoding) to determine the index represented in the bitstream. The decoder can then use each corresponding index to identify the location in the audio storage device of the matching speech segment corresponding to the input speech. The decoder can retrieve the speech segment from the audio storage device and can combine (e.g., concatenate or otherwise combine) the speech segments to generate a reconstructed or decoded speech that is similar to the input speech (e.g., sounds similar to the input speech).
[0037] In another illustrative example, the systems and techniques can be used for voice activation. In such examples, the representations stored in the audio representation storage device can represent keywords that, if recognized in the input speech, can be used to activate a device including a voice activation system to perform one or more functions (e.g., launch one or more applications, play music, set a timer, etc.). The voice activation system can receive a speech input, generate one or more representations (e.g., features or embedding vectors) of the speech input using a machine learning system, and determine whether the representation of the speech input matches one of the representations of the keywords stored in the audio representation storage device. If a match is determined, the voice activation system can activate the corresponding function of the device.
[0038] Other examples of applications for which the systems and techniques can be used include speech quality assessment (e.g., by measuring how well the representation of the speech input matches the stored representation of the stored speech), voice conversion (e.g., by outputting the stored speech corresponding to the stored representation of the speech that is determined to match the speech representation generated for the speech input), audio event detection (e.g., by identifying and outputting an indication of the stored speech corresponding to the stored representation of the speech that is determined to match the speech representation generated for the speech input), etc.
[0039] Aspects of the present disclosure will be described with reference to the figures.
[0040] Figure 1Illustrates an example embodiment of a system-on-chip (SoC) 100, which may include a central processing unit (CPU) 102 or a multi-core CPU configured to perform one or more of the functions described herein. Parameters or variables (e.g., neural signals and synaptic weights), system parameters associated with a computing device or model (e.g., a neural network model having parameters such as weights, biases, and / or other parameters), latencies, frequency bin information, task information, and other information may be stored in a memory block associated with the neural processing unit (NPU) 108, stored in a memory block associated with the CPU 102, stored in a memory block associated with the graphics processing unit (GPU) 104, stored in a memory block associated with the digital signal processor (DSP) 106, stored in the memory block 118, and / or may be distributed across multiple blocks. Instructions executed at the CPU 102 may be loaded from a program memory associated with the CPU 102 or may be loaded from the memory block 118.
[0041] The SoC 100 may also include additional processing blocks customized for specific functions, such as the GPU 104, the DSP 106, the connectivity block 110 (which may include fifth-generation (5G) connectivity, fourth-generation long-term evolution (4G LTE) connectivity, Wi-Fi connectivity, USB connectivity, Bluetooth connectivity, etc.), and the multimedia processor 112 that may, for example, detect and recognize gestures, speech, and / or other interactive user actions or inputs. In one embodiment, the NPU 108 is implemented in the CPU 102, the DSP 106, and / or the GPU 104. The SoC 100 may also include a sensor processor 114, one or more image signal processors (ISPs) 116, and / or an audio processing system 120. In some examples, the sensor processor 114 may be associated with or connected to one or more sensors for providing sensor inputs to the sensor processor 114. For example, the one or more sensors and the sensor processor 114 may be provided in the same computing device, coupled to the same computing device, or otherwise associated with the same computing device.
[0042] In some examples, the one or more sensors may include one or more microphones for receiving audio inputs. The audio inputs may include ambient sounds near a computing device associated with the SoC 100 and / or may include speech from a user of the computing device associated with the SoC 100 or one or more other users. In some cases, the computing device associated with the SoC 100 may additionally or alternatively be communicatively coupled to one or more peripheral devices (not shown) and / or be configured to communicate with one or more remote computing devices or external resources, for example, using a wireless transceiver and a communication network such as a cellular communication network.
[0043] In some cases, the audio input received by one or more microphones (and / or other sensors) can be processed by the SoC 100, CPU 102, DSP 106, NPU 108, and / or the audio processing system 120. For example, the audio processing system 120 can utilize a machine learning model (e.g., implemented using the CPU 102, DSP 106, and / or NPU 108) to process the stored audio to generate a representation of the stored audio that represents the characteristics of the audio data. The representation of the stored speech can be stored in an audio representation storage device (e.g., an audio representation database), which can be part of the audio processing system 120, part of the SoC 100, and / or external to the SoC 100 (e.g., on a separate chip, device, or system, such as a cloud server). The audio processing system 120 can utilize the same machine learning model to process the input audio received by one or more microphones (and / or other sensors) to generate a representation of the input audio that represents the characteristics of the input audio data. The audio processing system 120 can compare the representation generated for the input audio with the representation generated for the stored audio to determine one or more closest matching representations in the audio representation storage device. The audio processing system 120, the SoC 100, the applications communicating with the SoC 100, and / or other systems, devices, or components can perform one or more functions based on the matching representations from the audio representation storage device. Aspects regarding the audio processing system 120 will be discussed in more detail below with respect to Figure 2 and / or other figures.
[0044] Figure 2 is a diagram illustrating an example of an audio processing system 220 according to aspects of the present disclosure. The audio processing system 220 is Figure 1 an illustrative example of the audio processing system 120. As Figure 2 shown, the audio processing system 220 includes an audio storage device 224, an audio representation storage device 226, a representation generation engine 228, and an audio representation search and comparison engine 230.
[0045] The audio storage device 224 can store audio data, and the audio representation storage device 226 can store representations of the audio data stored in the audio storage device 224. As described below, the representations are generated by the representation generation engine 228. In some aspects, the audio storage device 224 can be a first database (e.g., an audio representation database), and the audio representation storage device 226 can be a second database separate from the first database (e.g., an audio representation database). In other examples, the audio storage device 224 and the audio representation storage device 226 can be parts of the same database, such as where the content of the audio storage device 224 is stored separately from the content of the audio representation storage device 226. Although the audio storage device 224 and the audio representation storage device 226 are shown as parts of the audio processing system 220, in some cases, the audio storage device 224 and / or the audio representation storage device 226 can be separate from the audio processing system 220.
[0046] The audio in the audio storage device 224 and the audio from the audio source can include speech (or voice), music, recorded sounds, any combination thereof, and / or other types of audio. For illustrative purposes, speech will be used to describe the examples herein. However, one of ordinary skill in the art will understand that the systems and techniques described herein apply to any type of audio information. For example, the audio data stored in the audio storage device 224 can include pulse code modulation (PCM) speech samples. In some cases, speech can be divided into segments (e.g., diphones including speech data from the middle of one phoneme of the speech to the middle of the next phoneme of the speech). In some aspects, the segments can have a fixed length, where all segments of the speech have the same (or common) length or duration. In some aspects, the segments can have a variable length, where the lengths of different segments can vary. In one illustrative example, the length of the segments can vary from 30 milliseconds (ms) to 150 ms. In the case where the audio is divided into variable length segments, the audio processing system 220 can resample the segments to convert the segments with variable lengths into audio segments with a fixed length (e.g., such that the speech segments have the same length or duration), as described below with respect to Figure 5 described.
[0047] The representation generation engine 228 can include a machine learning system (or machine learning model) that can be used to generate representations of the audio data from the audio stored in the audio storage device 224 and representations of the audio from the audio source 222. In some cases, the representation generation engine 228 can use an NPU (e.g., Figure 1implemented and / or operable in conjunction with an NPU (e.g., NPU 108), a DSP (e.g., DSP 106), a CPU (e.g., CPU 102), and / or other processors to apply a machine learning system to audio data. The machine learning system can be a neural network model trained to process audio data and generate one or more representations of the audio data. The representation can be a feature vector (also referred to as an embedding vector) representing the audio data from which the representation is generated. Illustrative examples of neural networks that can be used as part of the representation generation engine 228 include convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), multi-layer perceptron (MLP) neural networks, transformer neural networks, and the like.
[0048] In examples where the machine learning system includes a neural network, supervised learning, self-supervised learning, and / or other training techniques can be used to train the neural network. In some cases, a combination of supervised and self-supervised learning (and / or other training techniques) can be used to train the machine learning system. In cases where self-supervised learning is used to train the neural network, no labeled ground truth data is required. In such cases, the neural network can be trained on a large amount of speech (e.g., 50,000 hours of speech) to provide a very accurate and general speech model that can be stored in the audio representation storage device 226. New use cases will only require updating the database with new audio data, which can be done on the device (not offline), and there is no need to retrain the speech representation model stored in the audio representation storage device 226.
[0049] As described above, the representation generation engine 228 can use a machine learning model (e.g., a trained neural network) to process audio (e.g., a snippet of speech data) stored in the audio storage device 224 to generate a representation characterizing the stored audio data. In some cases, the representation can be a feature vector or embedding vector generated by the neural network. As described above, the representation of the stored speech can be stored in the audio representation storage device 226 (e.g., an audio representation database).
[0050] The representation generation engine 228 can also use a machine learning model (e.g., a trained neural network) to process input audio from the audio source 222 to generate one or more representations (e.g., feature vectors or embedding vectors) characterizing the input audio data. In an illustrative example, the audio source 222 can include a microphone (or microphones), and the input audio can include speech from a person captured by the microphone.
[0051] The audio representation search and comparison engine 230 may obtain (e.g., receive) a representation of the input audio from the representation generation engine 228 and may obtain (e.g., retrieve) the stored representations from the audio storage device 224. The audio representation search and comparison engine 230 may compare the representation of the input audio with the stored representations to determine one or more of the stored representations that are closest to the representation of the input audio. In some cases, the representation search and comparison engine 230 may use a distance calculation or metric (e.g., MSE, RMSE, SSE, or other distance metric) to compare the representation of the stored speech with the representation of the speech input to determine whether one or more of the representations of the stored speech match the representation of the speech input.
[0052] Figure 3 is an illustration Figure 2 of an example operation of the audio representation search and comparison engine 230 of the audio processing engine. As described above, the representations of the input audio and the stored audio may include feature vectors. For example, the feature vector 325 is a representation of the input speech 302, and the feature vectors in the audio representation storage device 326 are representations of the audio data from the audio storage device 224. Each of the feature vectors may include a specific number of values, such as 256 values, 512 values, or other numbers of values. The representation search and comparison engine 230 may use a distance calculation or metric (e.g., MSE, RMSE, SSE, or other distance metric) to determine the difference (or distance) between one or more of the values in the feature vector 325 representing the input speech 302 and one or more of the values in the feature vectors in the audio representation storage device 326. One or more of the feature vectors in the audio representation storage device 326 that produce the smallest (or smallest) distance value may be selected as the matching feature vectors.
[0053] In some cases, a specific number of representations may be generated for each segment of the stored audio and for each segment of the input audio, and that number of representations may be used by the audio representation search and comparison engine 230 to perform the comparison. For example, the audio processing system 220 may be divided into segments, where each segment has a length or duration of 120 ms. For an audio segment with a length of 120 milliseconds, the audio processing system 220 may generate features every 10 ms, resulting in 12 feature vectors for the segment. As described above, each feature vector may have a specific number of values (e.g., 256 values, 512 values, etc.). For example, referring to Figure 3, the set of feature vectors 327 representing segments of the input speech 322 may include 12 feature vectors, and each feature vector may include 512 values. Similarly, the set of feature vectors 329 representing segments of the stored speech in the audio representation storage device 326 may include 12 feature vectors, and each feature vector may include 512 values. The audio representation search and comparison engine 230 may use distance calculations (e.g., MSE, RMSE, SSE, or other distance metrics) to determine the difference between the values of the set of feature vectors 327 and the values of each set of feature vectors in the audio representation storage device 326. Based on the distance calculation (which may be considered a matching cost), the audio representation search and comparison engine 230 may determine that the set of feature vectors 327 and the set of feature vectors 329 match based on the difference between the values of the set of feature vectors 327 and the values of the set of feature vectors 329, thereby producing a minimum value from the determined differences.
[0054] As noted herein, the audio representation search and comparison engine 230 may select segments based on the best match to the input (e.g., based on the matching cost) for each segment. However, it may be beneficial to select segments based on the temporal evolution of the segments (e.g., by selecting segments to minimize discontinuities between segments). In some aspects, the audio representation search and comparison engine 230 may perform beam search and concatenation operations (e.g., Viterbi beam search and concatenation operations) to determine a match between one or more representations of the input speech and one or more representations stored in the audio representation storage device 226. For example, in addition to the above-mentioned matching cost (how much each output segment varies relative to the matching input segment), beam search and concatenation operations may be implemented by using a joint cost (e.g., indicating how different or "jumpy" each output segment is from the previous output segment). The joint cost may be calculated using any suitable technique, such as using a feature representation of deep learning (e.g., the distance between the feature representation at the end of one segment and the feature representation at the start of the next segment), which may be generated using a neural network model or other machine learning models.
[0055] Based on determining a matching representation (or set of representations) from the audio representation storage device 326, the audio processing system 220 may generate an output 232. The output 232 may include the best matching audio segment from the audio storage device 224 (e.g., for speech conversion, audio event detection, or other applications), a matching accuracy or score (e.g., for voice activation, speech quality assessment, or other applications) (such as a value between 0 and 1) based on the match between one or more representations of the input audio and one or more representations of the stored audio stored in the audio representation storage device 226, a bitstream (e.g., by Figure 4A(the index bitstream generated by the described speech encoder) and / or other types of outputs. For example, the best-matching audio segment can be an audio segment with a representation (or set of representations, such as a set of feature vectors 329) that is determined to match the representation (or set of representations, such as a set of feature vectors 327) of the input speech segment.
[0056] As described above, in some cases, the audio processing system 220 can be part of a speech encoder, a speech decoder, or a combined speech encoder-decoder (referred to as a "codec"). Figure 4A is a diagram illustrating an example of a speech encoder 440 and a speech decoder 450 according to aspects described herein. The speech encoder 440 includes the audio processing system 220 and can use the audio processing system 220 to generate an index bitstream 442. The index bitstream 442 represents an index that identifies (or points to) the location in the audio storage device 224 of the stored speech segment corresponding to the speech representation (in the audio representation storage device 226), where each speech representation is determined to match the corresponding representation of the input speech segment. The speech encoder can store the index bitstream or can send the index bitstream to a device that includes the speech decoder. The speech decoder 450 can receive the index bitstream 442 from the encoder.
[0057] The audio processing system 220 of the speech encoder 440 includes an audio storage device 224 and an audio representation storage device 226, as described with respect to Figure 2 The speech decoder 450 also includes an audio storage device 424, which includes the same audio data as the audio data stored in the audio storage device 224 of the audio processing system 220. The audio storage device 224 and the audio storage device 424 can store speech from multiple people. The representations (e.g., feature vectors, such as Figure 3 feature vectors) stored in the audio representation storage device 226 of the audio processing system 220 in the speech encoder 440 can be generated by a representation generation engine 228 to represent the speech (from multiple people) stored in the audio storage device 224. For example, a machine learning model (e.g., a neural network model) of the representation generation engine 228 in the audio processing system 220 can process the speech stored in the audio storage device 224 to generate a representation that represents the stored speech. For example, the stored speech can be divided into segments. The segments can be fixed-length segments or can be variable-length segments, as described with respect to Figure 5 and Figure 6As described. The machine learning model can generate one or more corresponding representations for each segment. In an illustrative example, the audio data can be partitioned such that each segment has a duration or length of 120 ms, and the machine learning model of the representation generation engine 228 can generate feature vectors for every 10 seconds of each segment (e.g., generate 12 feature vectors with 512 values for a 120 ms segment).
[0058] The speech encoder 440 can obtain (e.g., receive via one or more microphones, retrieve from a storage device, etc.) the input speech of the speaker. The machine learning model of the representation generation engine 228 processes the input speech to generate one or more representations (e.g., feature vectors, such as Figure 3 feature vectors). For example, similar to the speech stored in the audio storage device 224, the encoder can partition the input speech into segments (e.g., fixed-length or variable-length segments) and can generate one or more corresponding representations for each segment of the input speech.
[0059] The audio representation search and comparison engine 230 of the speech encoder 440 can search the audio representation storage device 226 for stored representations that match the representations generated for the input speech. For example, the audio representation search and comparison engine 230 can perform a distance calculation or metric (e.g., MSE, RMSE, SSE, or other distance metric) to determine whether any of the representations stored in the audio representation storage device 226 match the representations generated for the speech input. In one example, as described with respect to Figure 3 the representation search and comparison engine 230 can use a distance calculation (e.g., MSE, RMSE, SSE, or other distance metric) to determine the difference (or distance) between one or more of the values in the feature vectors 325 representing the input speech 302 and one or more of the values in the feature vectors in the audio representation storage device 326. One or more feature vectors in the audio representation storage device 326 that produce the minimum (e.g., smallest) distance value can be selected as the matching feature vectors. In another example, also as described with respect to Figure 3As described, the audio representation search and comparison engine 230 may use distance calculations (e.g., MSE, RMSE, SSE, or other distance metrics) to determine the differences between the values of various sets of features (e.g., set of feature vectors 327) generated for the input utterance 322 and the values of the various sets of feature vectors in the audio representation storage device 326. Based on the distance calculations, the audio representation search and comparison engine 230 may determine that the set of feature vectors 327 representing the input utterance 322 and the set of feature vectors 329 stored in the audio representation storage device 326 match based on the difference between the values of the set of feature vectors 327 and the set of feature vectors 329, thereby producing a minimum value from the determined differences between the values of the set of feature vectors 327 and all (or a subset) of the stored feature vectors.
[0060] Based on the comparison, the speech encoder 440 may determine a set of indices of various speech segments, each corresponding to a respective stored representation of a corresponding segment determined to match the speech input. Each index in the set of indices identifies the location in the audio storage device of a particular speech segment corresponding to one or more stored representations of one or more segments determined to match the speech input. In one illustrative example, referring to Figure 3 , the indices in the set of indices may indicate the location in the audio storage device 224 of a speech segment corresponding to the set of feature vectors 329 of the stored feature vectors that is determined to match a particular segment representing the input utterance 322 (e.g., the speech segment from which the set of feature vectors 329 was generated using the representation generation engine 228).
[0061] After determining the set of indices of the speech input, the speech encoder 440 may encode the set of indices into an index bitstream 442. In one illustrative example, the speech encoder 440 may quantize and / or entropy code (e.g., using an arithmetic encoder) the index values of the set of indices to encode the set of indices into the index bitstream 442. The encoder 440 may store the bitstream and / or may send the bitstream to the speech decoder 450. The speech encoder 440 may use any suitable technique to encode the indices, which in some cases may depend on the signal itself, transmission requirements, and / or other factors. In one illustrative example, the audio storage device 224 may store N segments of a fixed length L, in which case the speech encoder 440 may need to send B = log2(N) for each segment, for each fixed length L. Additional coding techniques may also be performed. For example, if one or more of the expected segments occur more frequently than the others, the speech encoder 440 may perform run-length encoding, which can represent the more frequently occurring segments with fewer bits and thus reduce the overall bit rate. In another example, if several segments are expected to be consecutive (e.g., obtaining a maximum length sequence from a database), the speech encoder 440 may again use run-length decoding, in which case a single bit may be used to signal to the speech decoder 450 whether consecutive segments in the database should be used or whether there should be a transition to another location in the database. In another example, such as if a fixed bit rate is desired, the speech encoder 440 may send the segment indices at a higher bit rate (which will still be very low compared to traditional speech decoders) per frame (e.g., every 20 ms). If a packet is lost during transmission, such examples may provide high error resilience because most packets will be recovered since they are the previous index plus 1.
[0062] In another illustrative example of decoding the indices, it may be assumed that with N bits, all frame indices in the database can be decoded. The value (2^N)-1 is reserved as a special character (SC). In such examples, various techniques may be performed by the speech encoder 440 to encode the sequence of database indices. For example, the speech encoder 440 may each use N bits to decode the sequence of frame indices, which can help with error resilience. Figure 4Bis a diagram illustrating such an example. As shown, N1, N2, N3, N4 of the matching audio segment 457 stored in the audio storage device 424 (for which the index will be encoded) are each decoded with N bits, resulting in a total of 4 * N bits. The resulting frame indices 459 are shown for each of N1, N2, N3, and N4 respectively. If the matching frame indices are decoded every 20 ms, decoding each index individually can be a high-quality option. In cases where encoding may be delayed, segment switching can be identified. In such cases, a method such as run-length decoding can be used to efficiently decode each matching segment, as described above. For example, N bits can be used to represent each frame in the matching segment (N1) with the first frame index. The speech encoder 440 can then perform run-length decoding on the index and can insert a special character (SC) between the run-length and the next index to identify the boundary. In some cases, the speech encoder 440 can decode each of N1, 4, and SC with N bits, for a total of 3 * N bits.
[0063] The speech decoder 450 can decode the index bitstream 442 to determine the indices represented in the bitstream. In some cases, the speech decoding engine 452 of the speech decoder 450 can perform inverse quantization and / or entropy decoding (e.g., using an arithmetic decoder) to decode the index bitstream 442. For example, the speech decoding engine 452 can perform the inverse operations of the encoding operations performed by the speech encoder 440, such as the inverse of the run-length decoding or other decoding techniques described above.
[0064] As described above, the audio storage device 424 of the speech decoder 450 stores the same audio data as that stored in the audio storage device 224 of the audio processing system 220. In this case, an index set can be used to identify the position of the speech segment determined to match the input speech in the audio storage device 224 (based on the comparison of the representation of the input speech with the representations stored in the audio representation storage device 326). Using the decoded index set, the segment retrieval engine 454 of the speech decoder 250 can retrieve speech segments from the audio storage device 424. The segment combination engine 456 of the speech decoder 250 can then combine the retrieved speech segments to generate a decoded speech 458 (also referred to as a reconstructed speech) that is similar to (e.g., sounds similar to) the input speech. In one example, the input speech can be a person saying "Hello, my name is Bob". By storing the speech of various users in the audio storage device 224 and the audio storage device 424, the segment retrieval engine 454 can retrieve audio segments from the audio storage device 424 that can be used to construct a phrase identical to the input speech. Thus, speech segments from a voice different from that of the person providing the speech input will be used to construct the decoded speech 458, but based on the comparison performed by the audio representation search and comparison engine 230, the voice will have speech characteristics similar to those of the human providing the input speech. In some cases, to combine the retrieved speech segments, the segment combination engine 456 can concatenate each retrieved segment to the subsequent segment to generate a complete speech output as the decoded speech 458. In some cases, the retrieved speech segments may overlap, in which case the segment combination engine 456 can align the segments so that the audio segments are properly aligned. For example, if the output is based on segments from different parts of the audio storage device 224, the segments may not be fully aligned (and thus may overlap), in which case concatenating the audio segments (e.g., the PCM signals of the audio segments) may result in glitches or other errors. In such cases, the segment combination engine 456 can align the segments (e.g., using the maximum correlation point) and can use overlap-and-add techniques to ensure that the transition between segments is as smooth as possible.
[0065] By using the audio processing system 220 to generate the index bitstream 442, the speech encoder 440 can greatly reduce the amount of data that needs to be provided to the speech decoder 450 for reconstructing the input speech. For example, the speech encoder 440 does not need to send the actual speech (or the decoded version of the actual speech) to the speech decoder 450. In an illustrative example, the audio storage device 224 and the audio storage device 424 can each store two and a half hours of speech, and the stored speech can be divided into segments, such as diphone segments. For example, two and a half hours of speech can correspond to approximately 96,000 different diphone segments. In some cases, as described herein, the length of the diphone segment can vary from, for example, 30 ms to 150 ms, in which case, on average, the length of each diphone segment is about 100 milliseconds. In other cases, the diphone segments are fixed-length segments (e.g., 20 ms, 40 ms, 60 ms, etc.). Based on the log of 96,000 2 , the speech encoder 440 can use 17 bits to represent a number between one and 96,000 (corresponding to 96,000 different diphone segments). Thus, the speech encoder 440 can use 17 bits to represent one diphone segment. Using the above variable-length example (where the average length of the segment is 100 ms), there will be 10 segments per second, and thus, on average, the speech encoder 440 can utilize only 170 bits per second (10 multiplied by 17 bits) for the index bitstream 442.
[0066] In some aspects, the audio processing system 220 can be part of a voice activation system. For example, in such aspects, the representations stored in the audio representation storage device 226 can represent keywords. The keywords can be used as part of voice activation. For example, if one of the keywords is recognized in the input speech from the audio source 222, the output 232 can be used to activate a device including the voice activation system to perform one or more functions (e.g., launch one or more applications, play music, set a timer, etc.). In one example, the representation generation engine 228 of the voice activation system can receive speech input from the audio source 222 and can use a machine learning system to generate one or more representations of the speech input (e.g., features or embedding vectors, such as Figure 3 the feature vector 325). The audio representation search and comparison engine 230 can then determine whether at least the representation of the speech input matches one of the representations of the keywords stored in the audio representation storage device 226. If a match is determined, the voice activation system can generate the output 232 to activate the corresponding functions of the device including the audio processing system 220 or a separate device.
[0067] Another example of an application for which the audio processing system 220 can be used is speech quality assessment. For example, the output 232 can indicate how well a representation of a speech input matches a stored representation of stored speech. In one example, the representation generation engine 228 of the voice activation system can receive a speech input from the audio source 222 and can use a machine learning system to generate one or more representations of the speech input (e.g., features or embedding vectors, such as Figure 3 feature vector 325). The audio representation search and comparison engine 230 can then determine whether at least the representation of the speech input matches one of the representations of keywords stored in the audio representation storage device 226. If a match is determined, the voice activation system can generate an output 232 that can indicate how well the representation of the speech input matches the representation stored in the audio representation storage device 226. In some cases, the output 232 can include a match score (such as a value between 0 and 1) that indicates how well the representation of the speech input matches the representation stored in the audio representation storage device 226. In some aspects, the match score can be compared to a match threshold (e.g., a value of 0.7, 0.8, 0.85, 0.9 or other value) to determine whether the input audio has high quality. In one example, the output 232 can include a match score of 0.8 and the match threshold can be a value of 0.75, in which case the input audio can be determined to have high quality (or sufficient quality for a given application). Such solutions can be used to determine the quality of a communication, such as a cellular phone conversation, a conference call using a conferencing tool, etc.
[0068] Other examples of applications for which the audio processing system 220 can be used include voice conversion (e.g., by outputting stored speech corresponding to a stored speech representation determined to match the speech representation generated for a speech input) and audio event detection (e.g., by identifying and outputting an indication of stored speech corresponding to a stored speech representation determined to match the speech representation generated for a speech input), etc.
[0069] Figure 5 FIG. 500 is a diagram illustrating an example of audio framing using fixed-length segments. As Figure 5 shown in the illustrative example of Figure 5As shown, for an audio segment 562 of the input speech, it is indicated that the generation engine 228 can generate feature vectors for each 10 - second duration of the audio segment 562, thereby producing 12 feature vectors. In Figure 5 the example of, each feature vector has 512 values. For example, a set of feature vectors 564 (including 12 feature vectors, where each feature vector has 512 values) can be generated for the audio segment 562. The representation stored in the audio representation storage device 526 may also include 12 feature vectors (where each feature vector has 512 values) for each speech segment with a length of 120 ms.
[0070] Using fixed - length speech segments allows the audio representation search and comparison engine 230 to perform difference or distance calculations (e.g., using MSE, etc.) and determine the output 232 using a set of representations with common dimensions (e.g., each set of feature vectors includes an array or matrix of size 12×512).
[0071] Figure 6 FIG. 600 is a diagram illustrating an example of audio framing using variable - length segments. For example, speech data stored in an audio storage device 624 (e.g., stored as PCM samples) can be divided into segments with variable lengths, for example, between 30 ms and 150 ms. Similarly, the input speech is also divided into segments with variable lengths (e.g., between 30 ms and 150 ms). As described above, if the input speech is divided into segments of the same length as the stored speech segments, the resulting representations can have common dimensions (e.g., each set of feature vectors representing the corresponding speech segment can be a matrix of dimension 12×512), and the distance or difference between the representation of the input speech and the representation of the stored speech can be easily determined. However, in Figure 6 the case of variable - length framing, it is necessary to compare sets of representations of segments of different lengths (e.g., it may be necessary to compare the set of representations of an input segment with a duration of 150 ms with the set of representations of a stored segment with a duration of 120 ms), which can be difficult due to the dimensions of different sets of representations having different dimensions.
[0072] To address such issues, the audio processing system 220 may resample one or more audio segments to convert the one or more variable-length audio segments into one or more fixed-length audio segments. Resampling may result in the representation of the input speech and the representation stored in the audio representation storage device 626 having fixed framing (with a fixed length), while the audio segments stored in the audio storage device 624 have variable lengths. For example, the representation of each speech segment may be resampled from several feature vectors generated for the speech segment (e.g., one feature vector every 10 ms duration) to k feature vectors, resulting in a set of resampled (e.g., downsampled) feature vectors of dimension k×N (where N is the number of values in each feature vector, such as Figure 6 the value of 512 as shown). The value of k may be set to any suitable value (e.g., values of 3, 5, 6, 7, etc.). In some cases, k may be a design parameter that can be adjusted in some cases to obtain a trade-off between quality and computational savings (e.g., a high value of k may result in high-quality matching, while a low value of k may result in low computational cost).
[0073] In one illustrative example, k may be equal to 3. As described herein, the audio processing system 220 may compute representations (e.g., feature vectors) at set time intervals (such as every 10 ms). For example, referring to Figure 6 , a speech segment 662 of the input speech may have a length or duration of 150 ms, and the audio processing system 220 may compute 15 representations (one representation every 10 ms of the speech segment 662). The audio processing system 220 may resample (e.g., downsample) the 15 representations by taking a subset of the representations from the complete set of representations (from the set of 15 representations) to produce a set of resampled feature vectors 664 (of dimension k×512) for the speech segment 662. In one example, using k = 3, the subset of representations selected by the audio processing system 220 may include the first representation, the last representation, and the middle representation. In another example, using k = 5, for an input speech segment that includes the word "hello" and is 250 ms long, the audio processing system 220 may generate 25 feature vectors (one feature vector every 10 ms). Using k = 5 as an illustrative example, the audio processing system 220 may downsample the set of 25 feature vectors by taking the first feature vector (e.g., vector 1), the feature vector (e.g., feature vector 25), the middle feature vector (e.g., vector 13), the feature vector at the quarter mark (e.g., feature vector 8), and the feature vector at the three-quarter mark (e.g., feature vector 18).
[0074] The audio processing system 220 may perform the same resampling for speech segments stored in the audio storage device 624 using the same k value used to resample the input speech. Then, the audio representation search and comparison engine 230 may perform the comparison using a set of resampled (e.g., downsampled) feature vectors of dimension k×N. For example, the set of resampled feature vectors 664 may be compared with the set of resampled feature vectors stored in the audio representation storage device 626 to determine the set of resampled feature vectors that is the closest match and determine the output 232.
[0075] As described above, the audio processing system 220 may utilize one or more machine learning systems or models to generate representations (e.g., features or embedding vectors). Machine learning (ML) may be considered a subset of artificial intelligence (AI). ML systems may include algorithms and statistical models that computer systems may use to perform various tasks by relying on patterns and inferences without using explicit instructions. An example of an ML system is a neural network (also referred to as an artificial neural network), which may include a group of interconnected artificial neurons (e.g., neuron models). Neural networks may be used in various applications and / or devices such as speech analysis, audio signal analysis, image and / or video decoding, image analysis and / or computer vision applications, Internet protocol (IP) cameras, Internet of Things (IoT) devices, autonomous vehicles, service robots, and the like.
[0076] Individual nodes in a neural network may simulate biological neurons by obtaining input data and performing simple operations on the data. The result of the simple operation performed on the input data is selectively passed to other neurons. Weight values are associated with each vector and node in the network, and these values limit the way the input data is related to the output data. For example, the input data for each node may be multiplied by the corresponding weight value, and the products may be summed. The sum of the products may be adjusted by an optional bias, and an activation function may be applied to the result, thereby producing an output signal or “output activation” (sometimes referred to as a feature map or activation map) of the node. The weight values may initially be determined by an iterative flow of training data through the network (e.g., the weight values are established during a training phase in which the network learns how to identify a particular class based on the characteristics of its typical input data).
[0077] There are different types of neural networks, such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), generative adversarial networks (GANs), multi-layer perceptron (MLP) neural networks, transformer neural networks, etc. For example, a convolutional neural network (CNN) is a feedforward artificial neural network. A convolutional neural network may include a collection of artificial neurons, each having a receptive field (e.g., a spatially local region of the input space) and together tiling an input space. RNNs work on the principle of saving the output of a layer and feeding that output back to the input to assist in predicting the result of that layer. A GAN is a generative neural network that can learn patterns in input data so that the neural network model can generate new synthetic outputs that plausibly could have come from the original dataset. A GAN may include two neural networks operating together, including a generative neural network that generates the synthetic output and a discriminative neural network that evaluates the authenticity of the output. In an MLP neural network, data may be fed into an input layer, and one or more hidden layers provide an abstraction level for the data. Predictions can then be made on an output layer based on the abstracted data.
[0078] Deep learning (DL) is an example of a machine learning technique and can be considered a subset of ML. Many DL methods are based on neural networks, such as RNNs or CNNs, and utilize multiple layers. Using multiple layers in a deep neural network allows for the gradual extraction of higher-level features from a given raw data input. For example, the output of the first layer of artificial neurons becomes the input to the second layer of artificial neurons, the output of the second layer of artificial neurons becomes the input to the third layer of artificial neurons, and so on. The layers located between the input and output of the entire deep neural network are typically referred to as hidden layers. The hidden layers learn (e.g., are trained) to transform the intermediate input from the previous layer into a slightly more abstract and composite representation that can be provided to the subsequent layer until the final or desired representation is obtained as the final output of the deep neural network.
[0079] As mentioned above, neural networks are examples of machine learning systems and may include an input layer, one or more hidden layers, and an output layer. Data is provided from the input nodes of the input layer, processed by the hidden nodes of one or more hidden layers, and an output is produced through the output nodes of the output layer. Deep learning networks typically include multiple hidden layers. Each layer of a neural network may include a feature map or activation map, which may include artificial neurons (or nodes). The feature map may include filters, kernels, etc. Nodes may include one or more weights that indicate the importance of the nodes in one or more of the layers. In some cases, a deep learning network may have a series of many hidden layers, where the early layers are used to determine simple and low-level characteristics of the input, and the later layers build a hierarchy of more complex and abstract characteristics.
[0080] Deep learning architectures learn hierarchical structures of features. For example, if presented with visual data, the first layer can learn to identify relatively simple features in the input stream, such as edges. In another example, if presented with auditory data, the first layer can learn to identify spectral power in specific frequencies. The second layer takes the output of the first layer as input and can learn to identify combinations of features, such as simple shapes in visual data or combinations of sounds in auditory data. For example, higher layers can learn to represent complex shapes in visual data or words in auditory data. Even higher layers can learn to identify common visual objects or spoken phrases. Deep learning architectures can perform particularly well when applied to problems with a natural hierarchical structure. For example, the classification of motorized vehicles can benefit from first learning to identify wheels, windshields, and other features. These features can be combined in different ways at higher levels to identify cars, trucks, and airplanes.
[0081] Figures 7A to 7C Illustrated is an example neural network that can be used for keyword detection in accordance with aspects of the present disclosure. The neural network can be designed to have multiple connection patterns. In a feedforward network, information passes from lower layers to higher layers, where each neuron in a given layer communicates with neurons in a higher layer. As described above, hierarchical representations can be constructed in successive layers of a feedforward network. The neural network can also have recurrent or feedback (also referred to as top-down) connections. In a recurrent connection, the output from a neuron in a given layer can be communicated to another neuron in the same layer. Recurrent architectures can help identify patterns that span more than one block of input data presented sequentially to the neural network. Connections from neurons in a given layer to neurons in lower layers are referred to as feedback (or top-down) connections. Networks with many feedback connections can be helpful when the identification of high-level concepts can assist in discerning specific low-level features of the input.
[0082] In some cases, the connections between the layers of a neural network can be fully connected or locally connected. Figure 7A Illustrated is an example of a fully connected neural network 702. In the fully connected neural network 702, neurons in the first layer can communicate their output to each neuron in the second layer, such that each neuron in the second layer will receive input from each neuron in the first layer. Figure 7BAn example of a locally connected neural network 704 is illustrated. In the locally connected neural network 704, neurons in the first layer can be connected to a limited number of neurons in the second layer. More generally, a locally connected layer of the locally connected neural network 704 can be configured such that each neuron in the layer will have the same or a similar connection pattern, but the connection strengths can have different values (e.g., 710, 712, 714, and 716). The locally connected connection pattern can result in spatially distinct receptive fields in higher layers because neurons in a given region of the higher layer can receive input that is tuned, through training, to the attributes of a restricted portion of the network's total input.
[0083] An example of a locally connected neural network is a convolutional neural network. Figure 7C An example of a convolutional neural network 706 is illustrated. The convolutional neural network 706 can be configured such that the connection strengths associated with the input to each neuron in the second layer are shared (e.g., 708). Convolutional neural networks can be well-suited for problems where the spatial location of the input is meaningful.
[0084] Figure 8 is a block diagram illustrating an example of a deep convolutional network (DCN) 850 according to aspects of the present disclosure. The DCN 850 can include multiple different types of layers based on connectivity and weight sharing. As Figure 8 shown, the DCN 850 includes convolutional blocks 854A, 854B. Each of the convolutional blocks 854A, 854B can be configured with a convolutional layer (CONV) 856, a normalization layer (LNorm) 858, and a max pooling layer (MAX POOL) 860.
[0085] The convolutional layer 856 can include one or more convolutional filters that can be applied to the input data 852 to generate a feature map. Although only two convolutional blocks 854A, 854B are shown, the present disclosure is not limited thereto, and any number of convolutional blocks (e.g., blocks 854A, 854B) can be included in the DCN 850 according to design preferences. The normalization layer 858 can normalize the output of the convolutional filters. For example, the normalization layer 858 can provide whitening or lateral inhibition. The max pooling layer 860 can provide spatially downsampled aggregation to achieve local invariance and dimensionality reduction.
[0086] For example, the parallel filter bank of the deep convolutional network can be loaded onto the CPU 102 or GPU 104 of the SOC 100 to achieve high performance and low power consumption. In an alternative embodiment, the parallel filter bank can be loaded onto the DSP 106 or ISP 116 of the SOC 100. Additionally, the DCN 850 can access other processing blocks that may be present on the SOC 100, such as the sensor processor 114 and keyword detection system 120 dedicated to sensors and navigation, respectively.
[0087] The deep convolutional network 850 may also include one or more fully connected layers, such as layer 862A (labeled "FC1") and layer 862B (labeled "FC2"). The DCN 850 may also include a logistic regression (LR) layer 864. Between each layer 856, 858, 860, 862A, 862B, 864 of the DCN 850 are weights (not shown) to be updated. The output of each of these layers (e.g., 856, 858, 860, 862A, 862B, 864) can serve as the input to a subsequent layer among these layers (e.g., 886, 858, 860, 862A, 862B, 864) of the deep convolutional network 850 to learn a hierarchical feature representation from the input data 852 (e.g., images, audio, video, sensor data, and / or other input data) supplied at the initial convolutional block 854A.
[0088] To adjust the weights, a learning algorithm can compute the gradient vector of the weights. The gradient can indicate the amount by which the error will increase or decrease when the weights are adjusted. At the top layer, the gradient can directly correspond to the value of the weight connecting the activated neurons in the penultimate layer and the neurons in the output layer. In the lower layers, the gradient can depend on the value of the weights and the error gradient computed in the higher layers. The weights can then be adjusted to reduce the error. This way of adjusting the weights can be referred to as "backpropagation" because it involves a "backward pass" through the neural network.
[0089] In practice, the error gradient of the weights can be computed over a small number of examples, such that the computed gradient is close to the true error gradient. This approximation method can be referred to as stochastic gradient descent. Stochastic gradient descent can be repeated until the achievable error rate of the entire system stops decreasing or until the error rate reaches a target level. After learning, new inputs can be presented to the DCN, and the forward pass through the network can produce an output 422 that can be considered an inference or prediction of the DCN.
[0090] The output of the DCN 850 is a classification score 866 for the input data 852. The classification score 866 can be a probability or a set of probabilities, where the probability is the probability that the input data includes features from the set of features that the DCN 850 is trained to detect.
[0091] Figure 9 is a flowchart illustrating an example of process 900 for encoding audio information in accordance with aspects of the present disclosure. In some aspects, process 900 may be performed by Figure 4A the speech encoder 440 and / or Figure 2 and / or Figure 4A the audio processing system 220. Operations of process 900 may be implemented as software components executed and run on one or more processors (e.g., Figure 11 the processor 1110 and / or other processors). Additionally, wireless communication devices may be enabled to transmit and receive signals in process 900, for example, via one or more antennas and / or one or more transceivers (e.g., wireless transceivers).
[0092] At operation 902, the speech encoder 440 may detect an input audio segment (e.g., an audio segment of audio obtained from the audio source 222). In some aspects, the input audio segment includes an input speech segment. At operation 904, the speech encoder 440 may process the input audio segment to generate a representation of the input audio segment. In some aspects, the representation of the input audio segment may include an embedding vector representing the input audio segment.
[0093] At operation 906, the speech encoder 440 may compare the representation of the input audio segment with a plurality of representations stored in at least one memory, the plurality of representations representing a plurality of audio segments. In some aspects, the plurality of audio segments includes a plurality of speech segments. In such cases, the plurality of representations may include a plurality of embedding vectors representing the plurality of audio segments.
[0094] At operation 908, the speech encoder 440 may determine one or more target representations of one or more target audio segments from the plurality of representations stored in at least one memory based on comparing the representation with the plurality of representations. In some aspects, to compare the representation of the input audio segment with the plurality of representations, the speech encoder 440 may determine a respective difference between the representation of the input audio segment and each corresponding representation of the plurality of representations. In some cases, the speech encoder 440 may determine one or more target representations from the plurality of representations based on one or more target representations having one or more minimum differences from the representation of the input audio segment. In some aspects, the speech encoder 440 may further determine one or more target representations based on search and cascade operations described herein.
[0095] In some aspects, one or more target audio segments have a fixed length. In some aspects, one or more target audio segments have a variable length. In such aspects, the speech encoder 440 may resample one or more target audio segments to convert the one or more target audio segments of variable length into one or more target audio segments of fixed length.
[0096] At operation 910, the speech encoder 440 may determine one or more indexes associated with one or more target audio segments. At operation 912, the speech encoder 440 may group the one or more indexes. At operation 914, the speech encoder 440 may transmit the grouped one or more indexes. In some aspects, the speech encoder 440 may encode the grouped one or more indexes as an audio bitstream. In some cases, to transmit the grouped one or more indexes, the speech encoder 440 may transmit the audio bitstream. In one illustrative example, the speech encoder 440 may transmit the audio bitstream at less than one thousand bits per second.
[0097] Figure 10 is a flowchart illustrating an example of a process 1000 for decoding audio information according to aspects of the present disclosure. In some aspects, the process 1000 may be performed by Figure 4A the speech decoder 450. The operations of the process 1000 may be implemented as software components executed and run on one or more processors (e.g., Figure 11 the processor 1110 and / or other processors). Additionally, wireless communication of signals in the process 1000 may be enabled, for example, by one or more antennas and / or one or more transceivers (e.g., a wireless transceiver).
[0098] At operation 1002, the speech decoder 450 may receive one or more grouped indexes associated with one or more target audio segments. In one illustrative example, the one or more target audio segments include one or more target speech segments. In some aspects, the speech decoder 450 may receive the one or more grouped indexes as an audio bitstream. In some cases, the one or more target audio segments have variable lengths. In one illustrative example, the audio bitstream is less than one thousand bits per second. At operation 1004, the speech decoder 450 may ungroup the one or more grouped indexes to generate one or more indexes associated with the one or more target audio segments.
[0099] At operation 1006, the speech decoder 450 may retrieve one or more target audio segments from at least one memory based on the one or more indexes. At operation 1008, the speech decoder 450 may combine the one or more target audio segments to generate decoded audio. In some aspects, to combine the one or more target audio segments, the speech decoder 450 may concatenate the one or more target audio segments to generate decoded audio. In some cases, the speech decoder 450 may output the decoded audio (e.g., via one or more speakers, store the decoded audio, send the decoded audio to another device, etc.).
[0100] In some aspects, the processes described herein (e.g., process 900, process 1000, and / or any other process described herein) may be performed by a computing device or apparatus. In one example, process 900, process 1000, and / or other techniques or processes described herein may be performed by Figure 2 the audio processing system 220. In another example, process 900, process 1000, and / or other techniques or processes described herein may be performed by Figure 11 the computing system 1100 shown in Figure 11 For example, a computing device having the computing device architecture of the computing system 1100 shown in Figure 2 may implement the audio processing system 220 to perform the operations of process 900 and / or the operations of process 1000.
[0101] The computing device may include any suitable device, such as a mobile device (e.g., a mobile phone), an extended reality (XR) device (e.g., a virtual reality (VR), augmented reality (AR), or mixed reality (MR) headset, AR or MR glasses, etc.), a wearable device (e.g., a network-connected watch or other wearable device), a vehicle (e.g., an autonomous or semi-autonomous vehicle) or a computing system or device of a vehicle, a desktop computing device, a tablet computing device, a server computer, a robotic device, a laptop computer, a network-connected television, a camera, and / or any other computing device having the resource capabilities to perform the processes described herein (including process 700, process 800, and / or any other process described herein). In some cases, the computing device or apparatus may include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other components configured to perform the steps of the processes described herein. In some examples, the computing device may include a display, a network interface configured to communicate and / or receive data, any combination thereof, and / or other components. The network interface may be configured to communicate and / or receive Internet Protocol (IP)-based data or other types of data.
[0102] The components of a computing device may be implemented in circuitry. For example, a component may include an electronic circuit or other electronic hardware and / or may be implemented using an electronic circuit or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, a graphics processing unit (GPU), a digital signal processor (DSP), a central processing unit (CPU), and / or other suitable electronic circuits), and / or may include computer software, firmware, or any combination thereof for performing the various operations described herein and / or may be implemented using computer software, firmware, or any combination thereof for performing the various operations described herein.
[0103] Process 900 and Process 1000 are illustrated as logic flowcharts, and the operations of these logic flowcharts represent sequences of operations that may be implemented in hardware, computer instructions, or a combination thereof. In the context of computer instructions, the operations represent computer-executable instructions stored on one or more computer-readable storage media that, when executed by one or more processors, perform the recited operations. In general, computer-executable instructions include routines, programs, objects, components, data structures, etc. that perform a particular function or implement a particular data type. The order in which the operations are described is not intended to be construed as a limitation, and any number of the described operations may be combined in any order and / or in parallel to implement the process.
[0104] Additionally, Process 900, Process 1000, and / or any other process described herein may be executed under the control of one or more computer systems configured with executable instructions and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that is executed jointly on one or more processors, implemented in hardware, or a combination thereof. As noted above, the code may be stored on a computer-readable or machine-readable storage medium, e.g., in the form of a computer program that includes a plurality of instructions capable of being executed by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.
[0105] Figure 11 An example of a computing system 1100 that may implement the various technologies described herein is shown. For example, computing system 1100 may implement the audio processing system 220 described with respect to Figure 2 Process 900 described with respect to Figure 9 Process 1000 described with respect to Figure 10 and / or any other audio processing operations described herein. The components of computing system 1100 communicate with each other using connection 1105. Connection 1105 may be a physical connection via a bus or a direct connection into processor 1110, such as in a chipset architecture. Connection 1105 may also be a virtual connection, a networking connection, or a logical connection.
[0106] In some embodiments, computing system 1100 is a distributed system, where the functions described in this disclosure can be distributed within one data center, multiple data centers, a peer-to-peer network, etc. In some embodiments, one or more of the described system components represent many such components that each perform some or all of the functions that the component is described for. In some embodiments, a component can be a physical device or a virtual device.
[0107] Example system 1100 includes at least one processing unit (CPU or processor) 1110 and connection 1105 that couples various system components including system memory 1115 (such as read-only memory (ROM) 1120 and random access memory (RAM) 1125) to processor 1110. Computing system 1100 may include a cache 1112 of high-speed memory that is directly connected to, in proximity to, or integrated as part of processor 1110. In some instances, computing system 1100 may copy data from memory 1115 and / or storage device 1130 to cache 1112 for quick access by processor 1110. In this way, the cache can provide a performance enhancement that avoids latency in processor 1110 while waiting for data. These modules and other modules can control or be configured to control processor 1110 to perform various actions. Other computing device memory 1115 may also be used. Memory 1115 may include various different types of memory with different performance characteristics.
[0108] Processor 1110 may include any general-purpose processor and hardware services or software services, such as services 11132, 11134, and 11136 stored in storage device 1130, which are configured to control processor 1110, as well as a dedicated processor in which software instructions are incorporated into the actual processor design. Processor 1110 can be substantially a complete stand-alone computing system that contains multiple cores or processors, buses, memory controllers, caches, etc. A multi-core processor can be symmetric or asymmetric.
[0109] To enable user interaction, computing system 1100 may also include an input device 1145, which may represent any number of input mechanisms, such as a microphone for speech, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, speech, etc. Computing system 1100 may also include an output device 1135, which may be one or more of a variety of output mechanisms known to those skilled in the art, such as a display, a projector, a television, a speaker device, etc. In some cases, a multimodal system may enable a user to provide multiple types of input / output to communicate with computing system 1100. Computing system 1100 may include a communication interface 1140, which generally may govern and manage user input and system output. The communication interface may perform or facilitate the reception and / or transmission of wired or wireless communications via wired and / or wireless transceivers, including using audio jack / plug, microphone jack / plug, universal serial bus (USB) port / plug, port / plug, Ethernet port / plug, fiber optic port / plug, dedicated wired port / plug, wireless signal transmission, low energy (BLE) wireless signal transmission, wireless signal transmission, radio frequency identification (RFID) wireless signal transmission, near field communication (NFC) wireless signal transmission, dedicated short range communication (DSRC) wireless signal transmission, 802.11 Wi-Fi wireless signal transmission, wireless local area network (WLAN) signal transmission, visible light communication (VLC), worldwide interoperability for microwave access (WiMAX), infrared (IR) communication wireless signal transmission, public switched telephone network (PSTN) signal transmission, integrated services digital network (ISDN) signal transmission, 3G / 4G / 5G / LTE cellular data network wireless signal transmission, ad hoc network signal transmission, radio wave signal transmission, microwave signal transmission, infrared signal transmission, visible light signal transmission, ultraviolet light signal transmission, wireless signal transmission along the electromagnetic spectrum, or some combination thereof. Communication interface 1140 may also include one or more global navigation satellite system (GNSS) receivers or transceivers for determining the location of computing system 1100 based on one or more signals received from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the United States' Global Positioning System (GPS), Russia's Global Navigation Satellite System (GLONASS), China's BeiDou Navigation Satellite System (BDS), and Europe's Galileo GNSS. There are no restrictions on operating on any particular hardware arrangement, and thus the underlying features here may be easily replaced to obtain improved hardware or firmware arrangements as they are developed.
[0110] The storage device 1130 can be a non-volatile and / or non-transitory and / or computer-readable memory device, and can be a hard disk or other types of computer-readable media that can store data accessible by a computer, such as a tape cartridge, a flash memory card, a solid-state memory device, a digital versatile disc, a cassette tape, a floppy disk, a flexible disk, a hard disk, a magnetic tape, a magnetic stripe / strip, any other magnetic storage medium, a flash memory, a memristor memory, any other solid-state memory, a compact disc read-only memory (CD-ROM) disc, a rewritable compact disc (CD) disc, a digital video disc (DVD) disc, a Blu-ray disc, a holographic disc, another optical medium, a secure digital (SD) card, a micro secure digital (microSD) card, a smart card chip, an Europay MasterCard and Visa (EMV) chip, a subscriber identity module (SIM) card, a mini / micro / nano / pico SIM card, another integrated circuit (IC) chip / card, a random access memory (RAM), a static RAM (SRAM), a dynamic RAM (DRAM), a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a flash EPROM (FLASHEPROM), a cache memory (L1 / L2 / L3 / L4 / L5 / L#), a resistive random access memory (RRAM / ReRAM), a phase change memory (PCM), a spin transfer torque RAM (STT-RAM), another memory chip or cartridge and / or combinations thereof.
[0111] The storage device 1130 can include software services (e.g., Service 1 1132, Service 2 1134, and Service 3 1136, and / or other services), servers, services, etc., which cause the system to perform functions when the code defining such software is executed by the processor 1110. In some embodiments, the hardware services that perform specific functions can include software components for performing the functions stored in a computer-readable medium connected to necessary hardware components such as the processor 1110, the connection 1105, the output device 1135, etc.
[0112] The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. The computer-readable medium can include non-transitory media, which can store data and do not include carrier waves and / or transient electronic signals propagated wirelessly or over a wired connection. Examples of non-transitory media can include, but are not limited to, magnetic disks or tapes, optical storage media (such as compact discs (CDs) or digital versatile discs (DVDs)), flash memory, memory, or memory devices. Code and / or machine-executable instructions can be stored on the computer-readable medium, and the code and / or machine-executable instructions can represent a process, function, subroutine, program, routine, subroutine, module, software package, class, or any combination of instructions, data structures, or program statements. By passing and / or receiving information, data, arguments, parameters, or memory contents, a code segment can be coupled to another code segment or hardware circuit. The information, arguments, parameters, data, etc. can be passed, forwarded, or sent via any suitable means, including memory sharing, message passing, token passing, network transmission, etc.
[0113] In some embodiments, computer-readable storage devices, media, and memories can include wired or wireless signals containing bitstreams, etc. However, when mentioned, non-transitory computer-readable storage media specifically exclude media such as energy, carrier signals, electromagnetic waves, and signals themselves.
[0114] Specific details are provided in the above description to provide a thorough understanding of the embodiments and examples provided herein. However, those of ordinary skill in the art will understand that the embodiments can be practiced without these specific details. For clarity of illustration, in some cases, the present technology can be presented as including separate functional blocks, including functional blocks containing devices, device components, steps or routines in a method embodied in software or a combination of hardware and software. Additional components other than those shown and / or described herein can be used. For example, circuits, systems, networks, processes, and other components can be shown in block diagram form as components to avoid obscuring these embodiments in unnecessary detail. In other cases, well-known circuits, processes, algorithms, structures, and technologies can be shown without necessary detail to avoid obscuring the embodiments.
[0115] Individual embodiments may be described above as processes or methods depicted as flowcharts, flow diagrams, data flow diagrams, structure diagrams, or block diagrams. Although a flowchart may describe operations as a sequential process, many of the operations may be performed in parallel or concurrently. In addition, the order of the operations may be rearranged. A process is terminated when its operations are completed, but a process may have additional steps not included in the figure. A process may correspond to a method, function, procedure, subroutine, subprogram, etc. When a process corresponds to a function, the termination of the process may correspond to the function returning to the calling function or the main function.
[0116] The processes and methods according to the above examples may be implemented using computer-executable instructions stored or otherwise obtained from a computer-readable medium. Such instructions may include, for example, instructions and data that cause a general-purpose computer, special-purpose computer, or processing device to configure the general-purpose computer, special-purpose computer, or processing device to perform a certain function or group of functions in other ways. Part of the computer resources used may be accessible through a network. The computer-executable instructions may be, for example, binary, intermediate format instructions such as assembly language, firmware, source code, etc. Examples of computer-readable media that may be used to store instructions, the information used, and / or the information created during the methods according to the described examples include magnetic or optical disks, flash memory, USB devices with non-volatile memory, networked storage devices, etc.
[0117] Devices implementing the processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description language, or any combination thereof, and may take any of a variety of form factors. When implemented in software, firmware, middleware, or microcode, the program code or code segments (e.g., computer program products) for performing the necessary tasks may be stored in a computer-readable or machine-readable medium. The processor may execute the necessary tasks. Typical examples of form factors include laptop computers, smart phones, mobile phones, tablet devices, or other small form factor personal computers, personal digital assistants, rack-mounted devices, stand-alone devices, etc. The functionality described herein may also be embodied in peripheral devices or plug-in cards. By additional example, such functionality may also be implemented on a circuit board between different chips or different processes executed on a single device.
[0118] Instructions, the media for conveying such instructions, the computing resources for executing them, and other structures for supporting such computing resources are example components for providing the functions described in this disclosure.
[0119] In the foregoing description, aspects of the present application have been described with reference to specific embodiments of the present application, but those skilled in the art will recognize that the present application is not limited thereto. Thus, although exemplary embodiments of the present application have been described in detail herein, it is to be understood that the inventive concept can be embodied and employed in various other ways, and the appended claims are intended to be construed to include such variations, unless limited by the prior art. The various features and aspects of the above applications can be used alone or in combination. Additionally, without departing from the broader spirit and scope of this specification, the embodiments can be used in any number of environments and applications beyond those described herein. Accordingly, the specification and drawings are to be regarded as illustrative rather than restrictive. For purposes of illustration, the methods are described in a particular order. It should be understood that in alternative embodiments, the methods can be performed in a different order than that described.
[0120] Those of ordinary skill in the art should understand that, without departing from the scope of this specification, the less than (“<”) and greater than (“>”) symbols or terms used herein can be replaced by the less than or equal to (“≤”) and greater than or equal to (“≥”) symbols, respectively.
[0121] In cases where a component is described as “configured to” perform certain operations, this configuration can be achieved, for example, by designing an electronic circuit or other hardware to perform the operations, by programming a programmable electronic circuit (e.g., a microprocessor or other suitable electronic circuit) to perform the operations, or any combination thereof.
[0122] The phrase “coupled to” means that any component is directly or indirectly physically connected to another component, and / or any component directly or indirectly communicates with another component (e.g., is connected to another component via a wired or wireless connection and / or other suitable communication interface).
[0123] Claim language or other language reciting “at least one of” a set and / or “one or more” in a set indicates that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, claim language reciting “at least one of A and B” means A, B, or A and B. In another example, claim language reciting “at least one of A, B, and C” means A, B, C, or A and B, or A and C, or B and C, or A and B and C. The language “at least one of” a set and / or “one or more” in a set does not limit the set to the items listed in the set. For example, claim language reciting “at least one of A and B” can mean A, B, or A and B, and can additionally include items not listed in the set of A and B.
[0124] The various illustrative logical blocks, modules, circuits, and algorithmic steps described in connection with the examples disclosed herein may be implemented as electronic hardware, computer software, firmware, or any combination thereof. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present application.
[0125] The techniques described herein may also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques may be implemented in any of a variety of devices, such as a general purpose computer, a wireless communication device handset, or an integrated circuit device with multiple uses, including applications in a wireless communication device handset and other devices. Any feature described as a module or component may be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques may be at least partially realized by a computer-readable data storage medium including program code, the program code including instructions that, when executed, perform one or more of the methods, algorithms, and / or operations described above. The computer-readable data storage medium may form part of a computer program product, which may include packaging materials. The computer-readable medium may include a memory or data storage medium, such as a random access memory (RAM) (such as a synchronous dynamic random access memory (SDRAM)), a read only memory (ROM), a non-volatile random access memory (NVRAM), an electrically erasable programmable read only memory (EEPROM), a flash memory, a magnetic or optical data storage medium, and the like. Additionally or alternatively, the techniques may be at least partially realized by a computer-readable communication medium that carries or conveys program code in the form of instructions or data structures that can be accessed, read, and / or executed by a computer, such as a propagated signal or wave.
[0126] The program code can be executed by a processor, which can include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such processors can be configured to perform any of the techniques described in this disclosure. A general-purpose processor can be a microprocessor; but in an alternative, the processor can be any conventional processor, controller, microcontroller, or state machine. The processor can also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. Thus, the term "processor" as used herein can refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or device suitable for implementing the techniques described herein.
[0127] Exemplary aspects of the present disclosure include:
[0128] Aspect 1. An apparatus for encoding audio information, the apparatus comprising: at least one memory; and at least one processor, the at least one processor coupled to the at least one memory and configured to: detect an input audio segment; process the input audio segment to generate a representation of the input audio segment; compare the representation of the input audio segment with a plurality of representations stored in the at least one memory, the plurality of representations representing a plurality of audio segments; based on comparing the representation with the plurality of representations, determine one or more target representations of one or more target audio segments from the plurality of representations stored in the at least one memory; determine one or more indexes associated with the one or more target audio segments; group the one or more indexes; and transmit the grouped one or more indexes.
[0129] Aspect 2. The apparatus according to aspect 1, wherein the representation of the input audio segment includes an embedding vector representing the input audio segment, and wherein the plurality of representations include a plurality of embedding vectors representing the plurality of audio segments.
[0130] Aspect 3. The apparatus according to any one of aspects 1 or 2, wherein the one or more target audio segments have a variable length.
[0131] Aspect 4. The apparatus according to aspect 3, wherein the at least one processor is configured to resample the one or more target audio segments to convert the one or more target audio segments of variable length into one or more target audio segments of fixed length.
[0132] Aspect 5. The apparatus according to any one of aspects 1 to 4, wherein: the at least one processor is configured to encode the grouped one or more indexes into an audio bitstream; and in order to transmit the grouped one or more indexes, the at least one processor is configured to transmit the audio bitstream.
[0133] Aspect 6. The apparatus according to aspect 5, wherein the at least one processor is configured to transmit the audio bitstream at less than one thousand bits per second.
[0134] Aspect 7. The apparatus according to any one of aspects 1 to 6, wherein: in order to compare the representation of the input audio segment with the plurality of representations, the at least one processor is configured to determine a respective difference between the representation of the input audio segment and each corresponding representation among the plurality of representations; and the at least one processor is configured to determine the one or more target representations from the plurality of representations based on the one or more target representations having one or more minimum differences from the representation of the input audio segment.
[0135] Aspect 8. The apparatus according to aspect 7, wherein the at least one processor is configured to further determine the one or more target representations based on search and concatenation operations.
[0136] Aspect 9. The apparatus according to any one of aspects 1 to 8, wherein the input audio segment includes an input speech segment, and wherein the plurality of audio segments includes a plurality of speech segments.
[0137] Aspect 10. An apparatus for decoding audio information, the apparatus comprising: at least one memory; and at least one processor, the at least one processor coupled to the at least one memory and configured to: receive one or more grouped indexes associated with one or more target audio segments; de-group the one or more grouped indexes to generate one or more indexes associated with the one or more target audio segments; retrieve the one or more target audio segments from the at least one memory based on the one or more indexes; and combine the one or more target audio segments to generate decoded audio.
[0138] Aspect 11. The apparatus according to aspect 10, wherein in order to combine the one or more target audio segments, the at least one processor is configured to concatenate the one or more target audio segments to generate the decoded audio.
[0139] Aspect 12. The apparatus according to any one of aspects 10 or 11, wherein the at least one processor is configured to output the decoded audio.
[0140] Aspect 13. The apparatus according to any one of aspects 10 to 12, wherein the one or more target audio segments have variable lengths.
[0141] Aspect 14. The apparatus according to any one of aspects 10 to 13, wherein the at least one processor is configured to receive the one or more packetized indexes as an audio bitstream.
[0142] Aspect 15. The apparatus according to aspect 14, wherein the audio bitstream is less than one thousand bits per second.
[0143] Aspect 16. The apparatus according to any one of aspects 10 to 15, wherein the one or more target audio segments include one or more target speech segments.
[0144] Aspect 17. A method for encoding audio information, the method comprising: detecting an input audio segment; processing the input audio segment to generate a representation of the input audio segment; comparing the representation of the input audio segment with a plurality of representations stored in the at least one memory, the plurality of representations representing a plurality of audio segments; based on comparing the representation with the plurality of representations, determining one or more target representations of one or more target audio segments from the plurality of representations stored in the at least one memory; determining one or more indexes associated with the one or more target audio segments; packetizing the one or more indexes; and transmitting the packetized one or more indexes.
[0145] Aspect 18. The method according to aspect 17, wherein the representation of the input audio segment includes an embedding vector representing the input audio segment, and wherein the plurality of representations include a plurality of embedding vectors representing the plurality of audio segments.
[0146] Aspect 19. The method according to any one of aspects 17 or 18, wherein the one or more target audio segments have variable lengths.
[0147] Aspect 20. The method according to aspect 19, the method further comprising resampling the one or more target audio segments to convert the one or more target audio segments of variable length into one or more target audio segments of fixed length.
[0148] Aspect 21. The method according to any one of aspects 17 to 20, the method further comprising: encoding the packetized one or more indexes as an audio bitstream; wherein transmitting the packetized one or more indexes includes transmitting the audio bitstream.
[0149] Aspect 22. The method according to aspect 21, the method further comprising transmitting the audio bitstream at less than one thousand bits per second.
[0150] Aspect 23. The method according to any one of aspects 17 to 22, wherein comparing the representation of the input audio segment with the plurality of representations includes determining a respective difference between the representation of the input audio segment and each respective representation of the plurality of representations, and further includes: determining the one or more target representations from the plurality of representations based on one or more of the target representations having one or more minimum differences from the representation of the input audio segment.
[0151] Aspect 24. The method according to aspect 23, the method further including determining the one or more target representations further based on search and concatenation operations.
[0152] Aspect 25. The method according to any one of aspects 17 to 24, wherein the input audio segment includes an input speech segment, and wherein the plurality of audio segments includes a plurality of speech segments.
[0153] Aspect 26. A method for decoding audio information, the method including: receiving one or more packetized indices associated with one or more target audio segments; depacketizing the one or more packetized indices to generate one or more indices associated with the one or more target audio segments; retrieving the one or more target audio segments from at least one memory based on the one or more indices; and combining the one or more target audio segments to generate decoded audio.
[0154] Aspect 27. The method according to aspect 26, wherein combining the one or more target audio segments includes concatenating the one or more target audio segments to generate the decoded audio.
[0155] Aspect 28. The method according to any one of aspects 26 or 27, the method further including outputting the decoded audio.
[0156] Aspect 29. The method according to any one of aspects 26 to 28, wherein the one or more target audio segments have variable lengths.
[0157] Aspect 30. The method according to any one of aspects 26 to 29, the method further including receiving the one or more packetized indices as an audio bitstream.
[0158] Aspect 31. The method according to claim 30, wherein the audio bitstream is less than one thousand bits per second.
[0159] Aspect 32. The method according to any one of aspects 26 to 31, wherein the one or more target audio segments include one or more target speech segments.
[0160] Aspect 33. The non-transitory computer-readable medium according to any one of aspects 17 to 25.
[0161] Aspect 34. An apparatus for encoding audio information, the apparatus comprising one or more components for performing the operations according to any one of aspects 17 to 25.
[0162] Aspect 35. The non-transitory computer-readable medium according to any one of aspects 26 to 32.
[0163] Aspect 36. An apparatus for decoding audio information, the apparatus comprising one or more components for performing the operations according to any one of aspects 26 to 32.
[0164] Aspect 37. The non-transitory computer-readable medium according to any one of aspects 17 to 25 and aspects 26 to 32.
[0165] Aspect 38. An apparatus for processing audio information, the apparatus comprising one or more components for performing the operations according to any one of aspects 17 to 25 and aspects 26 to 32.
Claims
1. An apparatus for encoding audio information, the apparatus comprising: at least one memory; and at least one processor, the at least one processor coupled to the at least one memory and configured to: detect an input audio segment; process the input audio segment to generate a representation of the input audio segment; compare the representation of the input audio segment with a plurality of representations stored in the at least one memory, the plurality of representations representing a plurality of audio segments; based on comparing the representation with the plurality of representations, determine one or more target representations of one or more target audio segments from the plurality of representations stored in the at least one memory; determine one or more indexes associated with the one or more target audio segments; group the one or more indexes; and transmit the grouped one or more indexes.
2. The apparatus according to claim 1, wherein the representation of the input audio segment comprises an embedding vector representing the input audio segment, and wherein the plurality of representations comprise a plurality of embedding vectors representing the plurality of audio segments.
3. The apparatus according to claim 1, wherein the one or more target audio segments have a variable length.
4. The apparatus according to claim 3, wherein the at least one processor is configured to resample the one or more target audio segments to convert the one or more target audio segments of variable length into one or more target audio segments of fixed length.
5. The apparatus according to claim 1, wherein: the at least one processor is configured to encode the grouped one or more indexes into an audio bitstream; and for transmitting the grouped one or more indexes, the at least one processor is configured to transmit the audio bitstream.
6. The apparatus according to claim 5, wherein the at least one processor is configured to transmit the audio bitstream at less than one thousand bits per second.
7. The apparatus according to claim 1, wherein: for comparing the representation of the input audio segment with the plurality of representations, the at least one processor is configured to determine a respective difference between the representation of the input audio segment and each corresponding representation among the plurality of representations; and the at least one processor is configured to determine the one or more target representations from the plurality of representations based on the one or more target representations having one or more minimum differences from the representation of the input audio segment.
8. The apparatus according to claim 7, wherein the at least one processor is configured to further determine the one or more target representations based on a search and concatenation operation.
9. The apparatus according to claim 1, wherein the input audio segment comprises an input speech segment, and wherein the plurality of audio segments comprise a plurality of speech segments.
10. An apparatus for decoding audio information, the apparatus comprising: at least one memory; and at least one processor, the at least one processor coupled to the at least one memory and configured to: Receive one or more packetized indices associated with one or more target audio segments; Depacketize the one or more packetized indices to generate one or more indices associated with the one or more target audio segments; Retrieve the one or more target audio segments from the at least one memory based on the one or more indices; And Combine the one or more target audio segments to generate decoded audio.
11. The apparatus according to claim 10, wherein, in order to combine the one or more target audio segments, the at least one processor is configured to concatenate the one or more target audio segments to generate the decoded audio.
12. The apparatus according to claim 10, wherein the at least one processor is configured to output the decoded audio.
13. The apparatus according to claim 10, wherein the one or more target audio segments have variable lengths.
14. The apparatus according to claim 10, wherein the at least one processor is configured to receive the one or more packetized indices as an audio bitstream.
15. The apparatus according to claim 14, wherein the audio bitstream is less than one thousand bits per second.
16. The apparatus according to claim 10, wherein the one or more target audio segments include one or more target speech segments.
17. A method for encoding audio information, the method comprising: Detect an input audio segment; Process the input audio segment to generate a representation of the input audio segment; Compare the representation of the input audio segment with a plurality of representations stored in the at least one memory, the plurality of representations representing a plurality of audio segments; Based on comparing the representation with the plurality of representations, determine one or more target representations of one or more target audio segments from the plurality of representations stored in the at least one memory; Determine one or more indices associated with the one or more target audio segments; Packetize the one or more indices; And Transmit the packetized one or more indices.
18. The method according to claim 17, wherein the representation of the input audio segment includes an embedding vector representing the input audio segment, and wherein the plurality of representations include a plurality of embedding vectors representing the plurality of audio segments.
19. The method according to claim 17, wherein the one or more target audio segments have variable lengths.
20. The method according to claim 19, the method further comprising resampling the one or more target audio segments to convert the one or more target audio segments of variable length into one or more target audio segments of fixed length.
21. The method according to claim 17, the method further comprising: Encode the packetized one or more indices as an audio bitstream; wherein transmitting the packetized one or more indices includes transmitting the audio bitstream.
22. The method according to claim 17, wherein comparing the representation of the input audio segment with the plurality of representations includes determining a respective difference between the representation of the input audio segment and each respective representation of the plurality of representations, and further includes: determining the one or more target representations from the plurality of representations based on the one or more target representations having one or more minimum differences from the representation of the input audio segment.
23. The method according to claim 22, the method further comprising determining the one or more target representations based further on search and concatenation operations.
24. The method according to claim 17, wherein the input audio segment includes an input speech segment, and wherein the plurality of audio segments includes a plurality of speech segments.
25. A method for decoding audio information, the method includes: receiving one or more packetized indexes associated with one or more target audio segments; de - packetizing the one or more packetized indexes to generate one or more indexes associated with the one or more target audio segments; retrieving the one or more target audio segments from at least one memory based on the one or more indexes; and combining the one or more target audio segments to generate decoded audio.
26. The method according to claim 25, wherein combining the one or more target audio segments includes concatenating the one or more target audio segments to generate the decoded audio.
27. The method according to claim 25, the method further comprising outputting the decoded audio.
28. The method according to claim 25, wherein the one or more target audio segments have variable lengths.
29. The method according to claim 25, the method further comprising receiving the one or more packetized indexes as an audio bitstream.
30. The method according to claim 25, wherein the one or more target audio segments include one or more target speech segments.