Selective processing of fragments of time series data based on fragment classification
By using a low-complexity machine learning model for feature extraction and classifiers, the problem of choosing appropriate operations in time series data segment processing is solved, achieving efficient classification and resource-saving processing results.
Patent Information
- Application Number
- CN202480048491.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-07-28
- Filing Date
- 2024-07-22
- Publication Date
- 2026-02-27
AI Technical Summary
When processing fragments of time series data, existing technologies struggle to effectively distinguish between different types of data content, leading to inappropriate processing operations, wasted resources, and computational burden. Furthermore, training complex classifiers requires a large amount of labeled data and resources.
A low-complexity machine learning model is used to generate latent space representations through a feature extractor and to classify segments using a classifier, thereby generating processing control signals to select appropriate downstream processing operations.
It enables efficient classification and appropriate processing of time series data, reduces computational resource consumption, improves data fidelity and processing efficiency, and reduces training data requirements.
Smart Images

Figure CN121586928A_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims priority to jointly owned U.S. non-provisional patent application No. 18 / 360,981, filed July 28, 2023, the entire contents of which are expressly incorporated herein by reference. Technical Field
[0003] This disclosure relates in general to the selective processing of time series data.
[0004] Related technical descriptions
[0005] Technological advancements have led to smaller and more powerful computing devices. For example, a wide variety of portable personal computing devices exist today, including small, lightweight, and easily portable cordless phones (such as mobile and smartphones, tablets, and laptops). These devices can transmit voice and data packets over wireless networks. Furthermore, many of these devices incorporate additional functionality, such as digital still cameras, digital camcorders, digital recorders, and audio file players. Moreover, these devices can process executable instructions, including software applications such as web browser applications that can be used to access the internet. Therefore, these devices can possess significant computing power.
[0006] A common application of such devices is processing various types of time-series data, such as sensor data, audio data, video data, and signals. Generally, time-series data is broken down into segments for processing, and different segments can include different types of content or data. In some cases, multiple different mechanisms are available for processing segments of time-series data, each with different advantages and disadvantages. Utilizing multiple different processing operations for different types of data within a single time-series data stream can be challenging if the content and / or data type of each segment of the time-series data is unknown in advance. Summary of the Invention
[0007] According to a particular embodiment, an apparatus includes a memory configured to store one or more segments of time-series data. The apparatus also includes one or more processors configured to use a feature extractor to generate a latent space representation of the segments of the time-series data. The one or more processors are also configured to provide one or more inputs to a classifier, including at least one input based on the latent space representation. The one or more processors are also configured to generate processing control signals for the segments based on the output of the classifier.
[0008] According to a particular implementation, a method includes using a feature extractor to generate a latent space representation of a segment of time series data. The method also includes providing one or more inputs to a classifier, the one or more inputs including at least one input based on the latent space representation. The method further includes processing control signals for generating the segment based on the output of the classifier.
[0009] According to a particular embodiment, a non-transitory computer-readable medium stores instructions that are executable by one or more processors to cause the one or more processors to use a feature extractor to generate a latent space representation of a segment of time-series data. These instructions also cause the one or more processors to provide one or more inputs to a classifier, the one or more inputs including at least one input based on the latent space representation. Furthermore, these instructions cause the one or more processors to generate processing control signals for the segment based on the output of the classifier.
[0010] According to a particular embodiment, an apparatus includes components for generating a latent space representation of a segment of time series data using a feature extractor. The apparatus also includes components for providing one or more inputs to a classifier, the one or more inputs including at least one input based on the latent space representation. The apparatus further includes components for generating a processing control signal for the segment based on the output of the classifier.
[0011] Other aspects, advantages, and features of this disclosure will become apparent upon reading the entire application, which comprises the following sections: description of the drawings, detailed description, and claims. Attached Figure Description
[0012] Figure 1A This is a block diagram illustrating specific exemplary aspects of a system capable of selectively processing time series data, based on some examples of this disclosure.
[0013] Figure 1B Based on some examples of this disclosure Figure 1A An illustrative diagram of the system.
[0014] Figure 2 Based on some examples of this disclosure Figure 1A The system is illustrated in the following diagram.
[0015] Figure 3 Based on some examples of this disclosure Figure 1A The system is illustrated in the following diagram.
[0016] Figure 4 Based on some examples of this disclosure Figure 1A The system is illustrated in the following diagram.
[0017] Figure 5Based on some examples of this disclosure Figure 1A The system is illustrated in the following diagram.
[0018] Figure 6A Examples of training feature extractors according to this disclosure are illustrated for use in... Figure 1A The system is used in all aspects.
[0019] Figure 6B Examples of trained classifiers according to this disclosure are illustrated for use in... Figure 1A The system is used in all aspects.
[0020] Figure 7 Examples of training a feature extractor and a classifier together according to some examples of this disclosure are illustrated for use in... Figure 1A The system is used in all aspects.
[0021] Figure 8 Based on some examples of this disclosure Figure 1A The diagram illustrates the illustrative aspects of the operation of the system's components.
[0022] Figure 9 Examples of integrated circuits capable of selectively processing time-series data according to some examples of this disclosure are illustrated.
[0023] Figure 10 This is an illustration of a mobile device capable of selectively processing time-series data, based on some examples of this disclosure.
[0024] Figure 11 This is an illustration of a headset capable of selectively processing time-series data, based on some examples of this disclosure.
[0025] Figure 12 This is an illustration of a wearable electronic device, according to some examples of this disclosure, capable of selectively processing time-series data.
[0026] Figure 13 The illustrations are of some examples of extended reality glasses devices (such as virtual reality, mixed reality, or augmented reality glasses) that are capable of selectively processing time-series data, according to the present disclosure.
[0027] Figure 14 This is an illustration of an earpiece that can operate to selectively process time series data, based on some examples of this disclosure.
[0028] Figure 15 This is a diagram illustrating a voice-controlled loudspeaker system capable of selectively processing time-series data, based on some examples of this disclosure.
[0029] Figure 16These are illustrations of cameras that are operable to selectively process time-series data, based on some examples of this disclosure.
[0030] Figure 17 This is an illustration of an extended reality headset (such as a virtual reality, mixed reality, or augmented reality headset) capable of selectively processing time-series data, according to some examples of this disclosure.
[0031] Figure 18 This is a diagram illustrating a first example of a vehicle capable of selectively processing time-series data, based on some examples of this disclosure.
[0032] Figure 19 This is a diagram illustrating a second example of a vehicle capable of selectively processing time-series data, based on some examples of this disclosure.
[0033] Figure 20 Based on some examples of this disclosure, it is possible to... Figure 1A A diagram illustrating a specific implementation of a method for selectively processing time-series data performed by a device.
[0034] Figure 21 This is a block diagram of a particular exemplary example of a device capable of selectively processing time-series data, according to some examples of this disclosure. Detailed Implementation
[0035] There are numerous applications for processing time-series data. Common examples include, but are not limited to, video processing, audio processing, motion processing, and sensor processing. In many of these applications, time-series data is broken down into segments for processing. For example, when time-series data is obtained via a data stream, the data stream can be sampled periodically or occasionally to form segments, or several such samples can be aggregated to form segments. For illustration, audio signals from a microphone can be sampled periodically to generate audio data samples, and multiple such audio data samples can be aggregated to form audio data segments (e.g., audio frames). Similar operations can be used to form video segments, sensor data segments, etc.
[0036] After segmentation, the time-series data fragments can undergo further processing. The specific downstream processing used for a fragment depends on the specific purpose for which the time-series data will be used. For example, when the time-series data includes audio or video data, the fragments may be encoded for transmission to another device (e.g., as part of a call or media stream) or compressed for storage.
[0037] In some cases, more than one type of downstream processing can be used for a specific type of time-series data. For example, there are many different compression schemes available for audio data. Similarly, different encoding schemes can be used for audio data. These different types of downstream processing can vary in terms of computational efficiency (e.g., processor cycles, power utilization), data efficiency (e.g., compression ratio), data fidelity, and many other factors. Furthermore, some types of downstream processing can be optimized for or otherwise well-suited for processing specific types of data. As an example, audio encoding schemes suitable for encoding audio data representing speech are designed for the range and sound of human vocalization. Therefore, such audio encoding schemes may be less suitable for encoding audio data representing music compared to encoding schemes tailored for music encoding.
[0038] Applying inappropriate downstream processing operations to specific time series data can be inefficient and may produce outputs that do not represent the time series data with the expected level of fidelity. This is particularly challenging when the characteristics or content of a set of time series data may change from time to time. For example, when the time series data includes audio data representing sounds captured by a microphone, the audio data may represent speech at a first time point, ambient sounds or noise at a second time point, and music at a third time point.
[0039] Therefore, one problem with processing time series data is that among several different available processing options, some are better suited than others for processing specific types of time series data content; however, the content of any single segment of time series data is often unknown beforehand. Thus, there is a problem in choosing the appropriate processing operation for a specific segment of time series data. One solution to this problem is to apply the same processing operation to all time series data. However, this solution leads to further problems. To illustrate, some segments may be processed inappropriately or inefficiently, resulting in waste, rework (e.g., reprocessing of such segments), etc.
[0040] Another solution to the above problem is to analyze each segment to determine its content type, allowing the segment to be routed to the appropriate processing operation. The problem with analyzing each segment is that performing such analysis is resource-intensive and may make it difficult to ensure that all possible content types have been considered. For example, audio data can represent a wide variety of different types of sounds, such as speech, music, alarms, birdsong, wind, waves, engine noise, etc. Sometimes algorithms are used to distinguish some of these various sound types (e.g., speech activity detection algorithms). However, such algorithms may not reliably distinguish all of these sound types. Machine learning models can be more accurate, but one or more classifiers capable of distinguishing all these various sound types would be very large, and training such classifiers would require a training dataset that includes labeled examples for each category. Furthermore, if it is later discovered that important sound types were omitted, new classifiers will need to be created and trained from scratch to account for the addition of new categories.
[0041] The aspects disclosed in this paper provide solutions to these and other problems by using low-complexity machine learning models to analyze and classify segments of time-series data, thereby efficiently routing each segment appropriately. As an example, compared to many other sound analysis models operating on the same data and hardware, the machine learning model described in this paper performs real-time analysis of audio data to distinguish speech sounds from all types of non-speech sounds using less than 100M floating-point operations (FLOPs) per classification result (e.g., approximately 20-50MFLOPs, such as 30MFLOPs), while these other sound analysis models would be expected to perform more than 100MFLOPs or even more than 1000MFLOPs, thus incurring a significant computational burden.
[0042] One way the solution described in this paper achieves this low complexity is by directly feeding the time-series data to downstream processing operations (e.g., decoding operations), rather than using the output of a segment classification operation for downstream processing. For example, in some implementations, the feature extractor is configured to receive data representing input segments to reduce the data in dimensionality and synthesize output data that approximately reconstructs the input segments. Since the synthesized output data is not used by downstream processing operations, it can be low-fidelity (e.g., low-dimensional relative to the dimensionality of the segments in the time-series data) without negatively impacting the output of downstream processing operations. In some implementations, further complexity reduction is achieved by performing segment classification based on a latent space representation having even lower dimensionality than the low-fidelity input to the feature extractor.
[0043] In a particular aspect, an apparatus configured to process time-series data includes a processing controller and one or more downstream processing components. The downstream processing components are configured to perform various processing operations on the time-series data, and the processing controller is configured to generate control signals based on analysis of segments of the time-series data to control which processing operations the downstream processing components use for each segment of the time-series data. As an example, the downstream processing components may include a first audio encoder (e.g., an encoder well-suited for encoding speech) and a second audio encoder (e.g., a general-purpose encoder), and processing control signals from the processing controller may cause the first audio encoder to be used to encode a first set of audio data segments of the time-series data (e.g., segments including speech), and may cause the second audio encoder to be used to encode a second set of audio data segments of the time-series data (e.g., segments not including speech). In this example, the first and second sets of audio data segments may be mixed in the time series. For example, the first set of audio data segments may represent audio data frames including speech, and the second set of audio data segments may represent audio frames not including speech (e.g., including non-speech sounds).
[0044] While audio data was used in the example above, in other examples, time-series data represents content other than sound or content that complements sound (e.g., video data, sensor data, etc.). In each of these other examples, time-series data can be described in terms of target data and non-target data, where target data corresponds to data optimized or otherwise well-suited for processing by a specific downstream processing component, and non-target data is all other data in the time series. In some implementations, the specific downstream processing component may include a target or dedicated processing component that is more efficient, less resource-intensive, or otherwise better suited for processing the target data than a general-purpose processing component.
[0045] One challenge of controlling downstream processing in the manner described above is that it is often impossible to know in advance which segments will be target data and which will be non-target data. In a particular aspect, the processing controller uses one or more machine learning models to classify each segment of time series data (e.g., as target or non-target data), and the processing control signals generated by the processing controller for a particular segment are based on the category assigned to that particular segment. For example, the processing controller includes a trained feature extractor and a trained classifier. In this context, "trained" indicates that at least some of the functionality of the components (e.g., the feature extractor or classifier) is based on the results of machine learning training techniques (e.g., functionality is learned). Training a machine learning model to classify segments of time series data can be challenging when there are a large number of possible categories that can be assigned to a particular type of time series data. For example, there are a large number of different types of sounds that audio data can represent, and training a machine learning model to identify each of these sound types is challenging because it would require sufficiently selecting labeled training data for each category. Furthermore, machine learning models trained in this way cannot reliably classify sound types not included in the training data.
[0046] To address these challenges, the feature extractor of the processing controller includes a machine learning model of the type of autoencoder, trained to synthesize time-series data of the target data type. Continuing with the audio example above, a feature extractor can be trained to reproduce audio data representing speech. An autoencoder trained in this way can accurately reproduce time-series data of the type represented in the training data, but will be less accurate in reproducing time-series data of types not represented in the training data. For example, an autoencoder trained using training data representing speech can accurately reproduce or synthesize audio data representing speech, but will be less accurate in reproducing or synthesizing audio data of all types of non-speech audio. Similarly, an autoencoder trained to reproduce target time-series data will reproduce the target time-series data more accurately than one trained to reproduce non-target time-series data. Furthermore, such an autoencoder can be trained using training data representing only the target data type, which significantly reduces training time and cost, and allows for the use of smaller (i.e., more memory and processor efficient) machine learning models.
[0047] In certain aspects, the classifier includes a machine learning model trained to determine, at least in part, whether a specific segment of time-series data belongs to a target data type or a non-target data type based on the output of a feature extractor. The output may include intermediate or final outputs of the feature extractor. For example, a specific segment of time-series data may be fed as input to the inference part of an autoencoder to generate a latent space representation of the specific segment, and the latent space representation may be fed as input to the generative network part of the autoencoder to generate synthetic data corresponding to a reproduced version of the specific segment. In this example, the latent space representation, synthetic data, or both may be used as the output of the autoencoder, which is then fed to the classifier or used to determine the input to the classifier. Additionally or alternatively, data representing the network states of one or more hidden layers of the inference part or the generative network part of the autoencoder may be used as the output.
[0048] In some implementations, the output of the feature extractor is further processed to generate input for a classifier. For example, an error value may be determined based on the difference between the synthetic data and a specific fragment used to generate the synthetic data. As another example, the error value may be determined based on the divergence between the latent space representation and the expected distribution representing the latent space representation of the target data.
[0049] In some implementations, in addition to the input based on the output of the feature extractor, other inputs may be provided to the classifier. For example, features extracted or determined from the segment may be included in the input, such as the segment's open-loop pitch, normalized correlation, spectral envelope, pitch stability, signal nonstationarity, linear prediction residual, spectral difference, spectral stationarity, or combinations thereof.
[0050] As another example, a pattern detector using a lightweight pattern detection process (compared to processing via feature extractors and classifiers) can also be used to process time-series data, and pattern indicators from the pattern detector can be provided as input to the classifier. For illustration, the pattern detector may include an audio pattern detector that performs statistical operations to determine whether the audio in the time-series data includes music. As another illustrative example, the pattern detector may be configured to determine whether an audio data frame includes voiced or unvoiced speech. As another illustrative example, the pattern detector may be configured to determine Enhanced Variable Rate Codec (EVRC) mode decisions based on segments. In such implementations, pattern detection may be based on assumptions that may not hold true for specific segments of the time-series data. For example, a pattern detector generating voiced / unvoiced pattern indicators might treat each segment of the time series as if it included speech, but this may not actually be the case. Therefore, the pattern indicators generated by the pattern detector are provided to a classifier, which uses the pattern indicators in conjunction with other data to classify the segments as target data or non-target data.
[0051] One benefit of the aspects disclosed herein is the ability to process both target and non-target time series data using different downstream processing operations without prior information about which segments of the time series include target data and which do not. This processing allows for more complex or optimized processing of some data (e.g., target data) and less complex processing of others (e.g., non-target data), thereby saving computational resources (e.g., processor time and memory) compared to using a general process for both target and non-target data, improving the processing of at least the target data in the time series (e.g., better compression, better data fidelity, etc.), and so on.
[0052] Another benefit of the aspects disclosed herein is the ability to distinguish target data within a time series that may include a wide range of non-target data, without requiring training data representing each predictable type of non-target data. For example, an autoencoder can be trained to reproduce a target data type using training data that includes only data of the target data type. Therefore, much less training data is required, and training time is reduced. The trained autoencoder can then be used as a feature extractor to generate feature data that is fed to a relatively simple classifier, such as a uniclass or binary classifier. Thus, both the feature extractor and the classifier can be relatively lightweight (e.g., compared to a classifier trained to distinguish between a large number of categories, such as each predictable type of non-target data plus the target data type), thereby saving computational resources (e.g., processor time and memory).
[0053] As used herein, the term “machine learning” should be understood to have any of its usual and conventional meanings within the fields of computer science and data science, including, for example, processes or techniques by which one or more computers can learn to perform certain operations or functions without being explicitly programmed to do so. As a typical example, machine learning can be used to enable one or more computers to analyze data to identify patterns in the data and generate results based on that analysis.
[0054] For some types of machine learning, the resulting output includes a data model (also known as a "machine learning model" or simply a "model"). Typically, a model is generated using a first dataset to facilitate analysis on a second dataset. For example, the first portion of a large dataset can be used to generate a model that can then be used to analyze the remaining portion of the large dataset. Similarly, a set of historical data can be used to generate a model that can be used to analyze future data.
[0055] Because a model can be used to evaluate a dataset different from the data used to generate the model, it can be considered a type of software (e.g., instructions, parameters, or both) automatically generated by a computer during the machine learning process. Therefore, the model can be transferable (e.g., it can be generated at a first computer and subsequently moved to a second computer for further training, use, or both). Additionally, the model can be combined with one or more other models to perform a desired analysis. For example, first data can be provided as input to a first model to generate first model output data, and the first model output data (alone, with the first data, or with other data) can be provided as input to a second model to generate second model output data indicative of the results of the desired analysis. Depending on the analysis and data involved, different combinations of models can be used to generate such results. In some examples, multiple models can provide model outputs that are input to a single model. In some examples, a single model provides model outputs as input to multiple models.
[0056] Because machine learning models are generated by computers based on input data, they can be discussed within at least two distinct time windows: the creation / training phase and the runtime phase. During the creation / training phase, the model is created, trained, adapted, validated, or otherwise configured by the computer based on the input data (often referred to as "training data" during this phase). It's important to note that a trained model corresponds to software that has been generated and / or refined during the creation / training phase to perform a specific operation (such as classification, prediction, encoding, or other data analysis or data synthesis operations). During the runtime phase (or "inference" phase), the model is used to analyze the input data to generate model output. The content of the model output depends on the type of model. For example, as a non-limiting example, a model can be trained to perform a classification task or a regression task.
[0057] In certain implementations, machine learning models can be trained and used on different computing devices. For example, a first computer or a first group of computers can be used to train the model, and after the model is trained, it can be executed by one or more different computers to analyze data. This transferability of machine learning models means that the computing device using the model (e.g., at runtime) does not need to also train the model. This separation of runtime computation and training provides several benefits. For example, a very large training dataset and a large number of iterations can be used to train the model, which consumes a lot of computing resources, and after training, the model can be moved to a computing device with more limited computing resources for runtime use. To illustrate, a high-end computing device, such as a server or a high-end desktop computer or a group of such computers, can be used to train the model, which has access to advanced processors, external power supplies, and large amounts of memory. This training is typically iterative and may require complex computations optimized for execution in parallel processing threads, as well as many memory input / output (I / O) operations. After training, the model can be moved to a more resource-constrained computing device, such as a smartphone or another device with fewer computing resources (e.g., fewer or fewer high-end processor cores, less memory, etc.), limited power (e.g., battery power), or other limitations (e.g., thermal limitations). Executing the model at runtime (e.g., during the inference phase after training) requires far fewer resources than training the model. Furthermore, the model can be trained on one device (or a set of devices) and subsequently used essentially as software that can be copied to any number of other devices for runtime use.
[0058] Training a model on a training dataset typically involves modifying the model's parameters with the goal of making the model's output possess specific characteristics based on the data input to the model. To distinguish it from model generation operations, model training may be referred to as optimization or optimization training in this paper. In this context, "optimization" refers to improving a metric, not necessarily finding an ideal value for that metric (e.g., a global maximum or minimum). Examples of optimization trainers include, but are not limited to, backpropagation trainers, derivative-free optimizers (DFO), and extreme learning machines (ELM). As an example of training a model, during supervised training of a neural network, input data samples are associated with labels. When input data samples are fed to the model, the model generates output data, comparing that output data with the labels associated with the input data samples to generate error values. The model's parameters are modified to attempt to reduce (e.g., optimize) the error values. As another example of training a model, during unsupervised training of an autoencoder, data samples are fed as input to the autoencoder, and the autoencoder reduces the dimensionality of the data samples (a lossy operation) and attempts to reconstruct the data samples into output data. In this example, the output data is compared with the input data samples to generate the reconstruction loss, and the parameters of the autoencoder are modified to attempt to reduce (e.g., optimize) the reconstruction loss.
[0059] Specific aspects of this disclosure are described below with reference to the accompanying drawings. In this description, common features are designated by common reference numerals. As used herein, various terms are used only for the purpose of describing particular embodiments and are not intended to be limiting of the embodiments. For example, the singular forms “a,” “an,” and “the” are intended to also include the plural forms unless the context clearly indicates otherwise. Furthermore, some features described herein are singular in some embodiments and plural in others. For illustration, Figure 1A Device 102 is depicted including one or more processors ("processor" 190 in FIG. 1). This indicates that in some embodiments, device 102 includes a single processor 190, and in other embodiments, device 102 includes multiple processors 190. For ease of reference herein, such features are generally introduced as "one or more" features and are subsequently referred to in the singular or optional plural form (as indicated by "(multiple)"), unless the aspect described relates to multiples of features.
[0060] In some figures, multiple instances of a particular type of feature are used. Although these features are physically and / or logically different, the same reference numerals are used for each feature, and these different instances are distinguished by adding letters to the reference numerals. Reference numerals are used without distinguishing letters when a feature is referenced herein as a group or of a type (e.g., when a specific feature among these features is not referenced). However, reference numerals are used with distinguishing letters when a specific feature among multiple features of the same type is mentioned herein. For example, referring to Figure 1, multiple operations are illustrated and associated with reference numerals 130A, 130B, and 130N. The distinguishing letter “A” is used when referring to a specific operation among these operations (such as the first operation 130A). However, reference numeral 130 is used without distinguishing letters when referring to any one of these operations or when referring to these operations as a group.
[0061] As used herein, the term "comprise" is used interchangeably with "include". Similarly, the term "wherein" is used interchangeably with "where". As used herein, "exemplary" indicates an example, embodiment, and / or aspect, and should not be construed as limiting or indicating a preference or preferred embodiment. As used herein, ordinal terms used to modify elements (such as structures, components, operations, etc.) (e.g., "first", "second", "third", etc.) do not themselves indicate any priority or order of that element relative to another element, but merely distinguish that element from another element with the same name (but using ordinal terms). As used herein, the term "set" refers to one or more specific elements among specific elements, while the term "multiple" refers to multiple (e.g., two or more) specific elements.
[0062] As used herein, “coupling” can include “communicationally coupled,” “electrically coupled,” or “physically coupled,” and may also (or alternatively) include any combination thereof. Two devices (or components) may be directly or indirectly coupled (e.g., communicationally coupled, electrically coupled, or physically coupled) via one or more other devices, components, wires, buses, networks (e.g., wired networks, wireless networks, or combinations thereof). As an illustrative, non-limiting example, two electrically coupled devices (or components) may be included in the same device or in different devices and may be connected via electronics, one or more connectors, or inductive coupling. In some embodiments, two communicationally coupled (such as electrically connected) devices (or components) may transmit and receive signals (e.g., digital or analog signals) directly or indirectly via one or more wires, buses, networks, etc. As used herein, “direct coupling” can include two devices coupled (e.g., communicationally coupled, electrically coupled, or physically coupled) without intermediate components.
[0063] In this disclosure, terms such as “determine,” “calculate,” “estimate,” “shift,” and “adjust” can be used to describe how one or more operations are performed. It should be noted that such terms should not be construed as restrictive, and similar operations can be performed using other techniques. Additionally, as mentioned herein, “generate,” “calculate,” “estimate,” “use,” “select,” “access,” and “determine” are used interchangeably. For example, “generating,” “calculating,” “estimate,” or “determining” a parameter (or signal) can refer to actively generating, estimating, calculating, or determining the parameter (or signal), or it can refer to using, selecting, or accessing a parameter (or signal) that has already been generated (e.g., by another component or device).
[0064] Referring to Figure 1, specific exemplary aspects of a system configured for selective processing of time series data are disclosed, and the system is generally designated as 100. System 100 includes a device 102 comprising one or more processors 190 and a memory 192 (e.g., one or more memory devices). Device 102 includes a processing controller 140 configured to control the selective processing of segments of time series data 110. For example, processing controller 140 is configured to generate a processing control signal 122 to control which of a set of downstream processing components 142 is used to process each segment of time series data 110.
[0065] In some embodiments, device 102 includes one or more input interfaces 108, and time-series data 110 is received via input interfaces 108. For example, in FIG. 1, microphone 104 is coupled to input interface 108 to receive sound 106. In this example, time-series data 110 may include audio data representing sound 106. In other embodiments, one or more cameras, one or more additional microphones, one or more sensors, or other devices that generate time-series data 110 may be coupled to input interface 108 to generate time-series data 110 or a portion thereof. For example, a light detection and ranging (LiDAR) system may be coupled to input interface 108, and time-series data 110 may represent a sequence of point clouds based on LiDAR returns. In still other examples, other active sensors or sensor systems (where “active” indicates that such sensors or systems generate information based on signal returns) or other passive sensors or systems (where “passive” indicates that such sensors or systems do not depend on signal returns) may be used to generate a sequence of sensor data corresponding to time-series data 110.
[0066] In some embodiments, device 102 includes modem 170, via which time-series data 110 can be received. For example, one or more devices 160 may transmit time-series data 110 to device 102 via one or more modulated signals, and modem 170 may demodulate the signals to generate time-series data 110 that is provided to processor 190.
[0067] In some implementations, time-series data 110 may be stored in memory 192 after being received via input interface 108 and / or modem 170. In such implementations, one or more segments 196 of time-series data 110 may be retrieved from memory 192 for processing by processing controller 140, by one or more downstream processing components 142 (e.g., in response to processing control signal 122), or both.
[0068] In the example illustrated in Figure 1, the processing controller 140 includes a preprocessor 112, a feature extractor 116, and a classifier 120. The preprocessor 112 is configured to prepare segment data 114 as input to the feature extractor 116, the classifier 120, or both. As an example, for some types of time-series data 110, the preprocessor 112 may be configured to segment the time-series data 110 (e.g., sample using a time window or frame it), and each segment or set of segments may be provided as part of the segment data 114. As another example, the preprocessor 112 may perform operations to modify the time-series data 110 (before or after segmentation). Examples of such modifications may include, but are not limited to, domain transformations (such as transforming time-domain data to frequency-domain data), resampling, normalization, filtering, signal enhancement, etc. In some embodiments, the preprocessor 112 generates at least a portion of the segment data 114 based on the analysis results of segments of the time-series data 110. For illustration, preprocessor 112 may perform one or more statistical analyses on one or more segments of time series data 110 and provide statistics associated with the segments as part of segment data 114. In some embodiments, preprocessor 112 may also perform operations to generate inputs for classifier 120. For example, preprocessor 112 may perform pattern detection operations to determine pattern indicators associated with one or more segments of time series data 110, and the pattern indicators may be provided to classifier 120 as one of a set of one or more inputs 118. As another example, preprocessor 112 may determine features of the segments (e.g., using an algorithm rather than a trained feature extraction process). For illustration, preprocessor 112 may determine the open-loop pitch, normalized correlation, spectral envelope, pitch stability, signal nonstationarity, linear prediction residual, spectral difference, spectral stationarity, or combinations thereof of the segments.
[0069] Feature extractor 116 and classifier 120 each include one or more machine learning models. Examples of machine learning models include, but are not limited to, perceptrons, neural networks, support vector machines, regression models, decision trees, Bayesian models, Boltzmann machines, adaptive neural fuzzy inference systems, and combinations, sets, and variations of these and other types of models. Variations of neural networks include, for example, but not limited to, prototype networks, autoencoders, transformers, self-focused networks, convolutional neural networks, deep neural networks, deep belief networks, etc. Variations of decision trees include, for example, but not limited to, random forests, boosted decision trees, etc.
[0070] In some embodiments, the feature extractor 116 includes or corresponds to an autoencoder. In such embodiments, the autoencoder includes an inference network portion, a bottleneck layer, and a generator network portion. In some contexts, the inference network portion of the autoencoder is referred to as the encoder network, and the generator network portion of the autoencoder is referred to as the decoder network; however, the inference network portion and the generator network portion are used herein to avoid potential confusion with other optional aspects of the downstream processing component 142 described below.
[0071] The inference network portion of the autoencoder is configured to receive input data samples (e.g., segment data 114 representing a segment of time series data 110) and reduce the dimensionality of the input data samples to the dimension of the bottleneck layer. In some embodiments, the autoencoder is a variational autoencoder, in which case the inference network portion maps the dimensionality-reduced representation of the input data samples to a probability distribution. The dimensionality-reduced representation of the input data samples in the bottleneck layer is also referred to herein as the "latent space representation". When the input to the autoencoder includes segment data 114 of a segment of time series data 110, the latent space representation at the bottleneck layer corresponds to or includes the latent space representation of the segment of time series data 110.
[0072] At least during the training of the autoencoder, the generative network part is configured to generate output data that attempts to reconstruct samples of the input data. For example, the generative network part may receive latent space representations of specific segments (or, in the case of a variational autoencoder, sampled latent space representations from a probability distribution) and perform a dimensionality expansion operation to generate output data with dimensions corresponding to the dimensions of the input to the inference network part. Thus, the latent space representations are trained to control properties of the output of the generative network part (e.g., properties of synthesized speech or other target audio), and the latent space representations of non-target audio (e.g., non-speech sounds) exhibit differences that can be used to classify segments of the audio data. For example, since the feature extractor 116 is trained to reconstruct data of the target data type, the latent space representation of the target data type tends to separate in latent space from the latent space representation of the non-target data type. This latent space separation between target and non-target data types can be detected by a classifier (e.g., in the latent space representations, or based on a reconstruction error metric associated with the output of the generative network part) to assign a classification to the segment. For example, if the input data samples include specific segments of time series data 110, the output of the generative network part includes data representing synthetic reconstructions of specific segments of time series data 110. In some implementations, the output of the generator network portion is not used at runtime (e.g., after the feature extractor has been trained and is processing time series data 110). For example, the operation of the generator network portion may be omitted (e.g., not performed by processor 190), or the output of the generator network portion may be discarded.
[0073] In some implementations, the output of the generative network portion (e.g., synthetic fragment data) is used to determine at least a portion of the input 118 of classifier 120. For example, the synthetic fragment data can be compared with fragment data 114 to compute an error metric, and the value of the error metric can be provided as part of the input 118. Since the dimensionality reduction performed by the inference network portion is a lossy operation, the output of the generative network portion will typically differ from the input data samples, and the error metric quantifies the difference between the output and the input. For example, the error metric can be viewed as vectors representing locations in a feature space, and the value of the error metric can be computed based on the distance between the vectors. As another example, the error metric can be computed based on a comparison of the probability distribution associated with the input (e.g., fragment data 114) and the probability distribution associated with the output (e.g., synthetic fragment data). For illustration, the error metric can be computed using the Itakura-Saito distance based on the input and output. Other examples of error metrics that can be used include, but are not limited to, scale-invariant signal distortion ratio or log-spectral distortion value.
[0074] In some implementations, the autoencoder of the feature extractor includes one or more recurrent layers, one or more dilated convolutional layers, and / or one or more other temporally dynamic layers. For example, the autoencoder may include one or more Long Short-Term Memory (LSTM) or Gated Recurrent Unit (GRU) layers. Such temporally dynamic structures enable the analysis of segments of time-series data 110 to take into account the context of the segments within the time series.
[0075] In some implementations, classifier 120 includes or corresponds to a neural network, decision tree, support vector machine, or another machine learning model or ensemble of machine learning models configured to generate a classification output (e.g., an output indicating whether input 118 is associated with a target data type). As an example, classifier 120 may include a neural network having an input layer configured to receive input 118, one or more hidden layers, and an output layer configured to generate an output indicating a determination of whether input 118 is associated with a target data type or a probability that input 118 is associated with a target data type. In some implementations, classifier 120 is a uniclass classifier or a binary classifier.
[0076] Processing controller 140 is configured to generate processing control signals 122 for specific segments of time-series data 110 based on the output of classifier 120. For example, if classifier 120 indicates that a specific segment is assigned to a target data type, processing controller 140 generates a first type of processing control signal 122, and if classifier 120 indicates that a specific segment is not assigned to a target data type, processing controller 140 generates a second type of processing control signal 122. As a non-limiting example, time-series data 110 may represent audio content, and the target data type may be audio data representing speech. In this example, classifier 120 generates an output indicating whether a specific segment of time-series data 110 includes speech, and processing controller 140 generates processing control signals 122 based on whether the segment includes speech.
[0077] Processing control signal 122 associated with the fragment is provided to one or more downstream processing components 142 to control which of the set of operations 130 (if any) is available to the downstream processing component 142 that performs the fragment.
[0078] The specific operations 130 that can be performed by the downstream processing component 142 depend on the nature of the time series data 110 and what the time series data 110 will be used for. Examples of operations 130 that can be performed by the downstream processing component 142 include, but are not limited to, encoding (e.g., for transmission), compression (e.g., for storage or transmission), rendering (e.g., for output), data integration (e.g., combining fragments of the time series data 110 with other data), etc.
[0079] In some implementations, certain operations in operation 130 are better suited than others in operation 130 for processing fragments that include the target data type. For example, a first operation 130A may be optimized for processing fragments of the target data type, a second operation 130B may be optimized for processing fragments that do not include the target data type, and an Nth operation 130N may include general operations that are not optimized for any particular data type. As another example, a first operation 130A may be optimized for processing fragments of the target data type, a second operation 130B may include general operations that are not optimized for any particular data type, and an Nth operation may be omitted.
[0080] As a particular, non-limiting example, when the time-series data 110 includes audio data representing sound, the first operation 130A may include a first audio encoder, and the second operation 130B may include a second audio encoder. In this example, the first audio encoder may include a speech encoder specifically configured to encode audio data representing speech, and the second audio encoder may include a general-purpose audio encoder. In this example, the speech encoder may be able to encode the speech audio frame in a manner that achieves high-fidelity reproduction of the speech frame, such as by not emphasizing portions of the audio frame in frequency bands outside of normal human speech, by emphasizing portions of the audio frame that are characteristic of normal human speech, or both. However, a speech encoder may be less efficient than a general-purpose audio encoder in terms of processing time, audio data compression, power utilization, or other factors. In this example, it may be advantageous to use the first operation 130A (in this example, the speech encoder) only on segments of the time-series data 110 representing speech. Therefore, the processing control signal 122 may instruct the downstream processing component 142, based on which segments include speech, which segments of the time-series data 110 should be processed using the first operation 130A. While the target data type in the example above is audio data representing speech, other types of target data are used in other examples. For illustration, when the time series data 110 includes video data, the target data type could include video frames representing motion, video frames including faces, blurred video frames, video frames with low lighting, etc.
[0081] Device 102 may include, correspond to, or be included within a device of various types of devices. For illustration, in various embodiments, processor 190 is integrated in at least one of the following: as referenced Figure 10 The described mobile phone or tablet computer device, as shown in the reference Figure 11 Further description of the head-mounted device, as in the reference Figure 12 The wearable electronic devices described, as in the reference Figure 13 The described augmented reality glasses, as reference Figure 14 The described earplugs, as referenced Figure 15 The described voice-controlled speaker system, as in the reference Figure 16 The camera equipment described, or as referenced Figure 17 The extended reality headset described. In another illustrative example, the processor 190 is integrated into a vehicle, such as a reference pod. Figure 18 and Figure 19 Further description.
[0082] Therefore, system 100 can selectively route specific segments of time series data 110 to appropriate downstream processing components 142 based on the content represented by the segments (e.g., whether the segments represent data of the target data type), without prior information about the content of the segments. The advantage of this selective routing and processing is that downstream processing operations that may be less efficient but are preferred for other reasons (such as speed or fidelity) can be performed on some segments, while other downstream processing operations that may be more efficient but are less preferred for other reasons (such as speed or fidelity) can be performed on other segments of the time series data. Additionally or alternatively, downstream processing operations that are particularly suitable for the target data type for reasons other than efficiency, speed, or fidelity can be used for segments that include the target data type, and other downstream processing operations (which may have other advantages) can be used for segments that do not include the target data type.
[0083] An additional benefit of the specific aspects disclosed herein is that when time series data 110 may include a large number of different types of data (of which only one type corresponds to the target data type), using an autoencoder as a feature extractor 116 and a uniclass or binary classifier as a classifier 120 simplifies training. For example, the autoencoder can be trained using only data of the target data type. In contrast, many other types of machine learning models require data that is more representative of the data to be processed after training (e.g., data including both the target data type and one or more non-target data types). Limiting the training data required to train the feature extractor reduces the cost of training the feature extractor (e.g., monetary costs and / or computational resource costs).
[0084] Similarly, a fairly limited dataset can be used to train a uniclass or binary classifier for use as classifier 120. For illustration, in some cases, the training data used to train a uniclass classifier may only include examples of the target data type. In such cases, a limited dataset containing only examples of the target data type can be used to train both the feature extractor and classifier 120, resulting in cost savings as described above. Generally, training a binary classifier requires training data that includes examples of both the target and non-target data types; however, when many types of non-target data exist, not all of the various non-target data types need to be represented and / or may be imbalanced. Therefore, the training data is far less constrained than the training data used to train a multi-class classifier, which is trained to distinguish various non-target data types.
[0085] although Figure 1AA microphone 104 coupled to device 102 is illustrated, but in other embodiments, microphone 104 may be integrated into device 102. In still other embodiments, other sensors, as alternatives to or supplements to microphone 104, may be coupled to or included within device 102 to generate time-series data 110. Additionally, although... Figure 1A Example 1 illustrates time-series data 110 received by processor 190 from input interface 108, but in other embodiments, time-series data 110 may be received via transmission from another device (e.g., one of the devices in device 160) and provided to processor 190 from modem 170. In the same or different embodiments, time-series data 110 received via input interface 108, modem 170, or both may be stored in memory 192 (e.g., as segment 196), and processor 190 may retrieve segment 196 from memory 192 for processing.
[0086] In a particular embodiment, processor 190 is configured to execute instructions 194 from memory 192 to perform one or more of the operations described above with reference to processing controller 140 or downstream processing component 142. For example, execution of instructions 194 may cause processor 190 to: generate a latent space representation of a segment of time series data 110 using feature extractor 116; provide input 118 to classifier 120, wherein input 118 includes at least one input based on the latent space representation; and generate a processing control signal 122 for the segment based on the output of classifier 120. As another example, execution of instructions 194 may cause processor to selectively perform a first operation 130A, a second operation 130B, or an Nth operation 130N using the segment based on processing control signal 122. For illustration, based on processing control signal 122, downstream processing component 142 may selectively perform a first encoding operation (e.g., first operation 130A) to encode the segment, or perform a second encoding operation (e.g., second operation 130B) to encode the segment. In this exemplary example, the first encoding operation may provide more efficient encoding of the target signal category (e.g., a signal representing the target data type) than the second encoding operation. Additionally or alternatively, the first encoding operation may provide higher quality encoding of the target signal category than the second encoding operation.
[0087] In some implementations, downstream processing component 142 includes different components associated with different operations 130. In such implementations, processing control signal 122 can control the activation and / or deactivation of various downstream processing components in downstream processing component 142. For example, a first operation 130A can be performed by a first subset of downstream processing components 142, and a second operation 130B can be performed by a second subset of downstream processing components 142. In this example, when processing control signal 122 associated with a specific segment of time series data 110 indicates that the first operation 130A should be performed, a first subset of downstream processing components 142 can be activated (e.g., powered on), a second subset of downstream processing components 142 can be deactivated (e.g., powered off), or both.
[0088] Figure 1B Based on some examples of this disclosure Figure 1A The system is illustrated in the diagram. Specifically, Figure 1B An example of an embodiment is illustrated, comprising a processing controller 140 and a downstream processing component 142, in which time-series data 110 includes audio data 158 and a target signal category corresponding to the audio data representing speech. Therefore, in Figure 1B In the illustrated example, the processing control signal 122 selects between a speech decoder 144 when a segment of audio data 158 is classified as representing speech, or one or more non-speech decoders 148 when a segment of audio data 158 is not classified as representing speech.
[0089] For example, in Figure 1B In the process, when a segment of audio data 158 is received, the processing controller 140 executes as described in the reference. Figure 1A The described operation assigns a segment category to the segment. In this example, the segment category indicates whether the segment represents speech. If the segment category indicates that the segment represents speech, the processing controller 140 generates a processing control signal 122 to select the speech decoder 144. Segments of audio data 158 are provided to the speech decoder 144, which in this example performs a linear prediction (LP) based decoding operation 146 to generate an output bitstream.
[0090] If the segment classification indicates that the segment does not represent speech (e.g., represents non-speech), then the processing controller 140 generates a processing control signal 122 to select the non-speech decoder 148. Segments of audio data 158 are provided to the non-speech decoder 148, and the non-speech decoder 148 performs appropriate decoding operations for the segments. Figure 1BIn the illustrated example, the non-speech decoder 148 is configured to perform frequency domain decoding operations (such as modified discrete cosine transform (MDCT) decoding operations, transform-decode-excited (TCX) decoding operations, etc.) and inactive signal decoding operations (such as comfort noise generation (CNG) operations). In some embodiments, Figure 1B The downstream processing component 142 corresponds to the decoding mode of the Enhanced Voice Service (EVS) codec. In this case, the downstream processing component 142 may include additional control and / or switching to select among various non-speech decoders 148 based on factors such as the signal-to-noise ratio or noise level of the audio data 158.
[0091] The outputs of the speech decoder 144 and the non-speech decoder 148 can be combined at the multiplexer (MUX) 154 to generate a bitstream 156 representing the audio data 158.
[0092] Figure 2 Based on some examples of this disclosure Figure 1A The system is illustrated in the diagram. Specifically, Figure 2 An example of the feature extractor 116 and classifier 120 of the processing controller 140 in Figure 1 is shown.
[0093] exist Figure 2 In the illustrated example, feature extractor 116 includes autoencoder 200. Autoencoder 200 includes inference network portion 204, bottleneck layer 206, and generative network portion 208. Inference network portion 204 is configured to receive input data samples 202(X), such as segment data representing fragments of time series data 110 of FIG1; and to reduce the dimension of input data samples 202 to the dimension of bottleneck layer 206 to form a latent space representation 212 of input data samples 202.
[0094] exist Figure 2 In the illustrated example, latent space representation 212 is provided as input to classifier 120. Classifier 120 is configured to generate a fragment classification 220 indicative of the input data sample 202 based on the latent space representation 212. For example, fragment classification 220 may indicate whether the input data sample 202 includes data of the target data type. For illustration, classifier 120 may include a uniclass classifier or a binary classifier, and fragment classification 220 may include a binary output having a first value (e.g., "1") if classifier 120 determines that the input data sample 202 includes the target data type, and a second value (e.g., "0") if classifier 120 determines that the input data sample 202 does not include the target data type.
[0095] Figure 1AThe processing controller 140 can output segment classification 220 as a processing control signal 122 for segments of time series data 110 corresponding to input data sample 202. Alternatively, Figure 1A The processing controller 140 can generate processing control signals 122 for segments of time series data 110 corresponding to input data sample 202 based on segment classification 220. For illustration, the processing controller 140 may include mapping data that maps a particular processing control signal 122 to a corresponding segment classification 220.
[0096] The generator network portion 208 is configured to receive the latent space representation 212 from the bottleneck layer 206 and generate a synthetic reconstruction 210 (X') of the input data sample 202 (X). In some embodiments, the synthetic reconstruction 210 of the input data sample 202 generated by the generator network portion 208 is not used at runtime (e.g., after the feature extractor 116 has been trained and is processing the time series data 110 of Figure 1). In such embodiments, the generator network portion 208 may be omitted from the feature extractor 116 during use. For example, the instructions and model parameters associated with the generator network portion 208 may not be loaded into the processor 190 during use. In other embodiments, the generator network portion 208 is present in the feature extractor 116 during use, but the output generated by the generator network portion 208 (e.g., the synthetic reconstruction 210 of the input data sample 202) is discarded, ignored, or used for purposes other than generating input to the classifier 120. Omitting or not executing the network generation section 208 at runtime provides the benefit of saving computational resources (e.g., processor time, memory I / O, power, etc.).
[0097] In some implementations, the inference network portion 204, the generation network portion 208, or both include one or more recursive layers, one or more dilated convolutional layers, and / or one or more other temporally dynamic layers, such as LSTM layers, GRU layers, etc. The presence of recursive and / or temporally dynamic layers enables the latent space representation 212, the synthetic reconstruction 210 of the input data sample 202, or both, to take into account other data samples of the time series that form the context of the input data sample 202.
[0098] Figure 3 Based on some examples of this disclosure Figure 1A The system is illustrated in the diagram. Specifically, Figure 3 Examples of the preprocessor 112, feature extractor 116, and classifier 120 of the processing controller 140 of Figure 1 are shown.
[0099] exist Figure 3 In the example, feature extractor 116 includes a reference Figure 2 The features and functions described are the same as those described. For example, Figure 3 Feature extractor 116 includes Figure 2 An autoencoder 200 is provided, comprising an inference network portion 204, a bottleneck layer 206, and a generator network portion 208. The inference network portion 204 is configured to receive input data samples 202(X) (such as segment data representing a segment of time series data 110 of FIG1) and reduce the dimension of the input data samples 202 to the dimension of the bottleneck layer 206 to form a latent space representation 212 of the input data samples 202. The generator network portion 208 is configured to receive the latent space representation 212 from the bottleneck layer 206 and generate a synthetic reconstruction 210(X') of the input data samples 202(X).
[0100] In addition, Figure 3 In the example, classifier 120 is configured to receive one or more inputs 118 and generate an output 220 indicating a segment classification 220 of the input data sample 202 based on the one or more inputs 118. As further described below, the one or more inputs 118 include at least one input based on the latent space representation 212 of the input data sample 202.
[0101] exist Figure 3 In this configuration, preprocessor 112 includes a fragment data generator 302. Fragment data generator 302 is configured to generate input data samples (e.g., input data sample 202) based on time series data 110. Specific operations performed by fragment data generator 302 depend on the nature of the time series data 110 (e.g., content and format). For example, if the time series data 110 includes one or more analog signals, fragment data generator 302 may perform framing operations to generate time window samples of the analog signals, where each time window sample or set of time window samples corresponds to a fragment of the time series data 110. Framing operations may also be performed in cases where the time series data includes a bitstream or one or more modulated signals. As other examples, depending on the format and content of the time series data 110, fragment data generator 302 may perform filtering operations, spectral analysis operations, domain transformation operations, data aggregation operations, statistical analysis operations, etc., to generate fragments of the time series data 110.
[0102] Optionally, the preprocessor 112 may include a pattern detector 304. In an embodiment where the preprocessor 112 includes a pattern detector 304, the pattern detector 304 is configured to analyze the time series data 110 to determine pattern indicators 306 associated with one or more segments of the time series data 110. In such an embodiment, the input 118 to the classifier 120 may include the pattern indicator 306.
[0103] Pattern detector 304 is configured to perform relatively lightweight and / or efficient analysis (compared to the operations performed by feature extractor 116 and classifier 120) on segments of time-series data to generate pattern indicators 306. For example, pattern detector 304 may use statistical techniques, pattern matching, etc. As an example, when time-series data 110 includes audio data, pattern detector 304 may include a speech pattern detector configured to distinguish between voiced and unvoiced speech. In this example, pattern indicators 306 for specific segments of time-series data 110 are generated based on the assumption that a particular segment includes speech (regardless of whether the segment actually includes speech). Thus, in this example, if the segment contains voiced speech, pattern indicator 306 indicates that the segment contains voiced speech, and if the segment does not contain voiced speech (e.g., includes unvoiced speech, silence, music, or any other sound), pattern indicator 306 has an uncertain value, such as a random value that may indicate either voiced or unvoiced speech. Therefore, in this example, the pattern indicator 306 cannot reliably indicate the fragment classification 220, but it does provide some data that the classifier 120 can use to assign fragment classification 220. As another example, the pattern indicator may include fragment-based enhanced variable rate codec (EVRC) pattern decisions.
[0104] In other examples, pattern detector 304 is configured to distinguish other sounds, such as differentiating speech from music or wind noise from other sounds. In still other examples, time series data 110 includes information other than or supplementing sound, such as video data or sensor readings, and pattern detector 304 performs pattern detection suitable for such other data. For illustration, when time series data 110 includes video data (e.g., a sequence of images), a pattern indicator 306 for a particular image (e.g., a segment of video data) can indicate whether the image includes a face.
[0105] exist Figure 3 In the classifier 120, the input 118 includes at least one input based on the latent space representation 212. For example, in... Figure 3 In the input 118, there are divergence values 322 and reconstruction error values 324, each of which is based on the latent space representation 212.
[0106] In certain respects, Figure 1A The processing controller 140 includes a divergence calculator 312 configured to determine a divergence value 322. The divergence value 322 is an error value indicating the divergence between the probability distribution of the latent space representation 212 and the expected distribution 310. In an embodiment where the divergence value 322 is in the input 118 to the classifier 120, the autoencoder 200 of the feature extractor 116 corresponds to a variational autoencoder (VAE). See reference... Figure 5 Further described, the latent space representation 212 of the VAE includes data representing a probability distribution (e.g., including mean and variance). In such embodiments, during training, the latent space representation 212 of the VAE is normalized to an expected probability distribution (e.g., expected distribution 310). The divergence value 322 during inference time (i.e., after training the VAE) indicates the divergence between the probability distribution of the latent space representation 212 and the expected distribution 310. The divergence value 322 may be calculated as, for example, but not limited to, the Kullback–Leibler divergence or another f-divergence indicating the difference between the two probability distributions.
[0107] In certain respects, Figure 1A The processing controller 140 includes a reconstruction error calculator 314 configured to determine a reconstruction error value 324. The reconstruction error value 324 indicates how closely the synthesized reconstruction 210 (based on the latent space representation 212 of the input data sample 202) matches the input data sample 202. Examples of calculations that can be used to determine the reconstruction error value 324 include, but are not limited to, cosine distance, Itakura-Saito distance, scale-invariant signal distortion ratio, and logarithmic spectral distortion value.
[0108] although Figure 3 The example illustrates three inputs 118 to classifier 120, but in other examples, inputs 118 to classifier 120 may include more than three or fewer inputs. As an example, pattern indicator 306 may be omitted from input 118. Additionally or alternatively, reconstruction error value 324 or divergence value 322 may be omitted from input 118. As another example, preprocessor 112 may include more than one pattern detector 304, in which case input 118 may include more than one pattern indicator 306. For illustration, preprocessor 112 may include a speech pattern detector and a music pattern detector, each of which generates a corresponding pattern indicator 306 provided as input to classifier 120.
[0109] Figure 4 Based on some examples of this disclosure Figure 1A The system is illustrated in the diagram. Specifically, Figure 4 Examples Figure 3 For example, time series data 110 includes audio data 402.
[0110] Figure 4 The illustrated examples include a preprocessor 112, a feature extractor 116, and a classifier 120, each of which can be adapted to operate on audio data 402 and / or audio segment data 410 based on audio data 402. For example, Figure 4The preprocessor 112 includes a fragment data generator 302 configured to generate audio fragment data 410 based on audio data 402. As another example, the feature extractor 116 includes an autoencoder 200 trained to take audio fragment data 410 as input to generate a latent space representation 212 of the audio fragment data 410, and to generate reconstructed audio fragment data 412 based on the latent space representation 212.
[0111] Optionally, Figure 4 The preprocessor 112 may also include an audio pattern detector 404, which is an example of pattern detector 304. The audio pattern detector 404 is configured to generate audio pattern indicators 408, such as a speech pattern indicator, a music pattern indicator, etc.
[0112] exist Figure 4 In this example, classifier 120 includes or corresponds to audio classifier 420. In this example, audio classifier 420 is configured to receive one or more inputs 118, such as references... Figures 1A to 3 The inputs described by any one of the inputs in the text are used to generate a segment classification 220, such as an audio category 422. As a non-limiting example, audio category 422 may indicate whether the audio segment data 410 includes data representing speech. In this example, when the audio segment data 410 includes data representing speech, Figure 1A The processing controller 140 may transmit a processing control signal 122 to the downstream processing component 142 to cause the downstream processing component 142 to perform a speech processing operation (e.g., a first operation 130A). The downstream processing component 142 may, in response to the processing control signal 122, obtain a portion of the audio data 402 associated with the processing control signal 122 (e.g., an audio segment corresponding to audio segment data) and use the speech processing operation to process the portion of the audio data 402.
[0113] Alternatively, in the example above, if the audio segment data 410 does not include data representing speech (as indicated by audio category 422), then Figure 1A The processing controller 140 may transmit a processing control signal 122 to the downstream processing component 142 to cause the downstream processing component 142 to perform a non-speech processing operation (e.g., a second operation 130B). The downstream processing component 142 may, in response to the processing control signal 122, obtain a portion of the audio data 402 associated with the processing control signal 122 (e.g., an audio segment corresponding to audio segment data) and process the portion of the audio data 402 using the non-speech processing operation.
[0114] Figure 5 Based on some examples of this disclosure Figure 1AThe system is illustrated in the diagram. Specifically, Figure 5 An example of feature extractor 116 is shown. Figure 5 In the illustrated example, feature extractor 116 includes a recursive VAE. As explained above, in other examples, feature extractor 116 includes different types of machine learning models. Figure 5 The depicted implementation illustrates a low-complexity non-autoregressive generative network portion 208 (e.g., such that the latent space contains most of the information), where simple priors are applied to gain information that is useful during inference (e.g., through classifier 120).
[0115] exist Figure 5 In this configuration, the feature extractor 116 includes an inference network portion 204, a bottleneck layer 206, and a generation network portion 208. The inference network portion 204 includes one or more fully connected (FC) layers 504 configured to receive fragment data 502 as input. The FC layers 504 are coupled to one or more LSTM layers configured to reduce the dimensionality of the fragment data 502, consider the context of the fragment data 502 in a time series, or both. The LSTM layers 506 are coupled to one or more FC layers 508 configured to perform dimensionality reduction (or further dimensionality reduction) of the fragment data 502.
[0116] FC layer 508 is coupled to linear layers 510 and 520, which correspond to bottleneck layer 206. Linear layer 510 is configured to generate an output representing the mean 512 of the latent space representation 212, and linear layer 520 is configured to generate an output representing the variance 522 of the latent space representation 212.
[0117] The generator network portion 208 is configured to sample the probability distribution of the latent space representation 212 to determine the sampled latent values 532 as input. Figure 5 In the illustrated example, the generator network portion 208 includes hidden layers that are mirror images of the hidden layers of the inference network portion 204. For example, the generator network portion 208 includes one or more FC layers 534 coupled to one or more LSTM layers 536, and one or more FC layers 538 coupled to one or more LSTM layers 536. The generator network portion 208 also includes a linear layer 540 as an output layer coupled to the FC layer 538. The linear layer 540 is configured to generate synthesized fragment data 542 as output.
[0118] As described above, in some implementations, the output of the generation network portion 208 is (e.g., Figure 5The synthesized fragment data 542 in the dataset is not used by the classifier. In such embodiments, the fidelity of the synthesized fragment data 542 to the fragment data 502 is not critical for the proper runtime (e.g., post-training) operation of the processing controller 140. Therefore, in some such embodiments, the hidden layers of the generating network portion 208 are not mirror images of the hidden layers of the inference network portion 204. For example, the generating network portion 208 may include fewer hidden layers than the inference network portion 204. As another example, the generating network portion 208 may include hidden layers different from those of the inference network portion 204. For illustration, the LSTM layer 528 of the generating network portion 208 may be omitted. While such differences between the generating network portion 208 and the inference network portion 204 may increase reconstruction error, reducing the number and / or complexity of the hidden layers in the generating network portion 208 saves computational resources.
[0119] Figure 6A Examples of training feature extractors according to this disclosure for use by the system of Figure 1 are illustrated, and Figure 6B Examples of trained classifiers according to this disclosure are illustrated for use in... Figure 1A The system is used in all aspects. Figure 6A and Figure 6B In the example shown, the feature extractor and classifier are trained separately.
[0120] refer to Figure 6A The training system 600 for training the feature extractor 608 includes training data 602, a preprocessor 112, the feature extractor 608, one or more error calculators 610, and an optimizer 614. The preprocessor 112 is optional and is used when it will be used after training (e.g., in the inference phase). The feature extractor 608 may include an autoencoder, such as... Figures 2 to 5 The autoencoder 200. Initially, the parameters of the feature extractor 608 (e.g., weights, biases, etc.) have not been tuned, such that the output (X') of the feature extractor 608 represents the input (X). The training process is iterative, and the parameters are gradually updated to reduce the mismatch between the input (X) and the output (X'), as further explained below.
[0121] In certain aspects, the training data 602 used to train the feature extractor 608 includes target data 604, and non-target data 606 may be omitted. For example, when the target data includes speech, the training data 602 may include audio data segments or files representing audio data streams, wherein each audio data segment includes speech, and no audio data segment does not include speech. The training data 602 typically includes a large number of data samples (e.g., segments) that are provided individually or in groups to the training system 600 to train the feature extractor.
[0122] During one iteration of the training system 600, a fragment of the target data 604 or a set of fragments of the target data 604 (e.g., a batch of consecutive fragments representing a time series) is provided to the preprocessor 112. The preprocessor 112 generates input (X) to the feature extractor 608, wherein the input (X) corresponds to or includes fragment data representing fragments of the training data 602, such as fragment data 114.
[0123] Feature extractor 608 is referenced above. Figures 2 to 5 The input (X) is processed as described in the autoencoder 200. For example, the feature extractor 608 passes data between multiple hidden layers to reduce the dimension of the input (X) to form a latent space representation, and processes the latent space representation to generate an output (X') based on the input (X).
[0124] The output (X') and input (X) are provided to the error calculator 610 to generate one or more error measures 612. The error measure 612 indicates the difference between the input (X) and the output (X'). For example, the error calculator 610 may include... Figure 3 The reconstruction error calculator 314. In some embodiments, the error calculator 610 also includes or alternatively includes Figure 3 A divergence calculator 312 is used. An error metric 612 is provided to an optimizer 614, which uses the error metric 612 to generate updated weights 616 for the feature extractor 608. The updated weights 616 are generated with the goal of reducing the error metric 612, for example, using a gradient descent backpropagation process or another machine learning optimization process.
[0125] After updating the feature extractor 608 based on the updated weights 616, the training system 600 may perform one or more additional iterations until a termination condition is met. For example, a termination condition may be met when a fixed (e.g., predetermined) number of iterations has been performed. As another example, a termination condition may be met when the error metric 612 meets a specified threshold, or when the rate of change of the error metric 612 (e.g., from one iteration to the next) meets a specified threshold.
[0126] When the termination condition is met, the training of feature extractor 608 is complete. Feature extractor 608 can then be validated (e.g., using validation data similar to training data 602) and / or used to train a classifier, as referenced below. Figure 6B As described.
[0127] refer to Figure 6BThe training system 650 for training classifier 620 includes portions of training system 600, such as training data 602, preprocessor 112, feature extractor 608, and error calculator 610. Training system 650 also includes classifier 620, one or more error calculators 626, and optimizer 632. The portions of training system 600 present in training system 650 are as described in reference [reference needed]. Figure 6A The operation is as described, except that both target data 604 and non-target data 606 are used during the training of classifier 620, and feature extractor 608 has already been trained. Therefore, the parameters of feature extractor 608 do not change during the training of classifier 620 using training system 650.
[0128] During one iteration of the training system 650, a fragment of training data 602 or a set of fragments of training data 602 (e.g., a batch of consecutive fragments representing a time series) is provided to the preprocessor 112. The preprocessor 112 generates input (X) to the feature extractor 608, where input (X) corresponds to or includes fragment data representing fragments of training data 602, such as fragment data 114. The training data 602 provided to the preprocessor 112 may include one or more data samples selected from target data 604, one or more data samples selected from non-target data 606, or both. As an example, batch training can be used, where batches representing target data 604 alternate with batches representing non-target data 606, but other batch training schemes may also be used.
[0129] Feature extractor 608 processes the input (X) as described above to generate a latent space representation 212 based on the input (X) and an output (X') based on the latent space representation 212. In some embodiments, the output (X') and the input (X) are provided to error calculator 610 to generate one or more error measures 612. For example, error calculator 610 may include... Figure 3 The reconstruction error calculator 314. In some embodiments, the latent space representation 212 is provided to the error calculator 610 to generate one or more error metrics 612. For example, the error calculator 610 may include... Figure 3 The divergence calculator 312.
[0130] At least one input based on the latent space representation 212 is provided to the classifier 620. For example, the input based on the latent space representation 212 may include the latent space representation 212, divergence values from the error calculator 610, reconstruction error values from the error calculator 610, or combinations thereof. In some embodiments, one or more other inputs 622 may also be provided to the classifier 620. For example, other inputs 622 may include one or more pattern indicators, as described above.
[0131] Based on the input, classifier 620 generates a classification output 624, such as... Figure 2 Fragment classification 220 or Figure 4 The audio category is 422. Classification outputs 624 for specific segments of training data 602 and labels 628 associated with those segments are provided to an error calculator 626. The error calculator 626 generates an error metric 630 (typically for a batch or group of training data 602) indicating the difference between the classification output 624 and the corresponding label 628, and the optimizer 632 determines the updated weights 634 of the classifier 620 based on the error metric 630.
[0132] After updating the classifier 620 based on the updated weights 634, the training system 650 may perform one or more additional iterations until a termination condition is met. For example, the termination condition may be met when a fixed (e.g., predetermined) number of iterations has been performed. As another example, the termination condition may be met when the error metric 630 meets a specified threshold, or when the rate of change of the error metric 630 (e.g., from one iteration to the next) meets a specified threshold.
[0133] When the termination condition is met, the training of classifier 620 is complete, and feature extractor 608 and classifier 620 can be deployed for runtime use as part of the processing controller 140 of Figure 1 (e.g., in the inference phase).
[0134] For reference Figure 6A and Figure 6B One advantage of training the feature extractor 608 and classifier 620 in separate training operations, as described, is that the training data 602 does not need to include a well-balanced selection of target data 604 and non-target data 606. For example, the feature extractor 608 can be trained using only target data 604. Training the classifier 620 typically requires the use of some non-target data 606; however, when the classifier 620 is a uniclass classifier, the target data 604 and non-target data 606 may be imbalanced (i.e., the set of target data 604 may be much larger than the set of non-target data 606). When the classifier 620 is a binary classifier, the target data 604 and non-target data 606 may need to be more balanced than required for training a uniclass classifier; however, the non-target data 606 does not need to include samples of every possible or every expected type of non-target data.
[0135] Figure 7 Examples of training feature extractors and classifiers according to this disclosure are illustrated for... Figure 1A The system's various aspects of use. (Refer to...) Figure 6A and Figure 6B Compared to the described training scheme, the feature extractor 608 and the classifier 620 are together... Figure 7 The training system 700 is trained. Training system 700 includes many of the same features as training systems 600 and 650, and, in addition to those noted below, such features are as described in the references. Figure 6A and Figure 6B It operates as described. For example, training system 700 includes training data 602 (which includes target data 604 and non-target data 606), preprocessor 112, feature extractor 608, error calculator 610, classifier 620, and error calculator 626. Training system 700 also includes one or more optimizers 732, which may include... Figure 6A Optimizer 614 Figure 6B The optimizer 632 or the combined optimizer 614 and optimizer 632 are both optimizers in various aspects.
[0136] During each iteration of the training system 700, a set of fragments of training data 602 (e.g., a batch of consecutive fragments representing a time series) is provided to the preprocessor 112. Generally, the batch of training data 602 used in each iteration includes only the target data 604 or only the non-target data 606.
[0137] For each segment in the batch, the preprocessor 112 generates input (X) to the feature extractor 608, where the input (X) corresponds to or includes segment data representing segments of the training data 602, such as segment data 114. The feature extractor 608 processes the input (X) as described above to form a latent space representation and a synthesized segment (X') based on the latent space representation. The latent space representation, the synthesized segment (X'), or both are provided as output 710 from the feature extractor 608.
[0138] The output 710 of the feature extractor 608 is provided to the error calculator 610, the classifier 620, or both. For example, for training... Figure 2 The classifier 120 outputs 710, which includes the latent space representation from the feature extractor 608, and this output 710 is provided to the classifier 620. In this example, the error calculator 610 can be omitted. As another example, for training... Figure 3 The classifier 120 outputs a latent space representation, a synthetic fragment (X'), or both. In this example, the latent space representation can be provided along with information describing the expected distribution (e.g., expected distribution 310) to a divergence calculator (e.g., Figure 3 A divergence calculator 312 is used to determine the divergence measure (e.g., divergence value 322). Additionally or alternatively, the synthesized fragment (X') may be provided along with the input (X) to a reconstruction error calculator (e.g., Figure 3The reconstruction error calculator 314 determines the reconstruction error metric (e.g., reconstruction error value 324). The error metric 612 from the error calculator 610 may include a divergence metric, a reconstruction error metric, or both.
[0139] At least one input based on the latent space representation from feature extractor 608 is provided to classifier 620. For example, the latent space representation-based input may include the latent space representation, divergence values from error calculator 610, reconstruction error values from error calculator 610, or combinations thereof. In some embodiments, other inputs 622 may also be provided to classifier 620. For example, other inputs 622 may include one or more pattern indicators, as described above.
[0140] Based on the input, classifier 620 generates a classification output 624, such as... Figure 2 Fragment classification 220 or Figure 4 The audio category. The classification output 624 of a specific segment of the training data 602 and the label 628 associated with the specific segment are provided to the error calculator 626. The error calculator 626 generates an error metric 630 (typically for a batch or group of training data 602) indicating the difference between the classification output 624 and the corresponding label 628.
[0141] Error metric 630 and possibly other data (such as error metric 612) are provided to optimizer 732. Optimizer 732 determines updated weights 616 for feature extractor 608, updated weights 634 for classifier 620, or both. In a particular aspect, when the training data 602 used in a particular iteration includes target data 604 (and omits non-target data 606), the updated weights 616 include weights (W) based on the target data. T The target data-based weights are applied to both the inference network portion 704 and the generation network portion 708 of the feature extractor 608. In contrast, when the training data 602 used in a particular iteration includes non-target data 606 (and the target data 604 is omitted), the updated weights 616 include weights based on the non-target data (W). NT The weights based on non-target data are applied only to the inference network portion 704 of the feature extractor 608. In this way, the generative network portion 708 is updated only based on the target data 604.
[0142] After updating feature extractor 608 and classifier 620, training system 700 may perform one or more additional iterations until a termination condition is met. For example, a termination condition may be met when a fixed (e.g., predetermined) number of iterations has been performed. As another example, a termination condition may be met when one or more of error metrics 612 or 630 meet a specified threshold, or when the rate of change of one or more of error metrics 612 or 630 (e.g., from one iteration to the next) meets a specified threshold.
[0143] When the termination condition is met, the training of feature extractor 608 and classifier 620 is complete, and feature extractor 608 and classifier 620 can be deployed for runtime use as part of the processing controller 140 of FIG1 (e.g., in the inference phase).
[0144] For reference Figure 7 One benefit of training the feature extractor 608 and classifier 620 together, as described, is that it enables improved separation between the latent space representation of the target data 604 and the latent space representation of the non-target data 606, which improves the classification accuracy of the classifier 620. For example, when the feature extractor 608 and classifier 620 are trained together, the inference network portion 704 of the feature extractor 608 is updated based on both the target data 604 and the non-target data 606, thus making the inference network portion 704 suitable (e.g., trained to) generate latent space representations of both the target data 604 and the non-target data 606. The generation network portion 708 is not trained to accurately synthesize the non-target data 606 because a large reconstruction error for the non-target data 606 helps the classifier 620 identify the non-target data 606.
[0145] The above references Figures 6A to 7 The training operations described are merely illustrative. Variations and / or different training operations described above can be used to train machine learning models to serve as feature extractors 116, classifiers 120, or both.
[0146] Figure 8 Figure 1 is an illustrative diagram illustrating aspects of the operation of components of the system according to some examples of this disclosure. Figure 8In this process, feature extractor 116 is configured to receive a sequence of data samples corresponding to time series data 110. For example, time series data 110 may include a sequence of consecutively captured frames of audio data, exemplified as a first frame (F1) 812, a second frame (F2) 814, and one or more additional frames including an Nth frame (FN) 816 (where N is an integer greater than two). Feature extractor 116 is configured to output a sequence 820 of feature datasets, including a first set 822, a second set 824, and one or more other sets including an Nth set 826. Each feature dataset may include or correspond to a latent space representation (e.g., Figures 2 to 5 The latent space representation of either of them (212), or the error value based on the latent space representation (such as the divergence value 322, the reconstruction error value 324, or both).
[0147] Classifier 120 is configured to receive a sequence 820 of feature data (and optionally other data) and output a sequence of fragment classifications 220 based on the feature data in sequence 820. For example, the sequence of fragment classifications 220 may include a first fragment classification (C1) 832 indicating a classification assigned to a first frame 812, a second fragment classification (C2) 834 indicating a classification assigned to a second frame 812, and an Nth fragment classification (CN) 836 indicating a classification assigned to an Nth frame 812.
[0148] Figure 9 Depicting Figure 1A The device 102 is an embodiment 900 of an integrated circuit 902 including one or more processors 190. The integrated circuit 902 also includes signal inputs 904 (such as one or more bus interfaces) to enable the reception of time-series data 110 for processing. Figure 9 In this embodiment, processor 190 includes a processing controller 140, downstream processing components 142, or both. Integrated circuit 902 also includes signal outputs 906, such as a bus interface, to enable the transmission of output data 908, such as the processing control signals 122 of FIG1. Figure 2 The fragment classification 220, or the result of processing fragments of time series data 110 by a selected subset of downstream processing components 142.
[0149] Figure 10An embodiment 1000, in which device 102 includes mobile device 1002 such as a telephone or tablet computer as an illustrative, non-limiting example, is depicted. Mobile device 1002 includes microphone 1004, camera 1006, and display screen 1008. Microphone 1004, camera 1006, or both may be configured to generate time-series data (e.g., time-series data 110 of FIG. 1). Additionally or alternatively, mobile device 1002 may include other sensors (e.g., inertial measurement unit, pressure sensor, temperature sensor, etc.) configured to generate time-series data. Components of processor 190 (including processing controller 140 and downstream processing components 142) are integrated into mobile device 1002 and are illustrated using dashed lines to indicate internal components that are not typically visible to the user of mobile device 1002.
[0150] In a particular example, processing controller 140 is configured to determine whether each segment of time-series data includes target data (e.g., data of a target data type) and generate a processing control signal based on that determination. Downstream processing component 142 is configured to perform specific operations on segments of time-series data in response to the processing control signal. For example, downstream processing component 142 may perform one or more operations in a first set for each segment including target data and one or more operations in a second set for each segment not including target data. For illustration, the target data type may include speech; in this case, downstream processing component 142 may use the first set of operations to process segments including speech and the second set of operations to process segments not including speech. In this illustrative example, the first set of operations may be more suitable than the second set of operations for processing speech segments (based on specific design objectives). As a non-limiting example, the first set of operations may include a speech encoder, and the second set of operations may include a general-purpose encoder. In this example, the speech encoder may be better than the general-purpose encoder at encoding speech (in terms of efficiency, fidelity, or some other metric).
[0151] Figure 11An embodiment 1100 in which device 102 includes a head-mounted device 1102 is depicted. The head-mounted device 1102 includes a microphone 1104 and a speaker 1106. The microphone 1104 may be configured to generate time-series data (e.g., time-series data 110 of FIG. 1). Additionally or alternatively, the head-mounted device 1102 may include other sensors (e.g., a camera, inertial measurement unit, pressure sensor, temperature sensor, etc.) configured to generate time-series data. Components of processor 190 (including processing controller 140 and downstream processing component 142) are integrated into the head-mounted device 1102. In a particular example, processing controller 140 is configured to determine whether each segment of the time-series data includes target data (e.g., data of a target data type) and generate a processing control signal based on that determination. Downstream processing component 142 is configured to perform a specific operation on the segment of time-series data in response to the processing control signal. For example, downstream processing component 142 may perform one or more operations in a first set for each segment that includes target data, and one or more operations in a second set for each segment that does not include target data. For illustration, the target data type may include speech; in this case, downstream processing component 142 may use the first set of operations to process segments that include speech, and may use the second set of operations to process segments that do not include speech.
[0152] Figure 12 An embodiment 1200 in which device 102 includes wearable electronic device 1202 (illustrated as a "smartwatch") is depicted. Wearable electronic device 1202 includes microphone 1204, sensor 1206, and display 1208. Microphone 1204, sensor 1206, or both may be configured to generate time-series data (e.g., time-series data 110 of FIG. 1). For example, sensor 1206 may include a camera, one or more physiological sensors (e.g., heart rate monitor, blood oxygen sensor), motion sensor, or one or more other types of sensors configured to generate time-series data. Components of processor 190, including processing controller 140 and downstream processing component 142, are integrated into wearable electronic device 1202. In a particular example, processing controller 140 is configured to determine whether each segment of time-series data includes target data (e.g., data of a target data type) and generate a processing control signal based on that determination. Downstream processing component 142 is configured to perform a specific operation on the segment of time-series data in response to the processing control signal. For example, when sensor 1206 includes one or more physiological sensors, downstream processing component 142 may perform one or more operations in a first set for each segment that includes target physiological data (e.g., target heart rate), and perform one or more operations in a second set for each segment that does not include target data.
[0153] Figure 13 An embodiment 1300 is depicted in which device 102 includes a portable electronic device corresponding to augmented reality glasses or mixed reality glasses 1302. Glasses 1302 include a holographic projection unit 1308 configured to project visual data onto the surface of lens 1310, or to reflect the visual data from the surface of lens 1310 onto the wearer's retina. Glasses 1302 also includes a microphone 1306 and a camera 1304. Microphone 1306, camera 1304, or both may be configured to generate time-series data (e.g., time-series data 110 of FIG. 1). Additionally or alternatively, glasses 1302 may include other sensors (e.g., inertial measurement unit, pressure sensor, temperature sensor, etc.) configured to generate time-series data.
[0154] Components of processor 190, including processing controller 140 and downstream processing component 142, are integrated into glasses 1302. In a particular example, processing controller 140 is configured to determine whether each segment of time-series data includes target data (e.g., data of a target data type) and generate processing control signals based on that determination. Downstream processing component 142 is configured to perform specific operations on segments of time-series data in response to the processing control signals. For example, camera 1304 may generate a series of images corresponding to the time-series data. In this example, when a particular image includes a face, downstream processing component 142 may perform operations to identify the face and transmit the output to holographic projection unit 1308 so that the name is displayed to the user of glasses 1302.
[0155] Figure 14 An embodiment 1400 is depicted in which device 102 includes a portable electronic device corresponding to a pair of earplugs 1410, the pair of earplugs including a first earplug 1402A and a second earplug 1402B. Although earplugs 1410 are described, it should be understood that the technology disclosed herein can be applied to other in-ear or over-ear playback devices.
[0156] The first earpiece 1402A includes a first microphone 1404, such as a high signal-to-noise ratio microphone positioned to capture the speech of the wearer of the first earpiece 1402A; an array of one or more other microphones configured to detect ambient sound and spatially distributed to support beamforming, exemplified as microphones 1422A, 1422B, and 1422C; an "inner" microphone 1424 located near the wearer's ear canal (e.g., to assist active noise cancellation); and a self-speech microphone 1426, such as a bone conduction microphone configured to convert sound vibrations from the wearer's ear bones or skull into audio signals. The second earpiece 1402B may be configured in a substantially similar manner to the first earpiece 1402A.
[0157] exist Figure 14 In this embodiment, the first earpiece 1402A includes components of a processor 190, including a processing controller 140 and a downstream processing component 142. The processing controller 140 is configured to process time-series data to generate processing control signals for the downstream processing component 142. The time-series data may correspond to or include data captured by any combination of microphones 1404, 1422A, 1422B, 1422C, 1424, and 1426 or other sensors of the earpiece 1410. The downstream processing component 142 is configured to perform specific operations on segments of the time-series data in response to the processing control signals. For example, if the target data type may include speech, in which case the downstream processing component 142 may use a first set of operations to process segments including speech and a second set of operations to process segments not including speech.
[0158] In some embodiments, earbuds 1402A and 1402B are configured to automatically switch between various operating modes, such as a pass-through mode in which ambient sounds are played via speaker 1406; a playback mode in which non-ambient sounds (e.g., streaming audio corresponding to telephone conversations, media playback, video games, etc.) are played back via speaker 1406; and an audio zoom mode or beamforming mode in which one or more ambient sounds are amplified and / or other ambient sounds are suppressed for playback at speaker 1406. In such embodiments, downstream processing component 142 may switch between these modes based on processing control signals from processing controller 140. For example, when speech is detected in the sound captured by one or more of the microphones 1404, 1422A, 1422B, 1422C, 1424, and 1426, the downstream processing component 142 can activate a first mode, and when no speech is detected in the sound captured by the microphones 1404, 1422A, 1422B, 1422C, 1424, and 1426, the downstream processing component 142 can activate a second mode.
[0159] Figure 15This is an embodiment 1500 in which device 102 includes a wireless speaker and a voice-activated device 1502. The wireless speaker and voice-activated device 1502 may have wireless network connectivity and is configured to perform auxiliary operations. The wireless speaker and voice-activated device 1502 includes a microphone 1504 and a speaker 1506. The microphone 1504 may be configured to generate time-series data (e.g., time-series data 110 of FIG. 1). Additionally or alternatively, the wireless speaker and voice-activated device 1502 may include other sensors (e.g., a camera, pressure sensor, temperature sensor, etc.) configured to generate time-series data. Components of processor 190 (including processing controller 140 and downstream processing component 142) are integrated into the wireless speaker and voice-activated device 1502.
[0160] In a specific example, processing controller 140 is configured to determine whether each segment of time-series data includes target data (e.g., data of a target data type) and generate a processing control signal based on that determination. Downstream processing component 142 is configured to perform specific operations on segments of time-series data in response to the processing control signal. For example, if the target data type may include speech, in which case downstream processing component 142 may use a first set of operations to process segments including speech and a second set of operations to process segments not including speech. For illustration, a keyword detector associated with voice-assisted operations may be activated only in response to the detection of speech in the audio captured by microphone 1504. For illustration, when processing controller 140 does not detect speech in the audio data from microphone 1504, the processing control signal from processing controller 140 may cause downstream processing component 142 to deactivate voice-assisted operations (including the keyword detector). However, when processing controller 140 detects speech in the audio data from microphone 1504, the processing control signal from processing controller 140 causes downstream processing component 142 to activate the keyword detector for voice-assisted operations.
[0161] Figure 16An embodiment 1600 is depicted in which device 102 includes a portable electronic device corresponding to camera device 1602. Camera device 1602 includes microphone 1104 and image sensor 1606. Microphone 1604, image sensor 1606, or both may be configured to generate time-series data (e.g., time-series data 110 of FIG. 1). Additionally or alternatively, camera device 1602 may include other sensors (e.g., inertial measurement unit, pressure sensor, temperature sensor, etc.) configured to generate time-series data. Components of processor 190 (including processing controller 140 and downstream processing component 142) are integrated into camera device 1602. In a particular example, processing controller 140 is configured to determine whether each segment of time-series data includes target data (e.g., data of a target data type) and generate a processing control signal based on that determination. Downstream processing component 142 is configured to perform a specific operation on the segment of time-series data in response to the processing control signal. For example, downstream processing component 142 may perform one or more operations in a first group on each segment that includes target data, and perform one or more operations in a second group on each segment that does not include target data.
[0162] Figure 17 An embodiment 1700 is depicted in which device 102 includes a portable electronic device corresponding to an extended reality (XR) headset 1702 (such as a virtual reality headset, a mixed reality headset, or an augmented reality headset). The XR headset 1702 includes a microphone 1704. The microphone 1704 can be configured to generate time-series data (e.g., time-series data 110 of FIG. 1). Additionally or alternatively, the XR headset 1702 may include other sensors (e.g., a camera, an inertial measurement unit, a pressure sensor, a temperature sensor, etc.) configured to generate time-series data. Components of processor 190 (including a processing controller 140 and a downstream processing component 142) are integrated into the XR headset 1702. In a particular example, the processing controller 140 is configured to determine whether each segment of the time-series data includes target data (e.g., data of a target data type) and generate a processing control signal based on that determination. The downstream processing component 142 is configured to perform a specific operation on the segment of time-series data in response to the processing control signal. For example, downstream processing component 142 may perform one or more operations in a first set for each segment that includes target data, and one or more operations in a second set for each segment that does not include target data. For illustration, the target data type may include speech; in this case, downstream processing component 142 may use the first set of operations to process segments that include speech, and may use the second set of operations to process segments that do not include speech.
[0163] Figure 18An embodiment 1800 is depicted in which device 102 corresponds to vehicle 1802 (exemplified as a manned or unmanned aerial device (e.g., a package delivery drone)) or is integrated within the vehicle. Vehicle 1802 includes microphone 1804 and camera 1806. Microphone 1804, camera 1806, or both may be configured to generate time-series data (e.g., time-series data 110 of FIG. 1). Additionally or alternatively, vehicle 1802 may include other sensors (e.g., inertial measurement unit, pressure sensor, temperature sensor, etc.) configured to generate time-series data. Components of processor 190 (including processing controller 140 and downstream processing component 142) are integrated into vehicle 1802. In a particular example, processing controller 140 is configured to determine whether each segment of time-series data includes target data (e.g., data of a target data type) and generate processing control signals based on that determination. Downstream processing component 142 is configured to perform specific operations on segments of time-series data in response to the processing control signals. For example, downstream processing component 142 may perform one or more operations in a first set for each segment that includes target data, and one or more operations in a second set for each segment that does not include target data. For illustration, user voice activity detection may be performed by downstream processing component 142 based on whether the image from camera 1806 includes a face, whether the sound captured by microphone 1804 includes speech, or both, determined by processing controller 140.
[0164] Figure 19 The diagram depicts a device 102 corresponding to a vehicle 1902 (illustrated as an automobile) or another embodiment 1900 integrated within such a vehicle. The vehicle 1902 includes a processing controller 140 and a downstream processing component 142. The vehicle 1902 also includes one or more microphones 1904, one or more sensors 1906 (e.g., cameras, LiDAR systems, internal measurement units, etc.) or combinations thereof configured to generate time-series data (e.g., time-series data 110 of FIG. 1). The processing controller 140 is configured to determine whether each segment of the time-series data includes target data and to generate a processing control signal based on that determination. The downstream processing component 142 is configured to perform specific operations on segments of the time-series data in response to the processing control signal. For example, the downstream processing component 142 may perform a first set of one or more operations on each segment including target data and a second set of one or more operations on each segment not including target data. For illustration, the target data type may include speech, in which case the downstream processing component 142 may use a first set of operations to process segments that include speech, and may use a second set of operations to process segments that do not include speech.
[0165] refer to Figure 20 This illustrates a specific embodiment of a method 2000 for selectively processing time-series data. In a particular aspect, one or more operations of method 2000 are performed by at least one of the processing controller 140, downstream processing component 142, processor 190, device 102, system 100, or combinations thereof of FIG1.
[0166] Method 2000 includes, at block 2002, using a feature extractor to generate a latent space representation of a segment of time series data. For example, processor 190 may execute instructions 194 to perform operations described by reference processing controller 140, such as generating segment data 114 based on time series data 110, and providing segment data 114 as input to feature extractor 116 to generate a latent space representation of segment data 114 (such as...). Figures 2 to 5 The latent space representation of any one of them (212).
[0167] In some implementations, the feature extractor includes an autoencoder comprising an inference network portion, a bottleneck layer, and a generative network portion. The inference network portion, the generative network portion, or both may include one or more recurrent layers, one or more dilated convolutional layers, one or more temporally dynamic layers, or any combination thereof. In some implementations, the autoencoder is a variational autoencoder, in which case the latent space representation includes the mean and standard deviation of the probability distribution.
[0168] An autoencoder for feature extractors can be trained to reproduce data segments from a target signal category. For example, the target signal category may include audio data with speech, in which case the autoencoder is trained to reproduce the speech data. In embodiments where time-series data represents audio data, segments of time-series data may include audio frames, and segment data may include a spectral representation of one or more audio data samples of the audio frames.
[0169] Method 2000 includes, at box 2004, providing one or more inputs to a classifier. The classifier may include, for example, a uniclass classifier or a binary classifier, and the output of the classifier indicates whether a segment is assigned to the target signal category.
[0170] One or more inputs to the classifier include at least one input based on the latent space representation from the feature extractor. For example, at least one of the one or more inputs provided to the classifier may include a latent space representation. As another example, at least one of the one or more inputs to the classifier may include an error value. For illustration, in some embodiments, the feature extractor includes a generative network portion, and inputs based on the latent space representation are provided to the generative network portion to generate synthetic segments of time-series data. In such embodiments, the error value may be determined based on a comparison of the segments and the synthetic segments. The error value may be determined, for example, based on the Itakura-Saito distance between the segments and the synthetic segments, the scale-invariant signal distortion ratio between the segments and the synthetic segments, the logarithmic spectral distortion value between the segments and the synthetic segments, etc. In embodiments where the feature extractor includes a VAE, the latent space representation defines a probability distribution that is normalized to a desired probability distribution during training. In such embodiments, the output of the generative network portion may include or be mapped to the probability distribution, and the error value may be determined based on the desired probability distribution and the probability distribution of the output.
[0171] In some implementations, one or more inputs to the classifier may also include pattern indicators associated with the segment. In such implementations, the pattern indicator may indicate, for example, whether the segment represents voiced speech. As another example, the pattern indicator may indicate whether the segment represents music. The pattern indicator can be determined using relatively lightweight operations, such as performing statistical analysis of time-series data and generating the pattern indicator based on the results of the statistical analysis. As another example, the pattern indicator may include segment-based enhanced variable rate codec (EVRC) pattern decisions.
[0172] In some implementations, one or more inputs to the classifier may also include one or more features associated with the segment, such as open-loop pitch, normalized correlation, spectral envelope, pitch stability, signal nonstationarity, linear prediction residual, spectral difference, spectral stationarity, or combinations thereof.
[0173] Method 2000 includes, at block 2006, generating a processing control signal for a segment based on the output of a classifier. The processing control signal indicates which of a set of available operations to perform on the segment of time-series data. For example, when the time-series data represents audio content, the classifier's output may indicate whether the segment includes an audio data type associated with a first audio encoder. In this example, the processing control signal may cause the segment to be selectively routed to one of two or more available audio encoders, such as routing to the first audio encoder if the segment includes an audio data type associated with the first audio encoder, or routing to a second audio encoder if the segment does not include an audio data type associated with the first audio encoder.
[0174] Therefore, based on the processing control signal, either a first encoding operation or a second encoding operation is performed to encode the segment. In this example, compared to the second encoding operation, the first encoding operation can provide higher quality encoding of the target signal category, more efficient encoding of the target signal category, or both.
[0175] Figure 20 Method 2000 can be implemented by a field-programmable gate array (FPGA) device, an application-specific integrated circuit (ASIC), a processing unit (such as a central processing unit (CPU)), a DSP, a controller, another hardware device, a firmware device, or any combination thereof. As an example, Figure 20 Method 2000 can be executed by a processor that executes instructions, such as reference Figure 21 As described.
[0176] refer to Figure 21 A block diagram illustrating a specific exemplary embodiment of the device is provided, and the device is generally designated as 2100. In various embodiments, device 2100 may have more than Figure 21 More or fewer components as illustrated. In an exemplary embodiment, device 2100 may correspond to device 102. In an exemplary embodiment, device 2100 may perform reference... Figures 1A to 20 One or more operations as described.
[0177] In a particular embodiment, device 2100 includes a processor 2106 (e.g., a central processing unit (CPU)). Device 2100 may include one or more additional processors 2110 (e.g., one or more DSPs). In a particular aspect, Figure 1A The processor 190 corresponds to processor 2106, processor 2110, or a combination thereof. Processor 2110 may include a voice and music decoder-decoder (codec) 2108, which includes a voice decoder (“vocoder”) encoder 2136, a vocoder decoder 2138, a processing controller 140, a downstream processing component 142, or a combination thereof.
[0178] Device 2100 may include memory 192 and codec 2134. Memory 192 may include instructions 152 that can be executed by one or more additional processors 2110 (or processor 2106) to implement the functionality described by reference processing controller 140, downstream processing component 142, or both. Memory 192 may also store segments 196 of the time-series data 110 of FIG1. Device 2100 may include a modem 170 coupled to antenna 2152 via transceiver 2150.
[0179] Device 2100 may include a display 2128 coupled to display controller 2126. Microphone 2190, speaker 2192, one or more sensors 2194, or a combination thereof may be coupled to codec 2134. Codec 2134 may include digital-to-analog converter (DAC) 2102, analog-to-digital converter (ADC) 2104, or both. In a particular embodiment, codec 2134 may receive analog signals from microphone 2190 and sensor 2194, or both, convert the analog signals to digital signals using ADC 2104, and provide the digital signals to processor 2110, processor 2106, or both (such as to voice and music codec 2108). The digital signals may be processed by processing controller 140 and at least a subset of downstream processing components 142. Voice and music codec 2108 may provide the digital signals to codec 2134. Codec 2134 may use ADC 2102 to convert the digital signals to analog signals and may provide the analog signals to speaker 2192.
[0180] In a particular embodiment, device 2100 may be included in a system-in-package or system-on-a-chip device 2122. In a particular embodiment, memory 150, processor 2106, processor 2110, display controller 2126, codec 2134, and modem 170 are included in the system-in-package or system-on-a-chip device 2122. In a particular embodiment, input device 2130 and power supply 2144 are coupled to the system-in-package or system-on-a-chip device 2122. Furthermore, in a particular embodiment, such as Figure 21 As illustrated, display 2128, input device 2130, speaker 2192, microphone 2190, sensor 2194, antenna 2152, and power supply 2144 are external to system-in-package or system-on-chip device 2122. In certain embodiments, each of display 2128, input device 2130, speaker 2192, microphone 2190, sensor 2194, antenna 2152, and power supply 2144 may be coupled to components of system-in-package or system-on-chip device 2122, such as an interface (e.g., input interface 108) or a controller.
[0181] Device 2100 may include smart speakers, soundbars, mobile communication devices, smartphones, cellular phones, laptops, computers, tablets, personal digital assistants, display devices, televisions, game consoles, music players, radios, digital video players, digital video disc (DVD) players, tuners, cameras, navigation devices, vehicles, head-mounted devices, augmented reality head-mounted devices, mixed reality head-mounted devices, virtual reality head-mounted devices, aircraft, home automation systems, voice-activated devices, wireless speakers and voice-activated devices, portable electronic devices, automobiles, computing devices, communication devices, Internet of Things (IoT) devices, extended reality (XR) devices, base stations, mobile devices, or any combination thereof.
[0182] In conjunction with the described embodiments, an apparatus includes components for generating latent space representations of segments of time-series data. For example, the components for generating latent space representations of segments of time-series data may correspond to system 100, device 102, processor 190, processing controller 140, feature extractor 116, autoencoder 200, inference network portion 204 and bottleneck layer 206, feature extractor 608, integrated circuit 902, one or more other circuits or components configured to generate latent space representations of segments of time-series data, or any combination thereof.
[0183] The device also includes components for providing one or more inputs to a classifier, wherein the one or more inputs include at least one input based on a latent space representation. For example, the components for providing one or more inputs to a classifier may correspond to system 100, device 102, processor 190, processing controller 140, feature extractor 116, preprocessor 112, autoencoder 200, inference network portion 204 and bottleneck layer 206, pattern detector 304, divergence calculator 312, reconstruction error calculator 314, audio pattern detector 404, feature extractor 608, error calculator 610, error calculator 626, integrated circuit 902, one or more other circuits or components configured to provide one or more inputs to a classifier, or any combination thereof.
[0184] The device also includes components for processing control signals to generate segments. For example, the components for processing control signals to generate segments may correspond to system 100, device 102, processor 190, processing controller 140, classifier 120, audio classifier 420, integrated circuit 902, one or more other circuits or components configured to generate processing control signals for segments, or any combination thereof.
[0185] In some embodiments, a non-transitory computer-readable medium (e.g., a computer-readable storage device, such as memory 150) includes instructions (e.g., instruction 152) that, when executed by one or more processors (e.g., one or more processors 2110, 2106, or 190), cause the one or more processors to use a feature extractor to generate a latent space representation of a segment of time-series data; provide one or more inputs to a classifier, the one or more inputs including at least one input based on the latent space representation; and processing control signals for generating segments based on the output of the classifier.
[0186] Specific aspects of this disclosure are described below in a collection of related embodiments:
[0187] According to Embodiment 1, an apparatus includes a memory configured to store one or more segments of time series data; one or more processors configured to use a feature extractor to generate a latent space representation of the segments of the time series data; provide one or more inputs to a classifier, the one or more inputs including at least one input based on the latent space representation; and processing control signals for generating the segments based on the output of the classifier.
[0188] Example 2 includes the device according to Example 1, wherein the classifier is a single-class classifier or a binary classifier, and the output indicates whether the segment is assigned to a target signal category.
[0189] Example 3 includes the device according to Example 1 or Example 2, wherein one or more processors are configured to selectively perform a first encoding operation to encode the segment or perform a second encoding operation to encode the segment based on the processing control signal, wherein the first encoding operation provides higher quality encoding of the target signal category than the second encoding operation.
[0190] Example 4 includes a device according to any one of Examples 1 to 3, wherein one or more processors are configured to selectively perform a first encoding operation to encode the segment or perform a second encoding operation to encode the segment based on the processing control signal, wherein the first encoding operation provides more efficient encoding of the target signal category than the second encoding operation.
[0191] Example 5 includes a device according to any one of Examples 1 to 4, wherein the time-series data represents audio content, and wherein the output indicates whether the segment includes an audio data type associated with a first audio encoder.
[0192] Example 6 includes a device according to any one of Examples 1 to 5, wherein one or more processors are configured to selectively route the segment to one of two or more audio decoders based on the processing control signal.
[0193] Example 7 includes a device according to any one of Examples 1 to 6, wherein the segment corresponds to a segment of audio data, and the input to the feature extractor includes a frequency domain representation of the segment of audio data.
[0194] Example 8 includes the device according to Example 7, wherein the frequency domain represents the power spectrum of the segment comprising audio data.
[0195] Example 9 includes the device according to any one of Examples 1 to 8, wherein the feature extractor includes an inference network portion and a generative network portion of an autoencoder.
[0196] Example 10 includes the device according to Example 9, wherein the autoencoder includes one or more recursive layers.
[0197] Example 11 includes the device according to Example 9 or Example 10, wherein the autoencoder includes one or more dilated convolutional layers.
[0198] Example 12 includes a device according to any one of Examples 9 to 11, wherein the automatic encoder includes one or more time-dynamic layers.
[0199] Example 13 includes a device according to any one of Examples 9 to 12, wherein the autoencoder is a variational autoencoder, and the latent space representation includes the mean and standard deviation of a probability distribution.
[0200] Example 14 includes a device according to any one of Examples 9 to 13, wherein the autoencoder is trained to reproduce data segments from a target signal category, and the classifier is configured to distinguish between data segments from the target signal category and data segments not from the target signal category based on the separation of latent space representations between the data segments from the target signal category and data segments not from the target signal category.
[0201] Example 15 includes the device according to any one of Examples 9 to 14, wherein the autoencoder is trained to reproduce speech data, and the classifier is configured to distinguish audio data segments that include speech from audio data segments that do not include speech.
[0202] Example 16 includes a device according to any one of Examples 9 to 15, wherein the one or more processors are configured to provide input to the generative network portion based on the latent space representation to generate a synthesized fragment of time series data; and to determine a reconstruction error value based on a comparison of the fragment with the synthesized fragment, wherein at least one of the one or more inputs provided to the classifier is based on the reconstruction error value.
[0203] Example 17 includes the device according to Example 16, wherein determining the error value includes calculating the Itakura-Saito distance based on the fragment and the synthesized fragment.
[0204] Example 18 includes the apparatus according to Example 16, wherein determining the error value includes calculating a scale-invariant signal distortion ratio based on the segment and the synthesized segment.
[0205] Example 19 includes the device according to Example 16, wherein determining the error value includes calculating a logarithmic spectral distortion value based on the segment and the synthesized segment.
[0206] Example 20 includes a device according to any one of Examples 9 to 19, wherein the one or more processors are configured to provide input to the generative network portion based on the latent space representation to generate a probability distribution; and to determine an error value based on the probability distribution, wherein at least one of the one or more inputs provided to the classifier is based on the error value.
[0207] Example 21 includes the device according to any one of Examples 2 to 20, wherein at least one of the one or more inputs provided to the classifier includes the latent space representation.
[0208] Example 22 includes a device according to any one of Examples 1 to 21, wherein the one or more processors are configured to determine a divergence metric indicating the difference between a probability distribution representing the latent space and a expected probability distribution, and wherein at least one of the one or more inputs provided to the classifier is based on the divergence metric.
[0209] Example 23 includes a device according to any one of Examples 1 to 22, wherein the one or more processors are configured to determine a pattern indicator associated with the segment, and wherein at least one of the one or more inputs provided to the classifier is based on the pattern indicator.
[0210] Example 24 includes the device according to Example 23, wherein the mode indicator indicates whether the segment represents voiced speech.
[0211] Example 25 includes the device according to Example 23 or Example 24, wherein the mode indicator indicates whether the segment represents music.
[0212] Example 26 includes a device according to any one of Examples 23 to 25, wherein the pattern indicator is based on statistical analysis of the time series data.
[0213] Example 27 includes a device according to any one of Examples 1 to 26, wherein the one or more processors are configured to determine one or more features associated with the segment, and wherein at least one of the one or more inputs provided to the classifier is based on the one or more features.
[0214] Example 28 includes the device according to Example 27, wherein one or more features include one or more of open-loop pitch, normalized correlation, spectral envelope, pitch stability, signal nonstationarity, linear prediction residual, spectral difference, and spectral stationarity.
[0215] Example 29 includes a device according to any one of Examples 1 to 28, wherein the one or more processors are configured to generate an Enhanced Variable Rate Codec (EVRC) mode decision based on the segment, and wherein at least one of the one or more inputs provided to the classifier is based on the EVRC mode decision.
[0216] Example 30 includes the device according to any one of Examples 1 to 29, wherein the feature extractor includes a first machine learning model trained on training data to synthesize time series data that approximates the input time series data.
[0217] Example 31 includes the device according to Example 30, wherein each segment of the training data includes speech.
[0218] Example 32 includes the device according to Example 30, wherein the training data includes segments representing speech and segments representing non-speech sounds.
[0219] Example 33 includes the device according to Example 30, wherein each segment of the training data includes data assigned to a target signal category.
[0220] Example 34 includes the device according to Example 30, wherein the training data includes data assigned to a target signal category and data not assigned to the target signal category.
[0221] Example 35 includes the device according to any one of Examples 30 to 34, wherein the classifier is trained using labeled training data based on the output of the feature extractor.
[0222] Example 36 includes the device according to any one of Examples 30 to 35, wherein the classifier includes a second machine learning model, and wherein the first machine learning model and the second machine learning model are jointly trained.
[0223] According to embodiment 37, a method includes using a feature extractor to generate a latent space representation of a segment of time series data; providing one or more inputs to a classifier, the one or more inputs including at least one input based on the latent space representation; and generating a processing control signal for the segment based on the output of the classifier.
[0224] Example 38 includes the method according to Example 37, wherein the classifier is a single-class classifier or a binary classifier, and the output indicates whether the segment is assigned to a target signal category.
[0225] Example 39 includes the method according to Example 37 or Example 38, and the method further includes selectively performing a first encoding operation to encode the segment or performing a second encoding operation to encode the segment based on the processing control signal, wherein the first encoding operation provides a higher quality encoding of the target signal category than the second encoding operation.
[0226] Example 40 includes the method according to any one of Examples 37 to 39, and the method further includes selectively performing a first encoding operation to encode the segment or performing a second encoding operation to encode the segment based on the processing control signal, wherein the first encoding operation provides more efficient encoding of the target signal category than the second encoding operation.
[0227] Example 41 includes the method according to any one of Examples 37 to 40, wherein the time-series data represents audio content, and wherein the output indicates whether the segment includes an audio data type associated with a first audio encoder.
[0228] Example 42 includes the method according to any one of Examples 37 to 41, and the method further includes selectively routing the segment to an audio encoder based on the processing control signal.
[0229] Example 43 includes the method according to any one of Examples 37 to 42, wherein the segment corresponds to an audio frame comprising a spectral representation of one or more audio data samples.
[0230] Example 44 includes the method according to any one of Examples 37 to 43, wherein the feature extractor includes an inference network portion and a generative network portion of an autoencoder.
[0231] Example 45 includes the method according to Example 44, wherein the autoencoder includes one or more recursive layers.
[0232] Example 46 includes the method according to Example 44 or Example 45, wherein the autoencoder includes one or more dilated convolutional layers.
[0233] Example 47 includes the method according to any one of Examples 44 to 46, wherein the autoencoder includes one or more time-dynamic layers.
[0234] Example 48 includes the method according to any one of Examples 44 to 47, wherein the autoencoder is a variational autoencoder, and the latent space representation includes the mean and standard deviation of the probability distribution.
[0235] Example 49 includes a method according to any one of Examples 44 to 48, wherein the autoencoder is trained to reproduce data segments from a target signal category, and the classifier is configured to distinguish between data segments from the target signal category and data segments not from the target signal category based on the separation of latent space representations between the data segments from the target signal category and data segments not from the target signal category.
[0236] Example 50 includes the method according to any one of Examples 44 to 49, wherein the autoencoder is trained to reproduce speech data, and the classifier is configured to distinguish audio data segments that include speech from audio data segments that do not include speech.
[0237] Example 51 includes the method according to any one of Examples 44 to 50, and the method further includes providing input to the generative network portion based on the latent space representation to generate a synthesized segment of time series data; and determining an error value based on a comparison of the segment with the synthesized segment, wherein at least one of the one or more inputs provided to the classifier is based on the error value.
[0238] Example 52 includes the method according to Example 51, wherein determining the error value includes calculating the Itakura-Saito distance based on the fragment and the synthesized fragment.
[0239] Example 53 includes the method according to Example 51, wherein determining the error value includes calculating the scale-invariant signal distortion ratio based on the segment and the synthesized segment.
[0240] Example 54 includes the method according to Example 51, wherein determining the error value includes calculating a logarithmic spectral distortion value based on the fragment and the synthesized fragment.
[0241] Example 55 includes the method according to any one of Examples 44 to 54, and the method further includes providing input to the generative network portion based on the latent space representation to generate a probability distribution; and determining an error value based on the segment and the probability distribution, wherein at least one of the one or more inputs provided to the classifier is based on the error value.
[0242] Example 56 includes the method according to any one of Examples 37 to 55, wherein at least one of the one or more inputs provided to the classifier includes the latent space representation.
[0243] Example 57 includes the method according to any one of Examples 37 to 56, and the method further includes determining a divergence measure indicating the difference between the probability distribution of the latent space representation and the expected probability distribution, and wherein at least one of the one or more inputs provided to the classifier is based on the divergence measure.
[0244] Example 58 includes the method according to any one of Examples 37 to 57, and the method further includes determining a pattern indicator associated with the segment, wherein at least one of the one or more inputs provided to the classifier is based on the pattern indicator.
[0245] Example 59 includes the method according to Example 58, wherein the mode indicator indicates whether the segment represents voiced speech.
[0246] Example 60 includes the method according to Example 58 or Example 59, wherein the mode indicator indicates whether the segment represents music.
[0247] Example 61 includes the method according to any one of Examples 58 to 60, and the method further includes performing statistical analysis on the time series data and generating the pattern indicator based on the statistical analysis.
[0248] Example 62 includes the method according to any one of Examples 37 to 61, and the method further includes determining one or more features associated with the segment, and wherein at least one of the one or more inputs provided to the classifier is based on the one or more features.
[0249] Example 63 includes the method according to Example 62, wherein one or more features include one or more of open-loop pitch, normalized correlation, spectral envelope, pitch stability, signal nonstationarity, linear prediction residual, spectral difference, and spectral stationarity.
[0250] Example 64 includes the method according to any one of Examples 37 to 63, and the method further includes generating an Enhanced Variable Rate Codec (EVRC) mode decision based on the segment, wherein at least one of the one or more inputs provided to the classifier is based on the EVRC mode decision.
[0251] Example 65 includes the method according to any one of Examples 37 to 64, wherein the feature extractor includes a first machine learning model trained on training data to synthesize time series data that approximates the input time series data.
[0252] Example 66 includes the method according to Example 65, wherein each segment of the training data includes speech.
[0253] Example 67 includes the method according to Example 65, wherein the training data includes segments representing speech and segments representing non-speech sounds.
[0254] Example 68 includes the method according to Example 65, wherein each segment of the training data includes data assigned to a target signal category.
[0255] Example 69 includes the method according to Example 65, wherein the training data includes segments assigned to: data assigned to a target signal category and data not assigned to the target signal category.
[0256] Example 70 includes the method according to Example 65, wherein the classifier is trained using labeled training data based on the output of the feature extractor.
[0257] Example 71 includes the method according to Example 65, wherein the classifier includes a second machine learning model, and wherein the first machine learning model and the second machine learning model are jointly trained.
[0258] According to Embodiment 72, a non-transitory computer-readable medium storage instruction is provided, the instruction being executable by one or more processors to cause the one or more processors to use a feature extractor to generate a latent space representation of a fragment of time series data; to provide one or more inputs to a classifier, the one or more inputs including at least one input based on the latent space representation; and a processing control signal to generate the fragment based on the output of the classifier.
[0259] Example 73 includes a non-transitory computer-readable medium according to Example 72, wherein the classifier is a single-class classifier or a binary classifier, and the output indicates whether the segment is assigned to a target signal category.
[0260] Example 74 includes a non-transitory computer-readable medium according to Example 72 or Example 73, wherein the instructions further cause the one or more processors to selectively perform a first encoding operation to encode the segment or perform a second encoding operation to encode the segment based on the processing control signal, wherein the first encoding operation provides a higher quality encoding of the target signal category than the second encoding operation.
[0261] Example 75 includes a non-transitory computer-readable medium according to any one of Examples 72 to 74, wherein the instructions further cause the one or more processors to selectively perform a first encoding operation to encode the segment or perform a second encoding operation to encode the segment based on the processing control signal, wherein the first encoding operation provides more efficient encoding of the target signal category than the second encoding operation.
[0262] Example 76 includes a non-transitory computer-readable medium according to any one of Examples 72 to 75, wherein the time-series data represents audio content, and wherein the output indicates whether the segment includes an audio data type associated with a first audio encoder.
[0263] Example 77 includes a non-transitory computer-readable medium according to any one of Examples 72 to 76, wherein the instructions further cause the one or more processors to selectively route the segment to an audio encoder based on the processing control signal.
[0264] Example 78 includes a non-transitory computer-readable medium according to any one of Examples 72 to 77, wherein the segment corresponds to an audio frame comprising a spectral representation of one or more audio data samples.
[0265] Example 79 includes a non-transitory computer-readable medium according to any one of Examples 72 to 78, wherein the feature extractor includes an inference network portion and a generative network portion of an autoencoder.
[0266] Example 80 includes a non-transitory computer-readable medium according to Example 79, wherein the autoencoder includes one or more recursive layers.
[0267] Example 81 includes a non-transitory computer-readable medium according to Example 79 or Example 80, wherein the autoencoder includes one or more dilated convolutional layers.
[0268] Example 82 includes a non-transitory computer-readable medium according to any one of Examples 79 to 81, wherein the autoencoder includes one or more time-dynamic layers.
[0269] Example 83 includes a non-transient computer-readable medium according to any one of Examples 79 to 82, wherein the autoencoder is a variational autoencoder, and the latent space representation includes the mean and standard deviation of a probability distribution.
[0270] Example 84 includes a non-transitory computer-readable medium according to any one of Examples 79 to 83, wherein the autoencoder is trained to reproduce data segments from a target signal category, and the classifier is configured to distinguish between data segments from the target signal category and data segments not from the target signal category based on the separation of latent space representations between the data segments from the target signal category and data segments not from the target signal category.
[0271] Example 85 includes a non-transitory computer-readable medium according to any one of Examples 79 to 84, wherein the autoencoder is trained to reproduce speech data, and the classifier is configured to distinguish audio data segments that include speech from audio data segments that do not include speech.
[0272] Example 86 includes a non-transitory computer-readable medium according to any one of Examples 79 to 85, wherein the instructions further cause the one or more processors to provide input to the generative network portion based on the latent space representation to generate a synthesized fragment; and to determine an error value based on a comparison of the fragment with the synthesized fragment, wherein at least one of the one or more inputs provided to the classifier is based on the error value.
[0273] Example 87 includes a non-transitory computer-readable medium according to Example 86, wherein determining the error value includes calculating the Itakura-Saito distance based on the fragment and the synthesized fragment.
[0274] Example 88 includes a non-transitory computer-readable medium according to Example 86, wherein determining the error value includes calculating a scale-invariant signal distortion ratio based on the segment and the synthesized segment.
[0275] Example 89 includes a non-transitory computer-readable medium according to Example 86, wherein determining the error value includes calculating a logarithmic spectral distortion value based on the segment and the synthesized segment.
[0276] Example 90 includes a non-transitory computer-readable medium according to any one of Examples 79 to 89, wherein the instructions further cause the one or more processors to provide input to the generative network portion based on the latent space representation to generate a probability distribution; and to determine an error value based on the segment and the probability distribution, wherein at least one of the one or more inputs provided to the classifier is based on the error value.
[0277] Example 91 includes a non-transitory computer-readable medium according to any one of Examples 72 to 90, wherein at least one of the one or more inputs provided to the classifier includes the latent space representation.
[0278] Example 92 includes a non-transitory computer-readable medium according to any one of Examples 72 to 91, wherein the instructions further cause the one or more processors to determine a divergence measure indicating the difference between the probability distribution of the latent space representation and the expected probability distribution, and wherein at least one of the one or more inputs provided to the classifier is based on the divergence measure.
[0279] Example 93 includes a non-transitory computer-readable medium according to any one of Examples 72 to 92, wherein the instructions further cause the one or more processors to determine a pattern indicator associated with the segment, and wherein at least one of the one or more inputs provided to the classifier is based on the pattern indicator.
[0280] Example 94 includes a non-transitory computer-readable medium according to Example 93, wherein the mode indicator indicates whether the segment represents voiced speech.
[0281] Example 95 includes a non-transitory computer-readable medium according to Example 93 or Example 94, wherein the mode indicator indicates whether the segment represents music.
[0282] Example 96 includes a non-transitory computer-readable medium according to any one of Examples 93 to 95, wherein the pattern indicator is based on statistical analysis of the time series data.
[0283] Example 97 includes a non-transitory computer-readable medium according to any one of Examples 72 to 96, wherein the instructions further cause the one or more processors to determine one or more features associated with the segment, and wherein at least one of the one or more inputs provided to the classifier is based on the one or more features.
[0284] Example 98 includes a non-transitory computer-readable medium according to Example 97, wherein one or more of the features include one or more of open-loop pitch, normalized correlation, spectral envelope, pitch stability, signal nonstationarity, linear prediction residual, spectral difference, and spectral stationarity.
[0285] Example 99 includes a non-transitory computer-readable medium according to any one of Examples 72 to 98, wherein the instructions further cause the one or more processors to generate an Enhanced Variable Rate Codec (EVRC) mode decision based on the segment, and wherein at least one of the one or more inputs provided to the classifier is based on the EVRC mode decision.
[0286] Example 100 includes a non-transitory computer-readable medium according to any one of Examples 72 to 99, wherein the feature extractor includes a first machine learning model trained on training data to synthesize time series data that approximates the input time series data.
[0287] Example 101 includes a non-transitory computer-readable medium according to Example 100, wherein each segment of the training data includes speech.
[0288] Example 102 includes a non-transitory computer-readable medium according to Example 100, wherein the training data includes segments representing speech and segments representing non-speech sounds.
[0289] Example 103 includes a non-transitory computer-readable medium according to Example 100, wherein each segment of the training data includes data assigned to a target signal category.
[0290] Example 104 includes a non-transitory computer-readable medium according to Example 100, wherein the training data includes segments assigned to: data assigned to a target signal category and data not assigned to the target signal category.
[0291] Example 105 includes a non-transitory computer-readable medium according to any one of Examples 100 to 104, wherein the classifier is trained using labeled training data based on the output of the feature extractor.
[0292] Example 106 includes a non-transitory computer-readable medium according to any one of Examples 100 to 105, wherein the classifier includes a second machine learning model, and wherein the first machine learning model and the second machine learning model are jointly trained.
[0293] According to embodiment 107, an apparatus includes components for generating a latent space representation of a segment of time series data; components for providing one or more inputs to a classifier, the one or more inputs including at least one input based on the latent space representation; and components for generating a processing control signal for the segment based on the output of the classifier.
[0294] Example 108 includes the apparatus according to Example 107, wherein the classifier is a single-class classifier or a binary classifier, and the output indicates whether the segment is assigned to a target signal category.
[0295] Example 109 includes the apparatus according to Example 107 or Example 108, the apparatus further including components for performing a first encoding operation to encode the segment; components for performing a second encoding operation to encode the segment; and components for selectively providing the segment to the components performing the first encoding operation or to the components performing the second encoding operation based on the processing control signal.
[0296] Example 110 includes an apparatus according to any one of Examples 107 to 109, wherein the time-series data represents audio content, and wherein the output indicates whether the segment includes an audio data type associated with a first audio encoder.
[0297] Example 111 includes the apparatus according to any one of Examples 107 to 110, and the apparatus further includes a component for selectively routing the segment to an audio encoder based on the processing control signal.
[0298] Example 112 includes the apparatus according to any one of Examples 107 to 111, wherein the segment corresponds to an audio frame comprising a spectral representation of one or more audio data samples.
[0299] Example 113 includes an apparatus according to any one of Examples 107 to 112, wherein the feature extractor includes an inference network portion and a generative network portion of an autoencoder.
[0300] Example 114 includes the apparatus according to Example 113, wherein the autoencoder includes one or more recursive layers.
[0301] Example 115 includes the apparatus according to Example 113 or Example 114, wherein the autoencoder includes one or more dilated convolutional layers.
[0302] Example 116 includes the apparatus according to any one of Examples 113 to 115, wherein the auto encoder includes one or more time-dynamic layers.
[0303] Example 117 includes the apparatus according to any one of Examples 113 to 116, wherein the autoencoder is a variational autoencoder, and the latent space representation includes the mean and standard deviation of a probability distribution.
[0304] Example 118 includes an apparatus according to any one of Examples 113 to 117, wherein the autoencoder is trained to reproduce data segments from a target signal category, and the classifier is configured to distinguish between data segments from the target signal category and data segments not from the target signal category based on the separation of latent space representations between the data segments from the target signal category and data segments not from the target signal category.
[0305] Example 119 includes an apparatus according to any one of Examples 113 to 118, wherein the autoencoder is trained to reproduce speech data, and the classifier is configured to distinguish audio data segments that include speech from audio data segments that do not include speech.
[0306] Example 120 includes an apparatus according to any one of Examples 113 to 119, the apparatus further comprising means for providing input to the generative network portion based on the latent space representation to generate a synthetic fragment of time series data; and means for determining an error value based on a comparison of the fragment with the synthetic fragment, wherein at least one of the one or more inputs provided to the classifier is based on the error value.
[0307] Example 121 includes the apparatus according to Example 120, wherein determining the error value includes calculating the Itakura-Saito distance based on the fragment and the synthesized fragment.
[0308] Example 122 includes the apparatus according to Example 120, wherein determining the error value includes calculating a scale-invariant signal distortion ratio based on the segment and the synthesized segment.
[0309] Example 123 includes the apparatus according to Example 120, wherein determining the error value includes calculating a logarithmic spectral distortion value based on the segment and the synthesized segment.
[0310] Example 124 includes an apparatus according to any one of Examples 113 to 123, the apparatus further comprising means for providing input to the generative network portion based on the latent space representation to generate a probability distribution; and means for determining an error value based on the segment and the probability distribution, wherein at least one of the one or more inputs provided to the classifier is based on the error value.
[0311] Example 125 includes an apparatus according to any one of Examples 107 to 124, wherein at least one of the one or more inputs provided to the classifier includes the latent space representation.
[0312] Example 126 includes the apparatus according to any one of Examples 107 to 125, and the apparatus further includes a component for determining a divergence metric indicating the difference between the probability distribution of the latent space representation and the expected probability distribution, and wherein at least one of the one or more inputs provided to the classifier is based on the divergence metric.
[0313] Example 127 includes the apparatus according to any one of Examples 107 to 126, and the apparatus further includes components for determining a pattern indicator associated with the segment, and wherein at least one of the one or more inputs provided to the classifier is based on the pattern indicator.
[0314] Example 128 includes the apparatus according to Example 127, wherein the mode indicator indicates whether the segment represents voiced speech.
[0315] Example 129 includes the apparatus according to Example 127 or Example 128, wherein the mode indicator indicates whether the segment represents music.
[0316] Example 130 includes the apparatus according to any one of Examples 127 to 129, and the apparatus further includes components for performing statistical analysis of the time series data to generate the pattern indicator.
[0317] Example 131 includes the apparatus according to any one of Examples 107 to 130, and the apparatus further includes components for determining one or more features associated with the segment, and wherein at least one of the one or more inputs provided to the classifier is based on the one or more features.
[0318] Example 132 includes the apparatus according to Example 131, wherein one or more features include one or more of open-loop pitch, normalized correlation, spectral envelope, pitch stability, signal nonstationarity, linear prediction residual, spectral difference, and spectral stationarity.
[0319] Example 133 includes the apparatus according to any one of Examples 107 to 132, and the apparatus further includes components for generating an Enhanced Variable Rate Codec (EVRC) mode decision based on the segment, and wherein at least one of the one or more inputs provided to the classifier is based on the EVRC mode decision.
[0320] Example 134 includes an apparatus according to any one of Examples 107 to 133, wherein the feature extractor includes a first machine learning model trained on training data to synthesize time series data that approximates the input time series data.
[0321] Example 135 includes the apparatus according to Example 134, wherein each segment of the training data includes speech.
[0322] Example 136 includes the apparatus according to Example 134, wherein the training data includes segments representing speech and segments representing non-speech sounds.
[0323] Example 137 includes the apparatus according to Example 134, wherein each segment of the training data includes data assigned to a target signal category.
[0324] Example 138 includes the apparatus according to Example 134, wherein the training data includes segments assigned to: data assigned to a target signal category and data not assigned to the target signal category.
[0325] Example 139 includes an apparatus according to any one of Examples 134 to 138, wherein the classifier is trained using labeled training data based on the output of the feature extractor.
[0326] Example 140 includes the apparatus according to any one of Examples 134 to 139, wherein the classifier includes a second machine learning model, and wherein the first machine learning model and the second machine learning model are jointly trained.
[0327] Those skilled in the art will also recognize that the various exemplary logic blocks, configurations, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein can be implemented as electronic hardware, computer software executed by a processor, or a combination of both. The various exemplary components, blocks, configurations, modules, circuits, and steps have been generally described above in terms of their functionality. Whether such functionality is implemented as hardware or processor-executable instructions depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each specific application, and such implementation decisions shall not be construed as departing from the scope of this disclosure.
[0328] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be directly implemented as hardware, a software module executed by a processor, or a combination of both. The software module can reside in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, hard disks, removable disks, compressed optical disc read-only memory (CD-ROM), or any other form of non-transitory storage medium known in the art. An exemplary storage medium is coupled to a processor such that the processor can read information from and write information to the storage medium. Alternatively, the storage medium may be integral with the processor. The processor and storage medium may reside in an application-specific integrated circuit (ASIC). The ASIC may reside in a computing device or user terminal. Alternatively, the processor and storage medium may reside as discrete components in a computing device or user terminal.
[0329] The prior description of the disclosed aspects is provided to enable those skilled in the art to make or use the disclosed aspects. Various modifications to these aspects will be apparent to those skilled in the art, and the principles defined herein can be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but should be granted the broadest scope that may be consistent with the principles and novel features as defined by the following claims.
Claims
1. An apparatus, the apparatus comprising: A memory configured to store one or more segments of time-series data; and One or more processors, said one or more processors being configured to: A feature extractor is used to generate latent space representations of segments of the time series data; Provide one or more inputs to the classifier, said one or more inputs including at least one input based on the latent space representation; and The processing control signal for the segment is generated based on the output of the classifier.
2. The device of claim 1, wherein the classifier is a single-class classifier or a binary classifier, and the output indicates whether the segment is assigned to a target signal category.
3. The device of claim 2, wherein the one or more processors are configured to selectively perform a first encoding operation to encode the segment or perform a second encoding operation to encode the segment based on the processing control signal, wherein the first encoding operation provides a higher quality encoding of the target signal category than the second encoding operation.
4. The device of claim 2, wherein the one or more processors are configured to selectively perform a first encoding operation to encode the segment or perform a second encoding operation to encode the segment based on the processing control signal, wherein the first encoding operation provides more efficient encoding of the target signal category than the second encoding operation.
5. The device of claim 1, wherein the time-series data represents audio content, and wherein the output indicates whether the segment includes an audio data type associated with the first audio encoder.
6. The device of claim 1, wherein the one or more processors are configured to selectively route the segment to one of two or more audio decoders based on the processing control signal.
7. The device of claim 1, wherein the segment corresponds to a segment of audio data, and the input to the feature extractor includes a frequency domain representation of the segment of audio data.
8. The device of claim 7, wherein the frequency domain represents the power spectrum of the segment comprising audio data.
9. The apparatus of claim 1, wherein the feature extractor comprises an inference network portion and a generation network portion of an autoencoder.
10. The device of claim 9, wherein the autoencoder is a variational autoencoder, and the latent space representation includes the mean and standard deviation of a probability distribution.
11. The apparatus of claim 9, wherein the autoencoder is trained to reproduce data segments from a target signal category, and the classifier is configured to distinguish between data segments from the target signal category and data segments not from the target signal category based on the separation of latent space representations between the data segments from the target signal category and data segments not from the target signal category.
12. The apparatus of claim 9, wherein the autoencoder is trained to reproduce speech data, and the classifier is configured to distinguish audio data segments that include speech from audio data segments that do not include speech.
13. The device of claim 9, wherein the one or more processors are configured to: The latent space representation is used to provide input to the generative network portion to generate synthesized fragments of time-series data; and A reconstruction error value is determined based on a comparison between the fragment and the synthesized fragment, wherein at least one of the one or more inputs provided to the classifier is based on the reconstruction error value.
14. The device of claim 9, wherein the one or more processors are configured to: The latent space representation is used to provide input to the generative network portion to generate a probability distribution; and An error value is determined based on the fragment and the probability distribution, wherein at least one of the one or more inputs provided to the classifier is based on the error value.
15. A method, the method comprising: Use feature extractors to generate latent space representations of fragments of time series data; Provide one or more inputs to the classifier, said one or more inputs including at least one input based on the latent space representation; and The processing control signal for the segment is generated based on the output of the classifier.
16. The method of claim 15, wherein the classifier is a single-class classifier or a binary classifier, and the output indicates whether the segment is assigned to a target signal category.
17. The method of claim 16, further comprising selectively performing a first encoding operation to encode the segment or performing a second encoding operation to encode the segment based on the processing control signal, wherein the first encoding operation provides a higher quality encoding of the target signal category than the second encoding operation.
18. The method of claim 16, further comprising selectively performing a first encoding operation to encode the segment or performing a second encoding operation to encode the segment based on the processing control signal, wherein the first encoding operation provides more efficient encoding of the target signal category than the second encoding operation.
19. The method of claim 15, wherein the time-series data represents audio content, and wherein the output indicates whether the segment includes an audio data type associated with a first audio encoder.
20. The method of claim 15, further comprising selectively routing the segment to an audio encoder based on the processing control signal.
21. The method of claim 15, wherein the segment corresponds to an audio frame comprising a spectral representation of one or more audio data samples.
22. The method of claim 15, wherein the feature extractor comprises an inference network portion and a generative network portion of an autoencoder.
23. The method of claim 22, wherein the autoencoder is trained to reproduce data segments from a target signal category, and the classifier is configured to distinguish the data segments from the target signal category from data segments not from the target signal category.
24. The method of claim 23, wherein the autoencoder is trained to reproduce speech data, and the classifier is configured to distinguish audio data segments that include speech from audio data segments that do not include speech.
25. The method according to claim 22, further comprising: The latent space representation is used to provide input to the generative network portion to generate synthesized fragments of time series data; as well as An error value is determined based on a comparison between the fragment and the synthesized fragment, wherein at least one of the one or more inputs provided to the classifier is based on the error value.
26. The method according to claim 22, further comprising: The latent space representation is used to provide input to the generative network portion to generate a probability distribution; as well as An error value is determined based on the fragment and the probability distribution, wherein at least one of the one or more inputs provided to the classifier is based on the error value.
27. The method of claim 15, further comprising determining a pattern indicator associated with the segment, wherein at least one of the one or more inputs provided to the classifier is based on the pattern indicator.
28. A non-transitory computer-readable medium storing instructions, the instructions being executable by one or more processors to cause the one or more processors to: Use feature extractors to generate latent space representations of fragments of time series data; Provide one or more inputs to the classifier, said one or more inputs including at least one input based on the latent space representation; and The processing control signal for the segment is generated based on the output of the classifier.
29. The non-transitory computer-readable medium of claim 28, wherein the feature extractor comprises an inference network portion and a generation network portion of an autoencoder, wherein the classifier is a single-class classifier or a binary classifier, and wherein the output indicates whether the segment is assigned to a target signal category.
30. An apparatus comprising: A component used to generate latent space representations of segments of time series data; A component for providing one or more inputs to a classifier, said one or more inputs including at least one input based on the latent space representation; and A component for generating processing control signals for the fragment based on the output of the classifier.