Energy efficient system and method for analyzing time series data using attentive power iteration
The Attentive Power Iteration model enables real-time time-series analysis on resource-constrained devices by using incremental PCA and matrix sketching, addressing the limitations of existing AI models and achieving high accuracy with reduced energy and time.
Patent Information
- Application Number
- JP2025019589
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-11-12
- Filing Date
- 2025-02-07
- Publication Date
- 2025-09-02
Smart Images

Figure 2025128033000001_ABST
Abstract
Description
[Technical Field]
[0001] TECHNICAL FIELD The present disclosure relates generally to systems and methods for analyzing time series data, and more particularly to systems and methods for analyzing time series data by performing attention weighting. [Background technology]
[0002] Resource-constrained devices, characterized by limited computational capacity and memory, are essential components of modern infrastructure. Devices such as edge devices (e.g., routers, smartphones) and distributed energy resources (DERs) like solar power units, turbines, and gas units all fall into this category. These devices accumulate vast amounts of time-series data streams over very short time intervals over time, making manual real-time analysis impractical. Artificial intelligence (AI) models, particularly those operating on time-series data, are gaining popularity due to their ability to autonomously identify underlying patterns. However, existing AI models, especially common neural networks, often require significant memory and computational energy, making them unsuitable for direct deployment on resource-constrained devices. A common workaround is to transfer collected data to a cloud server for analysis, but this method consumes bandwidth and time and raises concerns about data privacy and security during transmission. Furthermore, many AI models rely on batch learning, which processes entire time series simultaneously, rather than the incremental learning required for real-time analysis.
[0003] To address these challenges, there is a pressing need for streaming, cost-effective models that can analyze time-series data streams in real time directly on resource-constrained devices. Summary of the Invention
[0004] Aspects and advantages of the computer-implemented method and computing system according to the present disclosure will be set forth in part in the description that follows, or will be obvious from the description, or may be learned by practicing the techniques.
[0005] According to one embodiment, a computer-implemented method for analyzing time series data is provided. The method includes providing a sequence of time series batches of time series data to an Attentive Power Iteration (API) model. The method further includes generating, by the API model, a sequence of time series sketches based on the sequence of time series batches of time series data. The method further includes assigning a weight to a new time series batch in the sequence of time series batches based at least in part on a previous time series sketch in the sequence of time series sketches. The method further includes generating an output for each time series sketch in the sequence of time series sketches.
[0006] According to another embodiment, a computing system is provided. The computing system includes one or more processors. The computing system further includes one or more non-transitory computer-readable media, the one or more non-transitory computer-readable media collectively storing instructions that, when executed by the one or more processors, cause the computing system to perform operations. The operations include providing a sequence of time series batches of the time series data to an Attentive Power Iteration (API) model. The operations further include generating, by the API model, a sequence of time series sketches based on the sequence of time series batches of the time series data. The operations further include assigning a weight to a new time series batch of the sequence of time series batches based at least in part on a previous time series sketch in the sequence of time series sketches. The operations further include generating an output for each time series sketch in the sequence of time series sketches.
[0007] These and other features, aspects, and advantages of the computer-implemented method and computing system can be better understood with reference to the following description and claims. The drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments of the present technology and, together with the description, help to explain the principles of the technology. [Brief explanation of the drawings]
[0008] A full and enabling disclosure, to those skilled in the art, of computer-implemented methods and computing systems, including the best mode of making and using the systems and methods, is set forth in the specification, which makes reference to the following figures: [Figure 1] FIG. 1 is a block diagram of an exemplary computing system for performing real-time analysis of time-series data in accordance with an exemplary embodiment of the present disclosure. [Figure 2] 2 illustrates a block diagram of an exemplary streaming time series classification model that can be implemented by the computing system illustrated in FIG. 1 , according to an embodiment of the present disclosure. [Figure 3] 2 illustrates a block diagram of an exemplary streaming time series classification model that can be implemented by the computing system illustrated in FIG. 1 , according to an embodiment of the present disclosure. [Figure 4] 1 shows a flowchart of a computer-implemented method for analyzing time series data according to an embodiment of the present disclosure. [Figure 5] FIG. 1 illustrates a process according to an embodiment of the present disclosure. [Figure 6] 1 shows a flowchart of a computer-implemented method for analyzing time series data according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0009] Several embodiments of computer-implemented methods and computing systems are now described in detail. One or more examples of these embodiments are illustrated in the drawings. Each example is provided for illustrative purposes, not limiting, of the technology. Indeed, it will be apparent to those skilled in the art that modifications and variations can be made to the technology without departing from the scope or spirit of the claimed technology. For example, features illustrated or described as part of one embodiment can be used with another embodiment to yield a still further embodiment. Accordingly, the present disclosure is intended to cover such modifications and variations as come within the scope of the claims and their equivalents.
[0010] The word "exemplary" is used herein to mean serving as an example, instance, or illustration. In the description that follows, any implementation identified as "exemplary" should not necessarily be construed as recommended or advantageous over other implementations. Moreover, unless otherwise specified, all embodiments described herein should be considered exemplary.
[0011] The detailed description uses numbers and letters to refer to features in the drawings. Like or similar symbols in the drawings and the detailed description are used to indicate like or similar parts of the invention. In this specification, the terms "first," "second," and "third" may be used interchangeably to distinguish one component from another and are not intended to denote the location or importance of the individual components.
[0012] Approximate terms (such as "approximately," "about," "generally," and "substantially") are not intended to be limiting to the exact value specified. In at least some instances, approximate language may correspond to the precision of an instrument for measuring a value or the precision of a method or machine for constructing or manufacturing a component and / or system. In at least some instances, approximate language may correspond to the precision of an instrument for measuring a value or the precision of a method or machine for constructing or manufacturing a component and / or system. For example, approximate language may express values within a margin of 1, 2, 4, 5, 10, 15, or 20 percent for an individual value, a range of values, and / or any of the endpoints defining a range of values. When used in the context of an angle or direction, such terms include angles or directions within a range of 10 degrees greater or less than the stated angle or direction. For example, "approximately vertical" includes directions within 10 degrees in any direction (e.g., clockwise or counterclockwise) from vertical.
[0013] As used herein, the terms "comprises, comprising, includes, including, has, having, or variations of these words" are intended to indicate a non-exclusive inclusion. For example, a process, method, article, or apparatus that includes a list of features is not necessarily limited to only those features and may include other features not expressly listed, or other features inherent to such process, method, article, or apparatus. Furthermore, unless expressly stated to the contrary, "or" means an inclusive "or," not an exclusive "or." For example, a condition A or B is satisfied if any of the following cases exist: A is true (or exists) and B is false (or does not exist), A is false (or does not exist) and B is true (or exists), or both A and B are true (or exist).
[0014] Throughout this specification and claims, range limitations are combinable and interchangeable, and where such ranges are specified, unless the context or language indicates otherwise, such ranges are intended to include all subranges subsumed therein. For example, all ranges disclosed herein are inclusive of their endpoints, and the endpoints are independently combinable with each other.
[0015] This application relates to systems and methods for addressing the streaming time series classification problem, in which a categorical class label is predicted for each time series consisting of an ordered set of real-valued, often multivariate, attributes. Streaming time series classification is performed on time series streams to generate real-time labels while processing the data only once and utilizing limited storage. This disclosure includes neural networks for performing streaming time series classification, which are designed to operate with limited computational memory and energy resources while ensuring high processing speed.
[0016] For example, as data from the same time series arrive sequentially, the system can construct a real-time representation using incremental weighted principal component analysis (PCA) on a latent space trained with a supervised neural network. This representation is then employed for time series classification. While traditional PCA is an unsupervised technique for dimensionality reduction in non-time series data, the disclosed system and method uses weighted PCA on a latent space trained by considering both temporal aspects and supervised learning to accommodate the nature of time series data. Because multiple time series samples are not equally important, the disclosed system and method introduces weighted PCA. In weighted PCA, each time series sample is assigned a weight using a temporal attention model. With this weighted time series stream, the model can incrementally update its principal direction and form a distinctive real-time representation for classification. The complete system architecture encompasses latent space learning and incremental weighted PCA and is trained end-to-end using supervised learning methods. Specifically, the model can utilize streaming attentive power iteration (strAPI) to incrementally refine the dominant direction in a supervised latent space. The model can generate real-time representations and labels for each time series without needing to observe the entire sequence. The systems and methods of the present disclosure have been shown to consistently achieve superior classification accuracy, generally with minimal energy consumption and fast computational speed, when compared to other similar systems.
[0017] The disclosed systems and methods provide numerous technical effects and advantages. As an example, the systems and methods can be used to leverage machine learning models and attentive power iteration models for real-time processing of time-series data on resource-constrained devices. For example, the models continuously update a concise representation of the entire time series, improving classification (e.g., output) accuracy while saving energy and processing time. Notably, these models perform well in streaming scenarios and enable rapid decision-making without requiring access to the full time-series data. The models excel in classification accuracy and energy efficiency, achieving over 70% lower power consumption and three times faster task completion speeds compared to benchmarks. This research improves real-time responsiveness, energy savings, and operational efficiency in constrained devices, contributing to the optimization of various applications.
[0018] Referring now to the figures, exemplary embodiments of the present disclosure will be described in further detail.
[0019] 1 illustrates a block diagram of an exemplary computing system 100 for performing real-time analysis of time-series data in accordance with an exemplary embodiment of the present disclosure. The system 100 includes a user computing device 102, a server computing system 130, and a training computing system 150 communicatively coupled over a network 180.
[0020] The user computing device 102 can be any type of computing device (e.g., a control computing device for a machine (e.g., a control computing device for a wind turbine, water wheel, crane, etc.), a personal computing device (e.g., a laptop or desktop), a mobile computing device (e.g., a smartphone or tablet), a personal computing device (e.g., a laptop or desktop), a mobile computing device (e.g., a smartphone or tablet), a diagnostic computing device for diagnosing machine anomalies, a wearable computing device, an embedded computing device, an edge computing device, or any other type of computing device).
[0021] The user computing device 102 includes one or more processors 112 and memory 114. The one or more processors 112 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be a single processor or multiple operatively connected processors. The memory 114 may include one or more non-transitory computer-readable storage media (e.g., RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof). The memory 114 may store data 116 and instructions 118, which, when executed by the processor 112, cause the user computing device 102 to perform operations.
[0022] In some implementations, the user computing device 102 may store or include one or more machine-learned models 120. For example, the machine-learned models 120 may be or include various machine-learned models, such as neural networks (e.g., deep neural networks), or other types of machine-learned models, including nonlinear and / or linear models. Neural networks may include feedforward neural networks, recurrent neural networks (e.g., long-short-term memory recurrent neural networks), convolutional neural networks, temporal convolutional networks (TCNs), or other forms of neural networks. Some exemplary machine-learned models may utilize attention mechanisms (such as self-attention). For example, some exemplary machine-learned models may include multi-head self-attention models (e.g., Transformer models). Exemplary machine-learned models 120 are described with reference to FIGS. 2-3.
[0023] In some implementations, one or more machine-learned models 120 may be received from a server computing system 130 over a network 180, stored in a memory 114 of a user computing device, and then used or otherwise executed by one or more processors 112. In some implementations, a user computing device 102 may implement multiple parallel instances of a single machine-learned model 120.
[0024] Additionally or alternatively, one or more machine-learned models 140 may be included in or otherwise stored and executed by a server computing system 130 that communicates with the user computing device 102 according to a client-server relationship. For example, the machine-learned models 140 may be implemented by the server computing system 130 as part of a web service. Thus, one or more models 120 may be stored and implemented on the user computing device 102 and / or one or more models 140 may be stored and implemented on the server computing system 130.
[0025] User computing device 102 may also include one or more user input components 122 for receiving user input. For example, user input component 122 may be a touch-sensitive component (e.g., a touch-sensitive display screen or touchpad), a microphone, a conventional keyboard, or other means by which a user can provide user input.
[0026] The user computing device 102 may include and / or be communicatively connected to one or more sensors. The one or more sensors may be utilized to generate time series data. The time series data may describe machine performance and / or environmental signals over a period of time.
[0027] The server computing system 130 includes one or more processors 132 and memory 134. The one or more processors 132 may be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and may be a single processor or multiple operatively connected processors. The memory 134 may include one or more non-transitory computer-readable storage media (e.g., RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof). The memory 134 may store data 136 and instructions 138 that are executed by the processor 132 to cause the server computing system 130 to perform operations.
[0028] In some implementations, server computing system 130 includes or is implemented by one or more server computing devices. In examples where server computing system 130 includes multiple server computing devices, such server computing devices may operate according to a serial computing architecture, a parallel computing architecture, or some combination thereof.
[0029] As described above, the server computing system 130 may store or include one or more machine-learned models 140. For example, the models 140 may be or include various machine-learned models. Exemplary machine-learned models include neural networks or other multi-layer nonlinear models. Exemplary neural networks include feedforward neural networks, deep neural networks, recurrent neural networks, and convolutional neural networks. Some exemplary machine-learned models may utilize attention mechanisms (such as self-attention). For example, some exemplary machine-learned models may include multi-head self-attention models (e.g., Transformer models). Exemplary models 140 are described with reference to FIGS. 2-3.
[0030] User computing device 102 and / or server computing system 130 may train models 120 and / or 140 through interaction with a training computing system 150 that is communicatively coupled via network 180. Training computing system 150 may be a separate system from server computing system 130, may be part of server computing system 130, or may be part of user computing system 102.
[0031] The training computing system 150 includes one or more processors 152 and memory 154. The one or more processors 152 can be any suitable processing device (e.g., a processor core, a microprocessor, an ASIC, an FPGA, a controller, a microcontroller, etc.) and can be a single processor or multiple operably connected processors. The memory 154 can include one or more non-transitory computer-readable storage media (e.g., RAM, ROM, EEPROM, EPROM, flash memory devices, magnetic disks, etc., and combinations thereof). The memory 154 can store data 156 and instructions 158 that are executed by the processors 152 to cause the training computing system 150 to perform operations. In some implementations, the training computing system 150 includes or is implemented by one or more server computing devices.
[0032] The training computing system 150 may include a model trainer 160 that trains the machine-learned models 120 and / or 140 stored on the user computing device 102 and / or the server computing system 130 using various training or learning techniques, such as backpropagation of errors. For example, a loss function may be backpropagated through the model to update one or more parameters of the model (e.g., based on the gradient of the loss function). Various loss functions may be used, such as mean squared error, likelihood loss, cross-entropy loss, hinge loss, and / or various other loss functions. Gradient descent may be used to iteratively update the parameters over multiple training iterations.
[0033] In some implementations, performing backpropagation of errors may include performing truncated backpropagation through time. The model trainer 160 may implement several generalization techniques (e.g., weight decay, dropout, etc.) to improve the generalization ability of the model being trained. The model trainer 160 may include one or more teacher models for performing distilled training to generate a condensed model that can be implemented by the computing device 102.
[0034] Model trainer 160 contains computer logic used to provide the desired functionality. Model trainer 160 can be implemented in hardware, firmware, and / or software controlling a general-purpose processor. For example, in some implementations, model trainer 160 includes a program file stored on a storage device, loaded into memory, and executed by one or more processors. In other implementations, model trainer 160 includes one or more sets of computer-executable instructions stored on a tangible computer-readable storage medium (such as RAM, a hard disk, or an optical or magnetic medium).
[0035] Network 180 can be any type of communications network, such as a local area network (e.g., an intranet), a wide area network (e.g., the Internet), or a combination thereof, and can include any number of wired or wireless links. In general, communications on network 180 can be transmitted over any type of wired and / or wireless connection using a wide variety of communications protocols (e.g., TCP / IP, HTTP, SMTP, FTP), encodings or formats (e.g., HTML, XML), and / or protection schemes (e.g., VPN, Secure HTTP, SSL).
[0036] The machine-learned models described herein can be used in a variety of tasks, applications, and / or use cases.
[0037] In some implementations, inputs for the machine-learned models of the present disclosure may be sensor data, for example, from environmental sensors 182. The environmental sensors may be operably connected (or coupled) to a power generation system 184 (such as a gas turbine engine, a hydroelectric turbine, a wind turbine, or other power generation system).
[0038] The machine-learned model may process the sensor data to generate an output. As an example, the machine-learned model may process the sensor data to generate a recognition output. As another example, the machine-learned model may process the sensor data to generate a prediction output. As another example, the machine-learned model may process the sensor data to generate a classification output. As another example, the machine-learned model may process the sensor data to generate a segmentation output. As another example, the machine-learned model may process the sensor data to generate a visualization output. As another example, the machine-learned model may process the sensor data to generate a diagnostic output. As another example, the machine-learned model may process the sensor data to generate a detection output.
[0039] 1 illustrates an example of a computing system that can be used to implement the present disclosure. Other computing systems may be used. For example, in some implementations, a user computing device 102 may include a model trainer 160 and a training dataset 162. In such implementations, the model 120 may be both trained and used locally on the user computing device 102. In some implementations, the user computing device 102 may implement the model trainer 160 to personalize the model 120 based on user-specific data.
[0040] In many embodiments, computing system 100 may be a resource-constrained device (e.g., a memory-limited and / or processing-power-limited device), while in other embodiments, computing system 100 may be an edge computing device.
[0041] 2 illustrates a block diagram of an exemplary streaming time series classification model 200 that may be implemented by the computing system 100 described above with reference to FIG. 1, in accordance with an exemplary embodiment of the present disclosure. In some implementations, the model 200 is trained to receive a sequence of time series streams or batches 202. The model 200 may be further configured to perform a weighted principal component analysis (e.g., weighted PCA or WPCA) on the time series batches 202 via an Attentive Power Iteration (API) model (or weighting block) 204 to generate a sequence of time series sketches 206.
[0042] The API model 204 can incrementally refine the dominant direction in the supervised latent space. The API model 204 can find dominant eigenvalues and eigenvectors of a large number of time series batches 202 and incorporate an attention mechanism. The attention mechanism allows the API model 204 to heavily weight specific features or dimensions, thereby improving performance. Advantageously, the API model 204 can effectively capture temporal patterns of the time series batches 202 by iteratively updating a reduced representation of the time series batches 202 (i.e., by iteratively updating the sequence of time series sketches 206). The API model 204 can reliably highlight important temporal patterns and accurately classify them by assigning corresponding weights to the important temporal patterns. For example, important temporal patterns may include, but are not limited to, trends (e.g., long-term directional movements or tendencies), seasonality (e.g., periodic fluctuations or cycles), periodic patterns (e.g., non-seasonal fluctuations), or other patterns.
[0043] The model 200 can be further configured to provide real-time analysis of the sequence of time series sketches (e.g., prediction, classification, or other output) via a classification model 208. In some embodiments, as shown, the classification model 208 can generate a classification label 210 for each time series sketch 206A, 206B, 206C in the sequence of time series sketches 206.
[0044] As shown, model 200 may be supplied with a sequence of time series batches 202 (first time series batch 202A, second time series batch 202B, up to Nth time series batch 202C, etc.). Each batch in the sequence of time series batches 202 is a portion of a time series (or time series data) that may be reduced by model 200 before being assigned weights in API model 204. As a non-limiting example, given time series data collected over a 10-minute period, first time series batch 102A may be data for a first time (e.g., data for the time series collected from minute 0 to minute 1), second time series batch 102B may be data for a second time (e.g., data for the time series collected from minute 1 to minute 2), etc. The time series forming time series batch 202 may include any suitable time series data. For example, the time series may be sensor data, weather data, supply chain data, energy consumption data, stock price data, or other data collected over a period of time (e.g., periodically collected).
[0045] Advantageously, the API model 204 can effectively capture temporal patterns of the time series batch 202 by iteratively updating a reduced representation of the time series batch 202 (i.e., by iteratively updating the sequence of time series sketches 206). The API model 204 can reliably highlight significant temporal patterns by assigning corresponding weights to the significant temporal patterns, enabling accurate classification. For example, significant temporal patterns may include, but are not limited to, trends (e.g., long-term directional movements or tendencies), seasonality (e.g., periodic fluctuations or cycles), cyclical patterns (e.g., non-seasonal fluctuations), or other patterns.
[0046] The API model 204 can use weighted PCA for each time series batch 202. Weighted PCA is an extension of traditional PCA methods, in which the importance or significance of each time series batch 202A, 202B, 202C in the sequence of time series batches 202 is weighted. For example, in traditional PCA, all data points are treated equally (e.g., given equal weights). In contrast, in weighted PCA (i.e., WPCA), each time series batch 202A, 202B, 202C is assigned (or adjusted by) a weight that reflects its relative importance or contribution to the analysis. Weighted PCA can include multiple steps, such as, but not limited to, a data processing step, a weighting step, a covariance matrix calculation, an eigenvalue decomposition step, a principal component selection step, and a dimensionality reduction step.
[0047] How a time series batch is weighted by the API model 204 can be based on the importance of the temporal patterns. For example, the model 200 can be configured to weight certain trends identified in the time series batch 202 more heavily than other trends.
[0048] Model 200 integrates matrix sketching techniques within a neural network framework and is specifically tailored for real-time time series classification. As shown in Figure 2, this approach involves incorporating unsupervised matrix sketching within a supervised context. That is, matrix sketching can involve creating smaller, approximate representations (e.g., sketches or matrix sketches). Matrix sketching can be used to reduce the dimensionality of data while preserving important structural information. By incorporating matrix sketching within a supervised context, we can continuously update the sketch linked to a sequence of weighted time series streams. The dynamically updated sketch serves as a compact representation of the input time series, allowing for easy real-time classification into known time series classes at any timestamp. This innovative fusion of unsupervised matrix sketching and supervised neural networks addresses the demands of real-time processing while maintaining accuracy and efficiency for classification tasks.
[0049] 3 illustrates a block diagram of a streaming time-series classification model 300 that can be implemented by the computing system 100 described above with reference to FIG. 1 , according to an exemplary embodiment of the present disclosure. As illustrated, the model 300 can be provided with time-series data 302 as input and can generate an output 304 (e.g., a classification label, a prediction, or other output). The time-series data 302 can be raw time-series data obtained from a sensor, for example. In an exemplary implementation, the time-series data 302 can be obtained from an environmental sensor connected to a turbine system (e.g., a gas turbine system or a wind turbine system). The time-series data can be a time-indexed sequence of data, such as temperature data, pressure data, velocity data (e.g., rotational speed data, translational speed data, and / or flow velocity data), or other data.
[0050] The time series data 302 may be provided to the model 300 in a sequence of time series batches 306 (shown in dashed boxes in FIG. 3 ). Each time series batch 306 may be a portion of the entire time series data 302. In some implementations, each time series batch 306 may be an aggregated or compiled portion of the raw time series data 302. The sequence of time series batches 306 may be provided to the model 300 sequentially, with each new time series batch continuing where the previous time series batch left off. For example, a first time series batch may be provided to the model 300, followed by a second time series batch, followed by a third time series batch, and so on. As an example, for 10 or more minutes of time series data, the first time series batch may contain 1 minute of time series data, the second time series batch may contain 2 minutes of data, and the third time series batch may contain 3 minutes of time series data.
[0051] In some implementations, the sequence of time series batches 306 can be provided to an API model 308. The API model 308 generates a sequence of time series sketches 316 (e.g., U1, U2, ... U) based on the sequence of time series batches 306 of the time series data 302. N ) can be generated. The sequence of time series batches 306 can be position-coded, reduced, and weighted by the API model 308 to generate a sequence of time series sketches 316. Each time series sketch 316 (e.g., U1, U2, ... U N ) is a compact summary of each time series batch 306 (e.g., a compact summary of a data aggregation, such as a compact summary of a time series batch 306 that aggregates a portion of the raw time series data 302). In such an implementation, required information can be obtained from each time series batch 306, and other information can be extracted. In this way, a set of time series sketches 316 (e.g., U1, U2, ... U N) to calculate a metric can be advantageous because it can be less expensive than calculating an exact value. In an exemplary embodiment, each time series sketch may be a matrix sketch. However, in other embodiments, each time series sketch may be a collection of vectors or a graphical representation of the time series data. For example, in an embodiment where the time series data describes temperatures within a gas turbine engine (e.g., the temperature of one of the compressor section, the combustion section, and the turbine section), the time series sketch may be a compact representation of the temperature data (e.g., a compact matrix or a compact graphical representation).
[0052] For example, in various implementations, the API model 308 may include a position encoding model 310 of the API model 308. The position encoding model 310 may be provided with the sequence of time series batches 306, and the position encoding model 310 may generate, as output, a sequence of encoded time batches 312. The sequence of encoded time batches 312 may then be provided to a projection model 314, which may project the sequence of encoded time batches 312 into a sketch space. In particular, the projection model 314 may implement matrix sketching to project the sequence of encoded time batches 312 into the sketch space. The projection model 314 may generate a sequence of time series sketches from the sequence of encoded time batches 312.
[0053] In many embodiments, the API model 312 can include an incremental update model 318. A weight can be assigned to the new time series batch 306 (i.e., the new time series batch 306 can be modified or adjusted by the weight). The weight assigned to the new time series batch 306 can be based at least in part on a previous time series sketch in the sequence of time series sketches 316. The weight assigned to the new time series batch 306 can be a weighted sum of two elements, where a first element approximates the covariance of the observed data and a second element represents the covariance of the new batch.
[0054] In many embodiments, the model 300 can generate an output 304 (e.g., a classification label, a prediction, or other output) for each time series sketch 316 in a sequence of time series sketches 316. For example, the API model can feed the sequence of time series sketches 316 to a classification block 320 to predict a class probability distribution. As shown, the classification block 320 can include a fully connected layer and a softmax function. In various implementations, the model 300 can generate a label 322 (in real time) for each time series sketch 316 as the output 304. In other implementations, the model 300 may generate a prediction (in real time) for each time series sketch 316 as the output 304.
[0055] As one non-limiting example, model 300 may be provided with sensor data indicative of one or more parameters associated with a power generation system (such as a wind turbine or gas turbine engine) in a time-indexed form (e.g., time-series data). The sensor data may be pressure data (e.g., data from a pressure sensor), temperature data (e.g., data from a temperature sensor), vibration data (e.g., data from a vibration sensor or accelerometer), and / or speed data. The output 304 of model 300 may be a classification label indicative of a type of failure mode of the power generation system. For example, in an embodiment in which the power generation system is a gas turbine engine, failure mode classification labels may be output from model 300 including a bearing vibration fault event (e.g., bearing vibration exceeds a threshold for a predetermined period of time), a temperature fault event (e.g., temperature prevailing in the exhaust section of the gas turbine engine exceeds a threshold for a predetermined period of time, thereby indicating a leak in the combustion section), and an overspeed fault event (e.g., a rotor of the gas turbine engine spins at a speed above a failure threshold for a predetermined period of time). In an embodiment where the power generation system is a wind turbine system, the failure mode classification labels output from model 300 may include a blade crack event or other failure event.
[0056] The time series data can be analyzed in real time by the model 300. For example, sensor data can be fed to the model 300 in real time, which can generate classification label output 304 in real time. This allows the model 300 to determine a fault event as it occurs (or shortly thereafter). Rapid determination of a fault event is advantageous in that the fault event can be addressed quickly, thereby minimizing downtime of the power generation system.
[0057] In some embodiments, the model 300 may include a temporal convolutional network or model (TCN) 324. The sequence of time series batches 306 may be fed to the TCN 324 to generate a sequence of TCN-generated embeddings (e.g., as an output of the TCN 324). In various implementations, the TCN 324 may employ a sliding window to extract nonlinear features from the sequence of time series batches 306 before feeding them to the API model 308. For example, the model 300 may provide the sequence of TCN-generated embeddings to the API model 308 as input. The embeddings generated by the TCN may be a trained representation of the input data (e.g., time series data 302) in which specific nonlinear features have been extracted from the time series data by the TCN 324.
[0058] FIG. 4 shows a flowchart of a computer-implemented method 400 for analyzing time series data. Method 400 can be performed, for example, by the computing system 100 previously illustrated and described with reference to FIG. 1. At 402, method 400 can include providing a sequence of time series batches of time series data to an Attentive Power Iteration (API) model. At 404, method 400 can include generating, by the API model, a sequence of time series sketches based on the sequence of time series batches of time series data. At 406, method 400 can further include assigning a weight to a new time series batch in the sequence of time series batches based at least in part on a previous time series sketch in the sequence of time series sketches. At 408, method 400 can further include generating an output for each time series sketch in the sequence of time series sketches.
[0059] 5 illustrates a process 500 representing the API model 308 (FIG. 3) according to an embodiment of the present disclosure. In particular, the API model 308 can generate a sequence of time-series sketches according to the process illustrated in FIG.
[0060] In the first line, for each batch (e.g., X1, ...X N ) is processed in order.
[0061] In line 2, positional encoding is applied according to the order of the time series. This differs from the permutation-invariant samples in a typical data stream. Positional information within the time series is crucial. Process 500 achieves this by using Continuous Augmented Positional Embeddings (CAPE). CAPE is advantageous because it is computationally efficient and robust when dealing with input lengths.
[0062] In line 3, the encoded batch is fed to the fully connected layer f v The trained sketches are transformed by and then activated by relu, which introduces nonlinearity and improves the expressiveness of the trained sketches.
[0063] In lines 4 to 8, for the first batch (e.g., X1), the initial sketch vector U1 is updated until the initial sketch vector converges. In lines 9 to 18, for subsequent batches (i.e., not the first batch), the sketch vector U i is updated until it converges. In line 10, the sketch U i is initialized with the previous sketch U.
[0064] Lines 11-15 incorporate an attention model that assigns weights to the new batch relative to the most recent sketches. The update in line 15 is a weighted sum of two components: the first component approximates the observed data covariance, and the second component represents the covariance of the new batch. These weights come from the attention model in lines 11-13.
[0065] Notably, the normalization steps in lines 7 and 16 of Algorithm 2 have two advantages: 1) they prevent the sketch size from becoming excessively large, improving the stability of convergence, and 2) they ensure the robustness of classification against the length of the time series. Without normalization, the sketch size is correlated with the length of the time series, which reduces the similarity between short and long time series of the same class.
[0066] Referring now to FIG. 6, a flowchart of a computer-implemented method 600 for analyzing time series data is shown. Method 600 can be performed, for example, by the computing system 100 shown and described above with reference to FIG. 1. Method 600 can include, at 602, obtaining or receiving time series data as input. The time series data can describe a series of data points based on time. Method 600 can further include, at 604, processing the time series data with a time convolution model to generate one or more time series embeddings. The time convolution model can be the TCN 324 described above with reference to FIG. 3. Method 600 can further include processing the one or more time series embeddings with a representation generation model to generate an output representation. The output representation can describe a graphical representation of the series of data points based on time. The representation generation model can convert the time series embedding into a compact, informative embedding representation (e.g., by dimensionality reduction). In many implementations, the output representation can be a matrix sketch. The matrix sketch can reduce the size of the time series embedding while preserving key properties. Method 600 can further include (608) processing the output representation with a classification model to generate classification labels for the time series data. In many embodiments, the time series data is generated using one or more sensors associated with a power generation system (such as a gas turbine engine). Additionally, in various implementations, the classification labels include anomaly detection classifications. Anomaly detection classifications can identify trends or data points that deviate significantly from a norm of the time series data. For example, in time series data related to one or more parameters of a gas turbine engine, anomaly detection classifications can identify various trips (or fault events) of the gas turbine engine, such as overspeed, high exhaust section temperature, etc.
[0067] The computing system 100, models 200, 300, method 400, process 500, and method 600 described above with reference to Figures 1-6 offer numerous advantages over known models. For example, known artificial intelligence (AI) models (particularly common neural networks) often require significant memory and computational energy, making them unsuitable for deployment on resource-constrained devices. The models 200, 300 described above are designed to operate with limited computational memory and energy resources while maintaining high processing speeds. Configurations incorporating the models described herein are capable of performing complex data processing tasks on devices with limited computational resources, and can perform the tasks with reduced latency compared to conventional techniques. Experimental results demonstrate the superior performance of models 200, 300, demonstrating not only excellent accuracy but also significantly reduced energy consumption, increased execution speed, and reduced computational complexity.
[0068] The above-described models 200 and 300 advantageously utilize API models 204 and 308, enabling real-time processing on resource-constrained devices. Models 200 and 300 continuously update a compact representation of the entire time series, improving classification (e.g., output) accuracy while saving energy and processing time. In particular, models 200 and 300 excel in streaming scenarios and can make rapid decisions without requiring access to the entire time series data. Models 200 and 300 excel in classification accuracy and energy efficiency, consuming over 70% less power and completing tasks three times faster than benchmarks. This research advances real-time responsiveness, energy savings, and operational efficiency in constrained devices, contributing to the optimization of various applications.
[0069] This specification uses examples to disclose the invention, including the best mode, and also to enable any person skilled in the art to practice the invention (e.g., to make and use devices or systems, and to perform incorporated methods). The patentable scope of the invention is defined by the claims, and may include other examples that occur to those skilled in the art. Such other examples are intended to be within the scope of the claims if they contain structural elements that do not differ from the claim language, or if they contain equivalent structural elements that do not differ substantially from the claim language.
[0070] Other aspects of the invention are provided by the subject matter of the following embodiments.
[0071] [Embodiment 1] 1. A computer-implemented method for analyzing time series data, the method comprising: obtaining, by a computing system including one or more processors, a sequence of time-series batches of time-series data; generating, by said computing system and API (attentive power iteration) model, a sequence of time series sketches based on the sequence of time series batches of said time series data; assigning, by the computing system, a weight to a new time series batch in the sequence of time series batches based at least in part on a previous time series sketch in the sequence of time series sketches; and generating, by the computing system, an output for each time series sketch in the sequence of time series sketches. A computer-implemented method comprising: [Embodiment 2] 2. The computer-implemented method of claim 1, further comprising providing a sequence of time series batches to a temporal convolutional network (TCN) to generate a sequence of embeddings generated by the TCN, wherein the TCN extracts nonlinear features from the sequence of time series batches by performing sliding window processing on the time series data. [Embodiment 3] 3. The computer-implemented method of claim 1 or 2, further comprising providing a sequence of embeddings generated by a TCN to the API model as input. [Embodiment 4] A computer-implemented method according to any one of embodiments 1 to 3, further comprising supplying a sequence of time series batches to a position encoding model of the API model, the position encoding model generating a sequence of encoded time batches from the sequence of time series batches. [Embodiment 5] A computer-implemented method according to any one of embodiments 1 to 4, further comprising supplying a sequence of encoded time series batches to a projection model, the projection model generating a sequence of time series sketches from the sequence of encoded time series batches. [Embodiment 6] A computer-implemented method according to any one of embodiments 1 to 5, further comprising supplying a sequence of time series sketches to a classification block to predict class probability distributions. [Embodiment 7] A computer-implemented method according to any one of embodiments 1 to 6, wherein generating an output further comprises generating a classification label for each time series sketch in the sequence of time series sketches. [Embodiment 8] Producing the output further comprises: A computer-implemented method according to any one of embodiments 1 to 7, which generates a prediction for each time series sketch in a sequence of time series sketches. [Embodiment 9] A computer-implemented method according to any one of embodiments 1 to 8, wherein the API model generates a sequence of time-series sketches according to the process shown in FIG. 5. [Embodiment 10] 1. A computing system for analyzing time series data, the system comprising: one or more processors; and one or more non-transitory computer-readable media, the one or more non-transitory computer-readable media collectively storing instructions that, when executed by the one or more processors, cause the computing system to perform operations, the operations including: obtaining a sequence of time-series batches of said time-series data; generating a sequence of time series sketches based on a sequence of time series batches of time series data by an Attentive Power Iteration (API) model; assigning a weight to a new time series batch in the sequence of time series batches based at least in part on a previous time series sketch in the sequence of time series sketches; and Producing an output for each time series sketch in a sequence of time series sketches a computing system including: [Embodiment 11] 11. The computing system of claim 10, further comprising providing a sequence of time series batches to a temporal convolutional network (TCN) to generate a sequence of embeddings generated by the TCN, wherein the TCN extracts nonlinear features from the sequence of time series batches using a sliding window. [Embodiment 12] 12. The computing system of claim 10 or 11, further comprising providing the API model with a sequence of embeddings generated by a TCN as input. [Embodiment 13] A computing system described in any one of embodiments 10 to 12, further comprising supplying a sequence of time series batches to a position encoding model of the API, wherein the position encoding model generates a sequence of encoded time batches from the sequence of time series batches. [Embodiment 14] A computing system described in any one of embodiments 10 to 13, further comprising supplying a sequence of encoded time series batches to a projection model, wherein the projection model generates a sequence of time series sketches from the sequence of encoded time series batches. [Embodiment 15] A computing system described in any one of embodiments 10 to 14, further comprising supplying a sequence of time series sketches to a classification block to predict class probability distributions, the classification block including a fully connected layer and a softmax function. [Embodiment 16] Producing the output further comprises: A computing system according to any one of embodiments 10 to 15, comprising generating a classification label for each time series sketch in a sequence of time series sketches. [Embodiment 17] Producing the output further comprises: A computing system according to any one of embodiments 10 to 16, which generates a prediction for each time series sketch in a sequence of time series sketches. [Embodiment 18] A computing system described in any one of embodiments 10 to 17, wherein the API model generates a sequence of time-series sketches according to the process shown in FIG. 5. [Embodiment 19] A method according to the embodiments shown and described herein. [Embodiment 20] A computing system according to the embodiments shown and described herein. [Embodiment 21] 1. A computing system for time-based data classification, the system comprising: one or more processors; and 1. A system comprising: one or more non-transitory computer-readable media, the one or more non-transitory computer-readable media collectively storing instructions that, when executed by the one or more processors, cause the computing system to perform operations comprising: obtaining time series data, the time series data describing a series of data points based on time; processing the time series data with a temporal convolution model to generate one or more time series embeddings; processing the one or more time series embeddings with a representation generation model to generate output representations, the output representations describing a graphical representation of the series of data points based on time; and processing the output representations with a classification model to generate classification labels for the time series data. [Embodiment 22] A system described in any one of embodiments 10 to 21, wherein the output representation includes a matrix sketch. [Embodiment 23] A system described in any of embodiments 10 to 22, wherein the representation generation model includes one or more multi-head self-attention models and one or more power iteration models, and the one or more power iteration models are configured to process input data to generate one or more eigenvectors. [Embodiment 24] A system described in any one of embodiments 10 to 23, wherein the time series data is generated using one or more sensors associated with the turbine, and the classification labels include an anomaly detection classification. [Explanation of symbols]
[0072] 100 Computing Systems 102 User Computing Device 112 processors 114 memory 116 Data 118 Command 120 machine learning models 122 User Input Components 132 processors 134 memory 136 Data 138 Command 150 Training Computing System 152 processors 154 memory 156 Data 158 Command 160 Model Trainer 162 training dataset 180 Network 182 Environmental Sensors 184 Power Generation System 206 Time Series Sketch 208 Classification Model 210 Classification Labels 304 Output 308 API Model 310 Positional Coding Model 314 Projection Model 316 Time Series Sketch 318 Incremental Update Model 320 Classification Block 322 Label 324 Model (TCN) 400 ways 500 processes 600 ways
Claims
1. 1. A computer-implemented method for analyzing time series data, the method comprising: obtaining, by a computing system including one or more processors, a sequence of time-series batches of time-series data; generating, by said computing system and API (attentive power iteration) model, a sequence of time series sketches based on the sequence of time series batches of said time series data; assigning, by the computing system, a weight to a new time series batch in the sequence of time series batches based at least in part on a previous time series sketch in the sequence of time series sketches; and generating, by the computing system, an output for each time series sketch in the sequence of time series sketches. A computer-implemented method comprising:
2. 2. The computer-implemented method of claim 1, further comprising providing the sequence of time-series batches to a temporal convolutional network (TCN) to generate a sequence of embeddings generated by the TCN, wherein the TCN extracts nonlinear features from the sequence of time-series batches by performing sliding window processing on the time-series data.
3. The computer-implemented method of claim 2 , further comprising providing a sequence of embeddings generated by a TCN as input to the API model.
4. 2. The computer-implemented method of claim 1, further comprising: supplying the sequence of time series batches to a position encoding model of the API model, the position encoding model generating a sequence of encoded time batches from the sequence of time series batches.
5. 5. The computer-implemented method of claim 4, further comprising: feeding the sequence of encoded time series batches to a projection model, wherein the projection model generates a sequence of time series sketches from the sequence of encoded time series batches.
6. The computer-implemented method of claim 1 , further comprising feeding the sequence of time-series sketches to a classification block to predict class probability distributions.
7. Producing the output further comprises: The computer-implemented method of claim 1 , comprising generating a classification label for each time series sketch in the sequence of time series sketches.
8. Producing the output further comprises: The computer-implemented method of claim 1 , further comprising generating a prediction for each time series sketch in the sequence of time series sketches.
9. 1. A computing system for analyzing time series data, the system comprising: one or more processors; and one or more non-transitory computer-readable media, the one or more non-transitory computer-readable media collectively storing instructions that, when executed by the one or more processors, cause the computing system to perform operations, the operations including: obtaining a sequence of time-series batches of said time-series data; generating a sequence of time series sketches based on a sequence of time series batches of time series data by an Attentive Power Iteration (API) model; assigning a weight to a new time series batch in the sequence of time series batches based at least in part on a previous time series sketch in the sequence of time series sketches; and Producing an output for each time series sketch in a sequence of time series sketches a computing system including:
10. 10. The computing system of claim 9, further comprising providing the sequence of time-series batches to a temporal convolutional network (TCN) to generate a sequence of embeddings generated by the TCN, wherein the TCN extracts nonlinear features from the sequence of time-series batches using a sliding window.
11. The computing system of claim 9 , further comprising providing the API model with a sequence of embeddings generated by a TCN as input.
12. 10. The computing system of claim 9, further comprising: providing a sequence of time-series batches to a position encoding model of the API, the position encoding model generating a sequence of encoded time batches from the sequence of time-series batches.
13. 13. The computing system of claim 12, further comprising: feeding the sequence of encoded time series batches to a projection model, the projection model generating a sequence of time series sketches from the sequence of encoded time series batches.
14. 10. The computing system of claim 9, further comprising: feeding the sequence of time series sketches to a classification block to predict class probability distributions, the classification block including a fully connected layer and a softmax function.
15. Producing the output further comprises: The computing system of claim 9 , further comprising generating a classification label for each time series sketch in the sequence of time series sketches.