Predicting device usage with motif and combinatorial modeling of univariate time series datasets
By grouping and combining the univariate time series data sets, and generating mockup sequence diagrams and directed graphs, the problem of difficult-to-predict equipment power requirements in the data center is solved, and the optimization of equipment operation efficiency and resource management is achieved.
Patent Information
- Application Number
- CN202311293097.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2023-03-10
- Filing Date
- 2023-10-08
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2043-10-08
AI Technical Summary
It is difficult to effectively predict equipment power requirements in data center management, resulting in improper power utilization, increasing costs and reducing equipment efficiency and reliability.
Group univariate time series data sets through unsupervised machine learning models, generate motifs, use motif sequence diagrams and directed graphs for combined modeling, predict equipment usage, and dynamically adjust power supply.
It realizes dynamic forecasting of equipment power requirements, improves equipment operation efficiency, extends equipment life, reduces carbon footprint, and optimizes resource management of data centers.
Smart Images

Figure CN118626242B_ABST
Abstract
Description
Background Art
[0001] As computer systems increase in complexity, size, and processing power, the processing performed by these systems continues to grow. Monitoring systems have become increasingly popular in an attempt to manage the applications executed by computer systems and improve their overall efficiency. However, this is a difficult task. Data is being created at an ever-increasing rate, making it difficult to review. When data review is passed to a third party, the data received by the third party may not have access to all environmental data in the computer system.
[0002] Furthermore, data room infrastructure is often complex and difficult to track. Proper energy metrics are needed to determine the power requirements of servers and the broader data center. Standard input power generation is often a one-size-fits-all approach, increasing costs and reducing data center performance and health. Such issues require pragmatic solutions to help data center managers more efficiently manage their resources and reduce the overall costs associated with operating infrastructure. BRIEF DESCRIPTION OF THE DRAWINGS
[0003] The present disclosure according to one or more embodiments is described in detail with reference to the accompanying drawings, which are provided for purposes of illustration only and merely depict typical or example embodiments.
[0004] Figure 1 Computer systems, monitored devices, and user devices according to some examples of the present disclosure are shown.
[0005] Figure 2 A process for determining forecast predictions using topic and combination modeling according to some examples of the present disclosure is shown.
[0006] Figure 3 Examples of time series data according to some examples of the present disclosure are provided.
[0007] Figure 4 Examples of minimum and maximum values of time series data according to some examples of the present disclosure are provided.
[0008] Figure 5 Examples of time series data with identified outliers according to some examples of the present disclosure are provided.
[0009] Figure 6 Examples of time series data with identified clusters according to some examples of the present disclosure are provided.
[0010] Figure 7 A process for generating one or more cluster data according to some examples of the present disclosure is shown.
[0011] Figure 8Examples of directed graphs generated from motifs according to some examples of the present disclosure are provided.
[0012] Figure 9 is an example of a graph comparing actual time series data to forecasted data according to some examples of the present disclosure.
[0013] Figure 10 are example computing components that may be used to implement various features of the embodiments described in this disclosure.
[0014] Figure 11 Depicted is a block diagram of an example computer system in which various embodiments described herein may be implemented.
[0015] The drawings are not exhaustive and do not limit the disclosure to the precise forms disclosed. DETAILED DESCRIPTION
[0016] Data center managers are often faced with the challenge of deciding when and how often to introduce new workloads, either once or on a regular basis. Effective scheduling of new workloads can help save operating costs by avoiding peak usage periods and contribute to the long-term operating efficiency of servers and their peripherals. For example, overloading a power supply unit can cause overheating, which can adversely affect the functionality of the server. Conversely, underutilization of a power supply unit at different times can cause inefficiencies and may place undue stress on the power supply unit during other times. Limiting the amount of input power can help reduce waste and avoid unnecessary costs in running a server or data center. Conventional approaches to solving these problems are less than optimal than expected. For example, in one conventional approach, the user must provide the configuration of the server. This raises technical issues that limit the dynamic forecasting of the server's power requirements. Therefore, power requirements cannot be dynamically forecasted based on the historical patterns of power usage of the server, and the use of the equipment is more restricted.
[0017] Through the present disclosure, the power consumption of each device can be determined and used to collect the power usage requirements of the device. This data can be used to determine the capacity of the infrastructure required for the device or data center as a whole. Such predictions can also help determine the efficiency of the power supply. With the features of the present disclosure, the power requirements of the device can be dynamically predicted based on the historical pattern of its power usage. Devices that consume more power or less power (for example, than the optimal value) can be marked so that appropriate action can be taken. The device can operate more efficiently, which can result in a longer service life and reliability of the device and its components. The carbon footprint of the device or data center as a whole can also be reduced.
[0018] Forecasts can be based on the identification of multiple motifs in univariate time series data and combinatorial modeling of devices or combinations of devices in a data center. Motifs represent similarities across multiple clusters in time series data, which can be combined based on data signature similarities. The combined patterns identified in the similarities across clusters are called "motifs."
[0019] The identified similarities of consecutive data points across multiple clusters can be determined using various methods. For example, the system can implement an unsupervised machine learning model that is trained to identify and group consecutive sets of data points of a time series dataset into a first cluster. In some examples, the unsupervised machine learning model is trained to identify similar data signatures in each cluster and cross-match the data signatures of the clusters to form multiple motifs. In either sense, the system can extract motifs from a univariate time series dataset (e.g., using a customized K-Means algorithm or other clustering algorithm for extraction). The extraction can identify and cluster similar consecutive data points and data patterns in a univariate time series dataset and output motifs (e.g., representing repeating patterns and subsequences).
[0020] Multiple motifs can be used to generate data definitions, motif sequence diagrams, directed graphs, or other combinations of data points. These data points can be combined with other data points generated by a second machine learning model through a summation process. The output of the summation process can be used to predict device usage or other forecasts for monitored devices in the data center.
[0021] Technical advantages are realized throughout the application. The system described herein can more efficiently plan existing workloads or introduce new workloads at more optimal time periods. For example, this can prevent overloading equipment that may cause overheating, or mark equipment in the data center that consumes less power so that appropriate action can be taken. Overall cost savings and improved health and lifespan of equipment can be achieved. The solution of the present disclosure can be implemented in a variety of ways, including as software as a service (SaaS), as platform as a service (PaaS), in a cloud-based or cloud computing environment using underlying hardware components, or in equipment in an IT data center.
[0022] Figure 1Computer systems, monitored devices, and user devices according to some examples of the present disclosure are shown. In this example, the computer system, shown as a dedicated computer system, is labeled computer system 100. Computer system 100 is configured to interact with (multiple) monitored devices 140 using a processor 104, a memory 105, and a machine-readable storage medium 106. (Multiple) monitored devices 140 can interact with one or more user devices 142. Computer system 100 can be implemented as a cloud system, a data center computer server, a service, etc., although not every embodiment of the present disclosure requires such limitation.
[0023] Processor 104 may include a general or special purpose processing engine such as, for example, a microprocessor, controller, or other control logic. Processor 104 may be connected to a bus, although interaction with other components of computer system 100 or communication with the outside world may be facilitated using any communication medium.
[0024] Memory 105 may include random access memory (RAM) or other dynamic storage for storing information and instructions to be executed by processor 104. Memory 105 may also be used to store temporary variables or other intermediate information during execution of instructions to be executed by processor 104. Memory 105 may also include read-only memory (“ROM”) or other static storage devices coupled to the bus for storing static information and instructions for processor 104.
[0025] The machine-readable storage medium 106 is configured as a storage device that may include one or more interfaces, circuits, and modules for implementing the functionality discussed herein. The machine-readable storage medium 106 may carry one or more sequences of one or more instruction processors 104 for execution. Such instructions embodied on the machine-readable storage medium 106 may enable interaction with (multiple) monitored devices 140 to perform features or functions of the disclosed technology as discussed herein. For example, the interfaces, circuits, and modules of the machine-readable storage medium 106 may include, for example, a data processing component 108, a time series centroid component 110, a distance component 112, a clustering component 114, a detection component 116, a model training component 118, a data point extraction component 120, a summation component 122, a consumption prediction component 124, and a monitored device action component 126.
[0026] The data processing component 108 is configured to receive a time series data set and store the data in the time series data store 130. The time series data set may be received from sensors of the monitored device 140 and include various data signatures at different times. In some examples, the data received from the monitored device 140 is limited to data that a third party responsible for the monitored device 140 is willing to provide. Therefore, when the time series data set is received, the complete data history or composition of the data may not be available.
[0027] Illustrative examples of time series data sets are provided. For example, the time series data set may correspond to the monitored device 140 and its input power, where the input power is provided to the monitored device 140 and is sampled in hours. Another example of a time series data set may correspond to the processor utilization of the monitored device 140 when the processor is executing instructions for multiple software applications, such as Figure 3 shown.
[0028] The time series centroid component 110 is configured to determine each centroid of the time series dataset in anticipation of grouping a set of consecutive data points in the data into a first cluster of a plurality of clusters. For example, an unsupervised machine learning model can group a set of consecutive data points of the time series dataset into a first cluster by algorithmically determining the centroid of each consecutive grouping.
[0029] The centroids can be initialized using a variety of methods, including random selection (e.g., a random selection of data points from a received dataset), a trained machine learning model, a K-Means clustering algorithm or K-Means++ (e.g., by assigning each data point to the cluster with the value closest to the mean rather than the minimum / maximum value), or by initializing the centroids using outliers in a customized K-Means clustering algorithm.
[0030] In some examples, a clustering algorithm (e.g., a customized K-Means algorithm) can be configured to analyze small subsets of a univariate time series data set to find local minima and maxima of the subset of the time series data set, which can then be associated with the nearest centroid using a skewed Euclidean distance function. This creates a certain number of clusters, and then the positions of those nearest centroids are calculated again until the data associated with the centroid no longer changes, but remains consistent. These ultimately become the clusters that are analyzed to identify "motifs." As discussed herein, motifs represent similarities across multiple clusters that can be combined based on data signature similarities across the computer system 100.
[0031] The customized K-Means clustering algorithm can take into account the linear manner in which the time series data set is recorded when determining each cluster centroid (e.g., consecutive data points). In this sense, the time series centroid component 110 can group data points that are close to each other in time, so that they are more likely to be grouped with other readings from the same application, which are stored in the detected grouped data store 132. Clusters can include centroids, outliers, and local maxima and local minima. Local maxima and local minima can be determined to separate clusters into smaller clusters, where each smaller cluster can include a significant minimum or maximum point.
[0032] In order to find local minima and maxima on a very small subset of the time series data, the time series centroid component 110 can use an outlier centroid initialization method or other distance algorithm. For example, first, the time series centroid component 110 can assume that the number of recently found centroids is "k". Next, each data point can be associated with the nearest centroid (e.g., using a skewed Euclidean distance function). This can divide the data points into k clusters. Then, using the clusters, the time series centroid component 110 can recalculate the position of the centroid using a skewed Euclidean distance function or other distance algorithm. Any of these steps can be repeated until there is no longer a change in the membership of the data point, including dividing the data point into "k" clusters and calculating the position of the centroid using a skewed Euclidean distance function until there is no longer a change. The time series centroid component 110 can provide or output data points with cluster membership.
[0033] In some examples, the time series centroid component 110 can determine the number of clusters in an automatic and dynamic process. The clusters may not be manually defined. In this example, the number of clusters can be based on the number of maximum and minimum points in the time series data set, with the maximum and minimum points each being grouped as a set of consecutive data points in the time series data set. Illustrative examples of minimum and maximum values are shown in Figure 4 is provided.
[0034] In some examples, the minimum and maximum values determined by the time series centroid component 110 can correspond to actual values from the time series data set. The centroid of the cluster may not be selected from values not included in the data set (e.g., mean or average). Instead, the minimum and maximum values can correspond to actual data points received from the monitored device 140.
[0035] The distance component 112 is configured to determine the distance between each data point in the cluster and the centroid relative to a time series constraint (e.g., along a linear time series). This can improve upon standard distance functions that may incorrectly cluster data points without considering the linear nature of time series data.
[0036] The distance formula may determine that a change on the time axis (e.g., x-axis) is weighted less than the same change on the performance metric axis (e.g., y-axis) so that data points may be clustered along the time axis. An example of the type of distance formula that may be used is:
[0037]
[0038] Where "4" can be replaced with an even "n" value greater than or equal to 4. This value can be configured and customized by the user operating the user device 142. This value can be an exponential multiplier of the "y" portion of the formula, typically set to a fractional value. When the fractional value is raised to a higher "n" power, the "y" portion of the formula may become smaller. More weights can be associated with the x-axis, which corresponds to time values and can group time series data along a linear path.
[0039] The clustering component 114 is configured to determine one or more clusters. For example, the clustering component 114 can receive the local minima and maxima of a subset of the time series data set and the nearest centroid (determined from the time series centroid component 110). The two data sets can be correlated to form multiple clusters, and then the position of the nearest centroid can be recalculated until the calculated centroid no longer changes but remains consistent. These become analyzed to identify clusters of motifs, where additional clusters can be grouped with existing motifs that already contain similar clusters.
[0040] The detection component 116 is configured to implement a dynamic time warping (DTW) process on the defined clusters to detect similarities between clusters (e.g., in terms of data points forming peaks, valleys, or other shapes within each cluster) and generate one or more motifs. The DTW process can calculate the best match between two time series data sets by measuring the similarity of the two data sets along a linear time sequence. The best match can meet various constraints, including, for example, that each index from the first sequence can match one or more indexes from the other sequence, and vice versa, that the first index from the first sequence can match the first index from the other sequence (but not necessarily its only match), that the last index of the first sequence can match the last index from the other sequence (but not necessarily its only match), or that the mapping of indices from the first sequence to indices from the other sequence can be monotonically increasing, and vice versa. The best match can also meet a minimum cost, where the cost is calculated as the sum of the absolute differences between the values of each matching pair of indices. Similar clusters can be grouped in a motif. Multiple motifs can be generated for each group of similar clusters.
[0041] In some examples, the detection component 116 may include a parameter "c" that determines a minimum threshold for accepting whether two subsequences are similar. The threshold "c" may be inversely proportional to the compression ratio. In other words, for higher compression ratios, the value of "c" will be lower.
[0042] The results can be stored as an in-memory object, for example, in the detected group data store 132. The dataset can include metadata for each motif and an index of similar subsequences for each motif. The metadata can include, for example, values along the x-axis (time) and y-axis (calculated value) and an index of the closest or consecutive data point members of the sequence for application at the monitored device. The compressed dataset representation can store this or other metadata, which can describe each of the multiple motifs and a time-based index of the clusters grouped into each corresponding motif.
[0043] The compressed dataset representation can be stored according to a data schema format (e.g., in JSON format or in the time series data store 130). Since motifs can define patterns in the data (e.g., corresponding to data signatures of applications at monitored devices 140), the patterns, rather than individual points of the data, can be stored in the dataset representation according to the data schema format. In other words, the dataset representation can represent the entire time series dataset in a compressed format to occupy a reduced memory capacity. Once the dataset representation for each motif is generated, the dataset representation can be used to generate a new, compressed dataset that uses less memory. The compressed dataset can be stored in place of the original univariate time series dataset, and in some examples, the original univariate time series dataset can be deleted to save memory space.
[0044] For each data point in a motif (e.g., a cluster of data points), the properties of the data schema may include a unique identifier defined for the motif, a time start, a time stop, and the number of points. The data schema may also define the frequency of each data point in the motif and the closest member of the data point.
[0045] Multiple attributes or values can be stored in a data structure, such as an array data structure. The data structure can include a dataset representation for each motif and can be adjusted using an accuracy value. The accuracy value can be set by a system administrator to identify the number of parameters defined for the dataset representation for each motif. A larger accuracy value (e.g., greater than or exceeding a threshold accuracy value) can define more details in the dataset representation with higher accuracy between the repopulated time series dataset and the original time series dataset, which may result in more memory space required to store the dataset representation and the repopulated data. A smaller accuracy value (e.g., less than a threshold accuracy value) can define less details in the dataset representation with lower accuracy between the repopulated time series dataset and the original time series dataset, which may result in less memory space required to store the dataset representation and the repopulated data.
[0046] The model training component 118 is configured to receive a time series data set as input to a machine learning model. For example, the model training component 118 receives and processes or transforms data stored in the time series data store 130. This can include detecting and removing outliers in a time series data set (e.g., processor usage, input power periodicity data, or software application execution at the monitored device 140). This includes linearly interpolating any missing values in the time series data set and removing outliers in the data using the model training component 118. In some examples, a pre-trained machine learning model can implement anomaly detection and outlier removal (e.g., a one-class SVM). The data can be resampled, and only the average power consumed by the monitored device 140 can be captured at each interval (e.g., every hour or every two hours or any suitable interval). The refined time series data can be the base data set for training and validating the output generated by the ML model.
[0047] The model training component 118 is configured to determine outputs from a time series data set, including trend (e.g., long-term direction), seasonality (e.g., calendar-related movements), cyclic (e.g., systematic), and residual (e.g., non-systematic or short-term fluctuations) components (referred to as "seasonality"). In this example, existing algorithms predict a smooth curve that attempts to balance the trend and seasonality components, but may ignore small maxima and minima, which are considered under the residual component. In the case of the power and CPU data for the monitored device 140, the residual component can have more meaning and include information about the impact of the application or workload running on the monitored device 140. Therefore, when attempting to predict time series data, the outputs corresponding to the trend, seasonality, cyclic, and residual components can (ultimately) be combined with the above-mentioned model data.
[0048] In one non-limiting example, the past 70 days of data are used to train and validate the output generated by the ML model (e.g., a forecast of usage of the monitored device 140). In another non-limiting example, the past 12 months of data are used. Techniques such as mean absolute percentage error and root mean square error can be used to validate the output. For example, from the base data set, the first 70 days of data can be used to train the model for the server, and after successful training, the next 20 days of data can be used for validation. This process can be repeated on multiple monitored devices 140. In some examples, a parameter for the past "X" days of data can be set to the minimum time period of data required for processing. No data processing can be performed until "X" days of data are included.
[0049] Challenges in selecting and implementing training methods include long training times due to limited hardware resources. One training approach to overcome such challenges is to distribute training across multiple virtual machines (VMs) to generate models for each monitored device 140 and retrain using a warm-start strategy to shorten retraining time. Training can be optimized using scripts written to dynamically select the best set of hyperparameters for each model. In a non-limiting example, hyperparameters with the following combinations of parameters and their values can be used:
[0050] Changepoint_prior_scale: [0.001, 0.01, 0.05]. This parameter determines the flexibility of the trend, in particular how much the trend changes at trend change points. If too small, the trend may underfit the data and variance that should have been modeled using trend changes may instead be processed using a noise term. If too large, the trend may overfit and may model annual seasonality. The default value of 0.05 is initially recommended for many time series datasets and can be tuned over time. An example range could be around [0.001, 0.5].
[0051] Seasonality_prior_scale: [0.01, 0.1, 1.0, 10]. This parameter controls the flexibility of the seasonality. Again, large values allow the seasonality to adapt to large fluctuations, and small values reduce the amplitude of the seasonality. The default value is 10, which applies essentially no regularization. An example range for tuning this parameter might be around [0.01, 10]; when set to 0.01, one should find that the amplitude of the seasonality is forced to be very small.
[0052] Daily_seasonality: [True, False].
[0053] Growth: Logistics.
[0054] The model training component 118 is also configured to train an ML model using the dataset(s) to obtain predicted data for the monitored device 140 for the next or future time period. Various ML models or algorithms can be used for this purpose, including FBProphet, SARIMA / SARIMAX, Holt-Winter-ES, and Gated Recurrent Unit (GRU) networks. These models have appropriate training time and size to work with the techniques of this disclosure. In one example, the training-test split is 80%-20%.
[0055] FBProphet is open source software that can be implemented by the model training component 118. FBProphet is a process for forecasting time series data based on an additive model, where a nonlinear trend is accommodated with annual, weekly, and daily seasonality, along with holiday effects. FBProphet works well with time series datasets or any other dataset with strong seasonal effects and multi-seasonal historical data. FBProphet is robust to missing data and shifts in the trend and generally handles outliers well.
[0056] SARIMA is a class of statistical models for analyzing and predicting time series data. SARIMA is a generalization of another model called "autoregressive moving average" and adds the concept of integration. The parameters of the SARIMA model are defined as p, d, and q. "p" is the number of lagged observations included in the model, also known as the "lag order." "d" is the number of differences in the original observations, also known as the "degree of difference." "q" is the size of the moving average window, also known as the "order of the moving average." The python package "pmdarima" (Auto-ARIMA) can be used to find the correct set of parameters corresponding to (p, d, q) that exhibits a low AIC (Akaike Information Criteria) value for each monitored device 140.
[0057] Holt-Winters Exponential Smoothing (ES) can be used to forecast time series data that exhibit both trend and seasonal variation. Although Holt-Winters-ES is a relatively simple model, it can be a powerful forecasting algorithm. Holt-Winters-ES can handle seasonality in a data set simply by calculating the central value and then adding or multiplying the central value with the slope and seasonality, given the correct set of selected parameters for selection:
[0058] Level L t =α(y t -S t-s )+(1-α)(L t-1 +bt-1 );
[0059] Trend b t =β(L t -L t-1 )+(1-β)b t-1
[0060] Season S t =γ(y t -L t )+(1-γ)S t-s
[0061] Prediction F t+k =L t +kb t +S t+k-s
[0062] The Gated Recurrent Unit (GRU) network is a gating mechanism in a recurrent neural network (RNN) that uses connections through a sequence of nodes to perform memory-related machine learning tasks. GRUs can also solve the vanishing gradient problem posed by standard recurrent neural networks (RNNs). To address the vanishing gradient problem of standard RNNs, GRUs use update gates and reset gates. In some examples, the two gates are two vectors that determine which information should be passed to the output. The gates can be trained to retain information from long periods in the past without minimizing the details collected in the time series dataset or removing information that is not relevant to the forecast.
[0063] Using any of these or other models, the model training component 118 can generate forecast predictions as a dataset that can be stored as an in-memory object, for example, in the detected group data store 132. Forecast predictions for multiple motifs and compressed dataset representations can be stored therein.
[0064] In some examples, the forecast predictions can all be stored according to a data schema format (e.g., in JSON format). Because the forecast predictions can define patterns in the data (e.g., corresponding to a data signature of a seasonal trend at the monitored device 140), the patterns, rather than individual points of the data, can be stored in the data set representation according to the data schema format. Attributes of the data schema can include a unique identifier defined for each data signature, a time start, a time stop, and the number of points. The number of attributes or values can be stored in a data structure, such as an array data structure. The data structure can include a data set representation for each forecast prediction and can be adjusted using an accuracy value.
[0065] In some examples, the model training component 118 is configured to train a first machine learning model that generates motif forecast predictions and also train a second machine learning model that generates seasonal forecast predictions in parallel. In other words, the two machine learning models can be trained in parallel. The outputs from the first and second machine learning models can be provided to a summation process performed by the summation component 122.
[0066] Various machine learning models such as those described above can learn and forecast the power consumption or other metrics of the monitored device 140 for the next or future time period (e.g., per forecast prediction), as determined by the consumption forecast component 124. A check can be included to see if the forecast prediction was successful, and if so, the forecast prediction (e.g., in a non-limiting example, for the next 20 days or the next 30 days) can be passed to the consumption forecast component 124.
[0067] The data point extraction component 120 is configured to access the detected group data store 132 and determine a dataset representation of each of the plurality of motifs and a forecast prediction in a data schema format (eg, JSON format).
[0068] In some examples, the data point extraction component 120 is configured to generate a motif sequence graph for each motif in the compressed dataset representation (e.g., stored according to a data schema format). The motif sequence graph can be used to represent multiple sequences within multiple motifs using edges representing homology between segments. In some examples, multiple sequences can be represented by the same thread if there are multiple possible paths when traversing the thread in the sequence graph. In this way, a motif sequence graph can be created representing multiple motifs, where each motif corresponds to a path through the graph.
[0069] In some examples, the data point extraction component 120 is configured to generate a directed graph of motifs in a compressed dataset representation (e.g., stored according to a data schema format). The directed graph can assign weights to arrows, edges, and nodes to help identify the probability of traversing a particular sequence in the directed graph within a particular time interval. An illustrative directed graph utilizes Figure 8 is provided.
[0070] The summation component 122 is configured to aggregate the forecast predictions for each time point in the time series dataset using both the values generated from the first machine learning model (e.g., the motif forecast predictions) and the second machine learning model (e.g., the seasonal forecast predictions). Aggregation of the forecast prediction values can be performed during the summation process performed by the summation component 122. In providing the summation process, the summation component 122 can generate a compressed dataset representation that abstracts some degree of jaggedness generated in the original time series data and output by the first machine learning model, with additional seasonality generated by the second machine learning model. An illustrative aggregation of the forecast prediction values utilizes Figure 9 is provided.
[0071] The consumption prediction component 124 is configured to identify a time window (e.g., power consumption, application execution, etc.) within the forecast prediction. This can be accomplished by calculating an exponential moving average (EMA) of the forecast time series data set. Forecast predictions that exceed an overutilization threshold or fall below an underutilization threshold can be identified. The forecast prediction can determine the time of overutilization or underutilization for the monitored device 140. The forecast prediction can be used by a user or data management unit to schedule new workloads or modify existing workloads.
[0072] The monitored device action component 126 is configured to identify an action to be taken at the monitored device 140 based on the forecast prediction determined by the consumption prediction component 124. As an illustrative example, a user or data administrator unit can specify "t" hours (e.g., 5 hours in a non-limiting example) in order to find a "t" hour time window to run a new workload. In this case, if underutilization is identified, the monitored device 140 may be a good candidate for scheduling a new workload. In the event that overutilization is identified, the monitored device 140 may be a good candidate for moving the processing job to another machine or preventing any new processes from starting.
[0073] Additional details regarding the analysis operations, outputs, and actions corresponding to the time series datasets described herein are provided in U.S. patent application Ser. Nos. 17 / 991,500 (Docket No. P169908US; 61CT-361397) (Docket No. P169237US; 61CT-356847) and (Docket No. P170267US; 61CT-364847), which are incorporated herein by reference in their entirety for all purposes.
[0074] Figure 2 A process for determining forecast predictions using motifs and combined modeling is shown according to some examples of the present disclosure. Figure 1 The computer system 100 may be configured to execute machine-readable instructions to perform the process 200 described herein.
[0075] At block 210, a raw time series data set may be received. For example, the computer system 100 may receive a time series data set that includes data points corresponding to a monitored device or a monitored distributed system having multiple monitored devices. The time series data set may include CPU utilization, usage of sensors associated with the monitored devices, or other univariate time series data.
[0076] At block 215, a plurality of motifs may be generated. For example, the computer system 100 may group the time series dataset into a plurality of clusters, including identifying patterns of peaks or valleys in the original time series dataset and grouping each cluster as corresponding to the same activity at the monitored device. The original time series dataset may correspond to a plurality of these clusters across a time interval. As further described herein, the system may further identify similarities across the plurality of clusters and combine subsets of the plurality of clusters based on data signature similarities to form a plurality of motifs using an unsupervised machine learning model.
[0077] In some examples, multiple motifs can be pre-generated and clusters can be added to one or more motifs. For example, a continuous set of data points of an original time series dataset can be grouped into a first cluster. The grouping can be performed using any of the methods discussed herein, including using an unsupervised machine learning model. When the cluster is similar to other clusters of the first motif, the cluster can then be grouped into the first motif. The grouping can be performed using any of the methods discussed herein, including using a distance algorithm.
[0078] In some examples, a data schema can be accessed. For example, computer system 100 can access a data schema corresponding to a time series data set. Parameters of the data schema can include the type of data structure (e.g., an array of type "object"), attributes including data points, frequency, and nearest member, and whether specific attributes are required. Each data point attribute of the data schema can include a type (e.g., object), a unique identifier (e.g., an integer value), a time start (e.g., a string or integer value), a time stop (e.g., a string or integer value), and the number of points.
[0079] At block 220, data definitions may be generated for multiple motifs using a data schema. A compressed data set representation may identify a value for each attribute of the motif. For example, for a particular motif, data from a monitored device may typically include ten spikes within ten seconds. The data schema may include attributes for "spike" and "frequency," and a compressed data set representation for the particular motif may define values for those attributes (e.g., values for ten spikes within ten seconds). Using this compressed data set representation of the motif, the details of the particular motif may be abstracted and stored in a format corresponding to the data schema.
[0080] The level of detail represented by the compressed dataset can correspond to an accuracy value, discussed herein as parameter "c," which determines the minimum threshold for accepting two subsequences as similar. To this end, the accuracy value corresponds to the minimum threshold for accepting two subsequences as similar, and parameter "c" can be inversely proportional to the compression ratio. In other words, for higher compression ratios, parameter "c" will be lower. The accuracy value can be set by the system administrator.
[0081] In some cases, a larger accuracy value (exceeding a threshold accuracy value) may correspond to more detail in the compressed dataset representation and a greater degree of accuracy between the repopulated time series dataset and the original time series dataset, which may result in a larger memory space required to store the compressed dataset representation and the repopulated data. A smaller accuracy value (less than a threshold accuracy value) may correspond to less detail in the compressed dataset representation and a lower degree of accuracy between the repopulated time series dataset and the original time series dataset, which may result in a smaller memory space required to store the compressed dataset representation and the repopulated data.
[0082] A plurality of dataset representations may be generated, wherein each compressed dataset representation corresponds to each motif in the plurality of motifs. In other words, the original time series dataset used to generate the plurality of motifs may be compressed and represented as one or more dataset representations.
[0083] The process defined herein can generate a compressed dataset representation that defines a compressed dataset of an original time series dataset. A compressed time series dataset can be defined by the compressed dataset representation, which can be used to generate repopulated data having similar repeating data patterns as the original time series dataset but with fewer anomalies and distinctions found in the original time series dataset.
[0084] At block 225, a motif sequence graph can be generated. For example, a motif sequence graph can be generated for each motif in the compressed dataset representation to represent multiple sequences within the multiple motifs using edges representing homologies between segments. In some examples, multiple sequences can be represented by the same thread if there are multiple possible paths when traversing the thread in the sequence graph. In this way, a motif sequence graph can be created representing multiple motifs, where each motif corresponds to a path through the graph.
[0085] At block 230, a directed graph may be generated. For example, a directed graph may be generated. The directed graph may represent multiple motifs in the compressed dataset representation. The directed graph may assign weights to arrows, edges, and nodes to help identify the probability of traversing a particular sequence in the directed graph within a particular time interval.
[0086] At block 235, data points may be extracted from the motif and the directed graph. For example, the process may be accessed Figure 1 The detected group data store 132 in the plurality of motifs is used to determine a dataset representation of each motif in the plurality of motifs and a forecast prediction in a data schema format (eg, JSON format).
[0087] At box 250, a machine learning model can be trained. This may refer to a second machine learning model. For example, the process can receive a time series dataset as input to the second machine learning model. The process can detect and remove outliers in the time series dataset, including linearly interpolating any missing values in the time series dataset and removing outliers in the data. In some examples, a pre-trained machine learning model can implement anomaly detection and outlier removal (e.g., One-Class SVM). The data can be resampled. The refined time series data can be the base dataset for training and validating the output generated by the ML model.
[0088] In some examples, the second machine learning model is configured to determine outputs from a time series dataset, including trend (e.g., long-term direction), seasonality (e.g., movement related to the calendar), cycle (e.g., systematic), and residual (e.g., non-systematic or short-term fluctuations) components (referred to as "seasonality"). In this example, existing algorithms predict a smooth curve that attempts to balance the trend and seasonality components, but may ignore small maxima and minima, which are considered under the residual component. Algorithms may include, for example, FBProphet, SARIMA / SARIMAX, Holt-Winter-ES, and gated recurrent unit (GRU) networks. In this way, when attempting to predict time series data, the outputs corresponding to the trend, seasonality, cycle, and residual components can (ultimately) be combined with the motif data described in box 220.
[0089] At block 255, a data definition for the trained machine learning model may be generated. A second data definition may identify trend, seasonality, cyclic, and residual components. The data schema may include attributes of these components to represent the details of each of these influences on the time series dataset. The data may be abstracted and stored in a format corresponding to the data schema.
[0090] At block 260, data points may be extracted from the second trained machine learning model. For example, a second data definition of the trained machine learning model may be accessed, and data points forming the second data definition may be extracted corresponding to the trend, seasonality, cycle, and residual components in the compressed data format.
[0091] At block 270, a summation process may be performed on the data points using each of the first and second extracted data points. An aggregate of the forecast predictions may be generated. The aggregate may include values generated from both the first machine learning model (e.g., the motif forecast predictions) and the second machine learning model (e.g., the seasonal forecast predictions). The aggregation of the forecast prediction values may be performed during the summation process, which may generate some degree of jaggedness generated in the original time series data and abstracted by the compressed dataset representation output by the first machine learning model and with the added seasonality generated by the second machine learning model.
[0092] At block 275, the usage of the device or processor may be predicted. For example, the aggregated output from the first trained ML model and the second trained ML model may be used to obtain Figure 1 The monitored device 140 in the data center may be provided with predicted data for the next or future time period. The prediction may include, for example, a time window (e.g., power consumption, application execution, etc.) in the forecast prediction, so that an indication of exceeding an over-utilization threshold or falling below an under-utilization threshold may be identified. The forecast prediction may be used by a user or data management unit to schedule new workloads or modify existing workloads.
[0093] Figure 3 An example of time series data according to some examples of the present disclosure is provided. In example 300, processor time is identified for a monitored device running three different applications over a 15-second period, shown as a first application 310 (e.g., an antivirus scanner), a second application 320 (e.g., a web browser video), and a third application 330 (e.g., a notepad application). Each application runs with no overlapping processor time to illustrate the differences between signatures. In this example, applications 310, 320, and 330 are run independently, exclusively, and continuously to observe the differences in the time series signatures of each application.
[0094] Figure 4Examples of minimum and maximum values for time series data according to some examples of the present disclosure are provided. In example 400, an illustrative time series data set is provided with an illustrative maximum data point 410 and an illustrative minimum data point 420. For example, for each instance in which a point in the data set changes direction along the linear progression of the time series (e.g., progressing from increasing to decreasing, or from decreasing to increasing, etc.), a minimum point or a maximum point can be identified. In some examples, the time series centroid component 110 is configured to store the minimum or maximum value in the time series data store 130.
[0095] Figure 5 An example of time series data with identified outliers according to some examples of the present disclosure is provided. In example 500, the time series data is Figure 3 Example 300 is repeated, where processor time is identified for a monitored device running three different applications during a 15-second period, including a first application 510 (e.g., an antivirus scanner), a second application 520 (e.g., a web browser video), and a third application 530 (e.g., a notepad application). Additionally, an outlier is identified between the second and third applications, as shown by outlier 530.
[0096] Figure 6 An example of time series data with identified clusters according to some examples of the present disclosure is provided. In example 600, time series data is analyzed and grouped into multiple clusters, where data signature similarity is used to form multiple motifs using an unsupervised machine learning model. Illustrative clusters are provided, as shown in a first cluster 610 and a second cluster 620. These clusters can help create labeled time series data, where each cluster corresponds to a different tag (e.g., a different application or data signature, etc.). A cluster similar to the first cluster 610 is identified as a third cluster 630, which has a similar data signature determined using an unsupervised machine learning model. The model can be trained to identify similar data signatures in each cluster and match the data signatures.
[0097] In this context, each cluster of similar curves can be grouped into a motif. As discussed in this article, each motif can represent a recurring pattern and subsequence of data points that are grouped into each cluster. The distance algorithms discussed in this article can help find the similarity of curves between clusters. Figure 5 As shown, a plurality of similarly shaped data points in a first application 510 (eg, antivirus scanning) are each motif of the dataset. The plurality of motifs can be combined to form a shapelet, which can correspond to the entire dataset corresponding to the first application 510.
[0098] Clusters can be formed using a variety of processes (including Figure 7) to implement a custom K-Means clustering algorithm with outlier centroid initialization and skewed Euclidean distance function. In illustrative example 700, Figure 1 The clustering component 114 can perform one or more steps to determine a plurality of clusters.
[0099] At block 710, the inputs may include algorithm parameters, some of which may be determined by the operator Figure 1 The input may include, for example, an input data point D, an order parameter θ, and a time component N to generate a data point having cluster membership.
[0100] At block 715, a time series dataset may be received (e.g., via Figure 1 The time series dataset may be received from sensors of the monitored device 140 and include various data signatures at different times.
[0101] At block 720, the data may be processed, including normalizing the temporal features of the time series dataset. For example, feature normalization may scale individual data samples from the time series dataset to have a common and consistent unit of measurement.
[0102] At block 725, the data may be further processed, including implementing a scaler process on the time series dataset. The scaler process may help improve the wide variation in the data, create smaller feature standard deviations, and preserve zero entries in sparse data.
[0103] At block 730, Figure 1 The cluster component 114 can use the outlier centroid initialization method to receive local minima and maxima to obtain centroids c1, c2, ... ck. Figure 1 As discussed above, the distance component 112 performs the process in which the order parameters may be adjusted to determine the respective centroids.
[0104] At block 740, for each data point x i , use the skewed Euclidean distance function to find the nearest centroid (c1, c2, ... ck) and assign the point to the cluster. The skewed Euclidean distance function can include Figure 1 The distance component 112 is discussed in the distance formula.
[0105] At block 750, blocks 730 and 740 are repeated using different values of the order parameter θ and the time component N. These values are then analyzed to identify motifs and stored in Figure 1 The detected clusters in the packet data store 132 .
[0106] Figure 8Examples of directed graphs generated from motifs according to some examples of the present disclosure are provided. An illustrative directed graph 800 is provided, which is a graph consisting of a set of vertices connected by directed edges. The directed graph can illustrate the probability of traversing a particular sequence within a particular time interval. Using a generated motif sequence graph generated from multiple motifs in a compressed dataset representation, the directed graph can illustrate the progression of a sequence from the motif sequence graph.
[0107] Each point identified in the directed graph can correspond to a unique identifier of the compressed data set representation of the motif. The points identified in the directed graph can include patterns corresponding to each unique identifier, such as 1.0, 2.0, 17.0, 16.0, and so on. Thus, as shown in example 800, starting at point 810, a sequence of pattern "1.0" (from the compressed data set representation) is identified in the time series data set. The directed graph illustrates the probability of identifying the next sequence in the time series data, which is shown as point 820. At point 820, a sequence of pattern "2.0" is identified in multiple motifs. The next sequence is at point 830. At point 830, a sequence of pattern "17.0" is identified in multiple motifs. The next sequence is at point 840. At point 840, a sequence of pattern "16.0" is identified in multiple motifs, and so on.
[0108] In some examples, the sequence can have various starting points and can be based on the sequence specified by the system (e.g., Figure 1 The data point extraction component 120 of the computer system 100 in FIG. 10 is used to identify the most recent pattern identified for a particular time series. For example, if the most recently identified sequence corresponds to a sequence of pattern "17.0" at point 830, the next sequence likely to occur in the sequence is point 840 corresponding to the unique identifier "16.0." When a directed graph has been identified for a time series dataset, the predicted sequence can be overlaid on the time series dataset as a compressed dataset representation using the stored predictions.
[0109] In some examples, the directed graph 800 is also weighted using corresponding weights assigned to each node in the directed graph. The greater the weight for a sequence, the greater the probability that the next sequence in the time series data will be identified in the new time series data set.
[0110] Figure 9is an example of a graph comparing actual time series data with forecast data according to some examples of the present disclosure. In example 900, an illustrative graph is provided that shows an aggregation of forecast predictions for each time point in a time series dataset using both values generated from a first machine learning model (e.g., a motif forecast prediction) and a second machine learning model (e.g., a seasonal forecast prediction). Dataset 902 corresponds to the motif forecast predictions and data set 904 corresponds to the seasonal forecast predictions. Each of data sets 902, 904 can be output from a trained machine learning model. In example 910, the forecast predictions for a new time series dataset are provided with an overlay of the actual time series dataset values. Dataset 912 corresponds to the forecast predictions (after the summation process of the motif and seasonal forecast predictions), and data set 914 corresponds to the actual time series dataset values.
[0111] Figure 10 Example computing components that can be used to implement compression of time series datasets using motifs according to various embodiments are shown. Figure 10 , the computing component 1000 can be, for example, a server computer, a controller, or any other similar computing component capable of processing data. Figure 10 In an example implementation of , computing component 1000 includes a hardware processor 1002 and a machine-readable storage medium 1004 .
[0112] The hardware processor 1002 may be one or more central processing units (CPUs), semiconductor-based microprocessors, and / or other hardware devices suitable for fetching and executing instructions stored in the machine-readable storage medium 1004. The hardware processor 1002 may fetch, decode, and execute instructions (such as instructions 1006-1016) to control processes or operations for implementing a dynamically modular and customizable computing system. As an alternative or in addition to fetching and executing instructions, the hardware processor 1002 may include one or more electronic circuits that include electronic components for performing the functions of one or more instructions, such as a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), or other electronic circuits.
[0113] A machine-readable storage medium, such as machine-readable storage medium 1004, can be any electrical, magnetic, optical, or other physical storage device that contains or stores executable instructions. Thus, machine-readable storage medium 1004 can be, for example, random access memory (RAM), non-volatile RAM (NVRAM), electrically erasable programmable read-only memory (EEPROM), a storage device, an optical disk, or the like. In some embodiments, machine-readable storage medium 1004 can be a non-transitory storage medium, where the term "non-transitory" does not encompass transient propagating signals. As described in detail below, machine-readable storage medium 1004 can be encoded with executable instructions, such as instructions 1006-1016.
[0114] Hardware processor 1002 may execute instructions 1006 to receive a time series dataset. The time series dataset may be received from sensors of a monitored device. For example, the time series dataset may include various data signatures at different times. In some examples, the data received from the monitored device is limited to what a third party responsible for the monitored device is willing to provide and may be limited to a complete data history.
[0115] Hardware processor 1002 may execute instructions 1008 to group a set of consecutive data points of a time series dataset into clusters. The grouping may be implemented using an unsupervised machine learning model that is trained to group the consecutive set of data points into clusters. For example, the method may group the time series dataset into a first cluster of a plurality of clusters.
[0116] Hardware processor 1002 may execute instructions 1010 to group the first cluster into a first motif. For example, the method may group the first cluster into the first motif because the first cluster is similar to other clusters of the first motif. Similarity may be determined using any of the methods described herein, including by identifying similarities in data signatures. In some examples, the method may determine each centroid of the data set when determining similarities in the data and grouping the clusters into motifs.
[0117] The centroids can be initialized using various methods, including a customized K-Means clustering algorithm. The customized K-Means clustering algorithm can take into account the linear manner in which the time series dataset is recorded when determining each cluster centroid. This method can use outliers to initialize the centroid of each cluster to determine local maximum and local minimum points that correspond to the actual values from the time series dataset. In this way, the time series data can be divided into smaller cluster regions, where each smaller cluster can include a significant minimum or maximum point.
[0118] The method can also determine the distance between each data point and the centroid relative to a time series constraint (e.g., along a linear time series), which can help improve standard distance functions that may incorrectly cluster data points without considering the linear intrinsic nature of time series data. The distance formula can determine that a change on a time axis (e.g., the x-axis) is weighted less than the same change on a performance metric axis (e.g., the y-axis), so that data points can be clustered along the time axis (e.g., using the formula described herein).
[0119] Hardware processor 1002 may execute instructions 1012 to generate a compressed dataset representation using a plurality of motifs. The dataset representation may include metadata for the plurality of motifs according to a predefined data schema. In some examples, the dataset representation may correspond to a compressed dataset stored according to a data schema format (e.g., in JSON format or in a time series data store). Because motifs may define patterns in the data (e.g., corresponding to data signatures of an application at a monitored device), the patterns may be stored in the compressed dataset representation rather than individual points of the data.
[0120] Hardware processor 1002 may execute instructions 1014 to train a machine learning model. For example, instructions 1014 may identify similarities across multiple clusters and combine subsets of the multiple clusters based on data signature similarities to form multiple motifs using an unsupervised machine learning model. In another example, instructions 1014 may identify trend, seasonality, cycle, and residual components as second data definitions to represent details of each of these influences on the time series dataset. The data may be abstracted and stored in a format corresponding to the data schema.
[0121] An aggregate of forecast predictions can be generated. The aggregate can include values generated from both the first machine learning model (e.g., motif forecast predictions) and the second machine learning model (e.g., seasonal forecast predictions). Aggregation of forecast prediction values can be performed during a summation process that can generate some degree of jaggedness generated in the original time series data and abstracted by the compressed dataset representation output by the first machine learning model and with the added seasonality generated by the second machine learning model.
[0122] Hardware processor 1002 may execute instructions 1016 to schedule the monitored device for the next or future time period (e.g., Figure 1 The forecast prediction can be used by the user or data management unit to schedule new workloads or modify existing workloads.
[0123] Figure 11A block diagram of an example computer system 1100 is depicted in which various embodiments described herein may be implemented. The computer system 1100 includes a bus 1102 or other communication mechanism for communicating information, and one or more hardware processors 1104 coupled with the bus 1102 for processing information. The hardware processor(s) 1104 may be, for example, one or more general-purpose microprocessors.
[0124] The computer system 1100 also includes a main memory 1106, such as a random access memory (RAM), a cache, and / or other dynamic storage device, coupled to the bus 1102 for storing information and instructions to be executed by the processor 1104. The main memory 1106 may also be used to store temporary variables or other intermediate information during execution of instructions to be executed by the processor 1104. Such instructions, when stored in a storage medium accessible to the processor 1104, render the computer system 1100 as a special-purpose machine customized to perform the operations specified in the instructions.
[0125] The computer system 1100 also includes a read-only memory (ROM) 1108 or other static storage device coupled to the bus 1102 for storing static information and instructions for the processor 1104. A storage device 1110, such as a magnetic disk, optical disk, or USB thumb drive (flash drive), is provided and coupled to the bus 1102 for storing information and instructions.
[0126] The computer system 1100 may be coupled to a display 1112, such as a liquid crystal display (LCD) (or touch screen), via bus 1102 for displaying information to a computer user. An input device 1114, including alphanumeric and other keys, is coupled to bus 1102 for communicating information and command selections to processor 1104. Another type of user input device is a cursor controller 1116, such as a mouse, trackball, or cursor direction keys, for communicating direction information and command selections to processor 1104 and for controlling cursor movement on display 1112. In some embodiments, the same direction information and command selections as cursor control may be achieved by receiving touches on a touch screen without a cursor.
[0127] The computing system 1100 may include a user interface module for implementing a GUI, which may be stored in a mass storage device as executable software code executed by the computing device(s). By way of example, this module and other modules may include components such as software components, object-oriented software components, class components, and task components, processes, functions, properties, procedures, subroutines, program code segments, drivers, firmware, microcode, circuitry, data, databases, data structures, tables, arrays, and variables.
[0128] In general, the terms "component," "engine," "system," "database," "data store," and the like as used herein may refer to logic embodied in hardware or firmware, or to a collection of software instructions, which may have entry and exit points, written in a programming language such as, for example, Java, C, or C++. Software components may be compiled and linked into executable programs, installed in dynamic link libraries, or may be written in interpreted programming languages such as, for example, BASIC, Perl, or Python. It will be understood that software components may be called from other components or from themselves, and / or may be called in response to detected events or interrupts. Software components configured to execute on a computing device may be provided on a computer-readable or machine-readable storage medium, such as an optical disc, digital video disc, flash drive, magnetic disk, or any other tangible medium, or as a digital download (and may initially be stored in a compressed or installable format that requires installation, decompression, or decryption prior to execution). Such software code may be stored in part or in whole on a memory device of the executing computing device for execution by the computing device. The software instructions may be embedded in firmware, such as an EPROM. It will also be understood that hardware components may include connected logic units (such as gates and flip-flops), and / or may include programmable units (such as programmable gate arrays or processors).
[0129] Computer system 1100 can implement the techniques described herein using custom hardwired logic, one or more ASICs or FPGAs, firmware, and / or program logic that, in combination with the computer system, makes computer system 1100 a special-purpose machine or programs computer system 1100 to be a special-purpose machine. According to one embodiment, the techniques herein are performed by computer system 1100 in response to processor(s) 1104 executing one or more sequences of one or more instructions contained in main memory 1106. Such instructions may be read into main memory 1106 from another storage medium, such as storage device 1110. Execution of the sequences of instructions contained in main memory 1106 causes processor(s) 1104 to perform the process steps described herein. In alternative embodiments, hardwired circuitry may be used in place of or in combination with software instructions.
[0130] As used herein, the term "non-transitory media" and similar terms refer to any medium that stores data and / or instructions that cause a machine to operate in a specific manner. Such non-transitory media may include non-volatile media and / or volatile media. Non-volatile media include, for example, optical or magnetic disks, such as storage device 1110. Volatile media include dynamic memory, such as main memory 1106. Common forms of non-transitory media include, for example, floppy disks, diskettes, hard disks, solid-state drives, magnetic tape or any other magnetic data storage medium, CD-ROMs, any other optical data storage medium, any physical medium with a pattern of holes, RAM, PROMs and EPROMs, FLASH-EPROMs, NVRAMs, any other memory chips or cartridges, and networked versions thereof.
[0131] Non-transient media are distinct from, but can be used in conjunction with, transmission media. Transmission media participate in the transfer of information between non-transient media. For example, transmission media include coaxial cables, copper wire, and optical fiber, including the wires that comprise bus 1102. Transmission media can also take the form of acoustic or light waves, such as those generated during radio wave and infrared data communications.
[0132] Computer system 1100 also includes a communication interface 1118 coupled to bus 1102. Communication interface 1118 provides two-way data communication coupled to one or more network links connected to one or more local networks. For example, communication interface 1118 can be an integrated services digital network (ISDN) card, a cable modem, a satellite modem, or a modem that provides data communication connections to a telephone line of a corresponding type. As another example, communication interface 1118 can be a local area network (LAN) card to provide data communication connections with a compatible LAN (or with a WAN component of WAN communication). A wireless link can also be implemented. In any such implementation, communication interface 1118 sends and receives electrical, electromagnetic, or optical signals that carry digital data streams representing various types of information.
[0133] A network link typically provides data communication to other data devices through one or more networks. For example, a network link can provide a connection through a local network to a host computer or to data equipment operated by an Internet Service Provider (ISP). The ISP, in turn, provides data communication services through the global packet data communication network now commonly referred to as the "Internet." Both the local network and the Internet use electrical, electromagnetic, or optical signals that carry digital data streams. The signals through the various networks and the signals on the network link and through the communication interface 1118 are example forms of transmission media that carry digital data to and from the computer system 1100.
[0134] Computer system 1100 can send messages and receive data, including program code, through the network(s), network links, and communication interface 1118. In the Internet example, a server can send code requested by an application program through the Internet, an ISP, a local network, and the communication interface 1118 network.
[0135] The received code may be executed by processor 1104 as it is received, and / or stored in storage device 1110 or other non-volatile storage for later execution.
[0136] Each of the processes, methods, and algorithms described in the foregoing sections can be embodied in a code component executed by one or more computer systems or computer processors comprising computer hardware, and fully or partially automated by the code component. One or more computer systems or computer processors can also operate to support the performance of related operations in a "cloud computing" environment or as "software as a service" (SaaS). These processes and algorithms can be implemented in part or in whole in dedicated circuits. The various features and processes described above can be used independently of each other, or can be combined in various ways. Different combinations and sub-combinations are intended to fall within the scope of this disclosure, and certain method or process blocks may be omitted in some implementations. The methods and processes described herein are also not limited to any particular order, and the blocks or states associated therewith can be executed in other appropriate orders, or can be executed in parallel or in some other manner. Blocks or states can be added to or removed from the disclosed example embodiments. The performance of certain operations or processes can be distributed among computer systems or computer processors, not only residing within a single machine, but also deployed on multiple machines.
[0137] As used herein, circuit can utilize any form of hardware, software or its combination to realize.For example, one or more processors, controllers, ASIC, PLA, PAL, CPLD, FPGA, logic components, software routine or other mechanisms can be realized to form circuit.In realization, various circuits described herein can be realized as discrete circuits, or described function and feature can be partially or entirely shared between one or more circuits.Although the element of various features or functions can be described or required as separate circuits individually, these features and functions can be shared between one or more common circuits, and such description should not require or imply that separate circuits are needed to realize such features or functions.When circuit is realized in whole or in part using software, such software can be realized as and can perform the calculation or processing system (such as computer system 1100) of the function described about it and operate together.
[0138] As used herein, the term "or" may be interpreted as inclusive or exclusive. In addition, descriptions of resources, operations, or structures in the singular should not be construed to exclude the plural. Conditional language, such as "may," "could," "might," or "might," unless otherwise specifically stated or understood otherwise in the context of use, is generally intended to convey that some embodiments include, while other embodiments do not include, certain features, elements, and / or steps.
[0139] Unless expressly stated otherwise, the terms and phrases used in this document, and variations thereof, should be interpreted as open ended and not restrictive. Adjectives such as "conventional," "traditional," "normal," "standard," "known," and terms of similar meaning should not be interpreted as limiting the items described to a given time period or to items available at a given time, but should be understood to encompass conventional, traditional, normal, or standard technology that may be available or known at any time now or in the future. In certain cases, the presence of broad words and phrases such as "one or more," "at least," "but not limited to," or other similar phrases should not be understood as intending or requiring the use of a narrower context where such broad phrases may not be present.
[0140] It should be noted that the terms "optimize," "optimal," and the like as used herein may be used to mean enabling or achieving performance that is as efficient or perfect as possible. However, as one of ordinary skill in the art reading this document will recognize, perfection is not always achievable. Thus, these terms may also encompass enabling or achieving performance that is as good, efficient, or practical as possible under given circumstances, or enabling or achieving performance that is better than that achievable using other settings or parameters.
Claims
1. A system comprising: one or more processors; as well as a machine-readable storage medium storing instructions that, when executed by the one or more processors, cause the system to: Receive raw time series data from sensors of monitored equipment; Grouping the original time series data into a plurality of clusters; determining, using an unsupervised machine learning model, a data signature for a cluster in the plurality of clusters; combining subsets of the plurality of clusters based on data signature similarity to form a plurality of motifs; generating a first data definition using the plurality of motifs, wherein the data definition corresponds to a predefined data pattern and value for each motif in the plurality of motifs, wherein the data definition creates a compressed data set corresponding to the raw time series data; generating a second data definition using a seasonal prediction of the sensor of the monitored equipment; training a machine learning model using the first data definition and the second data definition, both received from the sensor of the monitored device, to obtain predicted power consumption data of the monitored device for a future time period; as well as Using the predicted power consumption data obtained by the trained machine learning model, schedule a workload to be executed at the monitored device during a portion of the future time period, wherein the predicted power consumption data for the portion of the future time period is below a power consumption threshold. 2 . The system of claim 1 , wherein the first data definition corresponds to a directed graph of the plurality of motifs, the directed graph comprising weighted edges between data points.
3. The system of claim 1 , wherein the first data definition and the second data definition are stored as JavaScript Object Notation (JSON) files, and the system further: Memory space is recovered by deleting the original time series data set and storing the first data definition and the second data definition in place of the original time series data. The system of claim 1 , wherein the monitored device is a processor in a server in an IT data center. 5 . The system of claim 1 , wherein the plurality of clusters comprises local minima and local maxima of a subset of the original time series data.
6. The system of claim 5, wherein the system further: A centroid of each of the plurality of clusters is determined using the local minimum data point and the local maximum data point of each of the plurality of clusters from the original time series data. 7 . The system of claim 6 , wherein the determining of the centroid of each of the plurality of clusters uses a customized K-means clustering algorithm that accounts for a linear approach to the original time series data.
8. The system of claim 1, wherein the workload scheduled to be executed at the monitored device during the portion of the future time period is a reduced workload. 9 . The system according to claim 1 , wherein the raw time series data comprises input power cycle data of the monitored equipment.
10. The system of claim 1 , wherein the first data definition is generated using a first machine learning model, the second data definition is generated using a second machine learning model, and the first machine learning model and the second machine learning model are different from the machine learning model trained to obtain the predicted power consumption data.
11. The system of claim 10, wherein the second machine learning model is one of: SARMA, FBProphet, Holt-Winter-ES, or a gated recurrent unit network.
12. A computer-implemented method comprising: The processor receives raw time series data from sensors of the monitored equipment; Grouping the original time series data into a plurality of clusters; determining, using an unsupervised machine learning model, a data signature for a cluster in the plurality of clusters; combining subsets of the plurality of clusters based on data signature similarity to form a plurality of motifs; generating a first data definition using the plurality of motifs, wherein the data definition corresponds to a predefined data pattern and value for each motif in the plurality of motifs, wherein the data definition creates a compressed data set corresponding to the raw time series data; generating a second data definition using a seasonal prediction of the sensor of the monitored equipment; training a machine learning model using the first data definition and the second data definition, both received from the sensor of the monitored device, to obtain predicted power consumption data of the monitored device for a future time period; as well as Using the predicted power consumption data obtained by the trained machine learning model, schedule a workload to be executed at the monitored device during a portion of the future time period, wherein the predicted power consumption data for the portion of the future time period is below a power consumption threshold.
13. The computer-implemented method of claim 12, wherein the first data definition corresponds to a directed graph of the plurality of motifs, the directed graph comprising weighted edges between data points.
14. The computer-implemented method of claim 12, wherein the first data definition and the second data definition are stored as JavaScript Object Notation (JSON) files, and the method further comprises: Memory space is recovered by deleting the original time series data set and storing the first data definition and the second data definition in place of the original time series data.
15. The computer-implemented method of claim 12, wherein the monitored device is a processor in a server in an IT data center.
16. The computer-implemented method of claim 12, wherein the plurality of clusters comprises local minima and local extremes of a subset of the original time series data.
17. The computer-implemented method of claim 16, wherein the method further comprises: A centroid of each of the plurality of clusters is determined using the local minimum data point and the local maximum data point of each of the plurality of clusters from the original time series data.
18. The computer-implemented method of claim 17, wherein the determining of the centroid of each cluster in a plurality of clusters uses a customized K-means clustering algorithm that accounts for a linear fashion of the original time series data.
19. The computer-implemented method of claim 12, wherein the workload scheduled to be executed at the monitored device during the portion of the future time period is a reduced workload.
20. The computer-implemented method of claim 12, wherein the raw time series data comprises input power cycle data of the monitored device.
Citation Information
Patent Citations
Unsupervised segmentation of a univariate time series dataset using motifs and shapelets
US12050626B2
Time series motif discovery method based on sub-sequence full join and maximum clique
CN109241118A
Time series analysis for predicting computational workloads
CN114430826A