Automatic forecast model provider
By introducing end-to-end systems into the forecast model system, combining edge computing and central computing resources, the forecast model is automatically trained, verified, selected and deployed, and the problem of deterioration in the prediction accuracy of forecast models in the existing technology is solved, achieving faster and more accurate response.
Patent Information
- Application Number
- CN202410669436.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-13
- Filing Date
- 2024-05-28
- Publication Date
- 2025-06-13
AI Technical Summary
The existing technology is difficult to respond quickly to changes in data trends, resulting in worsening the prediction accuracy of forecast models, which in turn causes financial losses.
The end-to-end system is adopted, combining edge computing resources and central computing resources to automatically train, verify, select and deploy forecast models to ensure that the forecast models are always adapted to rapidly changing data trends.
Quickly respond to data changes through automated processes, improving the prediction accuracy of forecast models, reducing financial losses, and improving corporate results.
Smart Images

Figure CN120146215A_ABST
Abstract
Description
Background Art
[0001] A machine learning model may include a computer program based on an algorithm that is trained to identify patterns in data and make predictions and / or classifications based on such learned pattern recognition.
[0002] As used herein, a forecasting model may refer to a machine learning model that processes historical and current / realtime information to predict a characteristic outcome. In many implementations, the forecasting model utilizes time series data (i.e., data points indexed in chronological order) to make such predictions. BRIEF DESCRIPTION OF THE DRAWINGS
[0003] The present disclosure is described in detail with reference to the following drawings according to one or more various examples. The drawings are provided for illustrative purposes only and depict examples only.
[0004] Figures 1A to 1B An example computing system is depicted for automatically training, validating, selecting, and deploying a forecasting model in response to rapidly changing data trends according to various examples of the techniques of the present disclosure.
[0005] Figure 2 An example computing resource is shown for automatically training, validating, and selecting a forecasting model according to various examples of the techniques of the present disclosure.
[0006] Figure 3 An example edge computing resource is shown for deriving a subset of time series data metrics, detecting a drift involving the subset of time series data metrics, and automatically deploying a new forecasting model trained to predict the (drifted) subset of time series data metrics according to various examples of the techniques of the present disclosure.
[0007] Figure 4 An example graph is depicted for comparing predictions of a forecasting model with actual / historical data for which the forecasting model is deployed to predict according to various examples of the techniques of the present disclosure.
[0008] Figure 5 A block diagram of an example computer system in which various examples described herein may be implemented is depicted.
[0009] The drawings are not exhaustive and do not limit the present disclosure to the exact forms disclosed. DETAILED DESCRIPTION
[0010] Forecast models have become key tools for many enterprises because they can use predictive analytics to address / overcome future uncertainties. However, the business process-related data used by forecast models for prediction often grows and changes rapidly. Changes in data trends (usually caused by unexpected events) may lead to deteriorated prediction accuracy for the forecast models, thereby causing financial losses for the enterprises that rely on them. For example, a forecast model can be deployed to predict certain operating parameters of machines in a manufacturing plant based on sensor data obtained from the machines. Unexpected events (such as changes in sensors on the machines, non-sensor-related function or structural modifications to the machines, etc.) may deteriorate the prediction accuracy of the forecast model. When certain business processes and / or business decisions are based on the increasingly inaccurate predictions of the forecast model, this may cause financial losses for the manufacturing plant.
[0011] In the above example, even if the deteriorated prediction accuracy of the forecast model is detected, many existing technologies still rely on human data scientists to manually select new forecast models that are more suitable for the new / changing data trends. Before deployment, data scientists usually also have to initiate the training and validation of new forecast models as well as other potential alternatives. This long manual process usually cannot keep up with the rapidly changing data trends - which means that less performant forecast models (i.e., forecast models with relatively low prediction accuracy) are usually deployed in inference / production for a longer period than the optimal ones.
[0012] Related to the above, many existing automation technologies in the field of predictive analytics only focus on individual parts of the forecast model pipeline (such as data acquisition, inference prediction, model training, model validation, model selection / deployment, etc.). The lack of cohesion between different systems may lead to inefficiencies (such as data latency between systems, dependence on manual data scientist input, etc.) - which may also cause a delay in response to rapidly changing data trends.
[0013] In this context, examples of the technology of the present disclosure provide an end-to-end system for automatically training, validating, selecting, and deploying a forecasting model in response to rapidly changing data trends. Such an end-to-end system includes edge computing resources (e.g., edge computing clusters associated with respective customers), which stream data directly from customer data sources (e.g., customer subsystems) and detect drifts in time series data derived from the streamed data. The end-to-end system also includes central computing resources (e.g., a centralized, cloud-based computer cluster), which respond to the drift detection by automatically training and validating instances of stored forecasting models (e.g., potentially numerous forecasting models stored in the Darts library) using fresh (i.e., recent / near real-time) time series data derived from the streamed data. The fresh time series data can be logically grouped into subsets of time series data metrics (i.e., multiple time series data metrics logically grouped according to heuristics). In some implementations, each subset of time series data metrics can be associated with a common customer subsystem (e.g., a subset of time series data metrics associated with a machine from the manufacturing plant example above). Here, logically grouping the time series data metrics into subsets can improve the efficiency of forecasting model training / validation.
[0014] For example, a system of the technology of the present disclosure can include central computing resources (e.g., a centralized, cloud-based computer cluster) and a series of edge computing resources (e.g., edge computing clusters), including a first edge computing resource. In response to detecting a drift in time series data (e.g., a subset of time series data metrics) for which a first forecasting model is deployed at the first edge computing resource to make predictions, the central computing resources can: (1) train an instance of a stored forecasting model using (fresh) training time series data derived from the time series data; (2) compare the predictions of the trained instance of the stored forecasting model with corresponding (fresh) validation time series data also derived from the time series data; (3) based on the comparison, determine that an instance of a second forecasting model among the trained instances of the stored forecasting models has the lowest prediction error for the predicted time series data; and (4) in response to determining that the trained instance of the second forecasting model has a lower prediction error for the predicted time series data than the first forecasting model, update the model registry database of the central computing resources with the parameters of the trained instance of the second forecasting model.
[0015] Correspondingly, in response to the central computing resource updating the model registry database with the parameters of the trained instance of the second forecasting model, the first edge computing resource may: (a) download the parameters of the trained instance of the second forecasting model from the model registry database; and (b) deploy the trained instance of the second forecasting model for predicting time series data. In various implementations, the first edge computing resource may also detect a drift involving the time series data - which triggers the forecasting model training, validation, and selection described in the previous paragraph. In some implementations, the first edge computing resource may detect a drift involving the time series data by comparing the historical predictions of the first forecasting model with the corresponding portions of the time series data over a time interval.
[0016] As described above, in certain implementations, the time series data may include a subset of time series data metrics that are logically grouped together for more efficient forecasting model training / validation. In these implementations, the first edge computing resource may execute machine-readable instructions of a data exchange process to: (i) stream data from a customer (e.g., data from the machines in the manufacturing plant example above); (ii) derive a subset of time series data metrics from the data streamed from the customer; (iii) upload the subset of time series data metrics to the time series database of the first edge computing resource; and (iv) upload the subset of time series data metrics to the time series database of the central computing resource. In some implementations, deriving a subset of time series data metrics from the data streamed from the customer may include: (a) indexing data points of the data in chronological order to generate time series data metrics; and (b) logically grouping the subset of time series data metrics together according to heuristics. In various implementations, the heuristics may include logically grouping together time series data metrics derived from a common customer subsystem (e.g., the respective machines of a customer manufacturing plant) as a subset of time series data metrics.
[0017] In the data exchange process using the first edge computing resource, the system can ensure that the time series data metrics used to train and validate the stored instances of the prediction model remain relatively fresh. Therefore, compared with potential alternative techniques that do not stream real-time (or near real-time) data to train and validate the prediction model, the system can respond faster to rapidly changing data trends. Correlatively, by (1) logically grouping the time series data metrics into corresponding subsets of time series data metrics, and (2) using these subsets of (logically grouped) time series data metrics to train, validate, and select instances of the prediction model - the system can train / validate / select the prediction model more efficiently in various customer subsystems / use cases. Other edge computing resources of the system can include similar data exchange processes that can be executed to stream real-time (or near real-time) data from customers and / or derive subsets of time series data metrics in the same / similar manner.
[0018] In some implementations, the central computing resource can execute machine-readable instructions of a model scheduler process that are operable to: (a) retrieve model training and validation tasks from a task database of the central computing resource, where the first model training and validation task utilizes time series data (e.g., a subset of time series data metrics); and (b) assign the model training and validation tasks to model listener processes, where the first model training and validation task is assigned to a first model listener process. Each model training and validation task can be associated with a corresponding customer subsystem and can utilize a corresponding subset of time series data metrics associated with the corresponding customer subsystem. Here, the central computing resource can execute the machine-readable instructions of the model scheduler process to assign the model training and validation tasks to the model listener processes in parallel to improve the speed / efficiency of prediction model training and validation across a wide variety of customer subsystems / use cases. The central computing resource can also execute the machine-readable instructions of the model listener processes (including the first model listener process). For example, the central computing resource can execute the machine-readable instructions of the first model listener process to perform the first model training and validation task.
[0019] It should be understood that in many practical implementations, the above-described system may include thousands (or more) of edge computing resources. These thousands of edge computing resources may be associated with thousands of customer subsystems / customer use cases (or more), and each subsystem / customer use case generates a large amount of data. Individually training and validating multiple instances of the prediction model to make predictions based on fresh time series data metrics derived from these thousands of customer subsystems / customer use cases is a daunting task. However, by leveraging a specialized computing architecture (e.g., a data exchange process of edge computing resources that can be executed to stream / process real-time (or near real-time) customer data, a model scheduler process of central computing resources that can be executed to parallelly schedule / assign thousands of model training and validation tasks to thousands of model listener processes, and thousands of model listener processes that can be executed to parallelly execute the assigned model training and validation tasks), the system can process this huge data processing task at a speed and efficiency that is several orders of magnitude higher than potential manual alternatives.
[0020] As described above, examples of the techniques of the present disclosure provide many advantages over the prior art. For example, by automatically training, validating, selecting, and deploying a prediction model in response to the occurrence of detected drift, the examples can respond to rapidly changing data trends faster than the prior art that requires manual / human data scientist input. Such a fast response time can improve the results of enterprises that rely on prediction model predictions. Correlatively, by providing an end-to-end system that manages all (or at least many) parts of the prediction model pipeline (e.g., data collection, inference prediction, model training, model validation, model selection / deployment, etc.), the examples can further improve the efficiency in the field of computational-based predictive analytics technology (e.g., reduce data latency between computing systems, reduce power consumption and processing time, etc.).
[0021] Examples of the techniques of the present disclosure are now described in conjunction with the accompanying drawings.
[0022] Figures 1A to 1B An example computing system 100 for automatically training, validating, selecting, and deploying a prediction model in response to rapidly changing data trends in accordance with various examples of the techniques of the present disclosure is depicted.
[0023] The computing system 100 includes edge computing resources 110 and central computing resources 120. Although not shown in Figures 1A to 1B , the computing system 100 may include additional (and in various implementations, many additional) edge computing resources with the same / similar architecture as the edge computing resources 110. For the sake of brevity, the central computing resources 120 are only partially depicted in Figure 1A . Similarly, the edge computing resources 110 are only partially depicted in Figure 1B .
[0024] Before describing the components of the computing system 100 in more detail, a high-level overview of the operation of the computing system 100 may be illuminating.
[0025] As shown, the edge computing resource 110 may execute a data exchange process 111 to: (i) stream data from a data source 102 (where the data source 102 may be associated with a first customer); (ii) derive a subset of time series data metrics from the streamed data; (iii) upload the subset of time series data metrics to a time series database 112 of the edge computing resource 110 (in some implementations, the time series database 112 may operate as a cache); and (iv) upload the subset of time series data metrics to a time series database 121 of the central computing resource 120. As described above, the edge computing resource 110 may execute the data exchange process 111 to derive a subset of time series data by: (a) indexing data points of the data streamed from the data source 102 in chronological order to generate time series data metrics; and (b) logically grouping together a subset of the time series data metrics according to heuristics. In various implementations, the heuristics may include logically grouping together time series data metrics derived from a common customer subsystem (e.g., corresponding machines in a manufacturing plant of the first customer) as a subset of the time series data metrics.
[0026] As shown, the edge computing resource 110 may execute a prediction engine process 113 to manage the deployment of one or more prediction models. In certain implementations, the corresponding prediction models may be used to make predictions based on corresponding subsets of the time series data metrics stored in the time series database 112. Such predictions may drive business processes and / or business decisions for a first customer associated with the edge computing resource 110. For example, the prediction engine process 113 may be executed to deploy a first prediction model to make predictions based on a first subset of the time series data metrics. Correlatively, the prediction engine process 113 may be executed to deploy a second prediction model to make predictions based on a second subset of the time series data metrics. The data exchange process 111 may be executed to derive a first subset of the time series data metrics from data streamed from a first customer subsystem. The first customer subsystem may include a machine in a manufacturing plant of the first customer, and the data streamed from the first machine may be related to the operating parameters of the machine. The data exchange process 111 may be executed to derive a second subset of the time series data metrics from data streamed from a second customer subsystem. The second customer subsystem may include a second machine in a manufacturing plant of the first customer, and the data streamed from the second machine may be related to the operating parameters of the second machine. The predictions made by the first and second prediction models may be used to drive manufacturing processes / manufacturing decisions at the manufacturing plant of the first customer.
[0027] In some implementations, before deploying a forecasting model for prediction, the edge computing resource 110 may execute a forecasting engine process 113 to verify that the time series data metrics stored in the time series database 112 are sampled from a common time interval (sometimes referred to as a "resampling operation"). Performing this resampling operation can ensure that the data used for prediction is appropriate in terms of time.
[0028] As shown, the edge computing resource 110 may execute a drift detection process 117 to detect drifts involving a first subset of the time series data metrics and / or a second subset of the time series data metrics. As used herein, "drift" may refer to a statistical property of a time series data metric (or time series data metrics) that changes in an unforeseen manner over time. These unforeseen changes may cause the prediction of a forecasting model deployed to predict the drifted time series data metric (or time series data metrics) to deteriorate. Thus, in some implementations, the drift detection process 117 may be executed to detect a drift involving a first subset of the time series data metrics by comparing, within a time interval, the historical predictions of a first forecasting model with corresponding portions of the first subset of the time series data metrics (it should be understood that in other implementations, such data drift may be detected using other techniques, such as using the Kolmogorov-Smirnov (K-S) test, using the population stability index, using the Page-Hinkley method, or using other drift detection techniques / algorithms). In some implementations, in response to a drift being detected, the drift detection process 117 may be executed to send a drift detection event to the central computing resource 120 (or more specifically, the model scheduler process 122 of the central computing resource 120). The drift detection event may indicate that a drift has been detected and that the central computing resource 120 (or more specifically, the model scheduler process 122 of the central computing resource 120) should initiate the training of instances of the stored forecasting models M_1–M_N to replace the first forecasting model.
[0029] Thus, in response to the detection of a drift involving a first subset of time series data metrics (e.g., in response to receiving a drift detection event indicating that a drift has been detected), the central computing resource 120 may execute a model scheduler process 122 to retrieve model training and validation tasks from the task database 123 of the central computing resource 120. Then the model scheduler process 122 may be executed to assign the model training and validation tasks to the model listener process 124(1) of the central computing resource 120. The model listener process 124(1) may be executed to perform the model training and validation tasks by: using the (fresh) training time series data derived from the first subset of time series data metrics to train an instance of the stored prediction models M_1–M_N (as shown, the model listener process 124(1) may be executed to retrieve the first subset of time series data metrics from the time series database 121 and split the first subset of time series data metrics into training time series data and validation time series data). Then the model listener process 124(1) may be executed to compare the predictions of the trained instances of the stored prediction models M_1–M_N with the corresponding (fresh) validation time series data also derived from the first subset of time series data metrics. Based on these comparisons, the model listener process 124(1) may be executed to determine that among the trained instances of the stored prediction models M_1–M_N, the trained instance of the stored prediction model M_2 has the lowest prediction error for predicting the first subset of time series data metrics. In response to determining that the trained instance of the stored prediction model M_2 has a lower prediction error for predicting the first subset of time series data metrics than the first prediction model, the model listener process 124(1) may be executed to update the model registry database 125 of the central computing resource 120 with the parameters of the trained instance of the stored prediction model M_2.
[0030] In response to the central computing resource 120 updating the model registry database 125 with the parameters of the trained instance of the stored prediction model M_2, the edge computing resource 110 may execute a download model scheduler process to download the parameters of the trained instance of the stored prediction model M_2 and upload such parameters to the model configuration database 115 of the edge computing resource 110. Then the prediction engine process 113 may be executed to retrieve the parameters of the trained instance of the stored prediction model M_2 from the model configuration database 115 and deploy the trained instance of the stored prediction model M_2 to predict the first subset of time series data metrics. In this example, the trained instance of the stored prediction model M_2 may replace the first prediction model that has become less suitable for predicting the first subset of time series data metrics after the drift.
[0031] The specific components of the computing system 100 are described in more detail below.
[0032] As described above, the computing system 100 includes a central computing resource 120 and one or more edge computing resources, including edge computing resource 110. As used herein, a computing resource (e.g., edge computing resource 110, central computing resource 120, etc.) can refer to one or more physical or logical / virtual computing devices. In certain implementations, a computing resource can include a computing cluster. For example, edge computing resource 110 can include an edge computing cluster, and central computing resource 120 can include a central computing cluster (e.g., a central, cloud-based computing cluster). As used herein, a computing cluster can refer to a group of physical devices that work together such that they can be regarded as a single computing system. The components of a computing cluster (sometimes referred to as computing nodes) can be interconnected via a network (e.g., a fast local area network (LAN)). Each computing node of a computing cluster can include a physical computing device that runs its own instance of an operating system. Examples of computing devices can include server computers, laptops, controllers, and Internet of Things (IoT) devices, or any other computing device capable of processing data. In some implementations, the computing nodes of a common computing cluster utilize the same / similar hardware and operating system, although this need not be the case. Generally speaking, a computing cluster can improve performance and availability compared to a single computing device. A computing cluster can also be more cost-effective than a single computing device with comparable speed / availability.
[0033] As shown in the figure, edge computing resource 110 and central computing resource 120 can perform certain computing processes. For example, edge computing resource 110 can perform a data exchange process 111, a prediction engine process 113, and a download model scheduler process 114. Similarly, central computing resource 120 can perform a model scheduler process 122 and model listener processes 124(1)-(n). As used herein, a "computing process" (sometimes abbreviated as "process") can refer to an instance of a computer program executed via one or more "computing threads". As used herein, a "computing thread" (sometimes abbreviated as "thread") can refer to the smallest sequence of programming instructions executed by a computing processor. Thus, edge computing resource 110 can execute the machine-readable instructions of each of its constituent computing processes. Similarly, central computing resource 120 can execute the machine-readable instructions of each computing process of its constituent computing processes.
[0034] As shown in the figure, the edge computing resource 110 and the central computing resource 120 may also include certain databases. For example, the edge computing resource 110 may include a time series database 112 (which may operate as a cache in some implementations) and a model configuration database. Similarly, the central computing resource 120 may include a time series database 121, a model registry database 125, and a task database 123. As used herein, a database may refer to an organized collection of data stored in a computing device. In many cases, the database is specifically organized for rapid search and retrieval by the computing device.
[0035] In various implementations, the edge computing resource 110 and the central computing resource 120 may include computing clusters implemented in a containerized environment. For example, the edge computing resource 110 may include a computing cluster that implements / executes its constituent processes (e.g., data exchange process 111, prediction engine process 113, and download model scheduler process 114) as microservices within a container. Similarly, the central computing resource 120 may include a computing cluster that implements / executes its constituent processes (e.g., model scheduler process 122 and model listener processes 124(1)-(n)) as microservices within a container. In these examples, the processes / microservices may be managed by a container orchestrator such as Kubernetes. The databases of the edge computing resource 110 and the central computing resource 120 may also be implemented using containers and may similarly be managed by the container orchestrator.
[0036] Now referring to the edge computing resource 110, the edge computing resource 110 may be associated with a first customer and may be implemented at the physical location of the first customer. For example (and as described above), the edge computing resource 110 may be implemented at a manufacturing plant of the first customer. In implementations where the computing system 100 includes additional edge computing resources, each additional edge computing resource may be associated with / implemented at the physical location of a corresponding additional customer.
[0037] As used herein, edge computing may refer to a distributed computing paradigm that brings data processing and / or data storage closer to the location where the data to be processed / stored is actually generated. For example (and as described above), the edge computing resource 110 may be implemented at a manufacturing plant of the first customer and may process / store data obtained from the machines of the manufacturing plant. In some implementations, the edge computing resource 110 may be incorporated / implemented in one or more of such machines.
[0038] Edge computing can reduce data latency (and associated inefficiencies) by bringing data processing / storage closer to the data source. Thus, by performing certain data processing steps at the edge computing resource 110 (e.g., deriving a subset of time series data metrics, deploying a forecasting model at inference time), the computing system 100 can improve efficiency by reducing data latency (and associated inefficiencies).
[0039] For example, the edge computing resource 110 can execute a data exchange process 111 to stream data from the data source 102. In various implementations, the data source 102 can be associated with one or more client subsystems (e.g., a first client subsystem, a second client subsystem, etc.). For example, in the example of a manufacturing plant, the data source 102 can be associated with one or more machines of the manufacturing plant, where each machine includes a corresponding client subsystem. In certain implementations, the edge computing resource 110 can execute the data exchange process 111 to stream data from multiple data sources (e.g., data source 103, data source 104, etc.) of a first client. These additional data sources can be associated with additional subsystems of the first client.
[0040] The data exchange process 111 can be executed to stream raw (or preprocessed) data from the data source 102. As described above, the data exchange process 111 (or another computing process of the edge computing resource 110) can be executed to convert the data streamed from the data source 102 into a more useful / compatible format for consumption by a forecasting model at inference time or during training. For example, the data exchange process 111 (or another computing process of the edge computing resource 110) can be executed to derive a subset of time series data metrics based on the data streamed from the data source 102. In some implementations, deriving a subset of time series data metrics based on the data streamed from the data source 102 can include: (a) indexing the data points of the data in chronological order to generate time series data metrics; and (b) logically grouping together a subset of the time series data metrics according to heuristics. In various implementations, the heuristics can include logically grouping together the time series data metrics derived from a common client subsystem (e.g., the respective machines of a manufacturing plant of a first client) as a subset of the time series data metrics.
[0041] As used herein, a time series data metric may refer to data points indexed in chronological order. Illustrative examples of time series data metrics may include: (1) temperature indexed by time in London; (2) temperature indexed by time in Berlin; (3) humidity indexed by time in London; (4) humidity indexed by time in Berlin; (5) voltage indexed by time of a first temperature sensor in a first machine; (6) voltage indexed by time of a first temperature sensor in a second machine; (7) voltage indexed by time of a second temperature sensor in the first machine; (8) voltage indexed by time of a second temperature sensor in the second machine; (9) data indexed by time obtained from a first CRM application; (10) data indexed by time obtained from a second CRM application; and so on.
[0042] As used herein, a subset of time series data metrics may include one or more time series data metrics logically grouped according to a heuristic (such a heuristic may be user-defined or automatically generated by computing system 100). Illustrative examples of subsets of time series data metrics may include: (1) a subset of time series data metrics that includes environmental parameters in London (e.g., a first time series data metric that includes temperature indexed by time in London, a second time series data metric that includes humidity indexed by time in London, a third time series data metric that includes SO2 level indexed by time in London, etc.); (2) a subset of time series data metrics that includes environmental parameters in Berlin (e.g., a first time series data metric that includes temperature indexed by time in Berlin, a second time series data metric that includes humidity indexed by time in Berlin, a third time series data metric that includes SO2 level indexed by time in Berlin, etc.); (3) a subset of time series data metrics that includes temperatures in major European cities (e.g., a first time series data metric that includes temperature indexed by time in London, a second time series data metric that includes temperature indexed by time in Berlin, a third time series data metric that includes temperature indexed by time in Rome, etc.); (4) a subset of time series data metrics that includes operating parameters of a customer's first machine (e.g., a first time series data metric that includes voltage indexed by time of a first temperature sensor in the first machine, a second time series data metric that includes voltage indexed by time of a second temperature sensor in the first machine, etc.); (5) a subset of time series data metrics that includes operating parameters of a customer's second machine (e.g., a first time series data metric that includes voltage indexed by time of a first temperature sensor in the second machine, a second time series data metric that includes voltage indexed by time of a second temperature sensor in the second machine, etc.); (6) a subset of time series data metrics that includes applications belonging to a customer's CRM domain (e.g., a first time series data metric that includes data indexed by time obtained from a first CRM application, a second time series data metric that includes data indexed by time obtained from a first CRM application, etc.).
[0043] As described above, by (1) logically grouping time series data metrics into corresponding subsets of time series data metrics, and (2) using these subsets of (logically grouped) time series data metrics to train, validate, and select instances of a forecasting model - computing system 100 can more efficiently train / validate / select a forecasting model across a wide variety of customer subsystems / use cases.
[0044] As shown in the figure, as Figures 1A to 1B shown, a data exchange process 111 can be performed to: (i) upload a subset of time series data metrics to the time series database 112 of the edge computing resource 110; and (ii) upload a subset of time series data metrics to the time series database 121 of the central computing resource 120. In some implementations, the time series database 112 can operate as a cache or similar short-term memory database. This may be the case because the time series data utilized within the edge computing resource 110 will typically be relatively fresh / current to improve prediction accuracy.
[0045] As shown in the figure, the edge computing resource 110 also includes a forecasting engine process 113. The forecasting engine process 113 (e.g., a Python computing process) can be executed to read / load a subset of time series data metrics from the time series database 112, and select a forecasting model stored in the model configuration database 115 that is most suitable for predicting the corresponding subset of time series data metrics. For example, the forecasting engine process 113 can be executed to determine that a first forecasting model stored in the model configuration database 115 is most suitable for predicting a first subset of time series data metrics. Correspondingly, the forecasting engine process 113 can be executed to determine that a second forecasting model stored in the model configuration database 115 is most suitable for predicting a second subset of time series data metrics, and so on. The forecasting engine process 113 can be executed to upload the parameters of the selected forecasting model(s) from the model configuration database 115 and deploy the selected forecasting model(s) to make predictions based on the subset of time series data metric data stored / cached in the time series database 112. In some implementations, the forecasting engine 113 can also be executed to write / upload the predictions of the selected forecasting model(s) to the time series database 112.
[0046] As described above, the drift detection process 117 can be executed to detect drifts involving a subset of time series data metrics stored in the time series database 112. For example, the drift detection process 117 can be executed to detect a drift involving a first subset of time series data metrics. The detected drift involving the first subset of time series data metrics can involve one or more of the constituent time series data metrics of the first subset of time series data metrics. In some examples, a drift involving a first time series data metric of a subset can suggest / anticipate a drift in other time series data metrics of the subset. As described above, "drift" can refer to the statistical properties of a time series data metric (or time series data metrics) that change in an unforeseen manner over time. These unforeseen changes can cause the prediction of a forecasting model deployed to predict the drift time series data metric (or time series data metrics) to deteriorate. Thus, in certain implementations, the drift detection process 117 can be executed to detect a drift involving a first subset of time series data metrics by comparing the historical predictions of a first forecasting model deployed to predict the first subset of time series data metrics with the corresponding portion of the first subset of time series data metrics over a time interval. However, in other implementations, such data drift can be detected using other techniques, such as using the Kolmogorov-Smirnov (K-S) test, using the population stability index, using the Page-Hinkley method, or using other drift detection techniques / algorithms. In some implementations, in response to a drift being detected, the drift detection process 117 can be executed to send a drift detection event to the central computing resource 120 (or more specifically, the model scheduler process 122 of the central computing resource 120). The drift detection event can indicate that a drift has been detected for the first subset of time series data metrics and that the central computing resource 120 (or more specifically, the model scheduler process 122 of the central computing resource 120) should initiate the training of instances of the stored forecasting models M_1–M_N to replace the first forecasting model. In certain implementations, detecting a drift for a single time series data metric alone may not be sufficient to trigger the sending of a drift detection event. For example, in these implementations, an offset may need to be detected for multiple time series data metrics within the first subset of time series data metrics before the drift detection process 117 is executed to send a drift detection event. In some implementations, each time series data metric within the first subset of time series data metrics can be assigned a weight that measures its relative importance level. Such weights can be compared with a threshold to determine whether to send a drift detection event. For example, detecting a drift for a single time series data metric with a weight (or other parameter) exceeding the threshold may be sufficient to trigger the sending of a drift detection event.In contrast, before a drift detection event is sent, the drift may need to be detected for multiple time series data metrics with weights below a threshold.
[0047] In response to the detection of a drift involving a first subset of time series data metrics (e.g., in response to receiving a drift detection event), the central computing resource 120 can train, validate, and select an instance of the prediction model M_2 to predict the (drifted) first subset of time series data metrics. After the trained instance of the prediction model M_2 has been selected and uploaded to the model registry database 125 of the central computing resource 120, the edge computing resource 110 can be notified. For example, the model registry database 125 can issue a notification callback to the download model scheduler process 114 to notify the edge computing resource 110 that a new profile has been uploaded to the model registry database 125, and this new profile stores the trained parameters of the prediction model M_2. In some examples, the notification callback can identify the trained instance of the prediction model M_2 that is selected to predict the (drifted) first subset of time series data metrics. Thus, the download model scheduler process 114 can be executed to download the new profile and upload the parameters of the trained instance of the prediction model M_2 to the model configuration database 115. Then the prediction engine process 113 can be executed to retrieve the parameters of the trained instance of the prediction model M_2 and deploy the trained instance of the prediction model M_2 to predict the (drifted) first subset of time series data metrics.
[0048] Now referring to the central computing resource 120, the central computing resource 120 includes a model scheduler process 122. The model scheduler process 122 can be executed to effectively manage / orchestrate prediction model training and validation in a manner similar to how the prediction engine process 113 is executed to manage / orchestrate the deployment of prediction models during inference.
[0049] For example, the edge computing resource 110 may notify the model scheduler process 122 (or more generally the central computing resource 120) that drift has been detected for a first subset of the time series data metrics. In response to such drift detection, the model scheduler process 122 may be executed to retrieve a first model training and validation task from the task database 123 of the central computing resource 120 (e.g., a relational database storing the relationship between the forecasting model and the validation task). The model scheduler process 122 may then be executed to assign the first model training and validation task to the model listener process 124(1). In a similar manner, the edge computing resource 110 may notify the model scheduler process 122 (or more generally the central computing resource 120) that drift has been detected for a second subset of the time series data metrics. Correlatively, another edge computing resource (not depicted) of the computing system 100 may notify the model scheduler process 122 (or more generally the central computing resource 120) that drift has been detected for a third subset of the time series data metrics. Although the examples provided herein involve two additional drift detections, it should be understood that there may be more simultaneous drift detections across many edge computing resources associated with many customers. As described above, the model scheduler process 122 may be executed to assign / distribute the model training and validation tasks corresponding to these various drift detections in parallel to the model listener processes 124(1)-(n). This parallel assignment / distribution of the model training and validation tasks across multiple model listener processes 124(1)-(n) may increase the overall training / validation speed for the computing system 100.
[0050] Now referring to the model listener process 124(1), after being assigned the first model training and validation task, the model listener process 124(1) may be executed to perform the first model training and validation task by: training an instance of the stored forecasting models M_1–M_N using (fresh) training time series data derived from the first subset of the time series data metrics. Here, the stored forecasting models M_1–M_N may include various types of forecasting models, including statistical models (such as AutoArima and Prophet), neural network models (such as long short-term memory (LSTM)), or any other type of machine learning model used for prediction.
[0051] In various implementations, the model listener process 124(1) can be executed to retrieve instances of the stored forecasting models M_1–M_N from a library / database such as the Darts library. Here, such a library (e.g., the Darts library) can provide storage for the forecasting models M_1–M_N. The library can also provide an API that can process time series data / time series data metrics. The use of the library can also facilitate easier comparison between the trained instances of the forecasting models using statistical indices such as the Mean Absolute Percentage Error (MAPE). The library can also support the construction / modification of instances of the stored forecasting models.
[0052] As shown, the model listener process 124(1) can be executed to retrieve a first subset of time series data metrics from the time series database 121 and split the first subset of time series data metrics into training time series data and validation time series data. In various implementations, the training time series data can be derived from a first time interval of the first subset of time series data metrics, and the validation time series data can be derived from a second time interval of the first subset of time series data metrics.
[0053] After using the training time series data derived from the first subset of time series data metrics to train the instances of the stored forecasting models M_1–M_N, the model listener process 124(1) can be executed to validate their performance. For example, the model listener process 124(1) can be executed to compare the predictions of the trained instances of the stored forecasting models M_1–M_N with the corresponding (fresh) validation time series data that is also derived from the first subset of time series data metrics. The model listener process 124(1) can be executed to perform this comparison using various statistical comparisons / indices such as the Mean Absolute Percentage Error (MAPE). For example, the model listener process 124(1) can be executed to calculate the MAPE index for each trained instance of the forecasting models M_1–M_N. Such a calculation can be performed for each constituent time series data metric of the first subset of time series data metrics. Then the model listener process 124(1) can be executed to compare the MAPE indices across the trained instances of the forecasting models M_1–M_N and across the time series data metrics.
[0054] Based on these comparisons, the model listener process 124(1) can be executed to determine that among the trained instances of the stored prediction models M_1–M_N, the trained instance of the stored prediction model M_2 has the lowest prediction error for a first subset of the predicted time series data metrics. In response to determining that the trained instance of the stored prediction model M_2 has a lower prediction error for the first subset of the predicted time series data metrics than the first prediction model (i.e., the prediction model currently deployed at the edge computing resource 110 for the first subset of the predicted time series data metrics), the model listener process 124(1) can be executed to update the model registry database 125 of the central computing resource 120 with the parameters of the trained instance of the stored prediction model M_2. As described above, this can include uploading a configuration file that stores the trained parameters of the instance of the stored prediction model M_2.
[0055] As described above, in response to the central computing resource 120 updating the model registry database 125 with the parameters of the trained instance of the stored prediction model M_2, the edge computing resource 110 can execute the download model scheduler process 114 to download the parameters of the trained instance of the stored prediction model M_2 and upload such parameters to the model configuration database 115 of the edge computing resource 110. Then the prediction engine process 113 can be executed to retrieve the parameters of the trained instance of the stored prediction model M_2 from the model configuration database 115 and deploy the trained instance of the stored prediction model M_2 to predict the first subset of the time series data metrics. In this example, the trained instance of the stored prediction model M_2 can replace the first prediction model that has become less suitable for predicting the drifted first subset of the time series data metrics.
[0056] Figure 2 An example computing resource 210 for automatically training, validating, and selecting a prediction model is shown in accordance with various examples of the techniques of the present disclosure. In some implementations, the computing resource 210 can be a central computing resource (e.g., a centralized cloud computing cluster) similar to Figures 1A to 1B the central computing resource 120. In these implementations, the computing resource 210 can include the central computing resource within a larger computing system that includes one or more edge computing resources (in some of these implementations, the computing resource 210 and the edge computing resources can be implemented using a containerized environment). However, in other implementations, the computing resource 210 need not be a central computing resource. For example, the computing resource 210 can be a computing resource deployed at the edge (i.e., an edge computing resource).
[0057] Now refer to Figure 2, the computing resource 210 can be, for example, a server computer, a controller, or any other similar computing component capable of processing data. In Figure 2 In an example implementation, the computing resource 210 includes a hardware processor 212 and a machine-readable storage medium 214.
[0058] The hardware processor 212 can be one or more central processing units (CPUs), semiconductor-based microprocessors, and / or other hardware devices suitable for fetching and executing instructions stored in the machine-readable storage medium 214. The hardware processor 212 can obtain, decode, and execute instructions such as instructions 216 - 222. As an alternative or supplement to fetching and executing instructions, the hardware processor 212 can include one or more electronic circuits that include electronic components for performing the functions of one or more instructions, such as a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), or other electronic circuits.
[0059] A machine-readable storage medium (such as the machine-readable storage medium 214) can be any electrical, magnetic, optical, or other physical storage device that contains or stores executable instructions. Thus, the machine-readable storage medium 214 can be, for example, random access memory (RAM), non-volatile RAM (NVRAM), electrically erasable programmable read-only memory (EEPROM), a storage device, an optical disc, etc. In some examples, the machine-readable storage medium 214 can be a non-transitory storage medium, where the term "non-transitory" does not cover transitory propagation indicators. As described in detail below, the machine-readable storage medium 214 can be encoded with executable instructions, such as instructions 216 - 222. Additionally, although Figure 2 the instructions shown in are in sequential order, the shown order is not the only order in which the instructions can be executed. Any instruction can be executed at any time in any order, can be repeatedly executed, and / or can be executed by any suitable one or more devices.
[0060] As shown, in response to detecting a drift in time series data for which a first prediction model is deployed to predict, the hardware processor 212 can execute instruction 216 to train an instance of the stored prediction model using training time series data derived from the time series data. In some implementations, the first prediction model can be deployed at an edge computing resource. In certain of these implementations, the edge computing resource can detect a drift in the time series data. In some implementations, the computing resource 210 and the edge computing resource can be part of a common computing system (such as Figures 1A to 1B the computing system 100). As described above, in various implementations, the time series data can include a subset of time series data metrics, although this need not be the case.
[0061] The hardware processor 212 may execute instructions 218 to compare the predictions of the trained instances of the stored prediction models with the corresponding validation time series data derived from the time series data. In some implementations, the training time series data and the validation time series data may be derived from different time intervals of the time series data.
[0062] Based on this comparison, the hardware processor 212 may execute instructions 220 to determine that among the trained instances of the stored prediction models, the trained instance of the second prediction model has the lowest prediction error for the predicted time series data.
[0063] In response to determining that the trained instance of the second prediction model has a lower prediction error for the predicted time series data than the first prediction model, the hardware processor 212 may execute instructions 222 to update the model registry of the computing resource 210 with the parameters of the trained instance of the second prediction model. As described above, in some implementations, the edge computing resource may download the parameters of the trained instance of the second prediction model from the model registry database and deploy the trained instance of the second prediction model to predict the time series data.
[0064] As described in connection with Figures 1A to 1B In some implementations, the hardware processor 214 may execute the machine-readable instructions of the model scheduler process to: (a) retrieve model training and validation tasks from the task database of the computing resource 210, where the first model training and validation task utilizes the time series data; and (b) assign the model training and validation tasks to the model listener processes, where the first model training and validation task is assigned to the first model listener process. In some implementations, the hardware processor 212 may execute the machine-readable instructions of the model scheduler process to assign the model training and validation tasks to the model listener processes in parallel. In some implementations, the hardware processor 212 may execute the machine-readable instructions of the first model listener process to perform the first model training and validation task.
[0065] Figure 3 An example edge computing resource 310 is shown for deriving a subset of time series data metrics, detecting drifts involving the subset of time series data metrics, and automatically deploying instances of new prediction models trained to predict the (drifted) subset of time series data metrics, according to various examples of the techniques of the present disclosure. The edge computing resource 310 may be similar in form and function to Figures 1A to 1B the edge computing resource 110. In these implementations, the edge computing resource 310 may be within a larger computing system that includes a central computing resource (e.g., Figures 1A to 1B the central computing resource 120, Figure 2 the computing resource 210, etc.) (e.g., Figures 1A to 1BThe computing system 100) is implemented.
[0066] Now refer to Figure 3 , the edge computing resource 310 can be, for example, a server computer, a controller, or any other similar computing component capable of processing data. In Figure 3 In an example implementation, the edge computing resource 310 includes a hardware processor 312 and a machine-readable storage medium 314.
[0067] The hardware processor 312 can be one or more central processing units (CPUs), semiconductor-based microprocessors, and / or other hardware devices suitable for fetching and executing instructions stored in the machine-readable storage medium 314. The hardware processor 312 can fetch, decode, and execute instructions, such as instructions 316 - 328. As an alternative or supplement to fetching and executing instructions, the hardware processor 312 can include one or more electronic circuits that include electronic components for performing the functions of one or more instructions, such as a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), or other electronic circuits.
[0068] A machine-readable storage medium (such as the machine-readable storage medium 314) can be any electrical, magnetic, optical, or other physical storage device that contains or stores executable instructions. Thus, the machine-readable storage medium 314 can be, for example, random access memory (RAM), non-volatile RAM (NVRAM), electrically erasable programmable read-only memory (EEPROM), a storage device, an optical disc, etc. In some examples, the machine-readable storage medium 314 can be a non-transitory storage medium, where the term "non-transitory" does not cover transitory propagation indicators. As described in detail below, the machine-readable storage medium 314 can be encoded with executable instructions, such as instructions 316 - 328. Additionally, although Figure 3 the instructions shown in are in sequential order, the shown order is not the only order in which the instructions can be executed. Any instruction can be executed at any time in any order, can be executed repeatedly, and / or can be executed by any suitable one or more devices.
[0069] As shown, the hardware processor 312 can execute instruction 316 to stream data from a customer.
[0070] The hardware processor 312 may execute instructions 318 to derive a subset of time series data metrics from the data streamed from the client. Deriving a subset of time series data metrics from the streamed client data may include: (a) indexing the data points of the streamed data in chronological order to generate time series data metrics; and (b) logically grouping together a subset of the time series data metrics according to heuristics. In some implementations, the heuristics may include logically grouping together time series data metrics derived from a common client subsystem as a subset of the time series data metrics.
[0071] The hardware processor 312 may execute instructions 320 to upload the subset of time series data metrics to the time series database of the edge computing resource. Correlatively, the hardware processor 312 may execute instructions 322 to upload the subset of time series data metrics to the time series database of a central computing resource (e.g., Figures 1A to 1B central computing resource 120, Figure 2 computing resource 210, etc.).
[0072] In some examples, the execution of instructions 318 - 322 by the hardware processor 312 may correspond to machine - readable instructions that execute a data exchange process of the edge computing resource 310.
[0073] The hardware processor 312 may execute instructions 324 to detect drift involving the subset of time series data metrics. In various implementations, this may include comparing, over a time interval, the historical predictions of a first prediction model deployed to predict the subset of time series data metrics with the corresponding portions of the (actual) subset of the time series data metrics. The magnitude of the increase in the deviation between the predicted value and the actual value (e.g., a deviation exceeding a threshold) may indicate the occurrence of data drift. As described above, the detection of such drift may trigger the central computing resource to train, validate, and select a second prediction model to predict the (drifted) subset of the time series data metrics.
[0074] In response to the central computing resource updating the model registry with the parameters of the trained instance of the selected second prediction model, the hardware processor 312 may execute instructions 326 to download the parameters of the trained instance of the second prediction model from the model registry.
[0075] The hardware processor 312 may then execute instructions 328 to deploy the trained instance of the second prediction model for predicting the (drifted) subset of the time series data metrics.
[0076] Figure 4 An example chart 400 is depicted that compares the predictions of a prediction model with the actual / historical data that the prediction model is deployed to predict, according to various examples of the techniques of the present disclosure.
[0077] The x - axis of chart 400 represents duration. The y - axis of chart 400 represents a scalar value for a particular parameter (e.g., the voltage of a temperature sensor in a machine, the temperature in London, etc.). Here, the scalar value of the time index for a particular parameter (i.e., the time - indexed data point that describes the particular parameter) can be referred to as a first time - series data metric.
[0078] The curve 410 in chart 400 (i.e., the curve depicted by a solid line) represents the actual / historical values for the first time - series data metric. The curve 420 in chart 400 (i.e., the curve depicted by a dashed line from time T1 to time T3) represents the prediction of a first forecasting model that is deployed to predict the first time - series data metric. The curve 422 in chart 400 (i.e., the curve depicted by a dashed line after time T3) represents the prediction of a second forecasting model that is deployed to predict the first time - series data metric after a drift is detected.
[0079] As shown, the prediction of the first forecasting model is very close to the actual / historical values for the first time - series data metric between times Tl and T2. However, starting at time T2, the prediction of the first forecasting model begins to deviate significantly from the actual / historical values for the first time - series data metric. As a visual representation of this deviation, curves 410 and 422 begin to separate in the drift window 450 (i.e., the shaded region between times T2 and T3). As described above, the prediction of the first forecasting model may begin to deteriorate at time T2 due to the occurrence of a drift. As used herein, “drift” (sometimes referred to as “data drift”) can refer to the statistical property of a time - series data metric (e.g., the first time - series data metric) that changes in an unforeseen way over time. These unforeseen changes can cause the prediction of a forecasting model (e.g., the first forecasting model) that is deployed to predict the time - series data metric with drift (e.g., the first time - series data metric) to deteriorate.
[0080] As an example, the first time - series data metric can be related to an operating parameter of a machine, such as the voltage of a temperature sensor. Changes to the voltage sensor, or other structural / functional modifications to the machine, may cause a drift. The first forecasting model may not be well - suited to predict the first time - series data metric after such a drift. Thus (and as described above), in response to detecting a drift involving the first time - series data metric, examples of the techniques of the present disclosure can automatically train, validate, and select an instance of a second forecasting model to predict the (drifted) first time - series data metric.
[0081] Examples can utilize various techniques to detect drift involving a first time series data metric. For example, an example can compare historical predictions of a first forecasting model with corresponding actual / historical values of the first time series data metric over a time interval (e.g., between times T2 and T3). If an example detects that the deviation exceeds a predetermined threshold, they can determine the occurrence of drift. In other implementations, examples can use other techniques to detect drift, such as using the Kolmogorov-Smirnov (K-S) test, using the population stability index, using the Page-Hinkley method, or using other drift detection techniques / algorithms.
[0082] After the second forecasting model is trained / validated, an example can deploy the second forecasting model (i.e., a trained instance of the second forecasting model) to predict the first time series data metric (with drift). As shown in chart 400, the second forecasting model is deployed at time T3. As can be seen in chart 400, the predictions of the second forecasting model track closer to the actual / historical values of the first time series data metric (with drift). Time T4 can represent the current time - after which only the predictions of the second forecasting model are available.
[0083] Figure 5 A block diagram depicts an example computer system 500 in which various embodiments described herein can be implemented. For example, Figures 1A to 1B edge computing resources 110 and central computing resources 120, Figure 2 computing resources 210, and Figure 3 edge computing resources 310 can be implemented using computing system 500. Computing system 500 includes a bus 502 or other communication mechanism for conveying information, and one or more hardware processors 504 coupled to bus 502 for processing information. The (multiple) hardware processors 504 can be, for example, one or more general-purpose microprocessors.
[0084] Computing system 500 also includes main memory 506, such as random access memory (RAM), cache, and / or other dynamic storage devices, coupled to bus 502 for storing information and instructions to be executed by processor 504. Main memory 506 can also be used to store temporary variables or other intermediate information during the execution of instructions to be executed by processor 504. When such instructions are stored in a storage medium accessible to processor 504, computing system 500 is presented as a special-purpose machine customized to perform the operations specified in the instructions.
[0085] The computer system 500 also includes a read-only memory (ROM) 508 or other static storage device coupled to the bus 502 for storing static information and instructions for the processor 504. A storage device 510, such as a magnetic disk, optical disk, or USB thumb drive (flash drive), is provided and coupled to the bus 502 for storing information and instructions.
[0086] The computer system 500 may be coupled via the bus 502 to a display 512, such as a liquid crystal display (LCD) (or touch screen), for displaying information to a computer user. An input device 514, including alphanumeric keys and other keys, is coupled to the bus 502 for communicating information and command selections to the processor 504. Another type of user input device is a cursor control 516, such as a mouse, trackball, or cursor direction keys, for communicating direction information and command selections to the processor 504 and for controlling cursor movement on the display 512. In some embodiments, the same direction information and command selections as for cursor control may be implemented via receiving touches on a touch screen without a cursor.
[0087] The computing system 500 may include a user interface module for implementing a GUI, which may be stored as executable software code executed by the (one or more) computing devices in a mass storage device. By way of example, this module and other modules may include components such as software components, object-oriented software components, class components, and task components, processes, functions, attributes, procedures, subroutines, program code segments, drivers, firmware, microcode, circuitry, data, databases, data structures, tables, arrays, and variables.
[0088] In general, terms such as "component", "engine", "system", "database", "data store", etc. as used herein may refer to logic embodied in hardware or firmware, or to a collection of software instructions that may have entry and exit points and are written in a programming language such as, for example, Java, C, or C++. Software components may be compiled and linked into an executable program, installed in a dynamic link library, or may be written in an interpreted programming language such as, for example, BASIC, Perl, or Python. It should be understood that software components may call from other components or from themselves, and / or may be called in response to detected events or interrupts. Software components configured to execute on a computing device may be provided on a computer-readable medium such as a compact disc, digital video disc, flash drive, magnetic disk, or any other tangible medium, or may be provided as a digital download (and may initially be stored in a compressed or installable format that requires installation, decompression, or decryption before execution). Such software code may be stored, in whole or in part, on the memory device of the executing computing device for execution by the computing device. Software instructions may be embedded in firmware such as an EPROM. It should also be understood that hardware components may include connected logic units such as gates and flip-flops, and / or may include programmable units such as programmable gate arrays or processors.
[0089] Computer system 500 may implement the techniques described herein using custom hardwired logic, one or more ASICs or FPGAs, firmware, and / or program logic, which in combination with the computer system causes or programs computer system 500 to be a special-purpose machine. According to one embodiment, the techniques herein are performed by computer system 500 in response to one or more sequences of one or more instructions contained in main memory 506 being executed by (one or more of) processors 504. Such instructions may be read into main memory 506 from another storage medium such as storage device 510. Execution of the instruction sequences contained in main memory 506 causes (one or more of) processors 504 to perform the processing steps described herein. In an alternative embodiment, hardwired circuitry may be used in place of or in combination with software instructions.
[0090] As used herein, the term "non-transitory medium" and like terms refer to any medium that stores data and / or instructions that cause a machine to operate in a particular manner. Such non-transitory media can include non-volatile media and / or volatile media. Non-volatile media includes, for example, optical or magnetic disks, such as storage device 510. Volatile media includes dynamic memory, such as main memory 506. Common forms of non-transitory media include, for example, floppy disks, flexible disks, magnetic disks, hard disks, solid state drives, magnetic tape, or any other magnetic data storage medium, CD-ROM, any other optical data storage medium, any physical medium with hole patterns, RAM, PROM, and EPROM, FLASH-EPROM, NVRAM, any other memory chip or cartridge, and network versions thereof.
[0091] Non-transitory media is different from transmission media, but can be used in combination with transmission media. Transmission media participates in the transfer of information between non-transitory media. For example, transmission media includes coaxial cables, copper wire, and fiber optics, including the wires that make up bus 502. Transmission media can also take the form of acoustic or light waves, such as those generated during radio wave and infrared data communications.
[0092] Computer system 500 also includes a communication interface 518 coupled to bus 502. Network interface 518 provides two-way data communication coupling to one or more network links connected to one or more local networks. For example, communication interface 518 can be an Integrated Services Digital Network (ISDN) card, cable modem, satellite modem, or a modem that provides a data communication connection to a corresponding type of telephone line. As another example, network interface 518 can be a Local Area Network (LAN) card to provide a data communication connection to a compatible LAN (or a WAN component for communicating with a WAN). A wireless link can also be implemented. In any such implementation, network interface 518 transmits and receives electrical, electromagnetic, or optical indicators that carry digital data streams representing various types of information.
[0093] Network links typically provide data communication to other data devices through one or more networks. For example, a network link can provide a connection through a local network to a host computer or to a data device operated by an Internet Service Provider (ISP). The ISP in turn provides data communication services through the global packet data communication network now commonly referred to as the "Internet". Both local networks and the Internet use electrical, electromagnetic, or optical indicators that carry digital data streams. The indicators through various networks and the indicators on network links and through communication interface 518 are example forms of transmission media that carry digital data to and from computer system 500.
[0094] The computer system 500 can send messages and receive data, including program code, via (a) network(s), network link, and communication interface 518. In an Internet example, a server can transmit request code for an application through the Internet, an ISP, a local network, and communication interface 518.
[0095] The received code can be executed by the processor 504 when it is received, and / or stored in the storage device 510 or other non-volatile storage for later execution.
[0096] Each of the processes, methods, and algorithms described in the foregoing sections can be embodied in code components executed by one or more computer systems or computer processors including computer hardware, and be fully or partially automated by the code components. The one or more computer systems or computer processors can also operate to support the execution of related operations in a “cloud computing” environment or as “software as a service” (SaaS). These processes and algorithms can be implemented in part or in whole in dedicated circuitry. The various features and processes described above can be used independently of each other, or can be combined in various ways. Different combinations and sub-combinations are intended to fall within the scope of this disclosure, and certain method or process blocks can be omitted in some implementations. The methods and processes described herein are also not limited to any particular order, and the blocks or states associated therewith can be executed in other appropriate orders, or can be executed in parallel or in some other manner. Blocks or states can be added to or removed from the disclosed example embodiments. The execution of some of the operations or processes can be distributed among computer systems or computer processors, not only residing within a single machine, but deployed across multiple machines.
[0097] As used herein, a circuit can be implemented using any form of hardware, software, or a combination thereof. For example, one or more processors, controllers, ASICs, PLAs, PALs, CPLDs, FPGAs, logic components, software routines, or other mechanisms can be implemented to constitute a circuit. In an implementation, the various circuits described herein can be implemented as discrete circuits, or the described functions and features can be partially or fully shared among one or more circuits. Although the elements of various features or functions can be described or claimed separately as separate circuits, these features and functions can be shared among one or more common circuits, and such description should not require or imply the need for separate circuits to implement such features or functions. In cases where the circuit is implemented in whole or in part using software, such software can be implemented to operate with a computing or processing system (such as computer system 500) capable of executing the functions described thereof.
[0098] As used herein, the term "or" may be construed in an inclusive or exclusive sense. Additionally, the description of a singular resource, operation, or structure should not be construed as excluding a plurality. Conditional language, such as "can," "could," "might," or "may," unless specifically stated otherwise or otherwise understood in the context in which it is used, is generally intended to convey that certain embodiments include, while other embodiments do not include, certain features, elements, and / or steps.
[0099] Unless otherwise expressly stated, the terms and phrases used in this document and their variants should be construed as open-ended rather than restrictive. Adjectives (such as "conventional," "traditional," "ordinary," "standard," "known") and terms of similar import should not be construed as limiting the items described to a given time period or to items available as of a given time; rather, they should be understood to encompass conventional, traditional, ordinary, or standard techniques that are available or known at any time, present or future. In some instances, the presence of broad words and phrases (such as "one or more," "at least," "but not limited to," or other similar phrases) should not be taken to imply that a narrower case is intended or required where such broad phrases might not be present.
Claims
1. A system comprising: A computing resource operable to execute machine-readable instructions to: In response to detecting a drift involving time series data, training a plurality of instances of stored forecast models using training time series data derived from the time series data, a first forecast model being deployed to forecast the time series data; as well as In response to determining that the trained instance of a second forecast model has a lower prediction error for predicting the time series data than the first forecast model, updating a model registry database of the computing resource with parameters of the trained instance of the second forecast model.
2. The system of claim 1 , further comprising an edge computing resource at which the first forecast model is deployed, wherein the edge computing resource is operable to execute machine-readable instructions to: detecting said drift involving said time series data; and In response to the computing resource updating the model registry database with the parameters of the trained instance of the second forecast model, downloading the parameters of the trained instance of the second forecast model from the model registry database and deploying the trained instance of the second forecast model for predicting the time series data.
3. The system of claim 2, wherein detecting the drift involving the time series data comprises: Historical predictions of the first forecast model are compared to corresponding portions of the time series data over a time interval.
4. The system of claim 1 , wherein the computing resource is further operable to execute machine-readable instructions to: It is determined that among the plurality of trained instances of the stored forecast models, the trained instance of the second forecast model has a lowest prediction error for predicting the time series data.
5. The system of claim 4, wherein determining that among the plurality of trained instances of the stored forecast models, the trained instance of the second forecast model has the lowest prediction error for predicting the time series data comprises: Predictions of the plurality of instances trained by the stored forecast model are compared to corresponding validation time series data derived from the time series data. 6 . The system of claim 5 , wherein the training time series data and the validation time series data are derived from different time intervals of the time series data.
7. The system of claim 2, wherein: The time series data includes a subset of time series data metrics; and The edge computing resource is operable to execute machine-readable instructions of a data exchange process to: Streaming data from customers; deriving said subset of time series data metrics from said data streamed from said client; Uploading the subset of time series data metrics to a time series database of the edge computing resource; as well as The subset of time series data metrics is uploaded to a time series database of the computing resource.
8. The system of claim 7, wherein deriving the subset of time series data metrics from the data streamed from the client comprises: Indexing data points of the data in time order to generate a time series data metric; as well as The subsets of time series data metrics are logically grouped together based on a heuristic.
9. The system of claim 8, wherein the heuristics include logically grouping together time series data metrics derived from common client subsystems as a subset of time series data metrics.
10. The system of claim 1, wherein: The computing resources are operable to execute machine readable instructions of a model scheduler process to: Retrieving model training and validation tasks from a task database of the computing resource, wherein a first model training and validation task utilizes the time series data; and The model training and validation tasks are assigned to a model listener process, wherein the first model training and validation task is assigned to a first model listener process; and the computing resource is operable to perform the first model training and validation task by executing the training and model registry database update steps of claim 1.
11. The system of claim 8, wherein the computing resources are operable to execute machine-readable instructions of the model scheduler process to assign the model training and validation tasks to the model listener process in parallel.
12. The system of claim 2, wherein the computing resources and the edge computing resources are implemented using a containerized computing environment.
13. A method comprising: In response to detecting a drift involving a subset of the time series data metrics, training a plurality of instances of the stored forecast models using training time series data derived from the subset of the time series data metrics, a first forecast model being deployed to predict the subset of the time series data metrics; as well as In response to determining that the trained instance of a second forecast model has a lower prediction error for the subset of forecasted time series data metrics than the first forecast model, updating a model registry database with parameters of the trained instance of the second forecast model.
14. The method according to claim 13, further comprising: It is determined that among the plurality of trained instances of the stored forecast models, the trained instance of the second forecast model has a lowest prediction error for the subset of forecasting time series data metrics.
15. The method of claim 14, wherein determining that among the plurality of trained instances of the stored forecast models, the trained instance of the second forecast model has the lowest prediction error for the subset of forecasted time series data metrics comprises: Predictions of the plurality of instances of the stored forecast model trained are compared to corresponding validation time series data derived from the subset of time series data metrics.
16. The method of claim 15, wherein the training time series data and the validation time series data are derived from different time intervals of the subset of time series.
17. The method according to claim 13, further comprising: Execute the Model Scheduler process to: Retrieving model training and validation tasks from a task database, wherein a first model training and validation task is associated with the subset of time series data metrics; and Assigning the model training and validation tasks to a model listener process, wherein the first model training and validation tasks are assigned to a first model listener process; and executing the first model listener process to perform the first model training and validation tasks by executing the training and model registry database update steps of claim 13.
18. The method of claim 17, wherein executing the model scheduler process to assign the model training and validation tasks to the model listener process comprises: The model scheduler process is executed to assign the model training and validation tasks to the model listener process in parallel.
19. A non-transitory computer readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to: In response to detecting a drift involving a subset of the time series data metrics, training a plurality of instances of the stored forecast models using training time series data derived from the subset of the time series data metrics, a first forecast model being deployed to predict the subset of the time series data metrics; and In response to determining that the trained instance of a second forecast model has a lower prediction error for the subset of forecast time series data than the first forecast model, updating a model registry database with parameters of the trained instance of the second forecast model.
20. The non-transitory computer readable medium storing instructions of claim 17, further comprising instructions to: It is determined that among the plurality of trained instances of the stored forecast models, the trained instance of the second forecast model has a lowest prediction error for the subset of forecasting time series data metrics.