Machine learning platform
The machine learning platform addresses the relevance gap by generating and monitoring models tailored to user tasks, ensuring accuracy and efficiency through data augmentation and automated updates, thus optimizing resource use.
Patent Information
- Application Number
- JP2022574853
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-02-18
- Filing Date
- 2021-02-18
- Publication Date
- 2025-07-23
- Estimated Expiration
- 2041-02-18
AI Technical Summary
Existing machine learning models may not be relevant to the specific problems of user applications, risking inadequate performance and resource inefficiency.
A machine learning platform that generates, deploys, and monitors models based on user-defined tasks, augments data, and adjusts learning data and task definitions to ensure model suitability, featuring automated parameter updates and backtesting for accuracy.
Enhances the relevance and accuracy of machine learning models, reducing resource waste by avoiding the use of outdated models and optimizing computational resources.
Smart Images

Figure 0007712299000001 
Figure 0007712299000002 
Figure 0007712299000003
Abstract
Description
Technical Field
[0001] The subject matter disclosed herein generally relates to methods, systems, and programs for a machine learning platform. Specifically, the present disclosure relates to systems, methods, and computer programs for generating and optimizing machine learning models.
Background Art
[0002] Machine learning is a field of study that gives computers the ability to learn (train) without being explicitly programmed. Machine learning explores the research and construction of algorithms, also called tools herein, that can learn from existing data and make predictions about new data. Such machine learning tools operate by constructing a model from example learning data.
Prior Art Documents
Non-Patent Documents
[0003]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] However, such models may not be relevant to the problems that the user is trying to solve in the user application, and thus there is a risk that they may not meet the requirements of the user application.
Means for Solving the Problem
[0005] To easily identify the discussion of particular elements or acts, one or more of the most significant digits of the reference numbers refer to the figure number in which the element is first introduced.
Brief Description of the Drawings
[0006]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Figure 16
Figure 17
Figure 18
[0007] In the following description, systems, methods, techniques, instruction sequences, and computing machine program products showing exemplary embodiments of the present subject matter will be described. In the following description, for the purpose of explanation, numerous specific details are set forth in order to provide an understanding of the various embodiments of the present subject matter. However, it will be apparent to those skilled in the art that the embodiments of the present subject matter may be practiced without some or all of these specific details. The examples are merely typifications of possible variations. Unless explicitly stated otherwise, the structure (e.g., structural components such as modules) is optional and may be combined or subdivided, and the operations (e.g., procedures, algorithms, or other functions) may be in a different order and may be combined or subdivided.
[0008] This application describes a machine learning platform for generating machine learning models. The machine learning platform generates a machine learning model based on user-defined tasks (e.g., problems to be solved or feature exploration), and learning data. Here, the learning data has been prepared (e.g., estimating missing data or filtering inaccurate or outlier data), and / or augmented with additional data from external dataset sources (e.g., additional data from users of the machine learning platform or from the market library of datasets). The machine learning platform deploys the machine learning model for access by applications existing externally to the machine learning platform. The machine learning platform monitors the performance (also referred to as health) of the deployed machine learning model and determines the suitability of the currently deployed machine learning model for the user-defined task. The machine learning platform adjusts the learning data and / or task definition based on the performance / health of the currently deployed machine learning model.
[0009] In one embodiment, a user of a machine learning platform uploads data to the data ingestion system of the machine learning platform. The uploaded data is augmented with additional data from a library of datasets or additional data from the user. The task system of the machine learning platform defines tasks that specify the problems to be solved (e.g., target columns, data exploration). The machine learning platform analyzes whether there are deficiencies or inconsistencies in the data and provides suggestions for better preparing the data for processing by machine learning algorithms. For example, the suggestions of the machine learning platform can include specific sections of the data for processing. This suggestion can affect the algorithmic approach to learning for the given data based on the provided partitioned data and the characteristics of the partitioned data (e.g., boundary values of the partitioned data, or the presence of missing data). The machine learning platform constructs a machine learning model based on the prepared data and evaluates the performance of the machine learning model. When the machine learning platform determines that the performance of the machine learning model is acceptable, it deploys the machine learning model by exporting the machine learning model to a deployment engine. In one example, an external application to the machine learning platform accesses the machine learning model. The machine learning platform monitors the usage and performance of the machine learning model. In one example, the machine learning platform determines that the performance of the machine learning model is no longer acceptable because the machine learning model no longer accurately solves the task. In such a situation, the machine learning platform recommends redefining the task. In another example, for the same task, the machine learning platform proposes different data partitions and different machine learning model search spaces as part of data preparation.In another example, the machine learning platform has a guidance or advice function that solicits additional clarification information from the user to clarify the uploaded data and provide context to the uploaded data.
[0010] In another exemplary embodiment, the machine learning platform provides a system for performing a time series forecast of the value of an item (e.g., stocks, commodities). Specifically, the machine learning platform has an automatic parameter update for an online forecasting model of the value of an item. For example, the machine learning platform uses historical data of the value of an item to generate a machine learning model. New data may have new statistical properties that are not compatible with the statistical properties of the historical data. In such a situation, the machine learning model becomes unsuitable for accurate forecasting. To overcome this situation, the system automatically updates the algorithm parameters based on the available new data. In a further embodiment, the machine learning platform has a backtest function for evaluating the accuracy of the prediction of the value of an item based on the machine learning model.
[0011] In another exemplary embodiment, a machine learning platform operating on a server is described. The machine learning platform accesses a dataset from a data store. A task is defined to identify the target of the machine learning algorithm from the machine learning platform. The machine learning algorithm forms a machine learning model based on the dataset and the task. The machine learning platform deploys the machine learning model and monitors the performance of the deployed machine learning model. The machine learning platform may update the machine learning model based on the monitoring.
[0012] As a result, one or more of the methodologies described herein facilitate the solution of technical problems of machine learning models that are outdated or inaccurate. Thus, one or more of the methodologies described herein can avoid the specific efforts or computing resources that would otherwise be involved in using outdated machine learning models. As a result, the resources used by one or more machines, databases, or devices (e.g., in an environment) may be reduced. Examples of such computing resources include processor cycles, network traffic, memory usage, data storage capacity, power consumption, network bandwidth, and cooling capacity.
[0013] FIG. 1 is a graphical representation of a network environment 100 in which some exemplary embodiments of the present disclosure may be implemented or deployed. One or more application servers 104 provide server-side functionality to network-connected user devices in the form of client devices 106 via a network 102. A web browser 110 (e.g., a browser) and a client application 108 (e.g., an “app”) are hosted and executed on the web browser 110. A user 130 operates the client device 106.
[0014] The Application Program Interface (API) server 118 and the web server 120 provide their respective program interfaces and web interfaces to the application server 104. A specific application server 116 hosts a machine learning platform 122 (which includes components, modules, and / or applications) and a service application 124. The machine learning platform 122 receives learning data (training data) from the client device 106, the third-party server 112, or the service application 124. The machine learning platform 122 generates a machine learning model based on the learning data. The machine learning platform 122 deploys the machine learning model and monitors the performance (e.g., accuracy) of the machine learning model.
[0015] In some exemplary embodiments, the machine learning platform 122 comprises a machine learning program (MLP), also referred to as a machine learning algorithm or tool. The machine learning program is utilized to perform operations related to predicting the value of an item at a future point in time, solving the values of a target column, or discovering features of the learning data.
[0016] Machine learning is a field of study that gives computers the ability to learn without being explicitly programmed. Machine learning explores the research and construction of algorithms, also referred to herein as tools, that can learn from existing data and make predictions about new data. Such machine learning tools operate by constructing a machine learning model from the learning data to make data-driven predictions or decisions, represented as outputs. Exemplary embodiments are presented with respect to several machine learning tools, but the principles presented herein may be applied to other machine learning tools.
[0017] In some exemplary embodiments, different machine learning tools may be used. For example, logistic regression (LR), naive Bayes, random forest (RF), neural network (NN), matrix factorization, and support vector machine (SVM) tools may be used to classify the attributes of the training data or to identify patterns in the training data.
[0018] Two general types of problems in machine learning are classification problems and regression problems. Classification problems, also called categorization problems, aim to classify an item into one of several categorical values (e.g., is this object an apple or an orange). Regression algorithms aim to quantify several items (e.g., by providing a value as a real number). In some embodiments, machine learning algorithms identify significant patterns in relation to other attributes in the training data. These algorithms utilize this training data to model such similar relationships that may affect the prediction results.
[0019] Service application 124 comprises a programmatic application accessed by client device 106. Examples of programmatic applications include a document authoring application, a communication application, a processing application, and an analysis application. Service application 124 exists external to machine learning platform 122. Service application 124 accesses a machine learning model generated in machine learning platform 122 and executes operations based on the machine learning model.
[0020] The web browser 110 communicates with the machine learning platform 122 via a web interface supported by the web server 120. Similarly, the client application 108 communicates with the machine learning platform 122 via a programmatic interface provided by the application program interface (API) server 118. In another example, the client application 108 communicates with the service application 124 via the application program interface (API) server 118.
[0021] The application server 116 is shown communicatively coupled to a database server 126 that facilitates access to an information storage repository or database 128. In an exemplary embodiment, the database 128 comprises a storage device that stores information (e.g., datasets, extended datasets, dataset marketplaces, machine learning model libraries) published and / or processed by the machine learning platform 122.
[0022] Furthermore, a third-party application 114 running on a third-party server 112 is shown having programmatic access to the application server 116 via a programmatic interface provided by the application program interface (API) server 118. For example, the third-party application 114 can support one or more features or functions on a website hosted by a third party by using information obtained from the application server 116. For example, the third-party application 114 provides a learning data marketplace to the machine learning platform 122.
[0023] Any system or machine shown in or related to FIG. 1 (e.g., a database, a device, a server) is modified (e.g., configured or programmed by software such as one or more software modules of an application, an operating system, firmware, middleware, or other programs) to perform one or more of the functions described herein for that system or machine, or is a special-purpose (e.g., special or otherwise non-general-purpose) computer equipped with such a special-purpose computer or can be implemented in such a special-purpose computer. For example, a special-purpose computer system capable of implementing any one or more of the methodologies described herein will be described later with respect to FIG. 16, and such a special-purpose computer may be a means for implementing any one or more of the methodologies described herein accordingly. In the technical field of such special-purpose computers, a special-purpose computer modified to perform the functions discussed herein by the structure discussed herein is technically improved compared to other special-purpose computers that lack the structure discussed herein or cannot perform the functions discussed herein. Therefore, a special-purpose machine configured according to the systems and methods discussed herein provides an improvement over the technology of similar special-purpose machines.
[0024] Furthermore, any two or more of the systems or machines (machines) illustrated in FIG. 1 can be combined into a single system or machine, and the functions described herein for any single system or machine can be subdivided among multiple systems or machines. Furthermore, any number and type of client devices 106 may be embodied within the network environment 100. Furthermore, some components or functions of the network environment 100 may be combined or arranged elsewhere in the network environment 100. For example, a part of the functions of the client device 106 may be embodied in the application server 116.
[0025] FIG. 2 illustrates a machine learning platform 122 according to an example embodiment. The machine learning platform 122 includes a dataset ingestion system 202, a machine learning system 204, a deployment system 206, a monitoring / assessment system 208, a task system 210, and an action system 212.
[0026] The dataset ingestion system 202 retrieves learning data for the machine learning system 204 from a data store 214 in a database 128. The data store 214 includes datasets provided by a client device 106, a service application 124, or a third-party application 114. In an example embodiment, the dataset ingestion system 202 annotates the learning data with statistical characteristics (e.g., mean, variance, n-th difference) and tags (e.g., part of speech for words in text data, day of the week for date-time values, anomaly flag for continuous data). In another exemplary embodiment, the dataset ingestion system 202 analyzes the learning data and determines whether additional learning data (related or complementary to the learning data) is available to further expand the learning data. In one example, the dataset ingestion system 202 requests that the client device 106 provide additional data. In another example, the dataset ingestion system 202 accesses a library of datasets in the data store 214 and expands the learning data with at least one dataset from the library of datasets. In yet another example, the dataset ingestion system 202 accesses a dataset marketplace (e.g., provided by a third-party application 114) to identify a dataset for expanding the learning data. For example, the dataset includes a column of postal codes. The dataset ingestion system 202 identifies the data as postal codes and proposes to expand the dataset by adding another dataset such as "average income" for each postal code from a library of other datasets (e.g., latitude, longitude, elevation, weather factors, social factors).
[0027] In another exemplary embodiment, the dataset ingestion system 202 comprises an advisor function that advises the client device 106 (which provides the dataset 216) on how to prepare the dataset 216 for processing by the machine learning system 204. For example, the dataset ingestion system 202 analyzes the structure of the dataset 216 and advises the client device 106 that the dataset contains missing values that should be corrected prior to processing by the machine learning system 204. In one example, the dataset ingestion system 202 estimates the missing values based on approximation.
[0028] The task system 210 defines tasks for the machine learning system 204. For example, the tasks identify goal parameters (e.g., the problem to be solved, the target sequence, data verification, and test methods, scoring metrics). The task system 210 receives the task definition from the client device 106, the service application 124, or a third-party application 114. In another example, the task system 210 receives an updated task from the action system 212. The task system 210 can also define non-machine learning tasks such as data transformation and analysis.
[0029] The machine learning system 204 uses machine learning algorithms to learn a machine learning model based on the data from the dataset ingestion system 202 and the tasks from the task system 210. In one example, the machine learning system 204 forms and optimizes a machine learning model to solve the tasks defined by the task system 210. Exemplary embodiments of the machine learning system 204 are further described below with respect to FIG. 3.
[0030] The deployment system 206 includes a deployment engine (not shown) that deploys a machine learning model to other applications (external to the machine learning platform 122). For example, the deployment system 206 defines an infrastructure such that the machine learning model exists in a queryable setting and can be used to make predictions in response to requests. An example of deployment includes uploading to the deployment system 206 a machine learning model or parameters for replicating such a model such that the deployment system 206 can then support the machine learning model and expose the associated functionality.
[0031] In another example, the deployment system 206 enables the service application 124 to access and use the machine learning model to generate forecasts and predictions regarding new data. The deployment system 206 stores the model in the model repository 218 of the database 128.
[0032] The monitoring / evaluation system 208 tests and evaluates the performance of the machine learning model (from the deployment system 206). For example, the monitoring / evaluation system 208 evaluates the algorithm and computational performance of the model by running tests and benchmarks against the model and facilitates comparison with other models. In another example, the monitoring / evaluation system 208 can receive verification of quality or performance from the users of the client device 106. In yet another example, the monitoring / evaluation system 208 is configured to enable the model to make predictions and take actions based on any trigger, rather than just application program interface (API) calls from other services, by tracking the usage of the model and monitoring its performance.
[0033] In another example, the deployment system 206 enables the service application 124 to access and use the machine learning model to generate forecasts and predictions regarding new data. In another embodiment, the deployment system 206 stores the model in the model repository 218 of the database 128.
[0034] The action system 212 triggers an external action (e.g., a call to the service application 124) based on predefined conditions. For example, the action system 212 detects that the deployment system 206 has deployed a machine learning model. In response to detecting the deployment of the machine learning model, the action system 212 notifies the service application 124 (e.g., by generating an alert of the deployment and communicating it to the service application 124). Other examples of actions from the action system 212 include re-training of the machine learning model, updating of model parameters, stopping of model functionality (fail-safe feature) if performance drops below a threshold, communication of alerts based on performance or usage (via an email / text / messaging platform).
[0035] The monitoring / evaluation system 208 monitors the deployment of the machine learning model. For example, the monitoring / evaluation system 208 continuously monitors the performance of the machine learning model (used by the service application 124) and provides feedback to the dataset ingestion system 202 and the task system 210 via the action system 212. For example, the service application 124 provides updated tasks to the task system 210 and the latest data to the dataset ingestion system 202. This process is sometimes referred to as meta-learning. In another example, the monitoring / evaluation system 208 monitors characteristics of the data such as the frequency of missing values or outliers and can also adopt different strategies to remedy these problems. In this way, the monitoring / evaluation system 208 narrows down which strategy to use for a given situation by learning which strategy is the most effective.
[0036] In one example embodiment, the monitoring / evaluation system 208 monitors the performance of the machine learning model. For example, the monitoring / evaluation system 208 intermittently evaluates the performance of the machine learning model as new data comes in such that an updated score representing the latest performance of the model can be derived. In another example, the monitoring / evaluation system 208 quantifies and monitors the sensitivity of the machine learning model to noise by perturbing the data and evaluating the impact on the model's score / prediction. After updating the machine learning model, the monitoring / evaluation system 208 can also test the machine learning model with a set of holdout data and confirm that the machine learning model is suitable for deployment (e.g., by comparing the performance of the new model with the performance of the previous model). Model performance can also be quantified in terms of computational time and required resources such that the user can be warned if changes in the frequency or type of data being ingested cause a decrease in efficiency or speed.
[0037] The monitoring / evaluation system 208 determines whether the performance / accuracy of the machine learning model is acceptable (e.g., whether it exceeds a threshold score). If the monitoring / evaluation system 208 determines that the performance / accuracy of the machine learning model is no longer acceptable, the action system 212 redefines the task in the task system 210 or proposes changes to the training data in the dataset ingestion system 202. For example, if the performance is no longer acceptable, the action system 212 issues a warning to the user 130 via communication means (e.g., email / text) and provides suggestions on the cause of the problem and improvement measures. Also, the action system 212 can update the model based on the latest data or stop the model from making predictions. In another exemplary embodiment, these action operations (action behaviors) can be defined by the user in a "if this then that" (IFTTT) manner.
[0038] Figure 3 shows a machine learning system 204 according to another exemplary embodiment. The machine learning system 204 includes a data segmentation (data partitioning) module 302, a task module 304, a model optimization system 306, and an optimization model learning system. The machine learning system 204 receives training data via a dataset ingestion system 202. The dataset ingestion system 202 provides the data to the data segmentation module 302. The data segmentation module 302 summarizes the data. For example, the data is summarized by calculating summary statistics and describing the distribution of the samples. Continuous values are binned and counted. Outliers and anomalies are flagged. The data segmentation module 302 further slices the summarized data into data slices such that the mathematical definition of the information contained in the original data is evenly distributed among the data slices. This is achieved by layering the data partitions and ensures that the data distribution between the slices matches as closely as possible. The data segmentation module 302 provides the data slices to the model optimization system 306.
[0039] The task system 210 provides user-defined tasks to the task module 304. The task module 304 includes a regression tool 310, a classification tool 312, and an unsupervised machine learning tool 314 as different types of machine learning tools. The task system 210 associates (maps) the user-defined task with any one of the machine learning tools. For example, if the user-defined task has the goal of predicting a categorical value, the task system 210 maps the task to the classification tool. The goal of predicting a continuous value will be mapped to the regression tool. If the user-defined task is to find underlying groupings in the data, it will be mapped to the clustering (unsupervised machine learning) tool. In one example, a lookup table is defined and a mapping (mapping) between different types of tasks and the types of machine learning tools is provided.
[0040] The model optimization system 306 learns a machine learning model based on the data slice and the type of machine learning tool. An exemplary embodiment of the model optimization system 306 will be further described below with respect to FIG. 4.
[0041] The optimization model learning system 308 receives the optimized machine learning model from the model optimization system 306 and provides the trained optimized machine learning model to the deployment system 206 by re-learning the model with all available and appropriate data.
[0042] Figure 4 shows a model optimization system according to one exemplary embodiment. The model optimization system 306 includes a model learning module 402, an optimization module 404, and a model performance estimator 406. The optimization module 404 proposes a specific model. The specific model is defined through a set of hyperparameters. These are a set (collection) of named values that together fully specify a particular model that is ready for model fitting to some learning data.
[0043] The model learning module 402 learns a specific model by using a plurality of data subsets. The model performance estimator 406 calculates a score representing the performance of the specific model. The optimization module 404 receives the score and proposes another specific model based on the score. For example, if a model as a random forest is given, the model is learned by using a plurality of data subsets. The performance can be calculated, for example, by using a loss function. If the score is below a threshold, the optimization module 404, for example, navigates the hyperparameter space according to the gradient of the loss function. A new set of values for the hyperparameters of the model is proposed.
[0044] Figure 5 shows a machine learning platform according to another exemplary embodiment. The machine learning platform 122 performs a time series forecast of the value of an item. The machine learning platform 122 includes a dataset ingestion system 502, a model learning system 504, a target specific feature discovery system 506, a prediction system 508, a backtest system 510, a target specific system 512, and an unsupervised machine learning tool 514.
[0045] The dataset capture system 502 receives data (e.g., a signal that changes (varies) over time) from a user of the client device 106. The user selects the signal that the user wishes to predict (from among the time-varying signals).
[0046] The target identification system 512 obtains a target (e.g., a future desired point for predicting the value of an attribute within the data) from the user or from the unsupervised machine learning tool 514. The target has a future value of a given time series based on the current point in time. The user uses the target identification system 512 to select a specific future value in the signal that the user wishes to predict.
[0047] The target identification feature discovery system 506 discovers features of the data. The features have derived signal properties that are temporally represented via a derived signal. Exemplary embodiments of the target identification feature discovery system 506 are further described below with respect to FIG. 6.
[0048] The model learning system 504 learns a machine learning model based on the features and the data. For example, a non-linear regressor is learned for historical points (features) of a plurality of time-varying signals.
[0049] The prediction system 508 generates a prediction of the value of an attribute in the data at a future desired point in time by using the machine learning model. For example, the prediction system 508 predicts the value at the next point in time in a given signal by using the machine learning model (e.g., given the annual data of stores A to F, what is the predicted revenue of store A next year).
[0050] The backtest system 510 evaluates the performance / accuracy of the machine learning model by performing a backtest on the data using the machine learning model. The predicted value output by the machine learning model is compared with the true value at that time period, and the performance of the model is quantified.
[0051] FIG. 6 shows a target specific feature discovery system 506 according to one exemplary embodiment. The target specific feature discovery system 506 includes a spectral signal embedding system 602, a user specific signal generation system 604, a feature reduction system 606, and a feature set optimization system 608.
[0052] The feature set includes a collection of signals and signal transformations. The feature set is provided to the spectral signal embedding system 602. The spectral signal embedding system 602 encapsulates the historical interdependencies of the signals via new features having simple short-term relationships. For example, a seasonal model (e.g., an ETS time series model) is a single and finite example. This is often used in the financial field where a signal is decomposed into multiple elements. Another example of historical interdependence involves complex signals that require more historical points and relationships to represent the complex signal. Exemplary embodiments of the spectral signal embedding system 602 are further described below with respect to FIG. 7.
[0053] For each future target value within the user-defined target set, the feature reduction system 606 measures the dependencies of the new features. The target set includes a collection of targets for a single time series. Exemplary embodiments of the feature reduction system 606 are further described below with respect to FIG. 8.
[0054] The feature set optimization system 608 generates a single feature set for each target within the target set. It is expected that the feature sets useful for short-term forecasts (e.g., the next day) will be very different from the feature sets useful for long-term forecasts (e.g., next year). The number of sales from a store yesterday is probably important for predicting the number of sales from the store tomorrow (probably much less important for several years ahead).
[0055] FIG. 7 is a diagram showing a spectral signal embedding system 602 according to one exemplary embodiment. The spectral signal embedding system 602 includes a signal scanning system 702, an information embedding system 704, and a collector 706.
[0056] For each signal in the feature set, the signal scanning system 702 detects the history points of the signal that encode the information contained in the signal. The information embedding system 704 shifts these points from the past to the present by generating a plurality of new signals. The collector 706 forms an extended feature set by collecting the new signals. For example, consider two signals. One is a very simple signal (e.g., a sine wave), and the other is a very complex signal. It is necessary to include the information from both signals in the signal set. However, from the complex signal, more history points are required to represent the complex signal than the sine wave. The collector 706 combines all these points into one feature set.
[0057] FIG. 8 shows a feature reduction system 606 according to an example embodiment. The feature reduction system 606 includes a first user-specific prediction target point 802 in the future, a correlation measurement or mutual information measurement and ranking 804, a second user-specific prediction target point 806 in the future, and a correlation measurement or mutual information measurement and ranking 808. For example, the forecast target is the sales one day ahead (one day in the future). The training data can be shifted one day backward to represent the future target according to our input signal. The feature reduction system measures the linear or non-linear relationship (e.g., mutual information) between the input signal and the shifted target signal, and cuts off the signals that do not have a high relationship with the shifted target. This process is repeated for each future target point (e.g., the sales one day in the future (one day later), the sales one week in the future (one week later)).
[0058] Figure 9 shows a method for deploying a machine learning model according to an exemplary embodiment. Method 900 may be performed by one or more computing devices as described below.
[0059] Note that in other embodiments, different orderings (sequencing), additional or fewer operations, and different nomenclatures or terms may be used to achieve similar functionality. In some embodiments, various operations may be performed in parallel with other operations in either a synchronous or asynchronous manner. The operations described herein are selected to describe some of the principles of operation in a simplified form.
[0060] In block 902, the dataset ingestion system 202 receives training data. In block 904, the dataset ingestion system 202 augments the training data with additional data from a library of datasets. In block 906, the task system 210 receives a task defined by a user of the machine learning platform 122. In block 908, the machine learning system 204 forms a machine learning model based on the training data and the task.
[0061] Figure 10 illustrates a method for deploying a machine learning model according to an example embodiment. Method 1000 may be performed by one or more computing devices as described below.
[0062] Note that in other embodiments, different orderings, additional or fewer operations, and different nomenclatures or terms can be used to achieve similar functionality. In some embodiments, various operations may be performed in parallel with other operations in either a synchronous or asynchronous manner. The operations described herein are selected to explain some of the principles of operation in a simplified form.
[0063] In block 1002, the machine learning system 204 forms a machine learning model (ML model) based on the learning data and tasks. In block 1004, the monitoring / evaluation system 208 tests the performance of the machine learning model. In block 1006, the deployment system 206 deploys the machine learning model to the service application 124. In block 1008, the monitoring / evaluation system 208 monitors the suitability (suitability, appropriateness) of the deployed machine learning model in the service application 124. The model is considered "suitable" if a pre-specified performance metric (such as log loss or confusion matrix) still ensures satisfactory performance across the tasks defined in block 906 (e.g., the performance exceeds a pre-set performance threshold). In decision block 1010, if the monitoring / evaluation system 208 determines that the machine learning model is "not suitable", the machine learning system 204 learns another machine learning model in block 1002. The monitoring / evaluation system 208 evaluates the performance of the machine learning model in block 1004. Subject to approval, the machine learning model is then deployed and continuously monitored by the monitoring / evaluation system 208 in block 1008, ensuring the continued suitability of the machine learning model.
[0064] FIG. 11 illustrates a method 1100 for generating an external action according to an exemplary embodiment. At block 1102, the dataset ingestion system 202 accesses data. At block 1104, the task system 210 accesses a task defined by a user of the service application 124. At block 1106, the machine learning system 204 forms a machine learning model based on the learning data and the task. At block 1108, the machine learning system 204 deploys the machine learning model at the application server 116. At block 1110, the action system 212 generates an external action based on the machine learning model. Examples of external actions are: sending an alert to an operator, performing a business action or process, marking a transaction as fraudulent.
[0065] FIG. 12 illustrates a method 1200 for monitoring the deployment of a machine learning model according to an example embodiment. At block 1202, the dataset ingestion system 202 accesses data. At block 1204, the task system 210 accesses a task defined by a user of the service application 124. At block 1206, the machine learning system 204 forms a machine learning model based on the learning data and the task. At block 1208, the machine learning system 204 provides the machine learning model to an application existing externally to the machine learning platform 122. At block 1210, the monitoring / evaluation system 208 monitors the deployment of the machine learning model at the application server 116.
[0066] FIG. 13 illustrates a method 1300 for updating a machine learning model according to an exemplary embodiment. In block 1302, the dataset ingestion system 202 accesses data. In block 1304, the task system 210 accesses a task defined by a user of the service application 124. In block 1306, the machine learning system 204 forms a machine learning model based on the learning data and the task. In block 1308, the machine learning system 204 provides the machine learning model to an application existing external to the machine learning platform 122. In block 1310, the monitoring / evaluation system 208 monitors the deployment of the machine learning model. The monitored parameters include, but are not limited to, execution time with new input data, distribution of the generated output, uncertainty of the generated output, memory usage, response time, and throughput. In block 1312, the action system 212 generates internal actions based on the monitored deployment in the external application system (e.g., the service application 124). Examples of external actions include sending an alert to an operator, performing a business action or process, and marking a transaction as fraudulent. In block 1314, the task system 210 updates the task based on the internal actions. In block 1316, the dataset ingestion system 202 accesses updated data from the external application system. In block 1318, the machine learning system 204 updates the machine learning model based on the updated data and the updated task.
[0067] FIG. 14 is a diagram showing a method 1400 for detecting data insufficiency according to an example embodiment. In block 1402, the dataset ingestion system 202 accesses data. In block 1404, the task system 210 accesses a task defined by a user of the service application 124. In block 1406, the machine learning system 204 forms a machine learning model based on the learning data and the task. In block 1408, the machine learning system 204 provides the machine learning model to an application existing externally to the machine learning platform 122. In block 1410, the monitoring / evaluation system 208 monitors the deployment of the machine learning model. In block 1412, the monitoring / evaluation system 208 detects data insufficiency based on the performance of the machine learning model. The data insufficiency may be in the form of missing data (missing data) or inaccurate data (e.g., possibly caused by malfunctioning or non-operating sensors). The feature importance of the predictions made from this data may be queried and compared to typical values in order to discover this. The performance metrics may be compared to typical / expected performance, as well as characteristics of the input data such as mean / standard deviation, quantiles over time, etc. In block 1414, the action system 212 generates internal actions based on the data insufficiency. In block 1416, the dataset ingestion system 202 accesses the missing data based on the internal actions. Given the missing data, it may be possible to find examples of similar data points that have been witnessed in the past. And the missing data may be approximated with values from these similar data points. Alternatively, alternative data can also be found by querying a library of datasets.
[0068] FIG. 15 is a diagram showing a method for forming a machine learning model based on a time series signal according to an example embodiment. The method 1500 may be executed by one or more computing devices as described below.
[0069] In block 1502, the dataset ingestion system 502 receives a dataset that includes a time-varying signal (time-variant signal). In block 1504, the target identification system 512 receives a request to predict a signal from the time-varying signal. In block 1506, the target identification system 512 receives the future time attribute of the signal. In block 1508, the target identification feature discovery system 506 receives a selection of a feature discovery mode. In block 1510, the model learning system 504 learns a machine learning model based on the discovered features. In block 1512, the backtest system 510 executes a backtest of the machine learning model. In block 1514, the prediction system 508 generates a prediction of the value of the signal at a future time attribute.
[0070] FIG. 16 shows a routine 1600 according to one embodiment. In block 1602, the routine 1600 accesses a dataset from a data store by one or more processors of a server. In block 1604, the routine 1600 receives a definition of a task to identify a target of a machine learning algorithm from a machine learning platform operating on the server. In block 1606, the routine 1600 forms a machine learning model based on the dataset and the task by utilizing a machine learning algorithm. In block 1608, the routine 1600 deploys the machine learning model by providing access to the machine learning model to an application that is external to the machine learning platform. In block 1610, the routine 1600 monitors the performance of the deployed machine learning model. In block 1612, the routine 1600 updates the machine learning model based on the monitoring.
[0071] FIG. 17 is a block diagram _1700 showing a software architecture 1704 that can be installed on any one or more of the apparatuses described in this specification. The software architecture 1704 is supported by hardware such as a machine 1702 that includes a processor 1720, a memory 1726, and I / O components 1730. In this example, the software architecture 1704 can be conceptualized as a stack of layers, each layer providing a particular function. The software architecture 1704 includes layers such as an operating system 1712, a library 1710, a framework 1708, and an application 1706. In operation, the application 1706 calls an application program interface API call 1732 through the software stack and receives a message 1734 in response to the application program interface API call 1732.
[0072] The operating system 1712 manages hardware resources and provides common services. The operating system 1712 includes, for example, a kernel 1714, services 1716, and drivers 1722. The kernel 1714 functions as an abstraction layer between the hardware and other software layers. For example, the kernel 1714 provides functions such as memory management, processor management (e.g., scheduling), component management, networking, and security configuration. The services 1716 can provide other common services to other software layers. The drivers 1722 are responsible for controlling or interfacing with the underlying hardware. For example, the drivers 1722 can include a display driver, a camera driver, a BLUETOOTH or BLUETOOTH Low Energy driver, a flash memory driver, a serial communication driver (e.g., a Universal Serial Bus (USB) driver), a WI-FI driver, an audio driver, a power management driver, and the like.
[0073] Library 1710 provides the low-level common infrastructure used by application 1706. Library 1710 can include a system library 1718 (e.g., the C standard library) that provides functions such as memory allocation functions, string manipulation functions, and mathematical functions. Further, library 1710 can include an application programming interface API library 1724 such as a media library (e.g., a library that supports the presentation and manipulation of various media formats such as Moving Picture Experts Group-4 (MPEG4), Advanced Video Coding (H.264 or AVC), Moving Picture Experts Group Layer-3 (MP3), Advanced Audio Coding (AAC), Adaptive Multi-Rate (AMR) audio codec, Joint Photographic Experts Group (JPEG or JPG), or Portable Network Graphics (PNG)), a graphics library (e.g., the OpenGL framework used to render in two dimensions (2D) and three dimensions (3D) in graphic content on a display), a database library (e.g., SQLite that provides various relational database functions), a web library (e.g., WebKit that provides web browsing functions). Library 1710 can also include a wide variety of other libraries 1728 to provide many other application programming interface APIs to application 1706.
[0074] Framework 1708 provides a high-level common infrastructure used by application 1706. For example, framework 1708 provides various graphical user interface (GUI) functions, high-level resource management, and high-level location services. Framework 1708 can provide a wide spectrum of other application program interfaces (APIs) that can be used by application 1706, some of which may be specific to a particular operating system or platform.
[0075] In an exemplary embodiment, application 1706 can include a machine learning platform 122, a service application 124, and a wide variety of other applications such as third-party application 114. Application 1706 is a program that executes functions defined in the program. Various programming languages can be employed to create one or more of application 1706, configured in various styles such as object-oriented programming languages (e.g., Objective-C, JAVA (registered trademark), or C++) or procedural programming languages (e.g., C or assembly language). In a specific example, third-party application 114 (e.g., an application developed by an entity other than the vendor of a particular platform using an ANDROID (registered trademark) or IOS (registered trademark) software development kit (SDK)) can be mobile software that runs on a mobile operating system such as IOS (registered trademark), ANDROID (registered trademark), WINDOWS (registered trademark) Phone, or other mobile operating systems. In this example, third-party application 114 can call application program interface API calls 1732 provided by operating system 1712 and can facilitate the functions described herein.
[0076] FIG. 18 is a pictorial representation of a machine 1800 on which instructions 1808 (e.g., software, program, application, applet, app, or other executable code) can be executed to cause the machine 1800 to perform any one or more of the methodologies discussed herein. For example, the instructions 1808 may cause the machine 1800 to perform any one or more of the methods discussed herein. The instructions 1808 transform a general, unprogrammed machine 1800 into a particular machine 1800 programmed to perform the functions described and illustrated in the described manner. The machine 1800 may operate as a stand-alone device or may be coupled (e.g., networked) to other machines. In a networked deployment, the machine 1800 may operate in the capacity of a server machine or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine 1800 may include, without limitation, a server computer, a client computer, a personal computer (PC), a tablet computer, a laptop computer, a netbook, a set-top box (STB), a PDA, an entertainment media system, a cellular phone, a smartphone, a mobile device, a wearable device (e.g., a smartwatch), a smart home device (e.g., a smart appliance), other smart devices, a web appliance, a network router, a network switch, a network bridge, or any machine capable of sequentially or otherwise executing instructions 1808 that specify actions to be taken by the machine 1800. Further, although only a single machine 1800 is illustrated, the term "machine" shall also be construed to include a collection of machines that individually or jointly execute instructions 1808 to perform any one or more of the methodologies discussed herein.
[0077] Machine 1800 may include a processor 1802, a memory 1804, and an I / O (input / output) component 1842, which may be configured to communicate with each other via a bus 1844. In an exemplary embodiment, the processor 1802 (e.g., a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a radio frequency integrated circuit (RFIC), another processor, or any suitable combination thereof) may include, for example, a processor 1806 and a processor 1810 that execute instructions 1808. The term “processor” is intended to include a multi-core processor having two or more independent processors (sometimes referred to as “cores”) that may execute multiple instructions contemporaneously. Although FIG. 18 shows multiple processors 1802, machine 1800 may include a single processor having a single core, a single processor having multiple cores (e.g., a multi-core processor), multiple processors having a single core, multiple processors having multiple cores, or any combination thereof.
[0078] Memory 1804 includes main memory 1812, static memory 1814, and storage 1816 that are accessible to processor 1802 via bus 1844. Main memory 1812, static memory 1814, and storage 1816 store instructions 1808 that embody any one or more of the methodologies or functions described herein. Instructions 1808 may also be present, fully or partially, within main memory 1812, within static memory 1814, within storage 1816, within machine-readable medium 1818 within at least one of processor 1802 (e.g., within a cache memory of the processor), or any suitable combination thereof, during execution of instructions 1808 by machine 1800.
[0079] The I / O component 1842 can include a variety of components for receiving inputs, providing outputs, generating outputs, transmitting information, exchanging information, capturing measurement values, and so on. The specific I / O components 1842 included in a particular machine depend on the type of the machine. For example, a portable machine such as a mobile phone can include a touch input device or other such input mechanism, while a headless server machine will probably not include such a touch input device. It will be understood that the I / O component 1842 can include many other components not shown in FIG. 18. In various exemplary embodiments, the I / O component 1842 can include an output component 1828 and an input component 1830. The output component 1828 can include a visual component (e.g., a display such as a plasma display panel (PDP), a light emitting diode (LED) display, a liquid crystal display (LCD), a projector, or a cathode ray tube (CRT)), an acoustic component (e.g., a speaker), a tactile component (e.g., a vibration motor, a resistive mechanism), other signal generators, and so on. The input component 1830 can include an alphanumeric input component (e.g., a keyboard, a touch screen configured to receive alphanumeric input, a photo-optical keyboard, or other alphanumeric input components), a point-based input component (e.g., a mouse, a touch pad, a trackball, a joystick, a motion sensor, or other pointing devices), a tactile input component (e.g., a physical button, a touch screen that provides the position and / or force of a touch or touch gesture, or other tactile input components), an audio input component (e.g., a microphone), and so on.
[0080] In further exemplary embodiments, the I / O component 1842 may include, among other broad components, a biometric component 1832, a motion component 1834, an environmental component 1836, or a location component 1838. For example, the biometric component 1832 may include components for detecting expressions (e.g., hand expressions, facial expressions, vocal expressions, body gestures, or eye tracking), measuring biometric signals (e.g., blood pressure, heart rate, body temperature, sweating, or brain waves), identifying a person (e.g., voice identification, retina identification, face identification, fingerprint identification, or brain wave-based identification), etc. The motion component 1834 may include, for example, an acceleration sensor component (e.g., an accelerometer), a gravity sensor component, a rotation sensor component (e.g., a gyroscope), etc. The environmental component 1836 may include, for example, a lighting sensor component (e.g., a photometer), a temperature sensor component (e.g., one or more thermometers for detecting ambient temperature), a humidity sensor component, a pressure sensor component (e.g., a barometer), an acoustic sensor component (e.g., one or more microphones for detecting background noise), a proximity sensor component (e.g., an infrared sensor for detecting nearby objects), a gas sensor (e.g., a gas detection sensor for detecting the concentration of dangerous gases for safety or measuring pollutants in the atmosphere), or other components that can provide a display, measurement, or signal corresponding to the surrounding physical environment. The location component 1838 may include a location sensor component (e.g., a GPS receiver component), an altitude sensor component (e.g., an altimeter or barometer for detecting the air pressure from which altitude can be derived), a direction sensor component (e.g., a magnetometer), etc.
[0081] Communication can be implemented by using a variety of technologies. The I / O component 1842 further includes a communication component 1840 that is operable to couple the machine 1800 to the network 1820 or the device 1822 via the coupling parts 1824 and 1826, respectively. For example, the communication component 1840 may include a network interface component or other suitable device for interfacing with the network 1820. In a further example, the communication component 1840 may include a wired communication component, a wireless communication component, a cellular communication component, a near field communication (NFC) component, a Bluetooth® component (e.g., Bluetooth® Low Energy), a Wi-Fi® component, and other communication components for providing communication via other modalities. The device 1822 may be either another machine or a variety of peripheral devices (e.g., a peripheral device coupled via USB).
[0082] Furthermore, the communication component 1840 may include a component operable to detect an identifier or to detect identifiers. For example, the communication component 1840 may include a radio frequency identification (RFID) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., an optical sensor for detecting one-dimensional barcodes such as universal product code (UPC) barcodes, quick response (QR) codes, Aztec codes, data matrix, data glyph, MaxiCode, PDF417, UltraCode, UCC_RSS-2D barcodes, and other optical codes, and multi-dimensional barcodes such as those mentioned above), or an acoustic detection component (e.g., a microphone for identifying tagged audio signals). Additionally, various information can be derived via the communication component 1840, such as location by Internet protocol (IP) geolocation, location by triangulation of Wi-Fi® signals, and location by detection of NFC beacon signals that may indicate a special location.
[0083] Various memories (e.g., memory 1804, main memory 1812, static memory 1814, and / or the memory of processor 1802) and / or storage device 1816 may store one or more sets of instructions and data structures (e.g., software) that are implemented or used by any one or more of the methodologies or functions described herein. These instructions (e.g., instruction 1808), when executed by processor 1802, cause various operations to implement the disclosed embodiments.
[0084] Instruction 1808 may be transmitted or received via network 1820, using a transmission medium, via a network interface device (e.g., the network interface component included in communication component 1840), using any one of a number of well-known transfer protocols (e.g., Hypertext Transfer Protocol (HTTP)). Similarly, instruction 1808 may be transmitted or received by device 1822, using a transmission medium via junction 1826 (e.g., a peer-to-peer connection).
[0085] Although embodiments have been described with reference to specific exemplary embodiments, it will be apparent that various modifications and changes can be made to these embodiments without departing from the broader scope of the disclosure. Accordingly, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense. The accompanying drawings, which form a part of this specification, illustrate specific embodiments in which the subject matter may be practiced, by way of example and not of limitation. The illustrated embodiments are described in sufficient detail to enable those skilled in the art to practice the teachings disclosed herein. Other embodiments may be utilized and derived therefrom without departing from the scope of the disclosure. Accordingly, this detailed description is not to be taken in a limiting sense, and the scope of the various embodiments is defined only by the appended claims, along with the full scope of equivalents to which such claims are entitled.
[0086] Such embodiments of the subject matter of the present invention may, for convenience only, be referred to herein individually and / or collectively by the term "invention", and there is no intention to arbitrarily limit the scope of the present application to a single invention if multiple inventions or inventive concepts are actually disclosed. Thus, while particular embodiments are illustrated and described herein, it should be understood that any arrangement calculated to achieve the same purpose may be substituted for the particular embodiments shown. The present disclosure is intended to cover any and all adaptations or variations of various embodiments. Combinations of the above embodiments and other embodiments not specifically described herein will be apparent to those skilled in the art upon review of the above description.
[0087] The abstract of the disclosure is provided to enable the reader to quickly grasp the nature of the technical disclosure. It is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. Further, in the foregoing detailed description, for the purpose of rationalizing the disclosure, it can be seen that various features are grouped in a single embodiment. This method of disclosure should not be construed as reflecting an intention that the claimed embodiments require more features than are expressly recited in each claim. Rather, as the following claims reflect, the inventive subject matter lies in less than all of the features of a single disclosed embodiment. Accordingly, the following claims are incorporated herein in a state where each claim exists as a separate embodiment in itself.
[0088] <Examples> Example 1 is a computer-implemented method comprising the following. A step of accessing a dataset from a data store by one or more processors of a server. A step of receiving, from a machine learning platform operating on the server, a definition of a task to identify a target of a machine learning algorithm. A step of forming a machine learning model based on the dataset and the task by utilizing a machine learning algorithm. A step of deploying the machine learning model by providing access to the machine learning model to an application that is external to the machine learning platform. A step of monitoring the performance of the deployed machine learning model. A step of updating the machine learning model based on the monitoring. The computer-implemented method comprises these steps.
[0089] Example 2 includes Example 1. The step of accessing the dataset further comprises a step of accessing a library of datasets from the data store, a step of identifying additional data from the library of datasets based on the task and the dataset, and a step of expanding the dataset with the additional data.
[0090] Example 3 includes any one of the above examples. The step of accessing the dataset further comprises a step of preparing the dataset for processing by partitioning and filtering the dataset based on the task. The machine learning model is based on the prepared dataset.
[0091] Example 4 includes any one of the above examples. The step of accessing the dataset further comprises a step of defining a model search space based on the dataset. The machine learning model is formed from the model search space.
[0092] Example 5 includes any one of the above examples. The monitoring process further includes a step of testing the machine learning model, a step of receiving a performance assessment of the machine learning model from an application, and a step of generating a performance indicator of the machine learning model based on the test and the performance assessment. The monitoring process further includes a step of determining that the performance indicator of the machine learning model exceeds the performance threshold of the machine learning model, and a step of updating the machine learning model in response to determining that the performance indicator of the machine learning model exceeds the performance threshold of the machine learning model. The update of the machine learning model is further characterized by including the following. That is, a step of updating the dataset. A step of updating the task definition based on the performance indicator of the machine learning model. A step of forming another machine learning model based on the updated definition of the task and the updated dataset by using a machine learning algorithm. These steps are further included in the update of the machine learning model.
[0093] Example 6 includes any one of the above examples. The step of updating the machine learning model further includes a step of detecting data deficit based on the performance of the machine learning model, a step of accessing additional data for remedying the data deficit from a data store, and a step of forming another machine learning model based on the additional data and the task by using a machine learning algorithm.
[0094] Example 7 includes any one of the above examples. The step of updating the machine learning model further includes a step of updating the task definition, a step of accessing additional data from a data store based on the updated definition of the task, and a step of forming another machine learning model based on the additional data and the updated definition of the task by using a machine learning algorithm. The target indicates at least one of the attributes of the dataset to be solved or the exploration of the feature quantities of the dataset.
[0095] Example 8 includes any one of the above examples. The computer-implemented method further includes a step of identifying feature quantities of a data set based on a target. The target indicates a future time point. The data set includes a plurality of time-varying signals. The computer-implemented method further includes a step of learning a machine learning model based on the identified feature quantities, and a step of generating a prediction of data of an attribute of the data set at a future time point by using the learned machine learning model.
[0096] Example 9 includes any one of the above examples. The computer-implemented method further includes a step of receiving a selection of a signal from a plurality of time-varying signals in a data set. The target indicates a future time point of the signal. The computer-implemented method further includes a step of performing a backtest of the selected signal by using a learned machine learning model, and a step of validating the learned machine learning model based on a result from the backtest.
[0097] Example 10 includes any one of the above examples. The computer-implemented method further includes a step of accessing a feature quantity set of a data set, and a step of forming an extended feature quantity set from the feature quantity set. The computer-implemented method further includes a step of measuring a dependency of the extended feature quantity set for each target in a target set, and a step of generating a single feature quantity for each target in the target set.
[0098] Example 11 includes any one of the above examples. The step of forming the extended feature quantity set further includes a step of scanning historical points of a corresponding signal for each signal in the data set, a step of shifting the historical points to the current time for each signal in the data set, a step of generating a new signal based on the shifted historical points for each signal in the data set, and a step of forming the extended feature quantity set based on the new signals from all the signals in the feature quantity set.
[0099] Example 12 is a computing device comprising a processor and a memory storing instructions that configure the computing device to perform an operation that, when executed by the processor, comprises: accessing a dataset from a data store by one or more processors of a server; receiving, from a machine learning platform operating on the server, a definition of a task to identify a target of a machine learning algorithm; further forming, based on the dataset and the task, a machine learning model by utilizing the machine learning algorithm; deploying the machine learning model by providing access to the machine learning model to an external application with respect to the machine learning platform; further monitoring the performance of the deployed machine learning model; and updating the machine learning model based on the monitoring.
[0100] Example 13 includes any one of the above examples. The step of accessing the dataset further comprises accessing a library of datasets from the data store, identifying additional data from the library of datasets based on the task and the dataset, and augmenting the dataset with the additional data.
[0101] Example 14 includes any one of the above examples. The step of accessing the dataset further comprises preparing the dataset for processing by partitioning and filtering the dataset based on the task. The machine learning model is based on the prepared dataset.
[0102] Example 15 includes any one of the above examples. The step of accessing the dataset further comprises defining a model search space based on the prepared dataset. The machine learning model is formed from the model search space.
[0103] Example 16 includes any one of the above examples. The monitoring process further includes a step of testing a machine learning model, a step of receiving a performance evaluation of the machine learning model from an application, and a step of generating a performance metric of the machine learning model based on the test and the performance evaluation. The monitoring process further includes a step of determining that the performance metric of the machine learning model exceeds a performance threshold of the machine learning model, and a step of updating the machine learning model in response to determining that the performance metric of the machine learning model exceeds the performance threshold of the machine learning model. The step of updating the machine learning model further includes a step of updating a dataset, a step of updating a task definition based on the performance metric of the machine learning model, and a step of forming another machine learning model based on the updated definition of the task and the updated dataset by using a machine learning algorithm.
[0104] Example 17 includes any one of the above examples. The step of updating the machine learning model further includes a step of detecting data insufficiency based on the performance of the machine learning model, a step of accessing additional data for improving the data insufficiency from a data store, and a step of forming another machine learning model based on the additional data and the task by using a machine learning algorithm.
[0105] Example 18 includes any one of the above examples. The step of updating the machine learning model further includes a step of updating a task definition, a step of accessing additional data based on the updated definition of the task from a data store, and a step of forming another machine learning model based on the additional data and the updated definition of the task by using a machine learning algorithm. The target indicates at least one of an attribute of a dataset to be solved or an exploration of feature quantities of the dataset.
[0106] Example 19 includes any one of the above examples. The operation further includes a step of identifying feature quantities of a data set based on a target. The target indicates a future time point. The data set includes a plurality of time-varying signals. The operation further includes a step of training a machine learning model based on the identified feature quantities, and a step of generating a prediction of data of an attribute of the data set at a future time point by using the trained machine learning model. The instruction is characterized in that it is configured to cause a computing device to perform such an operation.
[0107] Example 20 includes any one of the above examples. The operation further includes a step of receiving a selection of a signal from a plurality of time-varying signals in a data set. The target indicates a future time point of the signal. The operation further includes a step of performing a backtest on the selected signal by using a trained machine learning model, and a step of verifying the trained machine learning model based on the result from the backtest. The instruction is configured to cause a computing device to execute such an operation.
[0108] Example 21 includes any one of the above examples. The operation further includes a step of accessing a feature quantity set of a data set and a step of forming an extended feature quantity set from the feature quantity set. The operation further includes a step of measuring the dependency of the extended feature quantity set for each target in a target set, and a step of generating a single feature quantity for each target in the target set. The instruction is configured to cause a computing device to execute such an operation.
[0109] Example 22 is a computer-readable storage medium. The computer-readable storage medium comprises instructions that, when executed by a computer, cause the computer to perform operations comprising: accessing, by one or more processors of a server, a dataset from a data store; receiving, from a machine learning platform operating on the server, a definition of a task to identify a target of a machine learning algorithm; further forming, by utilizing the machine learning algorithm, a machine learning model based on the dataset and the task; deploying the machine learning model by providing access to the machine learning model to an application that is external to the machine learning platform; further monitoring the performance of the deployed machine learning model; and updating the machine learning model based on the monitoring.
Claims
1. A computer-implemented method, the computer-implemented method comprising: accessing, by one or more processors of a server, a dataset from a data store; receiving, from a machine learning platform operating on the server, a definition of a task to identify a target of a machine learning algorithm; forming, by using the machine learning algorithm, a machine learning model based on the dataset and the task; deploying the machine learning model by providing access to the machine learning model to an application existing external to the machine learning platform; monitoring the performance of the deployed machine learning model; updating the machine learning model based on the monitoring; identifying a feature of the dataset based on the target, wherein the target indicates a future time point and the dataset comprises a plurality of time-varying signals; learning the machine learning model based on the identified feature; generating a prediction of data of an attribute of the dataset at the future time point by using the learned machine learning model; A computer-implemented method comprising the above steps.
2. The step of accessing the dataset further comprises: accessing, from the data store, a library of the dataset; identifying additional data from the library of the dataset based on the task and the dataset; and extending the dataset with the additional data; The computer-implemented method according to claim 1, comprising the above steps.
3. The step of accessing the dataset further comprises preparing the dataset for processing by splitting and filtering the dataset based on the task, wherein the machine learning model is based on the prepared dataset; The computer-implemented method according to claim 1, comprising the above steps.
4. The step of accessing the dataset further comprises defining a model search space based on the dataset, wherein the machine learning model is formed from the model search space; The computer-implemented method according to claim 1, comprising the above steps.
5. The step of monitoring further comprises: The step of testing the machine learning model; The step of receiving, from the application, a performance evaluation of the machine learning model; The step of generating a performance metric of the machine learning model based on the test and the performance evaluation; The step of determining that the performance metric of the machine learning model exceeds a machine learning model performance threshold; The step of updating the machine learning model in response to determining that the performance metric of the machine learning model exceeds the machine learning model performance threshold; Comprising; The step of updating the machine learning model further comprises; The step of updating the dataset; The step of updating the definition of the task based on the performance metric of the machine learning model; and The step of forming another machine learning model based on the updated definition of the task and the updated dataset by using the machine learning algorithm; The computer-implemented method according to claim 1, comprising.
6. The step of updating the machine learning model further comprises; The step of detecting data insufficiency based on the performance of the machine learning model; The step of accessing, from the data store, additional data to improve the data insufficiency; and The step of forming another machine learning model based on the additional data and the task by using the machine learning algorithm; The computer-implemented method according to claim 1, comprising.
7. The step of updating the machine learning model further comprises; The step of updating the definition of the task; The step of accessing, from the data store, additional data based on the updated definition of the task; and The step of forming another machine learning model based on the additional data and the updated definition of the task by using the machine learning algorithm; Comprising; The target indicates at least one of the attributes of the dataset to be solved or the exploration of the feature quantities of the dataset; The computer-implemented method according to claim 1.
8. The computer-implemented method further comprises; The step of receiving a selection of a signal from a plurality of time-varying signals in the dataset, wherein the target indicates a future time point of the signal; The step of performing a backtest on the selected signal by using the trained machine learning model; A step of verifying the learned machine learning model based on the results of the backtest; The computer-implemented method according to claim 1, comprising the above.
9. The computer-implemented method further includes A step of accessing a feature set of the dataset; A step of forming an extended feature set from the feature set; For each target in the target set, a step of measuring the dependency of the extended feature set; and For each target in the target set, a step of generating a single feature; The computer-implemented method according to claim 1, comprising the above.
10. The step of forming the extended feature set further includes For each signal in the dataset, a step of scanning the historical points of the corresponding signal; For each signal in the dataset, a step of shifting the historical points to the present; For each signal in the dataset, a step of generating a new signal based on the shifted historical points; and Based on the new signals from all signals in the feature set, a step of forming an extended feature set; The computer-implemented method according to claim 9, comprising the above.
11. A computing device, the computing device includes A processor; and A memory storing instructions configured to cause the computing device to operate when executed by the processor; The operation includes A step of accessing a dataset from a data store by one or more processors of a server; A step of receiving a definition of a task of identifying a target of a machine learning algorithm from a machine learning platform operating on the server; A step of forming a machine learning model based on the dataset and the task by using the machine learning algorithm; A step of deploying the machine learning model by providing access to the machine learning model to an application existing outside the machine learning platform; A step of monitoring the performance of the deployed machine learning model; and A step of updating the machine learning model based on the monitoring step; A step of identifying feature amounts of the data set based on the target, wherein the target indicates a future time point, the data set includes a plurality of time-varying signals, and the step of identifying the feature amounts; A step of training the machine learning model based on the identified feature amounts; A step of generating a prediction of data of an attribute of the data set at the future time point by using the trained machine learning model; A computing device comprising the above.
12. The step of accessing the data set further comprises: A step of accessing a library of the data set from the data store; A step of identifying additional data from the library of the data set based on the task and the data set; and A step of expanding the data set with the additional data; The computing device according to claim 11, comprising the above.
13. The step of accessing the data set further comprises a step of preparing the data set for processing by dividing and filtering the data set based on the task, The machine learning model is based on the prepared data set; The computing device according to claim 11.
14. The step of accessing the data set further comprises a step of defining a model search space based on the prepared data set, The machine learning model is formed from the model search space; The computing device according to claim 11.
Citation Information
Patent Citations
Systems and techniques for predictive data analysis
JP2017520068A
Application development platform and software development kits that provide comprehensive machine learning services
WO2018217635A1