Automated artificial intelligence radial visualization

By showcasing the composition and performance of machine learning model pipelines through an interactive graphical user interface (GUI), this technology addresses the difficulty in evaluating and tracking model pipelines in existing technologies, thereby improving users' understanding of models and optimization efficiency.

CN114287012BActive Publication Date: 2025-12-19INTERNATIONAL BUSINESS MACHINE CORPORATION
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080060842.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-08-30
Filing Date
2020-08-25
Publication Date
2025-12-19
Estimated Expiration
2040-08-25

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively evaluate and track the structure, composition, overall architecture design, and performance of a large number of machine learning model pipelines, leaving users unable to fully understand their composition and performance.

Method used

It provides an interactive and visual graphical user interface (GUI) that generates a machine learning model pipeline by receiving machine learning tasks, transformers, and estimators, and extracts metadata for ranking and visualization, showcasing the composition, structure, overall architecture design, and performance of the model pipeline.

Benefits of technology

It enables visualization and ranking of machine learning model pipelines, helping users better understand and evaluate the structure and performance of models, and improving user experience and model optimization efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114287012B_ABST
    Figure CN114287012B_ABST
Patent Text Reader

Abstract

Methods, systems, and computer program products for providing automated machine learning visualizations are provided. Machine learning tasks, transformers, and estimators can be received into one or more machine learning component modules. The machine learning component modules generate one or more machine learning model pipelines. A machine learning model pipeline is a sequence of transformers and estimators, and an ensemble of machine learning pipelines is a collection of machine learning pipelines. A machine learning model pipeline, an ensemble of machine learning model pipelines, or a combination thereof, and corresponding metadata can be generated using the machine learning component modules. Metadata can be extracted from a machine learning model pipeline, an ensemble of machine learning model pipelines, or a combination thereof. An interactive visualization graphical user interface of a machine learning model pipeline, an ensemble of machine learning model pipelines, or a combination thereof, and the extracted metadata can be generated.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present invention relates generally to computing systems, and more particularly, to various embodiments for automated machine learning visualization by a processor. BACKGROUND

[0002] In today's society, consumers, business people, educators, and the like communicate in real-time, across great distances, and multiple times without boundaries or borders through various mediums. With the increase in use of computing networks such as the Internet, humans are currently inundated and overwhelmed with information available to them from a variety of structured and unstructured sources. Due to recent advances in information technology and the increasing popularity of the Internet, a wide variety of computer systems have been used in machine learning. Machine learning is a form of artificial intelligence that is used to allow computers to evolve behavior based on empirical data. SUMMARY

[0003] Various embodiments are provided for providing automated machine learning visualization by a processor. In one embodiment, a method for generating and constructing radial automated machine learning visualizations by a processor is provided, by way of example only. Machine learning ("ML") tasks, transformers, and estimators can be received into one or more machine learning component modules. The one or more machine learning component modules generate one or more machine learning models. A machine learning model pipeline, an ensemble of a plurality of machine learning model pipelines, or a combination thereof, and corresponding metadata can be generated using the one or more machine learning component modules. Metadata can be extracted from the machine learning model pipeline, the ensemble of the plurality of machine learning model pipelines, or the combination thereof. The extracted metadata and the machine learning model pipeline, the ensemble of the plurality of machine learning model pipelines, or the combination thereof can be ranked according to a metadata ranking criteria and a pipeline ranking criteria. An interactive visualization graphical user interface ("GUI") of the machine learning model pipeline, the ensemble of the plurality of machine learning model pipelines, or the combination thereof, and the extracted metadata can be generated according to the ranking. BRIEF DESCRIPTION OF DRAWINGS

[0004] In order that the advantages of the invention will be readily understood, a more particular description of the invention, briefly described above, will be rendered by reference to specific embodiments illustrated in the appended drawings. Understanding that these drawings depict only typical embodiments of the invention and are not therefore to be considered to be limiting of its scope, the invention will be described and explained with additional specificity and detail by the use of the accompanying drawings, in which:

[0005] Figure 1 is a block diagram that depicts an exemplary cloud computing node that is in accordance with an embodiment of the present invention;

[0006] Figure 2 is an additional diagram depicting an exemplary cloud computing environment in accordance with embodiments of the present application;

[0007] Figure 3 is an additional diagram depicting an abstraction model layer in accordance with embodiments of the present application;

[0008] Figure 4A is an additional diagram depicting a machine learning model in accordance with aspects of the present application;

[0009] Figures 4B-4C is an additional diagram depicting automated machine learning radial visualizations in a graphical user interface using a machine learning model in accordance with aspects of the present application;

[0010] Figures 5A-5J is an additional diagram depicting various views of automated machine learning radial visualizations in a graphical user interface using a machine learning model in accordance with aspects of the present application; and

[0011] Figure 6 is a flow diagram depicting an exemplary method for providing automated machine learning visualizations, in which aspects of the present application can likewise be implemented. DETAILED DESCRIPTION

[0012] Machine learning allows an automated processing system (“machine”), such as a computer system or a specialized processing circuit, to develop generalizations about a particular data set and use the generalizations to solve associated problems by, for example, classifying new data. Once the machine has learned a generalization from known attributes of input or training data (or has been trained with them), the machine can apply the generalization to future data to predict unknown attributes.

[0013] In one aspect, an automated artificial intelligence (“AI”) / machine learning system (“AutoAI system”) can generate a plurality (e.g., hundreds) of machine learning pipelines. An AutoAI tool can output ML models and the ranking of the ML models in a leaderboard, such as showing only the estimator plus parameters of the ML models in a drop-down list. However, this limited information detail and depth does not actually show how each machine learning model structure is created or how such machine learning model pipelines are created. When there are a large number of such machine learning model pipelines, a user cannot adequately assess the composition, make-up, and overall architectural design / development and performance of these structures (e.g., the user cannot keep track of their structure and performance in a concise manner).

[0014] Accordingly, various embodiments of the present application provide automated machine learning radial visualizations in a graphical user interface (“GUI”). Machine learning (“ML”) tasks, transformers, and estimators can be received into one or more machine learning component modules. The one or more machine learning component modules generate one or more machine learning models. The one or more machine learning component modules can be used to generate a machine learning model pipeline, an ensemble of multiple machine learning model pipelines, or a combination thereof and corresponding metadata. Metadata can be extracted from the machine learning model pipeline, the ensemble of multiple machine learning model pipelines, or the combination thereof. The extracted metadata and the machine learning model pipeline, the ensemble of multiple machine learning model pipelines, or the combination thereof can be ranked according to metadata ranking criteria and pipeline ranking criteria. An interactive visual graphical user interface (“GUI”) of the machine learning model pipeline, the ensemble of multiple machine learning model pipelines, or the combination thereof, and the extracted metadata can be generated according to the rankings.

[0015] In additional aspects, a machine learning model, ML tasks, selected data transformers, and selected data estimators can be input into one or more ML component modules. ML model pipelines, an ensemble of multiple ML model pipelines, or a combination thereof, and corresponding metadata can be generated from the one or more ML component modules. Metadata can be extracted from the ML model pipelines, the ensemble of multiple ML model pipelines, or the combination thereof. The ML model pipelines, the ensemble of multiple ML model pipelines, or the combination thereof, and their associated metadata components (e.g., data, estimators, transformers, component modules) can be ranked according to metadata and model pipeline ranking criteria. An interactive visual graphical user interface (“GUI”) of the ML model pipelines, the ensemble of multiple ML model pipelines, or the combination thereof can be generated according to the rankings.

[0016] In an additional aspect, ML tasks, transformers, and estimators can be received (as input data) into one or more machine learning constituent modules. The one or more machine learning constituent modules generate one or more machine learning models. A machine learning model pipeline is a sequence of transformers and estimators, and an ensemble of machine learning pipelines (e.g., a machine learning model pipeline that is a sequence of data transformers followed by an estimator algorithm) is an ensemble of machine learning pipelines. One or more machine learning constituent modules can be used to generate a machine learning model pipeline, an ensemble of multiple machine learning model pipelines, or a combination thereof, and corresponding metadata. Metadata can be extracted from a machine learning model pipeline, an ensemble of multiple machine learning model pipelines, or a combination thereof. An interactive visual graphical user interface (“GUI”) of a machine learning model pipeline, an ensemble of multiple machine learning model pipelines, or a combination thereof, and extracted metadata can be generated.

[0017] In one aspect, the present disclosure provides for the generation and visualization of machine learning models. Each machine learning model can be implemented as a pipeline or an ensemble of multiple pipelines. After any training operations, a machine learning model pipeline can receive test data and pass the test data through a sequence of data transformations (e.g., pre-processing, data cleaning, feature engineering, mathematical transformations, etc.), and can use an estimator operation of an estimator (e.g., logistic regression, gradient boosted trees, etc.) to produce a prediction on the test data.

[0018] It is understood that, although the present disclosure includes detailed descriptions of cloud computing, implementation of the teachings recited herein are not limited to a cloud computing environment. Rather, embodiments of the present disclosure are capable of being implemented in conjunction with any other type of computing environment now known or later developed.

[0019] Cloud computing is a model of service delivery for enabling convenient, on-demand network access to a shared pool of configurable computing resources (e.g. networks, network bandwidth, servers, processing, memory, storage, applications, virtual machines, and services) that can be rapidly provisioned and released with minimal management effort or interaction with a provider of the service. This cloud model can be composed of at least five characteristic that can be present in the cloud model: on-demand self-service, broad network access, resource pooling, rapid elasticity, measured service.

[0020] The characteristics are as follows:

[0021] On-demand self-service: a cloud consumer can unilaterally provision computing capabilities (e.g., server time and network storage), as needed automatically (without human interaction with the service's provider) from the cloud's provider.

[0022] Broad network access: capabilities are available over a network that is standard mechanism that promotes the use of different kinds of thin or thick client platforms (e.g., mobile phones, laptops, PDAs) as thin client platforms (e.g., mobile phones, laptops, PDAs) as

[0023] Resource pooling: the provider's computing resources are delivered as a service to multiple consumers using multi-tenant model, whereby different physical and virtual resources are dynamically assigned and reassigned according to consumer demand. There is a sense of location independence in that the consumer generally has no control or knowledge over the exact location of the provided resources but can be able to specify location at a higher level of abstraction (e.g., country, state, or data center).

[0024] Rapid elasticity: capabilities can be rapidly and elastically provisioned, in some cases automatically, to quickly scale out and rapidly released to quickly scale in. To the consumer, the

[0025] Measured service: cloud systems automatically control and optimize resource use by leveraging a metering capability at some level of abstraction appropriate to the type of service (e.g., storage, processing, bandwidth, and active user accounts). Resource usage can be monitored, controlled, and reported providing transparency for both the provider and consumer of the service.

[0026] The business model of a cloud system is as follows:

[0027] Software as a Service (SaaS): the capability provided to the consumer is to use the provider's applications running on a cloud infrastructure. The applications are accessible from various client devices through a thin client interface such as a web browser (e.g., web-based e-mail). The consumer does not manage or control the underlying cloud infrastructure including network, servers, operating systems, storage, or even individual application capabilities, with the possible exception of limited user-specific application configuration settings.

[0028] Platform as a Service (PaaS): the capability provided to the consumer is to deploy onto the cloud infrastructure consumer-created or acquired applications created using programming languages and tools supported by the provider. The consumer does not manage or control the underlying cloud infrastructure including networks, servers, operating systems, or storage, but has control over the deployed applications and possibly application hosting environment configurations.

[0029] Infrastructure as a Service (IaaS): This provides consumers with the capability to deploy and run any software, including operating systems and applications, on the underlying cloud infrastructure, including processing, storage, networking, and other basic computing resources. Consumers neither manage nor control the underlying cloud infrastructure, but they have control over the operating system, storage, and the applications deployed thereon, and may have limited control over chosen network components (such as host firewalls).

[0030] The deployment model is as follows:

[0031] Private cloud: The cloud infrastructure runs exclusively for a single organization. The cloud infrastructure can be managed by that organization or a third party and can exist inside or outside the organization.

[0032] Community cloud: A cloud infrastructure shared by several organizations that supports a specific community with common interests (such as mission, security requirements, policy, and compliance considerations). A community cloud can be managed by multiple organizations within the community or by third parties and can exist inside or outside the community.

[0033] Public cloud: Cloud infrastructure provided to the public or large industrial groups and owned by organizations that sell cloud services.

[0034] Hybrid cloud: A cloud infrastructure consisting of two or more cloud deployment models (private cloud, community cloud, or public cloud) that remain distinct entities but are bound together by standardized or proprietary technologies that enable data and applications to be ported together (such as cloud burst traffic balancing for load balancing between clouds).

[0035] Cloud computing environments are service-oriented, characterized by statelessness, loose coupling, modularity, and semantic interoperability. The core of cloud computing is its infrastructure, which comprises a network of interconnected nodes.

[0036] Now for reference Figure 1 The diagram illustrates an example of a cloud computing node. Cloud computing node 10 is merely one example of a suitable cloud computing node and is not intended to impose any limitation on the scope or functionality of the embodiments of the invention described herein. In any case, cloud computing node 10 can be implemented and / or perform any of the functions set forth above.

[0037] In cloud computing node 10 there is a computer system / server 12, which is operational with numerous other general purpose or special purpose computing system environments or configurations. Examples of well- known computing systems, environments, and / or configurations that can be suitable for use with computer system / server 12 include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, hand-held or laptop devices, multiprocessor systems, microprocessor-based systems, set top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments that include any of the above systems or devices, and the like.

[0038] Computer system / server 12 can be described in the general context of computer system-executable instructions, such as program modules, being executed by a computer system. Generally, program modules can include routines, programs, objects, components, logic, data structures, and so on that perform particular tasks or implement particular abstract data types. Computer system / server 12 can be practiced in distributed cloud computing environments with remote

[0039] As shown in Figure 1 Figure 1, computer system / server 12 in cloud computing node 10 is shown in the form of a general-purpose computing device. The components of computer system / server 12 can include, but are not limited to, one or more processors or processing units 16, a system memory 28, and a bus 18 that couples various system components including system memory 28 to processor 16.

[0040] Bus 18 represents one or more of any of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor or local bus using any of a variety of bus architectures. By way of example, and not limitation, such architectures include Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Enhanced ISA (EISA) bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.

[0041] Computer system / server 12 typically includes a variety of computer system readable media. Such media can be any available media that is accessible by computer system / server 12 and it includes both volatile and non-volatile media, removable and non-removable media.

[0042] The system memory 28 can include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. Computer system / server 12 can further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, a storage system 34 can be provided for reading from and writing to non-removable, non-volatile magnetic media (not shown and typically called a "hard drive"). Although not specifically shown, a magnetic disk drive can also be provided for reading from and writing to a removable, non-volatile magnetic disk (e.g., a "floppy disk"), and an optical disk drive can be provided for reading from or writing to a removable, non-volatile optical disk (such as a CD-ROM, DVD-ROM or other optical media). Each of these devices can be connected to bus 18 by one or more data media interfaces. As will be further depicted and described below, system memory 28 can include at least one program product having a set (e.g., at least one) of program modules that are configured to carry out the functions of embodiments of the application.

[0043] Program / utility 40 having a set (at least one) of program modules 42 can be stored in system memory 28 by way of example, and not limitation, as well as an operating system, one or more application programs, other program modules, and program data. Each of the operating system, one or more application programs, other program modules, and program data or some combination thereof, can include an implementation of a networking environment. Program modules 42 generally carry out the functions and / or methodologies of embodiments of the application as described herein.

[0044] Computer system / server 12 can also communicate with one or more external devices 14 such as a keyboard, a pointing device, a display 24, etc.; one or more devices that enable a user to interact with computer system / server 12; and / or any devices (e.g., network card, modem, etc.) that enable computer system / server 12 to communicate with one or more other computing devices. Such communication can occur via Input / Output (I / O) interface(s) 22. Still yet, computer system / server 12 can communicate with one or more networks such as a local area network (LAN), a general wide area network (WAN), and / or a public network (e.g., the Internet) via network adapter 20. As depicted, network adapter 20 communicates with the other components of computer system / server 12 via bus 18. It should be appreciated that although not shown, other hardware and / or software components could be used in conjunction with computer system / server 12. Examples, include, but are not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archival storage systems, etc.

[0045] Reference is now made to the drawings, wherein Figure 2The diagram illustrates an illustrative cloud computing environment 50. As shown, the cloud computing environment 50 includes one or more cloud computing nodes 10 to which local computing devices used by cloud consumers can communicate. These local computing devices include, for example, personal digital assistants (PDAs) or cellular phones 54A, desktop computers 54B, laptop computers 54C, and / or automotive computer systems 54N. The nodes 10 can communicate with each other. They can be physically or virtually grouped (not shown) in one or more networks, such as private clouds, community clouds, public clouds, or hybrid clouds, or combinations thereof, as described above. This allows the cloud computing environment 50 to provide infrastructure, platform, and / or software as a service, without requiring cloud consumers to maintain resources on their local computing devices. It should be understood that... Figure 2 The types of computing devices 54A-N shown are for illustrative purposes only, and computing node 10 and cloud computing environment 50 can communicate with any type of computerized device on any type of network and / or network-addressable connection (e.g., using a web browser).

[0046] Now for reference Figure 3 This demonstrates a cloud computing environment of 50 ( Figure 2 This provides a set of functional abstractions. It should be understood beforehand that... Figure 3 The components, layers, and functions shown are for illustrative purposes only, and embodiments of the invention are not limited thereto. As depicted, the following layers and corresponding functions are provided:

[0047] Device layer 55 includes physical and / or virtual devices embedded with and / or independent electronics, sensors, actuators, and other objects to perform various tasks in the cloud computing environment 50. Each device in device layer 55 integrates networking capabilities with other functional abstraction layers, enabling information obtained from the device to be provided to that device, and / or information from other abstraction layers to be provided to the device. In one embodiment, the various devices, including device layer 55, may be incorporated into a network of entities collectively referred to as the “Internet of Things” (IoT). As those skilled in the art will understand, such a network of entities allows data to communicate, be collected, and disseminated to achieve various purposes.

[0048] As shown in the figure, device layer 55 includes sensor 52, actuator 53, a "learning" thermostat 56 with integrated processing, sensors, and networked electronics, camera 57, controllable household socket / outlet 58, and controllable electrical switch 59, as shown. Other possible devices may include, but are not limited to, various additional sensor devices, networked devices, electronic devices (such as remote control devices), additional actuator devices, so-called "smart" appliances (such as refrigerators or washing machines / dryers), and a wide variety of other possible interconnected objects.

[0049] Hardware and software layer 60 includes hardware and software components. Examples of hardware components include: mainframes 61; RISC (Reduced Instruction Set Computer) architecture based servers 62; servers 63; blade servers 64; storage devices 65; and networks and networking components 66. In some embodiments, software components include network application server software 67 and database software 68.

[0050] Virtualization layer 70 provides an abstraction layer from which the following examples of virtual entities can be provided: virtual servers 71; virtual storage 72; virtual networks 73, including virtual private networks; virtual applications and operating systems 74; and virtual clients 75.

[0051] In one example, management layer 80 can provide the functions described below. Resource provisioning 81 provides dynamic procurement of computing resources and other resources that are utilized to perform tasks within the cloud computing environment. Metering and Pricing 82 provide cost tracking as resources are utilized within the cloud computing environment, and

[0052] Workloads layer 90 provides examples of functionality for which the cloud computing environment can be utilized. Examples of workloads and functions which can be provided from this layer include: mapping and navigation 91; software development and lifecycle management 92; virtual classroom education delivery; data analysis processing 94; transaction processing 95; and, in the context of the illustrated embodiments of the present application, various workloads and functions for providing radial automated machine learning visualizations 96. Additionally, the workloads and functions for providing radial automated machine learning visualizations 96 can include operations such as data analysis, data parsing, and notification functions as will be further described. Those of ordinary skill in the art will appreciate that the workloads and functions for providing radial automated machine learning visualizations 96 can also work in conjunction with other portions of the various abstraction layers, such as those in hardware and software 60, virtualization 70, management 80, and other workloads 90 (e.g., such as data analysis processing 94) to achieve the various purposes of the illustrated embodiments of the present application.

[0053] As previously described, the present application provides radial automated machine learning visualizations in a GUI, where the GUI is associated / communicates with an automated machine learning backend and frontend system. The automated machine learning backend system can receive and / or assemble one or more machine learning models, one or more machine learning tasks, one or more selected data transformers, and one or more selected data estimators into one or more machine learning component structures.

[0054] In another aspect, the present application provides automated machine learning radial visualizations to illustrate machine learning model and pipeline composition, constitution, and overall architectural design / development and performance. In one aspect, one or more machine learning tasks, candidate data transformers, and estimators can be received as input into an automated machine learning system (e.g., in a backend of a computing system such as, for example, an automated AI system). The machine learning tasks can each include a training dataset, an optimization metric, a label, and other defined tasks.

[0055] Machine learning models (and their composition, constitution, architecture, performance, training data, data transformers, and / or data estimators) can be combined into one or more machine learning component assemblies. One or more machine learning model pipelines composed of machine learning component assembly components that can include data transformers and estimators can be generated with their metadata and provided from a backend of a computing system.

[0056] A computing system (e.g., a front end of an automated machine learning system) can receive one or more machine learning model pipelines (e.g., one or more machine learning models and metadata) from an automated machine learning system (e.g., an automated AI backend). Machine learning model metadata can be excluded, such as data partition data (e.g., train / holdout, if applicable), machine learning model structure (e.g., data transformers and data estimators (e.g., as a sequence of transformers and estimators pipelines)), and parameters, data scores related to train / holdout metrics, and / or provenance data (e.g., a sequence of constituent algorithmic modules or “machine learning constituent structures”). Ranking of the machine learning pipelines can be determined based on the metadata, ranking criteria, and / or various metrics. For each metric, a minimum / maximum value can be determined. It should be noted that the metrics are part of the ML task that is input to the automated machine learning system (e.g., AutoAI system), and the models are generated based on these metrics. Example metrics can include, for example, accuracy, precision, ROC / AUC, root mean squared error, f1 score, etc. (for classification problems) and / or other regression problems. In one aspect, the ML task has at least one optimization metric, which can be used to create an optimized model pipeline based on these metrics and one or more evaluation metrics (for evaluating the resulting pipelines). Metrics that have an arithmetic value (e.g., accuracy) can also have a minimum and maximum value. The present invention determines / calculates the minimum and maximum metric values to show the pipeline nodes classified on the pipeline ring, and also to show annotations regarding these metrics. For example, if the accuracy metric is selected for ranking, the pipeline ring will show the minimum and maximum values, and then the pipeline nodes will be placed on the ring based on their accuracy scores.

[0057] The extracted and / or decomposed machine learning model metadata can be provided (e.g., placed) onto an interactive radial user interface (“UI”) that can include a plurality of concentric rings. For example, the concentric rings can include: 1) one or more data rings representing data folds for model training, 2) one or more estimator rings containing all available estimators, transformer rings containing all available transformers, 3) one or more provenance rings containing machine learning constituent modules (e.g., constituent algorithmic modules). In one aspect, the machine learning constituent modules can execute / run algorithms / operations that take as input the ML task (dataset, metrics, objectives, etc.), transformers, and estimators, and output a machine learning model pipeline.

[0058] In one aspect, examples of machine learning component modules include, but are not limited to: 1) pre-processing / data cleaning to convert raw data into digital format; 2) hyperparameter optimization (HPO) to identify / seek optimal hyperparameters for a given data set's pipeline; 3) model selection to select and rank a given set of machine learning pipelines for a given data set; 4) feature engineering to perform data transformations and add / remove new features to / from a data set; 5) synthesis to create an ensemble based on a set of machine learning pipelines and / or data sets; 6) automatic model generation modules based on a given ML task (e.g., training, hold-out / test data set, optimization / evaluation metric, target variable, etc.), transformers and estimators input to backend system (e.g., black-box automatic model generation modules using existing AutoAI frameworks); and / or 7) human-driven model generation modules (e.g., driven by a human user creating data based on a given training data set, transformers and estimators input to backend system).

[0059] Each machine learning model pipeline node in the radial UI (e.g., in the pipeline ring) is connected to its estimator, transformer, and ensemble module in the estimator ring, transformer ring, and ensemble ring, respectively. Each data transformer can be annotated with their order of appearance in the machine learning model pipeline.

[0060] Turning now to Figures 4A-4C , a block diagram illustrating example functional components 400, 415, and 425 depicting various mechanisms in accordance with the illustrated embodiments is shown. In one aspect, Figures 1-3 one or more of the components, modules, services, applications, and / or functions described in

[0061] As shown in Figure 4A , a machine learning pipeline structure 400 is depicted. In one aspect, the machine learning pipeline structure 400 can perform various computations, data processing, and other functions in accordance with various aspects of the present application. The machine learning pipeline structure 400 can be provided by the computer system / server 12 of Figure 1 .

[0062] As will be appreciated by one of ordinary skill in the art, the depiction of various functional units in the machine learning pipeline structure 400 is for illustrative purposes, as the functional units can be located within the machine learning pipeline structure 400 or elsewhere within and / or between distributed computing components.

[0063] For example, the machine learning pipeline architecture 400 can include one or more machine learning models 410. The machine learning models 410 can include one or more data transformers 420 and one or more estimators 430. In operation, raw data 402 can be provided to the machine learning models 410, in which one or more machine learning operations can be performed using the data transformers 420 and / or the estimators 430 to provide predictions 404.

[0064] Turning now to Figure 4B the present disclosure provides a backend automated machine learning system 440 and a frontend automated machine learning system 450 (e.g., a frontend / UI subsystem). In one aspect, the backend automated machine learning system 440 and / or the frontend automated machine learning system 450 can be internal and / or external to the computer system / server 12. Figure 1

[0065] The frontend automated machine learning system 450 can include a metadata extraction module 452 and / or a ranking module 454. The frontend automated machine learning system 450 can also be in communication with an interactive visualization GUI 460 (e.g., an interactive radial visualization GUI).

[0066] In one aspect, the backend automated machine learning system 440 can include Figure 4A one or more machine learning models 410, which can be referred to hereinafter as machine learning component structures 412A-412N. The backend automated machine learning system 440 can receive a training dataset 414, a set of candidate data transformers 420, and candidate estimators 430 as inputs.

[0067] The backend automated machine learning system 440 can assemble the training dataset 414, the set of data transformers 420 (e.g., candidate data transformers and / or selected data transformers), and the selected / candidate estimators 430 (and / or one or more machine learning models 410, one or more machine learning tasks) into one or more machine learning component structures, such as machine learning component structures 412A-N. In one aspect, the machine learning component structures 412A-N can include one or more machine learning models, one or more selected data transformers 420, and one or more candidate estimators 430.

[0068] ​In one aspect, machine learning constituent structures or combinations such as machine learning constituent structures 412A-N can be used to generate each model pipeline output by the backend automated machine learning system 440 (e.g., backend sub-system). In one aspect, examples of machine learning constituent structures 412A-N can include and / or perform one or more operations such as: 1) pre-processing / data cleaning, which converts raw data into a digital format, 2) hyperparameter optimization (“HPO”), which locates, finds, identifies one or more optimal hyperparameters for a pipeline for a given dataset, 3) machine learning model selection, which ranks a given set of machine learning pipeline models (training data, holdout / test dataset, metrics, target variable, etc.) for a given ML task, 4) feature engineering, which performs data transformations and adds / removes new features to / from a dataset, 5) synthesis, which creates an ensemble (e.g., machine learning model pipeline ensemble) based on a set of pipelines and / or datasets, 6) a black box automated model generation module (e.g., using an existing AutoAI framework) based on a given training dataset 414, transformer 420, and estimator 430 input to the backend automated machine learning system 440, and / or 7) one or more human driven model generation modules (driven by a user based on a given training dataset 414, transformer 420, and estimator 430 input to the backend automated machine learning system 440 to create data).

[0069] Accordingly, the backend automated machine learning system 440 can generate a machine learning model pipeline, an ensemble of multiple machine learning model pipelines, or a combination thereof, and corresponding metadata from one or more machine learning constituent structures. Accordingly, the backend automated machine learning system 440 can output a set of machine learning models that can be included in a machine learning model pipeline and / or an ensemble of multiple machine learning model pipelines and their associated metadata (including but not limited to metadata related to selected data transformers 420, selected data estimators 430, and / or performance / metric metadata, etc.). It should be noted that the provenance metadata for each generated model pipeline can be a list or sequence of machine learning constituent structures 412A-N that have been used to create the generated model pipeline and / or an ensemble of multiple generated model pipelines. It should be noted that as used herein, “provenance metadata” can refer to metadata that describes how a machine learning model pipeline was generated. For example, if a machine learning model pipeline is associated with a constituent module referred to as “ACME AutoAI,” the “ACME AutoAI algorithm” was used to find that machine learning model pipeline. The provenance can refer to what data or what ML task has been used to create a particular machine learning model pipeline.

[0070] In one aspect, the front-end automated machine learning system 450 can receive a machine learning model pipeline and / or an ensemble of multiple machine learning model pipelines (e.g., machine learning model and / or machine learning composition structures 412A-N) from the back-end automated machine learning system 440. The front-end automated machine learning system 450 can extract associated metadata from the ensemble of machine learning model pipelines and / or multiple machine learning model pipelines (e.g., machine learning composition structures 412A-N). The front-end automated machine learning system 450 can generate and / or produce one or more interactive visualizations of the ensemble of machine learning model pipelines and / or multiple machine learning model pipelines (e.g., machine learning model and / or machine learning composition structures 412A-N) and associated metadata attributes presented to a user based on various ranking criteria. The visualizations are dynamically updated as new models are input over time. During training or after training, the user can interact with the visualizations to discover a wide range of model pipeline attributes and the training task at hand.

[0071] In one aspect, the metadata extraction module 452 can extract metadata from the incoming ensemble of machine learning model pipelines and / or multiple machine learning model pipelines, which can be a single pipeline or an ensemble of pipelines.

[0072] In one aspect, the metadata of each machine model pipeline can include, but is not limited to: 1) machine learning model structure, which includes those data transformers 420 and estimators 430 and their associated parameters used in the machine learning model, 2) performance metadata, which can include scores for a number of metrics (e.g., if there are multiple metrics (e.g., ROC_AUC, accuracy, precision, recall, fl score, etc.) during training) for a training set (e.g., a holdout set, if applicable), the predictions of each generated machine learning model pipeline can be evaluated (scored) based on the above metrics. In addition, there can be a training set or test / holdout set that is not used to train and generate the pipeline. It is only used for evaluation, 3) one or more plots based on the scoring data, such as confusion matrix, receiver operating characteristic “ROC” / area under the curve (“AUC”) curve, etc., 4) data, such as the training data 414 (or a subset thereof) used to generate the particular machine learning model pipeline, 5) provenance data, 6) composition modules (e.g., machine learning composition structures 412A-N), which include the composition modules of the automated back-end machine learning system 440 (and their parameters, if any) that were used to generate the particular machine learning model, and / or 7) creation time data.

[0073] In an alternative aspect, the metadata for each ensemble of multiple machine learning model pipelines can include, but is not limited to: 1) machine learning model pipeline structure, 2) parameters, e.g., parameters that determine the method / way in which multiple machine learning model pipelines are integrated, 3) performance metadata, which can include scores for multiple metrics (e.g., accuracy, receiver operating characteristic "ROC" / area under curve ("AUC") curve, etc.) for a training set (e.g., holdout set, if applicable), 4) one or more plots, which are based on scoring data, e.g., for confusion matrix, ROC / AUC curve, etc., 5) data, e.g., training data 420 (or a subset thereof) that was used to generate a particular machine learning model pipeline ensemble, 5) provenance data, 6) constituent modules (e.g., machine learning constituent structures 412A-N), which include constituent modules (and their parameters, if any) of a backend automated machine learning system 440 that were used to generate a particular machine learning model, and / or 7) creation time data.

[0074] In an aspect, the ranking module 454 can receive, as input, the extracted metadata for existing machine learning model pipelines and incoming machine learning model pipelines, and output different rankings for models, transformers, estimators, or constituent modules. For each ranking with a numerical value, a minimum ranking value and a maximum ranking value are also computed.

[0075] The ranking module 454 can rank each of the machine learning model pipelines and / or ensembles of multiple machine learning model pipelines according to a ranking criterion. For example, in an aspect, the ranking criterion can include, e.g.: 1) no ranking (e.g., arbitrary), 2) creation time (which can be the default), 3) training (cross-validation) scores for different machine learning metrics (e.g., accuracy, precision, ROC / AUC, root mean squared error, etc.), 4) holdout (or test) scores for different machine learning metrics (e.g., accuracy, precision, ROC / AUC, root mean squared error, etc.).

[0076] In an aspect, the ranking module 454 can rank the transformers 420, data folds (e.g., during training, data can be partitioned into data folds, which are subsets of the original dataset, and the data folds can be displayed in an interactive GUI), estimators 430, and constituent modules (e.g., machine learning constituent structures 412A-N) according to additional ranking criteria. For example, in an aspect, the additional ranking criteria can include, e.g.: 1) no ranking (e.g., arbitrary), 2) alphabetical name order (which can be the default), 3) frequency of use in current pipelines, 4) visualization optimization criteria, 5) size (for data folds), and / or 5) average score (of the pipeline to which the additional ranking criteria are applied).

[0077] The automated machine learning frontend system 450 can also include default ranking criteria, where, for example, the user 480 can submit user input 482 and can change the ranking criteria by interacting with the automated machine learning frontend system 450. Each combination of ranking criteria selection 462 can result in a different ranking criteria view 456 that is presented to the user 480.

[0078] Given a ranking criteria (such as the ranking criteria selection 462), the metadata of the machine learning models, the machine learning models, the data estimators 430, the data transformers 420, and the constituent modules (e.g., the machine learning constituent structures 412A-N) can each be placed on a radial user interface (“UI”) 470 (e.g., a radial UI visualization view) of an interactive visualization GUI 460 (which can include one or more concentric rings, which is also depicted in Figure 4C FIG. 4B) according to their ranking criteria.

[0079] As Figure 4C shown, the radial UI visualization 470 is presented to the user. The order of the rings in the radial UI visualization 470 can vary according to the application, visualization optimization, or user preference. Once the model metadata is extracted / generated, the metadata can be placed as nodes (e.g., ring-like points on each ring, by way of example only) on the radial UI visualization 470, which includes multiple concentric rings. The concentric rings of the radial UI visualization 470 can include, for example: 1) a pipeline ring 475, which contains all available model pipelines and pipeline ensembles, 2) a data ring 472, which contains the training data partitions (and holdout data partitions, if available) used for training the pipelines, 3) an estimator ring 474, which contains all available estimators (which are input to the backend training subsystem), 4) a transformer ring 476, which contains all available transformers (which are input to the backend training subsystem), 5) an origin ring 478, which contains all available constituent modules / machine learning constituent structures 412A-N (of the automated machine learning backend system 440).

[0080] The radial UI visualization 470 can include a radial leaderboard. That is, the pipelines can be sorted according to different ranking criteria (e.g., training score, holdout score, creation time, etc.). A minimum and maximum value can be determined for each ranking criterion that can be indicated on the end of a ring. Likewise, the pipeline metadata (e.g., machine learning pipelines, machine learning models, data folds, estimators, transformers, constituent modules / machine learning constituent structures 412A-N) can be sorted from minimum to maximum according to the ranking criterion they are used in the view of the radial UI visualization 470. The minimum and maximum values for the current ranking criterion can be shown at the end of each ring. In this way, a user (e.g., user 480) is able to visualize in a concise manner how the machine learning model pipelines and the metadata of the machine learning model pipelines are ranked according to different criteria. As an alternative to a radial leaderboard, the ranked pipelines and metadata can be depicted on a linear leaderboard.

[0081] In additional aspects, the radial UI visualization 470 can be automatically updated when one or more events (e.g., trigger events) occur. For example, the trigger events can include one or more of the following.

[0082] A trigger event can occur when a new machine learning model pipeline is input to the automated machine learning front-end system 450. In this case, the machine learning model pipeline can be placed as a new node in the pipeline ring 475 and labeled by an identifier (“ID”) and optionally a value according to a ranking criterion. Additionally, one or more connections (e.g., connection lines, just as an example) can be depicted / shown to the metadata nodes (e.g., transformers, estimators, data folds, constituent modules) on their respective rings.

[0083] A trigger event can occur at the beginning of a training process, when training data enters the automated machine learning back-end system 440 and is split into training / holdout folds and the data ring is updated with the data folds. The data fold statistics can also be displayed on the radial view.

[0084] A trigger event can occur when a model selection module is present and executed in the automated machine learning back-end system 440, and the estimator ring and / or the pipeline ring can be updated by retaining (or highlighting) the top K machine learning model pipelines or estimators on the ring and deleting (or fading out) the rest.

[0085] A trigger event can occur when an HPO constituent module is present and executed in the automated machine learning back-end system 440, and the corresponding pipeline and pipeline ring exhibit an additional self-rotating ring around it to represent the layer of HPO optimization.

[0086] It should also be noted that the radial UI visualization 470 can update one or more of the multiple rings. For example, the data ring can be updated when new training data is input to the automated machine learning backend system 440. The transformer ring can be updated when new transformers are input to the automated machine learning backend system 440. The estimator ring can be updated when new estimators are input to the automated machine learning backend system 440. The provenance ring can be updated when new constituent modules are added to the (backend) system.

[0087] In additional aspects, the radial UI visualization 470 can be automatically updated when the user performs one or more events. For example, user events (e.g., user-initiated trigger events) can include one or more of the following.

[0088] In one aspect, the radial UI visualization 470 can be automatically updated when the user selects a ranking criterion. By reclassifying the pipelines and their metadata in the radial view, the entire radial UI visualization 470 is updated to reflect these criteria.

[0089] The radial UI visualization 470 can be automatically updated when the user selects, clicks, or hovers over a model pipeline node on the pipeline ring 475. The pipeline node can be labeled with its value (e.g., score, creation time, etc.) corresponding to the pipeline ranking criterion of the current view of the radial UI visualization 470. The connections of the pipeline node to its metadata nodes (e.g., estimators, transformers, constituent modules, and data folds) can be featured, delineated, and / or highlighted in the respective rings of the connections. Additionally, the metadata nodes can be labeled with their parameters for that particular pipeline. More in-depth information about this pipeline can be shown on a separate window or pop-up window, including ROC / AUC curves (e.g., binary classification problem), scatter plots of predicted values versus measured values (e.g., regression problem), scores for all supporting metrics about the problem type, etc.

[0090] The radial UI visualization 470 can be automatically updated when a user selects, clicks on, or hovers over a model pipeline ensemble node 477 such as a model pipeline ensemble 477 in the pipeline ring 475. A model pipeline ensemble node 477 (e.g., a pipeline ensemble node that can have connections to one or more pipelines that it includes) can be labeled with its values (scores, creation, parameters, time, etc.) corresponding to the pipeline ranking criteria of the current radial UI visualization 470 view. In one aspect, a metadata node is a node in a metadata ring such as, for example, a data ring, a transformer ring, an estimator ring, a composition module ring (e.g., essentially any node in the ring that is not a pipeline or ensemble). That is, a metadata node is a non-pipeline / ensemble node. For example, metadata nodes can include data folds (partitions), transformers, estimators, composition modules, etc.

[0091] Pipeline nodes on the pipeline ring 475 and / or data in the data ring 472 associated with a model pipeline ensemble node 477 can be illustrated, depicted, highlighted, and their connections to their metadata. Ensemble parameters can be provided that can specify the combination rules for their pipelines and data.

[0092] The radial UI visualization 470 can be automatically updated when a user selects, clicks on, or hovers over a metadata (e.g., an estimator, transformer, composition module, or data fold) (or group of metadata) in its corresponding ring. All instances of that metadata (or group of metadata) are shown on a new temporary instance ring and labeled with their occurrence statistics (number, frequency, or rate) relative to the metadata. The connections of that metadata node (or group) to all models in the pipeline ring that use that metadata (or group) can be illustrated.

[0093] The metadata nodes can be labeled with their occurrence statistics (e.g., number, frequency, or rate) on models in the pipeline ring. For example, when an estimator is selected in the estimator ring 474: 1) a new ring of all instances with that estimator name can be depicted, each labeled with the number of times it occurs. Selecting each such instance can show its detailed parameters, 2) all connections to models in the pipeline ring that use that estimator are shown, and / or 3) the estimator node in the estimator ring 474 is labeled with the number of times it occurs on models in the pipeline ring 475. The same type of visualization updates apply to other metadata such as transformers, composition modules, and data folds.

[0094] Thus, the radial UI visualization 470 can be automatically updated / switched between different metric views when the user 1) selects a pipeline to view details about it (e.g., enabling viewing of scores, transformer parameters, estimator parameters (e.g., upon hovering over a selected region of the radial UI), constituent modules), 2) selects a transformer to view the pipeline associated with it, 3) selects an estimator to view the pipeline associated with it, 4) selects a data partition to view the pipeline associated with it, and / or 5) selects a constituent module to view the pipeline associated with it.

[0095] In view of FIG. 4, Figures 5A-5J Various radial visualization GUI views 500, 515, 525, 535, 545, 555, 565, 575, 585, and 595 of various automated machine learning radial visualization components in a graphical user interface depicting the structure of machine learning models and pipelines are further depicted. In one aspect, Figure 1 One or more of the components, modules, services, applications, and / or functions described in -4 can be used in Figures 5A-5J For brevity, repeated descriptions of similar elements employed in other embodiments described herein (e.g., -4) are omitted. Figure 1 Preprocessing

[0096] As shown in FIG. 5, preprocessing operations can occur at the core node of the radial visualization GUI 500 (see also 470 of FIG. 4) with different arcs to represent data partitioning with additional segments to represent a“partitioning” layer of percentages of data for training and testing, e.g., 90% training data, 10% holdout data, etc. In other words, the visualization (e.g., radial UI visualization 470) builds and adds arcs, layers, and nodes gradually as it progresses. The core node of the radial visualization GUI 500 utilizes different rings (e.g., data ring 472 and estimator ring 474) to represent different subsets of data, with additional segments representing a“partitioning” additional layer of percentages of data for training, holdout, and testing purposes.

[0097] Figure 5A Model Selection Figure 4C

[0098] Model Selection

[0099] In Figures 5B-5C ​​In the middle, a model selection operation can be performed. For example, a machine learning model selection can be performed, where the machine learning model selection phase is represented by an arc around a data source node (e.g., "File_Name_1.CSV"). In one aspect, nodes can be added to the arc, each node representing a separate estimator or its associated pipeline. For example, a click node can cause a machine learning pipeline to be connected to an internal estimator node to indicate which estimator was used to generate that particular pipeline (e.g., Pipeline 1 with an ROC / AUC score of 0.843). That is, the core node (in the middle) can be clicked to cause the pipeline to be connected to the estimator node (on the right) to indicate which estimator was used to generate that particular pipeline (e.g., Pipeline 1 with an ROC / AUC score of 0.843). In one aspect, the core node (in the middle) can be clicked to cause the pipeline to be connected to the estimator node (on the left) to indicate which estimator was used to generate that particular pipeline (e.g., Pipeline 1 with an ROC / AUC score of 0.843). Figures 5A-5C In the middle, or "File_Name_1.CSV", the core node can be clicked to cause the pipeline to be connected to the estimator node (on the right) to indicate which estimator was used to generate that particular pipeline (e.g., Pipeline 1 with an ROC / AUC score of 0.843). In one aspect, the core node (in the middle) can be clicked to cause the pipeline to be connected to the estimator node (on the left) to indicate which estimator was used to generate that particular pipeline (e.g., Pipeline 1 with an ROC / AUC score of 0.843). Figure 5C It should be noted that the order of the rings can be interchanged. For example, "internal" can be used with reference to Figures 5B-5C but the rings can generally have a different order.

[0100] It should be noted that the "candidate estimators" (e.g., the small hollow circles above the GUI 517) can represent multiple estimators. The GUI 517 illustrates user interaction on one of the estimators (e.g., clicking or hovering over the estimator node) and displays details, such as estimator type, estimator name, and / or ROC_AUC with a score of 0.701. After one or more estimators have been selected, the GUI 519 depicts the selected estimators as the only estimators now being displayed. The GUI 519 illustrates user interaction on one of the selected estimators (e.g., clicking or hovering over the estimator node) and displays details, such as top ranked estimator, estimator name (e.g., decision tree), and / or ROC_AUC with a score of 0.842. It should be noted that the GUI 519 can selectively display (or not display) the estimators based on user configuration, application, or product. For example, in one aspect, the GUI 519 hides all unselected estimators while displaying all selected estimators. In additional aspects, the GUI 519 displays all unselected estimators but can more prominently display all selected estimators (e.g., highlight, flash, provide a spinning ring, etc.). Thus, the interactive GUI (e.g., the radial UI visualization 470) can selectively display, hide, highlight, or emphasize or de-emphasize one or more components, features, rings, or nodes according to user preference or technical capabilities of the computing / media display device.

[0101] As these nodes are used in the process, the visualization operation can refine over time to perform the best nodes (e.g., Figures 5A-5CPipelines 1 and 2 (e.g., ROC / AUC 0.834 and 0.830) are used until the maximum number of selected executors is determined (e.g., this number is defined by the user). Once the maximum number of executors is selected, another layer / arc can be added outside the previous arc. New node types can be added to the new arc, which has connections to its class (e.g., estimator).

[0102] Hyperparameter optimization

[0103] like Figures 5D-5E The diagram illustrates hyperparameter optimization operations, where the number of rings around a single pipeline node represents the layer of hyperparameter optimization being performed on that pipeline. Rotating rings (e.g., rotational motion on the rings to indicate some underlying activity / operation, such as optimization) are used to represent a pipeline being actively optimized (e.g., pipeline 2 with an ROC / AUC of 0.830). A new node (e.g., pipeline #2) can be replicated from the properties of a previous node (e.g., pipeline 1), with added rotating rings to indicate that parameters are being modified / optimized.

[0104] Feature Engineering

[0105] like Figure 5F As shown, feature engineering can be provided, for example, where the connection line between the transformer and the pipeline indicates that feature engineering was performed on the pipeline. Hovering over a pipeline node (e.g., pipeline node 546) (e.g., using a GUI trigger such as a mouse) can display the associated transformer, which has numbers to indicate the order in which features were applied to the pipeline.

[0106] As an example only, while the process is running, the connections between the transformers and the pipeline can be drawn, mapped, depicted, and / or sketched in the order they are attempted during the machine learning model generation process. In one aspect, the new node 548 is copied from the previous node, and a second rotation loop can be added around node 548 to represent additional modification / optimization layers.

[0107] In addition, in such Figure 5F As an additional aspect described herein, a new node 548 can be copied from the previous node, and connections are drawn between the inner and outer arc nodes in the order in which modifications are being applied. Information for these modifications may also provide numerical values ​​to specify the order in which the modifications are applied. For example... Figure 5G As shown, the same process is repeated based on the initial maximum number of executors until the process is complete, and the user is free to interact with the visual elements (at any time).

[0108] In addition, Figure 5Gdepict one or more user interactions (during completion or after). The dataset, as well as the percentage of training data compared to the holdout dataset, can be represented in a semi-circle or full-circle layout. One or more portions of participating in the radial UI visualization 565, such as selecting or hovering over any visualization node, displays a tooltip with contextual information for more details. Hovering also displays all direct associations to that node, represented with connecting lines. The number of rings around the node represents different layers of optimization / modification. It should be noted that the use of different estimators and transformers during training is used for illustrative purposes only, as an example.

[0109] Further, as Figures 5H-5J shown, any node outside of the core can be hovered over, which shows all directly connected nodes and associated labels. To illustrate the linking between estimators and transformers, a pipeline can be implemented to link them together. One or more portions of participating in the radial UI visualization view 575, 585, such as selecting or hovering over the core of the visualization, provides additional details about the core node. Hovering over a secondary view (e.g., progress graph 594) highlights the corresponding information in the visualization. Hovering over the legend 592 highlights the corresponding node / arc type represented in the visualization. Further, selecting / clicking on a node in the visualization causes the user to scroll to the corresponding table item below.

[0110] Figure 6 is an additional flowchart 600 depicting an additional exemplary method for automated machine learning visualization, in which aspects of the application can likewise be implemented. The functionality 600 can be implemented as a method that is executed as instructions on a machine, where the instructions are included on at least one computer readable medium or one non-transitory machine readable storage medium. The functionality 600 can begin in block 602.

[0111] Machine learning tasks, transformers, and estimators can be received into one or more machine learning component modules, as in block 604. In one aspect, receiving one or more machine learning tasks further comprises receiving training data, holdout data, test data, an optimization metric, an evaluation metric, a target variable, one or more transformers, and one or more estimators into one or more machine learning component modules.

[0112] The one or more machine learning component modules generate one or more machine learning models. The one or more machine learning component modules can be used to generate a machine learning model pipeline, an ensemble of multiple machine learning model pipelines, or a combination thereof, and corresponding metadata, as in block 606. Metadata can be extracted from the machine learning model pipeline, the ensemble of multiple machine learning model pipelines, or the combination thereof, as in block 608. The extracted metadata and the machine learning model pipeline, the ensemble of multiple machine learning model pipelines, or the combination thereof can be ranked according to the metadata ranking criteria and the pipeline ranking criteria, as in block 610. An interactive visual graphical user interface (“GUI”) of the machine learning model pipeline, the ensemble of multiple machine learning model pipelines, or the combination thereof, and the extracted metadata can be generated according to the ranking, as in block 612. The function 600 can end, as in block 614.

[0113] In one aspect, in conjunction with and / or as part of at least one block of Figure 6 The operations of the method 600 can include each of the following, in one aspect. The operations of the method 600 can define metadata for a machine learning model pipeline to include data about which of the one or more selected data transformers and the one or more selected data estimators are included in the one or more machine learning component structures, performance data, parameters, and metric data related to the one or more machine learning models and training data, and data related to the creation of the one or more machine learning component structures. The operations of the method 600 can define metadata for an ensemble of multiple machine learning model pipelines to include structural data related to which of the combination of multiple machine learning model pipelines are included, performance data related to the one or more machine learning models and training data, and parameters and metric data.

[0114] In one aspect, the operations of the method 600 can define metadata for a machine learning model pipeline to include structural metadata, performance metadata, and provenance metadata. The structural metadata for the machine learning model pipeline includes data transformers and estimators of the machine learning model pipeline and associated parameters and hyperparameters. The performance metadata includes scores of optimization and evaluation metrics for the machine learning task, a confusion matrix, a ROC / AUC curve, and the like, of the machine learning model pipeline. The provenance metadata for the machine learning model pipeline includes the machine learning task based on which the machine learning model pipeline was created. The machine learning task (for the machine learning model pipeline) also includes training data (or a subset thereof) used to train the machine learning model pipeline, test or holdout data used to evaluate the machine learning model pipeline, optimization and evaluation metrics, target variables, and / or constituent modules (e.g., machine learning constituent modules used to generate the machine learning model pipeline and their parameters, if any), time of creation, and / or resources (compute time, memory, and the like) spent in creating.

[0115] In an additional aspect, the operations of the method 600 can define metadata for an ensemble of multiple machine learning model pipelines to include structural metadata, performance metadata, and provenance metadata. The structural metadata for the ensemble of multiple machine learning model pipelines includes all machine learning model pipelines of the ensemble. The performance metadata includes parameters that determine the manner in which the pipelines are integrated. The performance metadata includes scores of optimization and evaluation metrics for the machine learning task, a confusion matrix, a ROC / AUC curve, and the like, of the ensemble.

[0116] The provenance metadata for the ensemble of multiple machine learning model pipelines includes the machine learning task based on which the machine learning model pipeline was created. The machine learning task (for the machine learning model pipeline and / or the ensemble of multiple machine learning model pipelines) also includes training data (or a subset thereof) used to train the ensemble of machine learning model pipelines, test or holdout data used to evaluate the machine learning model pipeline, optimization and evaluation metrics, target variable(s) and / or constituent modules (e.g., machine learning constituent modules used to generate the ensemble of machine learning model pipelines and their parameters, if any), time of creation, and / or resources (compute time, memory, and the like) spent in creating.

[0117] The operations of the method 600 can display the interactive visual GUI as a radial structure with a plurality of concentric rings, in which one or more nodes are displayed, wherein the plurality of concentric rings includes at least a machine learning pipeline ring, a data ring, an estimator ring, a transformer ring, and a component module ring, wherein the one or more nodes represent a machine learning model pipeline, an ensemble of a plurality of machine learning model pipelines, or a combination thereof, data, one or more estimators, one or more transformers, and machine learning component modules used to generate the machine learning model pipeline, the ensemble of the plurality of machine learning model pipelines, or the combination thereof, wherein the one or more nodes in the plurality of concentric rings are sequentially displayed based on different ranking criteria.

[0118] The operations of the method 600 can associate the one or more nodes with one or more of the plurality of concentric rings based on operations to: associate and display details related to the machine learning model pipeline or the ensemble of the plurality of machine learning model pipelines when the user interacts with one or more of the plurality of concentric rings; or display each of the machine learning model pipeline, the ensemble of the plurality of machine learning model pipelines, or the combination thereof associated with one or more selected data transformers, one or more data estimators, one or more machine learning component modules, one or more data partitions, or a combination thereof when the user interacts with one or more of the plurality of concentric rings.

[0119] The operations of the method 600 can automatically update the interactive visual GUI when one or more triggering events occur, and / or automatically update the interactive visual when the user performs operations to: 1) select ranking criteria for the machine learning model pipeline, the ensemble of the plurality of machine learning model pipelines, or the combination thereof, and ranking criteria for data, transformers, estimators, component modules for visualization in their corresponding rings, 2) select or interact with one or more nodes located within the interactive visualization of the machine learning model pipeline, the ensemble of the plurality of machine learning model pipelines, or the combination thereof, and / or 3) select or interact with one or more nodes of the metadata rings (transformer, estimator, data, component module rings) within the interactive GUI visualization.

[0120] The operations of the method 600 can associate one or more nodes with one or more of the plurality of concentric rings based on the following operations: associating and displaying details related to the machine learning model pipeline or the collection of the plurality of machine learning model pipelines when the user interacts with one or more of the plurality of concentric rings; and / or displaying each of the machine learning model pipeline, the collection of the plurality of machine learning model pipelines, or a combination thereof associated with one or more selected data transformers, one or more selected data estimators, one or more machine learning component modules, one or more data partitions, or a combination thereof when the user interacts with one or more of the plurality of concentric rings.

[0121] The operations of the method 600 can automatically update the interactive visualization GUI when one or more triggering events occur, or automatically update the interactive visualization when the user performs the following operations: selecting a ranking criteria for the machine learning model pipeline, the collection of the plurality of machine learning model pipelines, or a combination thereof; selecting or interacting with one or more nodes located within the interactive visualization of the machine learning model pipeline, the collection of the plurality of machine learning model pipelines, or a combination thereof; and / or selecting metadata within the interactive visualization GUI.

[0122] The present disclosure can be a system, a method, and / or a computer program product. The computer program product can include a computer readable storage medium (or media) having computer readable program instructions stored therein (or thereon) for causing a processor to carry out aspects of the present disclosure.

[0123] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer readable storage medium includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable), or electrical signals transmitted through a wire.

[0124] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions to storage media within the respective computing / processing device for execution by a processor.

[0125] Computer readable program instructions for carrying out operations of the present application can be assembly instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present application.

[0126] Aspects of the present application are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer readable program instructions.

[0127] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein comprises an article of manufacture including

[0128] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0129] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0129] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

Claims

1. A method for providing automated machine learning visualizations by one or more processors, comprising: receiving one or more machine learning tasks, one or more transformers, and one or more estimators into one or more machine learning component modules; generating, using the one or more machine learning component modules, a machine learning model pipeline, an ensemble of a plurality of machine learning model pipelines, or a combination thereof, and corresponding metadata, wherein a machine learning model pipeline is a sequence of transformers and estimators, and an ensemble of machine learning pipelines is an ensemble of machine learning pipelines; extracting metadata from the machine learning model pipeline, the ensemble of a plurality of machine learning model pipelines, or a combination thereof; and generating an interactive visualization graphical user interface (GUI) of the machine learning model pipeline, the ensemble of a plurality of machine learning model pipelines, or a combination thereof, the extracted metadata, or a combination thereof; displaying the interactive visualization GUI as a radial structure having a plurality of concentric rings in which one or more nodes are displayed, wherein the plurality of concentric rings include a machine learning pipeline ring, a data ring, an estimator ring, a transformer ring, and a component module ring, and wherein the one or more nodes represent the machine learning model pipeline, the ensemble of a plurality of machine learning model pipelines, or a combination thereof, data, the one or more estimators, the one or more transformers, and the machine learning component modules used to generate the machine learning model pipeline, the ensemble of a plurality of machine learning model pipelines, or a combination thereof.

2. The method of claim 1, wherein, receiving the one or more machine learning tasks further comprises receiving training data, holdout data, test data, optimization metrics, evaluation metrics, target variables, the one or more transformers, and the one or more estimators into the one or more machine learning component modules.

3. The method of claim 1, further comprising: defining metadata for the machine learning pipeline to include structural metadata, performance metadata, provenance metadata, or a combination thereof, related to the machine learning pipeline, wherein the structural metadata related to the machine learning pipeline includes data transformers and estimators of a machine learning model pipeline and associated parameters and hyperparameters, the performance metadata related to the machine learning pipeline includes scores of optimization and evaluation metrics for a machine learning task of a machine learning model pipeline, confusion matrices, ROC / AUC curves, and the provenance metadata related to the machine learning pipeline includes a machine learning task based on which the machine learning task was created; or defining metadata for the ensemble of a plurality of machine learning model pipelines to include structural metadata, performance metadata, provenance metadata, or a combination thereof, related to the ensemble of a plurality of machine learning model pipelines.

4. The method of claim 1, further comprising: rank the extracted metadata and the machine learning model pipeline, the ensemble of the plurality of machine learning model pipelines, or a combination thereof according to a pipeline ranking criteria and a metadata ranking criteria; rank the structural metadata, the performance metadata, the provenance metadata, or a combination thereof for the machine learning pipeline according to the metadata ranking criteria; or rank the structural metadata, the performance metadata, the provenance metadata, or a combination thereof for the ensemble of the plurality of machine learning model pipelines according to the metadata ranking criteria.

5. The method of claim 1, wherein, the one or more nodes in the plurality of concentric rings are sequentially displayed based on different ranking criteria.

6. The method of claim 5, further comprising: associating one or more nodes with one or more of the plurality of concentric rings based on: associating and displaying details related to the machine learning model pipeline or the ensemble of the plurality of machine learning model pipelines when a user interacts with the one or more of the plurality of concentric rings; or displaying each of the machine learning model pipeline, the ensemble of the plurality of machine learning model pipelines, or a combination thereof associated with one or more selected data transformers, the one or more data estimators, the one or more machine learning component modules, one or more data partitions, or a combination thereof when the user interacts with the one or more of the plurality of concentric rings.

7. The method of claim 4, further comprising: automatically updating the interactive visualization GUI when one or more triggering events occur, or automatically updating the interactive visualization when a user: selects the pipeline ranking criteria and the metadata ranking criteria, the one or more transformers, the one or more estimators, the one or more machine learning component modules, or a combination thereof of the machine learning model pipeline, the ensemble of the plurality of machine learning model pipelines, or a combination thereof for visualization in one or more corresponding rings of the interactive visualization; selects or interacts with one or more nodes located within the interactive visualization of the machine learning model pipeline, the ensemble of the plurality of machine learning model pipelines, or a combination thereof; or selects or interacts with one or more nodes of one or more metadata rings within the interactive visualization GUI, wherein the plurality of metadata rings include a transformer ring, an estimator ring, a data ring, and a component module ring.

8. A system for providing automated machine learning visualizations, comprising: one or more computers having executable instructions that, when executed, cause the system to: receive one or more machine learning tasks, one or more transformers, and one or more estimators into one or more machine learning component modules; ​ generating, using the one or more machine learning component modules, a machine learning model pipeline, an ensemble of a plurality of machine learning model pipelines, or a combination thereof, and corresponding metadata, wherein the machine learning model pipeline is a sequence of transformers and estimators, and the ensemble of machine learning pipelines is an ensemble of machine learning pipelines; extracting metadata from the machine learning model pipeline, the ensemble of a plurality of machine learning model pipelines, or a combination thereof; and generating an interactive visual graphical user interface (GUI) of the machine learning model pipeline, the ensemble of a plurality of machine learning model pipelines, or a combination thereof, the extracted metadata, or a combination thereof; displaying the interactive visual GUI as a radial structure having a plurality of concentric rings in which one or more nodes are displayed, wherein the plurality of concentric rings include a machine learning pipeline ring, a data ring, an estimator ring, a transformer ring, and a component module ring, and wherein the one or more nodes represent the machine learning model pipeline, the ensemble of a plurality of machine learning model pipelines, or a combination thereof, data, the one or more estimators, the one or more transformers, and the machine learning component modules used to generate the machine learning model pipeline, the ensemble of a plurality of machine learning model pipelines, or a combination thereof.

9. The system of claim 8, wherein, The executable instructions for receiving the one or more machine learning tasks further include receiving training data, holdout data, test data, optimization metrics, evaluation metrics, target variables, the one or more transformers, and the one or more estimators into the one or more machine learning component modules.

10. The system of claim 8, wherein, The executable instructions further: define metadata for the machine learning pipeline to include structural metadata, performance metadata, provenance metadata, or a combination thereof related to the machine learning pipeline, wherein the structural metadata related to the machine learning pipeline includes data transformers and estimators of the machine learning model pipeline and associated parameters and hyperparameters, the performance metadata related to the machine learning pipeline includes scores of optimization and evaluation metrics for the machine learning task, confusion matrices, ROC / AUC curves of the machine learning model pipeline, and the provenance metadata related to the machine learning pipeline includes machine learning tasks based on which the machine learning task was created; or define metadata for the ensemble of a plurality of machine learning model pipelines to include structural metadata, performance metadata, provenance metadata, or a combination thereof related to the ensemble of a plurality of machine learning model pipelines.

11. The system of claim 8, wherein, The executable instructions further: rank the extracted metadata and the machine learning model pipeline, the ensemble of a plurality of machine learning model pipelines, or a combination thereof according to pipeline ranking criteria and metadata ranking criteria; rank the structural metadata, performance metadata, provenance metadata, or a combination thereof for the machine learning pipeline according to the metadata ranking criteria; or or rank the structural metadata, the performance metadata, the provenance metadata, or a combination thereof for the ensemble of the plurality of machine learning model pipelines based on the metadata ranking criteria.

12. The system of claim 8, wherein, the one or more nodes in the plurality of concentric rings are sequentially displayed based on different ranking criteria.

13. The system of claim 8, wherein, the executable instructions further: associate one or more nodes with one or more of the plurality of concentric rings based on: associate and display details related to the machine learning model pipeline or the ensemble of the plurality of machine learning model pipelines as the user interacts with the one or more of the plurality of concentric rings; or display each of the machine learning model pipeline, the ensemble of the plurality of machine learning model pipelines, or a combination thereof associated with one or more selected data transformers, the one or more data estimators, the one or more machine learning component modules, one or more data partitions, or a combination thereof as the user interacts with the one or more of the plurality of concentric rings.

14. The system of claim 11, wherein, the executable instructions further: automatically update the interactive visualization GUI upon occurrence of one or more triggering events, or automatically update the interactive visualization upon the user: selecting the pipeline ranking criteria and the metadata ranking criteria of the machine learning model pipeline, the ensemble of the plurality of machine learning model pipelines, or a combination thereof, the one or more transformers, the one or more estimators, the one or more machine learning component modules, or a combination thereof for visualization in one or more corresponding rings of the interactive visualization; selecting or interacting with one or more nodes located within the interactive visualization of the machine learning model pipeline, the ensemble of the plurality of machine learning model pipelines, or a combination thereof; or selecting or interacting with one or more nodes for one or more metadata rings of a plurality of metadata rings within the interactive visualization GUI, wherein the plurality of metadata rings include a transformer ring, an estimator ring, a data ring, and a component module ring.

15. A computer program product for providing automated machine learning visualizations by a processor, the computer program product comprising a non-transitory computer-readable storage medium having computer-readable program code portions stored therein, the computer-readable program code portions comprising: receiving one or more machine learning tasks, one or more transformers, and one or more estimators into an executable portion of one or more machine learning component modules; generating, using the one or more machine learning component modules, an executable portion of a machine learning model pipeline, an ensemble of a plurality of machine learning model pipelines, or a combination thereof and corresponding metadata, wherein a machine learning model pipeline is a sequence of transformers and estimators, and an ensemble of machine learning pipelines is an ensemble of machine learning pipelines; extracting an executable portion of metadata from the machine learning model pipeline, the ensemble of the plurality of machine learning model pipelines, or a combination thereof; and generating an executable portion of an interactive visual graphical user interface (GUI) of the machine learning model pipeline, the ensemble of the plurality of machine learning model pipelines, or a combination thereof, the extracted metadata, or a combination thereof; displaying the interactive visual GUI as a radial structure having a plurality of concentric rings in which one or more nodes are displayed, wherein the plurality of concentric rings include a machine learning pipeline ring, a data ring, an estimator ring, a transformer ring, and a component module ring, wherein the one or more nodes represent the machine learning model pipeline, the ensemble of the plurality of machine learning model pipelines, or a combination thereof, data, the one or more estimators, the one or more transformers, and the machine learning component modules used to generate the machine learning model pipeline, the ensemble of the plurality of machine learning model pipelines, or a combination thereof.

16. The computer program product of claim 15, wherein, receiving the executable portion of the one or more machine learning tasks further receives, in the one or more machine learning tasks, training data, holdout data, test data, optimization metrics, evaluation metrics, target variables, the one or more transformers, and the one or more estimators into the one or more machine learning component modules.

17. The computer program product of claim 15, further comprising an executable portion that performs the following operations: define metadata for the machine learning pipeline to include structural metadata, performance metadata, provenance metadata, or a combination thereof, related to the machine learning pipeline, wherein, the structural metadata related to the machine learning pipeline includes data transformers and estimators of the machine learning model pipeline and associated parameters and hyperparameters, the performance metadata related to the machine learning pipeline includes scores of optimization and evaluation metrics, confusion matrices, ROC / AUC curves of the machine learning model pipeline for the machine learning task, the provenance metadata related to the machine learning pipeline includes the machine learning task based on which the machine learning pipeline was created; or defining metadata for the ensemble of the plurality of machine learning model pipelines to include structural metadata, performance metadata, provenance metadata, or a combination thereof related to the ensemble of the plurality of machine learning model pipelines.

18. The computer program product of claim 15, further comprising an executable portion that performs the following operations: ranking the extracted metadata and the machine learning model pipeline, the ensemble of the plurality of machine learning model pipelines, or a combination thereof according to pipeline ranking criteria and metadata ranking criteria; ranking the structural metadata, the performance metadata, the provenance metadata, or a combination thereof for the machine learning pipeline according to the metadata ranking criteria; or ranking the structural metadata, the performance metadata, the provenance metadata, or a combination thereof for the ensemble of the plurality of machine learning model pipelines according to the metadata ranking criteria.

19. The computer program product of claim 15, wherein, The one or more nodes in the plurality of concentric rings are sequentially displayed based on different ranking criteria.

20. The computer program product of claim 15, further comprising an executable portion that: associates one or more nodes with one or more of the plurality of concentric rings based on: displays details related to the machine learning model pipeline or the ensemble of the plurality of machine learning model pipelines in association with the one or more of the plurality of concentric rings as the user interacts with the one or more of the plurality of concentric rings; or displays each of the machine learning model pipeline, the ensemble of the plurality of machine learning model pipelines, or a combination thereof in association with one or more selected data transformers, the one or more data estimators, the one or more machine learning component modules, one or more data partitions, or a combination thereof as the user interacts with the one or more of the plurality of concentric rings.

21. The computer program product of claim 18, further comprising an executable portion that: automatically updates the interactive visualization GUI upon occurrence of one or more triggering events, or automatically updates the interactive visualization upon the user: selecting the pipeline ranking criteria and the metadata ranking criteria of the machine learning model pipeline, the ensemble of the plurality of machine learning model pipelines, or a combination thereof, the one or more transformers, the one or more estimators, the one or more machine learning component modules, or a combination thereof for visualization in one or more corresponding rings of the interactive visualization; selecting or interacting with one or more nodes located within the interactive visualization of the machine learning model pipeline, the ensemble of the plurality of machine learning model pipelines, or a combination thereof; or selecting or interacting with one or more nodes of one or more of a plurality of metadata rings within the interactive visualization GUI, wherein, the plurality of metadata rings comprises a transformer ring, an estimator ring, a data ring, and a component module ring.

Citation Information

Patent Citations

  • Parallel rendering and visualization method and system based on data flow diagram

    CN103679789A

  • Data orchestration platform management

    JP2019133610A