Enhancing a machine learning pipeline corpus to synthesize new machine learning pipelines

CN115796298BActive Publication Date: 2026-09-04FUJITSU LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211064368.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-03-31
Filing Date
2022-09-01
Publication Date
2026-09-04
Estimated Expiration
2042-09-01

AI Technical Summary

Technical Problem

因此,用于ML流水线的自动生成的当前技术可能不能生成准确的ML流水线并且可能需要大量的计算时间和资源

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115796298B_ABST
    Figure CN115796298B_ABST
Patent Text Reader

Abstract

Operations can include receiving an ML item stored in an ML corpus database. Operations can also include causing a first ML pipeline of a first set of ML pipelines associated with the received ML item to transition to determine a second set of ML pipelines. The transition of the first ML pipeline can correspond to replacing a first ML model associated with the first ML pipeline with a second ML model associated with one of a predefined set of ML pipelines. Operations can also include selecting one or more ML pipelines from the second set of ML pipelines based on a performance score associated with each ML pipeline of the determined set of ML pipelines. Operations can also include augmenting the ML corpus database to include the selected one or more ML pipelines and the first set of ML pipelines.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references / reference additions to related applications

[0002] This application claims priority to U.S. Provisional Patent Application Serial No. 63 / 260,908, filed September 3, 2021, entitled “Using Data Augmentation With Learning From Human-Written Pipelines To Generate High-Quality MLPipelines,” the entire contents of which are incorporated herein by reference. Technical Field

[0003] The implementations discussed in this disclosure relate to enhancing machine learning pipeline corpora to synthesize new machine learning pipelines. Background Technology

[0004] Advances in artificial intelligence (AI) and machine learning (ML) have led to the application of AI / ML algorithms across various fields. Typically, ML pipelines are manually created by data scientists for a given dataset. Manual creation of ML pipelines can be a time-consuming task, requiring significant effort from expert users such as data scientists. Recently, some techniques have been developed for the automated generation of ML pipelines for datasets. Current techniques for the automated generation of ML pipelines generally follow an exploratory approach, where a vast space of possible ML pipelines is iteratively searched based on the instantiation and testing of multiple candidate ML pipelines to find the optimal pipeline for a given dataset. Therefore, current techniques for the automated generation of ML pipelines may not produce accurate ML pipelines and may require substantial computational time and resources.

[0005] The subject matter claimed in this disclosure is not limited to implementations that address any shortcomings or operate only in environments such as those described above. Rather, this background information is provided merely to illustrate an example technical field in which some of the implementations described in this disclosure can be practiced. Summary of the Invention

[0006] According to one aspect of the implementation, the operation may include receiving ML projects from a plurality of ML projects stored in a machine learning (ML) corpus database. Each ML project in this document may include a dataset and a set of ML pipelines applicable to the dataset. The operation may also include mutating a first ML pipeline in a first ML pipeline set associated with the received ML project based on a predefined set of ML pipelines to determine a second ML pipeline set. In this document, the mutation of the first ML pipeline may correspond to replacing a first ML model associated with the first ML pipeline with a second ML model associated with a predefined ML pipeline in the predefined ML pipeline set. The operation may also include selecting one or more ML pipelines from the determined second ML pipeline set based on performance scores associated with each second ML pipeline in the determined second ML pipeline set. The operation may also include augmenting the ML corpus database to include the selected one or more ML pipelines and the first ML pipeline set associated with the received ML project.

[0007] The objectives and advantages of the implementation will be realized and accomplished, at least by means of the elements, features and combinations specifically pointed out in the claims.

[0008] Both the foregoing general description and the following detailed description are given by way of example and are illustrative rather than limiting of the claimed invention. Attached Figure Description

[0009] Example embodiments will be described and illustrated with additional features and details using the accompanying drawings, in which:

[0010] Figure 1 This is a graph representing an example environment related to enhancements to a machine learning pipeline corpus;

[0011] Figure 2 This is a block diagram of a system for enhancing machine learning pipeline corpora;

[0012] Figure 3 This is a flowchart illustrating an example method for enhancing a machine learning pipeline corpus to synthesize new machine learning pipelines;

[0013] Figure 4 This is a flowchart illustrating an example method for transforming a first ML pipeline in an ML pipeline set associated with a received ML project to determine a second ML pipeline set.

[0014] Figure 5 This is a diagram illustrating an exemplary scenario for enhancing a machine learning pipeline corpus based on a predefined machine learning pipeline;

[0015] Figure 6A This is a diagram illustrating an exemplary scenario of identifying a code snippet of a first ML model associated with an exemplary first ML pipeline;

[0016] Figure 6B This is a diagram illustrating an exemplary scenario for instantiating a second ML model selected from a set of predefined models associated with a predefined ML pipeline set;

[0017] Figure 7 This is a flowchart illustrating an example method for instantiating a second ML model selected from a set of predefined models associated with a predefined ML pipeline set;

[0018] Figure 8 This is a flowchart illustrating an example method for identifying one or more statements associated with a first ML model;

[0019] Figure 9 This is a flowchart illustrating an example method for obtaining model slices to identify code snippets of the first ML model;

[0020] Figure 10 This is a flowchart illustrating an example method for training a meta-learning model; and

[0021] Figure 11 An exemplary scenario for training a meta-learning model is shown;

[0022] All of these figures are based on at least one embodiment described in this disclosure. Detailed Implementation

[0023] Some embodiments described in this disclosure relate to methods and systems for enhancing a machine learning pipeline corpus to synthesize new machine learning pipelines. In this disclosure, ML projects can be received from a plurality of ML projects stored in a machine learning (ML) corpus database. Furthermore, a first ML pipeline in a first set of ML pipelines associated with the received ML projects can be transformed based on a predefined set of ML pipelines to determine a second set of ML pipelines. Herein, the transformation of the first ML pipeline can correspond to replacing the first ML model associated with the first ML pipeline with a second ML model associated with a predefined ML pipeline in the predefined set of ML pipelines. Thereafter, one or more ML pipelines can be selected from the determined second ML pipeline set based on performance scores associated with each second ML pipeline in the determined second ML pipeline set. Furthermore, the ML corpus database can be enhanced to include the selected one or more ML pipelines and the first set of ML pipelines associated with the received ML projects.

[0024] According to one or more embodiments of this disclosure, the field of artificial intelligence (AI) / machine learning (ML) technology can be improved by configuring a computing system in a manner that allows the computing system to enhance an ML pipeline corpus to synthesize new ML pipelines. The computing system can receive ML items from a plurality of ML items stored in an ML corpus database. Each ML item in the plurality of ML items may include a dataset and a set of ML pipelines applicable to that dataset. The computing system can determine a second ML pipeline set by transforming a first ML pipeline in a first ML pipeline set associated with the received ML items based on a predefined set of ML pipelines. In this document, the transformation of the first ML pipeline may correspond to replacing a first ML model associated with the first ML pipeline with a second ML model associated with a predefined ML pipeline in the predefined set of ML pipelines. Furthermore, the first ML pipeline set may correspond to each ML pipeline associated with the received ML items. A first ML pipeline can be selected from the first ML pipeline set. Subsequently, the computing system can select one or more ML pipelines from the determined second ML pipeline set based on the performance score associated with each second ML pipeline in the determined second ML pipeline set. Furthermore, the computing system can enhance the ML corpus database to include the selected one or more ML pipelines and the first ML pipeline set associated with the received ML projects.

[0025] Traditional methods for generating ML pipelines may require an explicit search of a large space of possible ML pipelines to determine the optimal ML pipeline for a given ML project. Therefore, conventional techniques for automatically generating ML pipelines may fail to produce accurate pipelines and may require significant computational time and resources. On the other hand, the disclosed technique (performed by a computing system) may include transforming a first ML pipeline in a first set of ML pipelines associated with a received ML project to determine a second set of ML pipelines. In this paper, the transformation of the first ML pipeline may correspond to replacing the first ML model associated with the first ML pipeline with a second ML model associated with a predefined ML pipeline in a predefined set of ML pipelines. Furthermore, the disclosed technique may include selecting one or more ML pipelines from the determined second set of ML pipelines based on performance scores associated with each second ML pipeline in the determined second set of ML pipelines. An ML corpus database may be enhanced to include the selected one or more ML pipelines and the first set of ML pipelines. Therefore, the disclosed technique can select such a transformed ML pipeline that may have high performance scores, thereby ensuring good accuracy of the final ML pipeline set in the enhanced ML corpus database.

[0026] The electronic device of this disclosure can follow a generative approach, whereby a meta-learning model can learn from a corpus of existing ML pipelines created by data scientists for other datasets and use it to efficiently and optimally synthesize ML pipelines for new datasets. This disclosure can substantially address key challenges of automated machine learning methods based on generative learning. The electronic device of this disclosure can use data augmentation techniques to augment an ML corpus database by systematically transforming a given ML pipeline by replacing the ML model used in an existing ML pipeline in the corpus with other feasible options to create a new ML pipeline population. The new ML pipeline population can be used to provide higher quality and more consistent learning features to the meta-learning model, enabling the meta-learning model to learn from the augmented ML corpus database and subsequently synthesize new, higher-quality ML pipelines for user datasets. The transformation can employ a novel abstract syntax tree (AST) level analysis of the human-written pipeline to extract the necessary procedural elements and transform them in a grammatically well-formed manner, taking into account all the grammatical and stylistic variations that human-written programs such as ML pipelines may typically contain. Therefore, new ML pipelines can be synthesized effectively.

[0027] The embodiments of the present disclosure will be described with reference to the accompanying drawings.

[0028] Figure 1This is a diagram illustrating an example environment related to the enhancement of a machine learning pipeline corpus according to at least one embodiment described in this disclosure. (Refer to...) Figure 1 An environment 100 is shown. Environment 100 may include electronic device 102, server 104, client device 106, database 108, and communication network 110. Electronic device 102, database 104, and client device 106 may be communicatively coupled to each other via communication network 110. Electronic device 102 may include a meta-learning model 102A. Figure 1 The diagram also shows a user 112 who can associate with or operate electronic device 102 (or user terminal device 106). Database 108 may include multiple ML projects. For example, the multiple ML projects may include "n" ML projects, such as ML project-1114, ML project-2116, ..., and ML project-n118. Each of the multiple ML projects may include a dataset and a set of ML pipelines applicable to that dataset. For example, ML project-1114 may include dataset 114A and ML pipeline set 114B. Similarly, ML project-2116 may include dataset 116A and ML pipeline set 116B. Likewise, ML project-n118 may include dataset 118A and ML pipeline set 118B.

[0029] Figure 1 The "n" ML items shown are merely illustrative. Without departing from the scope of this disclosure, multiple ML items can include only two ML items or more than "n" ML items. For the sake of brevity, ... Figure 1 Only “n” ML items are shown in this document. However, in some implementations, more than “n” ML items may exist without limiting the scope of this disclosure.

[0030] Electronic device 102 may include appropriate logic, circuitry, and interfaces. Electronic device 102 may be configured to receive ML items (e.g., ML item-1 114) from a plurality of ML items stored in a machine learning (ML) corpus database (e.g., database 108). Electronic device 102 may also be configured to transform a first ML pipeline in a first ML pipeline set associated with the received ML item to determine a second ML pipeline set based on a predefined set of ML pipelines. Electronic device 102 may also be configured to select one or more ML pipelines from a determined second ML pipeline set based on performance scores associated with each second ML pipeline in the determined second ML pipeline set. Electronic device 102 may also be configured to augment the ML corpus database to include the selected one or more ML pipelines and the first ML pipeline set associated with the received ML item. Examples of electronic device 102 may include, but are not limited to, computing devices, smartphones, cellular phones, mobile phones, gaming devices, mainframes, servers, computer workstations, and / or consumer electronics (CE) devices.

[0031] The meta-learning model 102A may include appropriate logic, circuitry, interfaces, and / or code. The meta-learning model 102A can be configured to use meta-learning algorithms to generate predictive models based on previously trained models (e.g., ML pipelines) and data frames or features. The meta-learning model 102A can learn from the outputs of other learning algorithms. For example, for prediction, the meta-learning model 102A can learn based on the outputs of other learning algorithms. In another example, parameters of other ML models (e.g., neural network models, multinomial regression models, random forest classifiers, logistic regression models, or ensemble learning models) and data frames / features corresponding to each ML algorithm can be fed to the meta-learning model 102A. The meta-learning model 102A can learn meta-features and meta-heuristics based on the parameters and data frames / features associated with each of the input ML models. In the example, once the meta-learning model 102A is trained, it can be used to generate ML pipelines based on input features or data frames associated with ML projects.

[0032] Server 104 may include appropriate logic, circuitry, and interfaces and / or code. Server 104 may be configured to receive ML projects from a plurality of ML projects stored in a machine learning (ML) corpus database (e.g., database 108). Server 104 may also be configured to transform a first ML pipeline in a first ML pipeline set associated with the received ML projects to determine a second ML pipeline set. Server 104 may also be configured to select one or more ML pipelines from the determined second ML pipeline set based on performance scores associated with each second ML pipeline in the determined second ML pipeline set. Server 104 may also be configured to augment the ML corpus database to include the selected one or more ML pipelines and the first ML pipeline set associated with the received ML projects. Server 104 may be implemented as a cloud server and may perform operations via web applications, cloud applications, HTTP requests, repository operations, file transfers, etc. Other example implementations of server 104 may include, but are not limited to, database servers, file servers, web servers, media servers, application servers, mainframe servers, or cloud computing servers.

[0033] In at least one embodiment, server 104 can be implemented as multiple distributed cloud-based resources using several techniques known to those skilled in the art. Those skilled in the art will understand that the scope of this disclosure is not limited to implementing server 104 and electronic device 102 as two separate entities. In some embodiments, without departing from the scope of this disclosure, the functionality of server 104 may be wholly or at least partially integrated into electronic device 102. In some embodiments, server 104 may host database 108. Alternatively, server 104 may be decoupled from database 108 but may be communicatively coupled to database 108.

[0034] Client device 106 may include appropriate logic, circuitry, interfaces, and / or code. Client device 106 may be configured to store real-time applications, in which new ML pipelines may be synthesized based on enhancements to an ML corpus. In some embodiments, client device 106 may receive first user input from a user (e.g., a data scientist, such as user 112) and generate a predefined set of ML pipelines based on the received first user input. In another embodiment, client device 106 may receive one or more second user inputs from a user (e.g., a data scientist, such as user 112) and generate a set of ML pipelines associated with each of multiple ML projects based on one or more second user inputs. Additionally, client device 106 may receive multiple datasets associated with multiple projects from various sources (e.g., online dataset repositories, code repositories, and online open-source projects). Client device 106 may be configured to upload the predefined set of ML pipelines, along with the multiple ML pipelines and multiple datasets associated with the multiple projects, to server 104. Multiple uploaded ML pipelines and datasets can be stored as multiple ML projects in the ML corpus database in database 108. Uploaded predefined sets of ML pipelines can also be stored in database 108 along with multiple ML projects. Examples of client devices 106 may include, but are not limited to, mobile devices, desktop computers, laptop computers, computer workstations, computing devices, mainframes, servers (e.g., cloud servers), and server groups.

[0035] Database 108 may include appropriate logic, interfaces, and / or code. Database 108 may be configured to store multiple ML projects, wherein each ML project may include a dataset and a set of ML pipelines applicable to that dataset. Database 108 may also store a predefined set of ML pipelines. Database 108 may be derived from data in relational or non-relational databases or from a set of comma-separated value (CSV) files in a general or big data storage system. Database 108 may be stored or cached on a device such as a server (e.g., server 104) or electronic device 102. The device storing database 108 may be configured to receive queries from electronic device 102 for ML projects from multiple machine learning (ML) projects. In response, the device storing database 108 may be configured to retrieve and provide the queried machine learning ML project, including the dataset for the queried ML project and the set of ML pipelines applicable to that dataset, based on the received query, to electronic device 102.

[0036] In some implementations, database 108 may be hosted on multiple servers located in the same or different locations. The operation of database 108 may be performed using hardware including processors, microprocessors (e.g., to perform or control the execution of one or more operations), field-programmable gate arrays (FPGAs), or application-specific integrated circuits (ASICs). In some other instances, database 108 may be implemented using software.

[0037] Communication network 110 may include a communication medium through which electronic device 102, server 104, and user terminal device 106 can communicate with each other. Communication network 110 may be either a wired or wireless connection. Examples of communication network 110 may include, but are not limited to, the Internet, cloud networks, cellular or wireless mobile networks (e.g., LTE and 5G new radio), Wi-Fi networks, personal area networks (PANs), local area networks (LANs), or metropolitan area networks (MANs). Various devices in environment 100 may be configured to connect to communication network 110 according to various wired and wireless communication protocols. Examples of such wired and wireless communication protocols may include, but are not limited to, at least one of Transmission Control Protocol and Internet Protocol (TCP / IP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), File Transfer Protocol (FTP), ZigBee, EDGE, IEEE 802.11, Li-Fi, 802.16, IEEE 802.11s, IEEE 802.11g, multi-hop communication, wireless access point (AP), device-to-device communication, cellular communication protocols, and Bluetooth (BT) communication protocols.

[0038] In operation, electronic device 102 can receive ML projects (e.g., ML project-1 114) from multiple ML projects stored in an ML corpus database (e.g., database 108). Each of the multiple ML projects can include a dataset and a set of ML pipelines applicable to that dataset. ML projects help applications perform tasks such as prediction tasks (e.g., classification or regression) without being programmed to do so. The dataset can include historical data corresponding to a specific ML task defined on that dataset. An ML pipeline, script, or program can be a sequence of operations for training an ML model for a specific ML prediction task. It is understood that data scientists or developers can generate several ML pipelines and datasets and upload such ML pipelines and datasets as knowledge bases to various online source code and ML repositories on the Internet. The knowledge base can be downloaded from the Internet and stored in database 108. Electronic device 102 can receive ML projects from multiple ML projects in database 108. For example, electronic device 102 can receive ML project-1 114. In this paper, electronic device 102 can receive dataset 114A and ML pipeline set 114B corresponding to ML project-1 114. For example, in Figure 5 The document also provides details on several ML projects.

[0039] It can be noted that human-written ML pipelines, which may constitute the training corpus of learning-based machine learning methods, may not contain the best representative ML pipeline solution for each dataset. Worse still, the pipeline may have varying quality, and therefore, may lack any learnable patterns regarding which components can be used for which types of datasets. This problem can be particularly severe when using ML models as the choice of ML model in an ML pipeline, and can significantly impact the accuracy of the ML pipeline. Human-written ML pipelines can be instantiated in many ways, such as through cross-validation, hyperparameter optimization, etc. Safe refactoring of syntactically and semantically correct human-written ML pipelines can be critical and challenging. Therefore, human-written ML pipelines may need to be transformed.

[0040] Electronic device 102 can transform a first ML pipeline in a set of ML pipelines associated with a received ML project based on a predefined set of ML pipelines to determine a second set of ML pipelines. In this document, the transformation of the first ML pipeline can correspond to replacing the first ML model associated with the first ML pipeline with a second ML model associated with a predefined ML pipeline in the predefined set of ML pipelines. It can be noted that this transformation can systematically change the ML model of the original ML pipeline to improve the accuracy of that ML pipeline. As discussed, the first ML pipeline may be human-written and may not be optimal for various reasons. In one example, the data scientist who wrote the set of ML pipelines associated with the received ML project may not be an expert, and therefore may not be able to use the optimal ML model. In another example, the set of ML pipelines associated with the received ML project may not be optimal because the optimal model for the dataset is not available. Therefore, the set of ML pipelines applicable to the dataset corresponding to the received ML project may need to be transformed to improve its accuracy. Human-written ML pipelines may differ syntactically and semantically. Therefore, a universal pattern for replacing an original ML pipeline with another may not exist. The electronic device 102 of this disclosure can transform a first ML pipeline associated with a received ML item based on a predefined set of ML pipelines to determine a second set of ML pipelines. In one embodiment, the first ML pipeline can be transformed to determine the second ML pipeline. In an alternative embodiment, the first ML pipeline can be transformed to determine the second set of ML pipelines. In another embodiment, the electronic device 102 can transform each ML pipeline or a selected group of ML pipelines in the set of ML pipelines associated with the received ML item. For example, Figure 3 The document also provides details of the transformation of the first ML pipeline.

[0041] Electronic device 102 can select one or more ML pipelines from the determined second ML pipeline set based on performance scores associated with each second ML pipeline in the determined second ML pipeline set. Each second ML pipeline in the second ML pipeline set can be ranked based on the performance scores associated with each second ML pipeline in the determined second ML pipeline set. Performance scores can be F1 scores, R2 scores, etc., associated with the corresponding ML pipeline in the second ML pipeline set. For a given dataset and original ML pipelines, electronic device 102 can execute each transformed ML pipeline on the training dataset in a manner similar to how a data scientist executes ML pipelines on the dataset. Since the purpose of this disclosure is to enhance machine learning pipeline corpora with robust ML pipelines, only the best transformed ML pipelines can be retained along with existing ML pipelines for meta-learning model training. One or more ML pipelines selected from the determined second ML pipeline set can be the best ML pipelines, which can be retained for meta-learning model training. For example, in Figure 3 The document also provides details on selecting one or more ML pipelines.

[0042] Electronic device 102 can enhance the ML corpus database in database 108 to include one or more selected ML pipelines and a first set of ML pipelines associated with the received ML item (e.g., ML item-1114). Enhancement can be a standard technique for improving the quality of an ML corpus database. Systematic data augmentation can be used to improve an ML corpus database. Once one or more ML pipelines are selected, electronic device 102 can add the selected one or more ML pipelines to database 108 to improve the quality of database 108. In one embodiment, electronic device 102 can add the selected one or more ML pipelines together with the first set of ML pipelines to database 108. In another embodiment, electronic device 102 can replace the first set of ML pipelines in database 108 with the selected one or more ML pipelines. For example, in Figure 5 The document also provides details on the enhancements to the ML corpus database.

[0043] Electronic device 102 can augment an ML corpus database to help improve the training of meta-learning model 102A by using a high-quality ML pipeline that performs well on a given dataset. Meta-learning model 102A can generate an abstract version of the ML pipeline. Therefore, if the accuracy of meta-learning model 102A is below a certain quality threshold, the generated ML pipeline may not be optimal. On the other hand, if the quality of meta-learning model 102A is good or above the certain quality threshold, the generated ML pipeline may be better. Therefore, if the training of meta-learning model 102A is based solely on the raw ML pipeline that can be obtained directly from the Internet, the quality of meta-learning model 102A may be unacceptable, despite precautions taken during the download of the raw ML pipeline.

[0044] Without departing from the scope of this disclosure, Figure 1 Modifications, additions, or omissions may be made. For example, environment 100 may include more or fewer elements than those shown and described in this disclosure. For example, in some embodiments, environment 100 may include electronic device 102 but not database 108. Additionally, in some embodiments, the functionality of each of database 108 and server 104 may be incorporated into electronic device 102 without departing from the scope of this disclosure.

[0045] Figure 2 This is a block diagram of a system for enhancing a machine learning pipeline corpus, according to at least one embodiment described in this disclosure. (Combined with...) Figure 1 To explain using elements Figure 2 . Reference Figure 2 The diagram 200 shows a block diagram of a system 202 including an electronic device 102. The electronic device 102 may include a processor 204, a memory 206, a meta-learning model 102A, an input / output (I / O) device 208 (including a display device 208A), and a network interface 210.

[0046] Processor 204 may include appropriate logic, circuitry, and interfaces, and may be configured to execute a set of instructions stored in memory 206. Processor 204 may be configured to execute program instructions associated with different operations to be performed by electronic device 102. For example, some operations may include receiving machine learning (ML) projects, switching a first ML pipeline, selecting one or more ML pipelines from a determined set of second ML pipelines, and enhancing an ML corpus database. Processor 204 may be implemented based on a variety of processor technologies known in the art. Examples of processor technologies may include, but are not limited to, a central processing unit (CPU), x86-based processors, reduced instruction set computing (RISC) processors, application-specific integrated circuit (ASIC) processors, complex instruction set computing (CISC) processors, graphics processing units (GPUs), and other processors.

[0047] Despite Figure 2 While shown as a single processor, processor 204 may include any number of processors configured to individually or collectively perform or direct the execution of any number of operations of electronic device 102 as described herein. Additionally, one or more processors may reside on one or more different electronic devices, such as different servers. In some embodiments, processor 204 may be configured to interpret and / or execute program instructions stored in memory 206 and / or process data. After program instructions are loaded into memory 206, processor 204 may execute the program instructions. Some examples of processor 204 may be a graphics processing unit (GPU), a central processing unit (CPU), a reduced instruction set computer (RISC) processor, an ASIC processor, a complex instruction set computer (CISC) processor, a coprocessor, and / or combinations thereof.

[0048] Memory 206 may include appropriate logic, circuitry, and interfaces, and may be configured to store one or more instructions to be executed by processor 204. The one or more instructions stored in memory 206 may be executed by processor 204 to perform different operations of processor 204 (and electronic device 102). Memory 206 may be configured to store multiple ML items, a predefined set of ML pipelines, a first set of ML pipelines, a second set of ML pipelines, and one or more selected ML pipelines. Examples of memory implementations may include, but are not limited to, random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), hard disk drive (HDD), solid-state drive (SSD), CPU cache, and / or secure digital card (SD card).

[0049] I / O device 208 may include appropriate logic, circuitry, and interfaces, and may be configured to receive input from user 112 and provide output based on the received input. For example, I / O device 208 may receive user input associated with the generation of an ML pipeline, a dataset associated with an ML pipeline, or an ML project from user 112. Furthermore, I / O device 208 may present a determined second set of ML pipelines, one or more selected ML pipelines, and / or the predicted output of meta-learning model 102A, which may be trained based on an enhanced ML corpus database. I / O device 208 may include various input and output devices and may be configured to communicate with processor 204. Examples of I / O device 208 may include, but are not limited to, a touchscreen, keyboard, mouse, joystick, microphone, display device (e.g., display device 208A), and speaker.

[0050] Display device 208A may include suitable logic, circuitry, and interfaces, and may be configured to display a second set of ML pipelines and / or one or more selected ML pipelines. Display device 208A may be a touchscreen that allows a user (e.g., user 112) to provide user input via display device 208A. The touchscreen may be at least one of resistive, capacitive, or thermal touchscreens. Display device 208A may be implemented using several known technologies, such as, but not limited to, liquid crystal display (LCD), light-emitting diode (LED) display, plasma display, and / or organic LED (OLED) display technologies and / or at least one of other display devices. According to embodiments, display device 208A may refer to a display screen of a head-mounted device (HMD), a smart glass device, a see-through display, a projection-based display, an electrochromic display, or a transparent display.

[0051] Network interface 210 may include appropriate logic, circuitry, and interfaces, and may be configured to facilitate communication between processor 204, server 104, user terminal device 106 (or any other device in environment 100) via communication network 110. Network interface 210 may be implemented using various known technologies to support wired or wireless communication between electronic device 102 and communication network 110. Network interface 210 may include, but is not limited to, antennas, radio frequency (RF) transceivers, one or more amplifiers, tuners, one or more oscillators, digital signal processors, encoder-decoder (CODEC) chipsets, subscriber identification module (SIM) cards, or local buffer circuitry. Network interface 210 may be configured to communicate wirelessly with networks such as the Internet, intranets, or wireless networks (e.g., cellular telephone networks, wireless local area networks (LANs), and metropolitan area networks (MANs)). Wireless communication can be configured to use one or more of a number of communication standards, protocols and technologies, such as Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (W-CDMA), Long Term Evolution (LTE), Fifth Generation (5R) New Radio (NR), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Wi-Fi (e.g., IEEE 802.11a, IEEE 802.11b, IEEE 802.11g or IEEE 802.11n), Voice over Internet Protocol (VoIP), Li-Fi Lightweight, Wi-MAX, Email Protocol, Instant Messaging and Short Message Service (SMS).

[0052] Modifications, additions, or omissions may be made to the example electronic device 102 without departing from the scope of this disclosure. For example, in some embodiments, the example electronic device 102 may include any number of other components that may not be explicitly shown or described for brevity purposes.

[0053] Figure 3 This is a flowchart illustrating an example method for enhancing a machine learning pipeline corpus to synthesize new machine learning pipelines, according to embodiments of this disclosure. (Combined with...) Figure 1 and Figure 2 To describe the elements Figure 3 . Reference Figure 3 Flowchart 300 is shown. The method shown in flowchart 300 can begin at block 302 and can be performed by any suitable system, device, or apparatus, such as by... Figure 1 Example electronic device 102 or Figure 2Execution 204. Although shown in discrete boxes, depending on the specific implementation, the steps and operations associated with one or more boxes in flowchart 300 may be divided into additional boxes, combined into fewer boxes, or removed.

[0054] At box 302, ML projects (e.g., ML project-1 114) from multiple ML projects stored in a machine learning (ML) corpus database (e.g., database 108) can be received. In this document, each of the multiple ML projects may include a dataset and a set of ML pipelines applicable to that dataset. Processor 204 can be configured to receive ML projects from the multiple ML projects stored in the ML corpus database (e.g., database 108). ML projects can help applications perform tasks such as prediction tasks (e.g., classification or regression) without being programmed to do so. The dataset may include historical data corresponding to a specific ML task defined on that dataset. The ML pipeline may include sequences of operations that can be used to train ML models for a specific ML task. It is understood that data scientists or developers may generate several ML pipelines and datasets and upload such ML pipelines and datasets as knowledge bases to various online source code and ML repositories on the Internet. The knowledge base can be downloaded from the Internet and can be stored in database 108. Electronic device 102 can receive ML projects from the multiple ML projects in database 108. For example, electronic device 102 can receive ML item-1 114. In this document, electronic device 102 can receive dataset 114A and ML pipeline set 114B corresponding to ML item-1 114.

[0055] At box 304, a first ML pipeline in a first ML pipeline set associated with a received ML project can be transformed based on a predefined ML pipeline set to determine a second ML pipeline set. In this document, the transformation of the first ML pipeline can correspond to replacing the first ML model associated with the first ML pipeline with a second ML model associated with a predefined ML pipeline in the predefined ML pipeline set. Furthermore, the first ML pipeline set can correspond to each ML pipeline associated with the received ML project. A first ML pipeline can be selected from the first ML pipeline set. Processor 204 can be configured to transform the first ML pipeline associated with the received ML project based on a predefined ML pipeline in the predefined ML pipeline set to determine the second ML pipeline set. It can be noted that this transformation can systematically change the ML model of the original ML pipeline to improve the accuracy of that ML pipeline. As discussed, the first ML pipeline can be human-written and may not be optimal for various reasons. In one example, the data scientist who might write the first set of ML pipelines associated with the received ML project might not be an expert, and therefore might not be able to use the optimal ML model. In another example, the first set of ML pipelines associated with the received ML project might not be optimal because the optimal model for the dataset is not available. Therefore, the first set of ML pipelines associated with the received ML project may need to be transformed to improve its accuracy. Human-written ML pipelines may be syntactically and semantically different. Therefore, there may not be a universal pattern that can be used to replace the original ML pipeline with another ML pipeline. Processor 204 can transform the first ML pipeline associated with the received ML project based on a predefined set of ML pipelines to determine a second set of ML pipelines. In one implementation, the first ML pipeline can be transformed to determine a second ML pipeline. In an alternative implementation, the first ML pipeline can be transformed to determine a second set of ML pipelines that includes multiple second ML pipelines. In one implementation, electronic device 102 can transform each ML pipeline in the set of ML pipelines associated with the received ML project individually. For example, in Figure 5 The document also provides details of the transformation of the first ML pipeline.

[0056] At box 306, one or more ML pipelines can be selected from the determined second ML pipeline set based on a performance score associated with each of the second ML pipelines in the determined second ML pipeline set. Processor 204 can be configured to select one or more ML pipelines from the determined second ML pipeline set based on a performance score associated with each of the second ML pipelines in the determined second ML pipeline set. Each second ML pipeline in the determined second ML pipeline set can be ranked based on a performance score associated with each of the second ML pipelines in the determined second ML pipeline set. For a given dataset and original ML pipelines, electronic device 102 can automatically execute each transformed ML pipeline on the training dataset in a manner similar to how a data scientist executes ML pipelines on the dataset. Since the purpose of this disclosure is to enhance machine learning pipeline corpora with robust ML pipelines, only the best transformed ML pipelines can be retained along with existing ML pipelines. One or more ML pipelines selected from the determined second set of ML pipelines may be the optimal ML pipelines, which can be reserved for training the meta-learning model.

[0057] In implementation, the performance score associated with each of the second ML pipelines in the defined set of second ML pipelines may correspond to, but is not limited to, the F1 score or R2 score associated with the corresponding ML pipeline. In the example, the performance score of the ML pipeline may correspond to the ratio of the F1 score to the R2 score of the ML pipeline. It is understood that the F1 score can be used to determine the accuracy of a given ML model for a given dataset based on the harmonic mean of the precision and recall of a given ML model. In this document, the precision of the model may be determined based on the ratio of the number of true positive results to the total number of positive results including both true positives and false positives. The recall or sensitivity of a given ML model may be determined based on the ratio of the number of true positive results to the total number of true positives and false negatives. The F1 score may be determined based on equation (1), as follows:

[0058]

[0059] The F1 score can range from "0" to "1". In this document, the F1 score is likely to be close to "1" when the precision and recall of a given ML model are likely to be close to "1". Similarly, the F1 score can be "0" when the precision or recall of a given ML model is likely to be close to "0". In an implementation, an F1 score can be determined associated with each of the second ML pipelines in the determined set of second ML pipelines to select one or more ML pipelines. In an example, the determined set of ML pipelines may include five ML pipelines, such as ML pipeline-1 with an F1 score of "0.2", ML pipeline-2 with an F1 score of "0.27", ML pipeline-3 with an F1 score of "0.55", ML pipeline-4 with an F1 score of "0.65", and ML pipeline-5 with an F1 score of "0.72". In this document, ML pipeline-5 may be the most accurate of the five ML pipelines. Therefore, ML Pipeline-5 can be selected as one or more ML pipelines.

[0060] The R² score (also known as the R-squared score) can be a coefficient that indicates how well a given ML model fits a dataset. The R² score describes how changes in the independent variables can affect the dependent variable of a given ML model. The R² score can be determined based on the ratio of the sum of squares (SSR) of the regression or residuals to the total sum of squares (TSS). In this context, the SSR can be the total variation of the predicted values ​​from the mean of all dependent variables. The TSS can be the total variation of the actual values ​​from the mean. Similar to the F1 score, the R² score can have values ​​between "0" and "1". When the R² score is "1", the variation in the dependent variable can be fully described by the variation in the independent variables. That is, when the R² score is close to "1", the given ML model fits the dataset better. In an implementation, an R² score can be determined associated with each of the second ML pipelines in a defined set of second ML pipelines to select one or more ML pipelines. In the example, the determined second set of ML pipelines may include five ML pipelines, such as ML pipeline-1 with an R2 score of "0.4", ML pipeline-2 with an R2 score of "0.5", ML pipeline-3 with an R2 score of "0.75", ML pipeline-4 with an R2 score of "0.65", and ML pipeline-5 with an R2 score of "0.6". In this paper, since ML pipeline-3 can have the highest R2 score among the five ML pipelines, it may be the most suitable one. Therefore, ML pipeline-3 can be selected as one or more ML pipelines.

[0061] In an implementation, one or more ML pipelines selected from the determined second ML pipeline set may include, but are not limited to, ML pipelines associated with the maximum performance score (from the determined second ML pipeline set), a first group of ML pipelines that may correspond to performance scores above a threshold (from the determined second ML pipeline set), or a second group of ML pipelines that may correspond to a predefined number of ML pipelines based on performance scores (from the determined second ML pipeline set).

[0062] For example, the selected one or more ML pipelines may include ML pipelines (from the determined second set of ML pipelines) associated with the highest performance score (e.g., F1 score or R2 score). In this document, the ML pipeline with the highest performance score may be selected as one or more ML pipelines from the determined second set of ML pipelines. In the example, the determined set of ML pipelines may include seven ML pipelines, such as ML pipeline-1 with an R2 score of "0.4", ML pipeline-2 with an R2 score of "0.5", ML pipeline-3 with an R2 score of "0.75", ML pipeline-4 with an R2 score of "0.65", ML pipeline-5 with an R2 score of "0.6", ML pipeline-6 ​​with an R2 score of "0.8", and ML pipeline-7 with an R2 score of "0.49". In this document, ML pipeline-6 ​​has the highest R2 score of 0.8. Therefore, ML Pipeline-6 ​​can be selected as one or more ML pipelines.

[0063] In another example, the selected one or more ML pipelines may include a first set of ML pipelines (from the determined second set of ML pipelines) that can correspond to performance scores above a threshold. Here, the threshold can be a value that can be used to select one or more ML pipelines (e.g., F1 score or R2 score). Consider a previous scenario involving a determined second set of ML pipelines comprising seven ML pipelines and a threshold of “0.7”. Here, ML pipeline-3 (with an R2 score of “0.75”) and ML pipeline-4 (with an R2 score of “0.8”) can be selected as one or more ML pipelines that can have performance scores (e.g., R2 scores) greater than the threshold (e.g., “0.7”).

[0064] In another example, the selected one or more ML pipelines may include a second set of ML pipelines (from the determined second ML pipeline set) that may correspond to a predefined number of ML pipelines based on performance scores. Processor 204 may select the top K (e.g., the top 3) ML pipelines based on the performance of each second ML pipeline in the determined second ML pipeline set. Consider a previous scenario with a determined second ML pipeline set including seven ML pipelines and a predefined value (i.e., K) of three. In such a case, the top three ML pipelines (based on their respective performance scores) may be selected as one or more ML pipelines. For example, the selected one or more ML pipelines may include ML pipeline-2 with an R2 score of "0.69", ML pipeline-3 with an R2 score of "0.75", and ML pipeline-4 with an R2 score of "0.8".

[0065] At block 308, the ML corpus database (e.g., database 108) can be enhanced to include one or more selected ML pipelines and a first set of ML pipelines associated with received ML items. Processor 204 can be configured to enhance the ML corpus database (i.e., database 108) to include one or more selected ML pipelines and a first set of ML pipelines associated with received ML items. Once one or more ML pipelines are selected, electronic device 102 can add the selected one or more ML pipelines to database 108 to improve the quality of the ML corpus database stored in database 108. In one embodiment, electronic device 102 can store the selected one or more ML pipelines together with the first set of ML pipelines in database 108. In another embodiment, electronic device 102 can replace the first set of ML pipelines in database 108 with the selected one or more ML pipelines. Control can be passed to the end.

[0066] Although flowchart 300 is shown as discrete operations, such as 302, 304, 306, and 308, in some embodiments, without departing from the essence of the disclosed embodiments, such discrete operations may be divided into other operations, combined into fewer operations, or eliminated, depending on the specific implementation.

[0067] Figure 4 This is a flowchart illustrating an example method, according to an embodiment of the present disclosure, for transforming a first ML pipeline in a set of ML pipelines associated with a received ML project to determine a second ML pipeline set. (In conjunction with...) Figure 1 , Figure 2 and Figure 3 To describe the elements Figure 4 . Reference Figure 4 Flowchart 400 is shown. The method shown in flowchart 400 can begin at 402 and can be performed by any suitable system, device, or apparatus, such as by... Figure 1 Example electronic device 102 or Figure 2 Execution 204. Although shown in discrete boxes, depending on the specific implementation, the steps and operations associated with one or more boxes in flowchart 400 may be divided into additional boxes, combined into fewer boxes, or removed.

[0068] At box 402, a code snippet of the first ML model associated with the first ML pipeline can be identified. In this document, a code snippet can be a portion of the first ML pipeline's code that implements the first ML model. Processor 204 can be configured to identify the code snippet of the first ML model associated with the first ML pipeline. Since only a portion of the code in the first ML pipeline can implement the first ML model, a specific portion can be identified and transformed to obtain a second ML model. For example, in Figure 9 The document also provides details on the identification of code snippets from the first ML model.

[0069] At box 404, one or more input parameters associated with the identified code segment can be determined. Processor 204 can be configured to determine one or more input parameters associated with the identified code segment. To maintain the functionality of the second ML model identical to that of the first ML model, one or more input parameters corresponding to the identified code segment of the first ML model can be fed into the second ML model. Thus, one or more input parameters of the identified code segment can be determined and later fed into the second ML model to obtain an equivalent output, which can be determined based on a transformation of the first ML model. For example, in Figure 6A The document also provides details of the identification of one or more input parameters associated with the identified code snippet.

[0070] In an implementation, the determined input parameters associated with the identified code snippet may include at least one of the following: a training dataset, a test dataset, and a set of hyperparameters associated with the first ML model. Herein, the training dataset may be used to train the first ML model based on updates to the weights associated with it. The test dataset may be used to test the trained first ML model to check whether its accuracy is within permissible limits. The set of hyperparameters may correspond to a set of parameters that can control the learning of the first ML model. For example, the learning rate, the number of neural network layers, the number of neurons in each neural network layer, etc., may correspond to the set of hyperparameters. The processor 204 may collect all read accesses in one or more input parameters of each application programming interface (API) in the identified code snippet.

[0071] At box 406, a second ML model can be selected from a set of predefined models associated with a predefined ML pipeline set. Processor 204 can be configured to select the second ML model from the set of predefined models associated with the predefined ML pipeline set. The set of predefined models can be ML models that may have already been created. For example, the set of predefined models may correspond to human-written ML models, which can be created as template ML models for certain application scenarios and datasets. The set of predefined models can be stored in database 108. Processor 204 can receive the set of predefined models associated with the predefined ML pipeline set from database 108. Alternatively, the set of predefined models can be pre-stored in memory 206, and processor 204 can retrieve the set of predefined models from memory 206. Once the set of predefined models can be received / retrieved, processor 204 can select the second ML model from the set of predefined models. For example, in Figure 5 The document also provides details on selecting a second ML model.

[0072] At box 408, the selected second ML model can be instantiated based on replacing the first ML model with the selected second ML model in the first ML pipeline. Processor 204 can be configured to instantiate the selected second ML model based on replacing the first ML model with the selected second ML model in the first ML pipeline. In this document, the selected second ML model can replace the identified code snippet. One or more input parameters associated with the identified code snippet can be provided to the selected second ML model. Since only the identified code snippet corresponding to the first ML model in the first ML pipeline can be replaced by the selected second ML model, and one or more input parameters can be the same, the functionality of the instantiated second ML model can remain the same as that of the first ML model. For example, in Figure 7The documentation also provides details on instantiating the second ML model. Control can be passed to the end.

[0073] Although flowchart 400 is shown as discrete operations such as 402, 404, 406, and 408, in some embodiments, such discrete operations may be divided into other operations, combined into fewer operations, or eliminated, depending on the specific implementation, without departing from the essence of the disclosed embodiments.

[0074] Figure 5 This is a diagram illustrating an exemplary scenario for enhancing a machine learning pipeline corpus based on a predefined machine learning pipeline, according to at least one embodiment described in this disclosure. Combined with... Figure 1 , Figure 2 , Figure 3 and Figure 4 To describe the elements Figure 5 . Reference Figure 5 An exemplary scenario 500 is illustrated. Exemplary scenario 500 may include a database 108, a first ML pipeline 502, a code snippet recognizer 504, a code snippet 506, an input parameter lookup unit 508, a first ML model 510, a predefined model set 512, an instantiation block 514, a second ML pipeline set 516, an ML pipeline evaluator 518, one or more ML pipelines 520, and an enhancement block 522. Database 108 may include multiple ML projects, including "n" ML projects, such as ML project-1 114, ML project-2 116, ..., and ML project-n 118. Each of the multiple ML projects may include a dataset and an ML pipeline set applicable to that dataset. ML project-1 114 may include dataset 114A and ML pipeline set 114B, and ML project-2 116 may include dataset 116A and ML pipeline set 116B. Similarly, ML project-n118 can include dataset 118A and ML pipeline collection 118B.

[0075] Figure 5 The "n" ML items shown are for illustrative purposes only. Without departing from the scope of this disclosure, multiple ML items may include only two ML items or more than "n" ML items. For the sake of brevity, Figure 5 Only “n” ML items are shown in this document. However, in some implementations, more than “n” ML items may exist without limiting the scope of this disclosure.

[0076] Processor 204 can be configured to receive machine learning ML projects from multiple ML projects stored in an ML corpus database (e.g., database 108). Each of the multiple ML projects can include a dataset and a set of ML pipelines applicable to that dataset. Machine learning ML projects help applications perform tasks such as prediction without being programmed to do so. Datasets can include training datasets, validation datasets, and test datasets for the corresponding machine learning (ML) project. Training datasets can be used to train an ML model corresponding to the machine learning (ML) project. Validation datasets can be used to validate the trained ML model corresponding to the machine learning (ML) project. Validation datasets can also improve the accuracy of the training dataset and prevent overfitting or underfitting of a given ML model. Test datasets can be used to test the trained ML model to check if the accuracy of the trained ML model is within permissible limits. ML pipelines can be scripts or programs that include a sequence of operations for training an ML model for a specific ML prediction task. It is understood that data scientists or developers can generate several ML pipelines and datasets and upload such ML pipelines and datasets as knowledge bases to various online source code and ML repositories on the Internet. A knowledge base can be downloaded from the Internet and stored in database 108. Electronic device 102 can receive multiple ML projects from database 108. For example, electronic device 102 can receive ML project-1 114. In this document, electronic device 102 can receive a dataset 114A and an ML pipeline set 114B corresponding to ML project-1 114. Each ML pipeline in ML pipeline set 114B can be transformed one after another. For example, the first ML pipeline 502 can be selected for transformation.

[0077] Processor 204 can be configured to identify code fragment 506 of a first ML model associated with a first ML pipeline 502. Herein, code fragment 506 can be a portion of the code in the first ML pipeline 502 that implements the first ML model. Since only a portion of the code in the first ML pipeline 502 can implement the first ML model, that particular portion can be transformed to obtain a second ML model. To identify code fragment 506 of the first ML model associated with the first ML pipeline 502, the first ML pipeline 502 can be provided to code fragment recognizer 504. For example, in… Figure 6A The document also provides details on the identification of code snippet 506 from the first ML model.

[0078] Processor 204 can be configured to determine one or more input parameters associated with the identified code snippet 506 using input parameter lookup 508. Input parameter lookup 508 can be fed the identified code snippet 506 to determine one or more input parameters associated with it. The one or more input parameters may include training datasets, test datasets, and hyperparameters. To maintain the functionality of the second ML model identical to that of the first ML model, one or more input parameters corresponding to the identified code snippet 506 of the first ML model can also be fed into the second ML model. Therefore, one or more input parameters of the identified code snippet 506 can be determined. Figure 5 In this context, a first ML model with one or more input parameters (i.e., an ML model with input parameters) is denoted as first ML model 510. For example, in... Figure 6A The document also provides details of the identification of one or more input parameters associated with the identified code snippet 506.

[0079] Processor 204 can be configured to select a second ML model from a predefined model set 512 associated with a predefined ML pipeline set. The predefined model set 512 can be an already created ML model. The predefined model set 512 associated with the predefined ML pipeline set can be stored in database 108 or pre-stored in memory 206. Therefore, if the predefined model set 512 is stored in database 108, processor 204 can receive the predefined model set 512 from database 108; alternatively, if the predefined model set 512 is pre-stored in memory 206, processor 204 can retrieve the predefined model set 512 from memory 206.

[0080] Processor 204 can be configured to instantiate a selected second ML model using instantiation block 514 based on replacing the first ML model 510 with the selected second ML model in the first ML pipeline 502. The selected second ML model can be instantiated based on replacing the first ML model 510 with the selected second ML model in the first ML pipeline 502. Herein, a code snippet of the selected second ML model can replace the identified code snippet 506. One or more input parameters associated with the identified code snippet 506 can be provided to the selected second ML model. Since only the identified code snippet 506 corresponding to the first ML model 510 in the first ML pipeline 502 can be replaced by the selected second ML model, the functionality of the first ML pipeline 502 after the transformation can remain the same. After replacing the first ML model 510 with the second ML model, a second ML pipeline set 516 can be determined. For example, refer to... Figure 5 The second ML pipeline set 516 can include three ML pipelines. For example, in Figure 9 The document also provides details on instantiating the second ML model.

[0081] Processor 204 can be configured to select one or more ML pipelines from the determined second ML pipeline set 516 based on a performance score associated with each second ML pipeline in the determined second ML pipeline set 516. ML pipeline evaluator 518 can determine a performance score associated with each second ML pipeline in the determined second ML pipeline set 516. The performance score associated with each second ML pipeline in the determined second ML pipeline set 516 can correspond to at least one of an F1 score or an R2 score associated with the corresponding ML pipeline. In an implementation, one or more ML pipelines selected from the determined second ML pipeline set 516 may include, but are not limited to, one of the following: ML pipelines from the determined second ML pipeline set 516 associated with the maximum performance score; a first group of ML pipelines from the determined second ML pipeline set 516 corresponding to performance scores above a threshold; or a second group of ML pipelines from the determined second ML pipeline set 516 corresponding to a predefined number of ML pipelines based on performance scores. See also... Figure 5 The determined second ML pipeline set 516 comprises three ML pipelines. The selected one or more ML pipelines 520 may include two ML pipelines (e.g., the first two ML pipelines) selected from the determined second ML pipeline set 516 based on the performance score of each ML pipeline. For example, in Figure 3 The document also describes selecting one or more ML pipelines from the determined second ML pipeline set 516.

[0082] Processor 204 can use enhancement block 522 to enhance database 108 to include one or more selected ML pipelines 520 and a first ML pipeline set 502 associated with the received ML item-1 114. Once one or more ML pipelines 520 are selected, processor 204 can add the selected one or more ML pipelines 520 to database 108 to improve the quality of database 108. In one embodiment, electronic device 102 can add the selected one or more ML pipelines 520 together with the first ML pipeline set 502 to database 108. In another embodiment, electronic device 102 can replace the first ML pipeline 502 in database 108 with the selected one or more ML pipelines 520. Similarly, each of the first ML pipelines associated with each of multiple ML projects (e.g., ML project-1 114, ML project-2 116, and ML project-n 118) can be transformed, and one or more of their corresponding selected ML pipelines can be added to enhance database 108. (See reference...) Figure 5 After enhancement, the ML pipeline set 114B corresponding to ML item-1 114 can be changed to ML pipeline set 524, the ML pipeline set 116B corresponding to ML item-2 116 can be changed to ML pipeline set 526, and the ML pipeline set 118B corresponding to ML item-n 118 can be changed to ML pipeline set 528. The enhancement of database 108 can then be used to provide higher quality and more consistent learning pipelines / features to the meta-learning model 102A for learning from the corpus, and subsequently synthesize new, higher quality ML pipelines for other datasets.

[0083] It should be noted that Figure 5 Scenario 500 is for illustrative purposes and should not be construed as limiting the scope of this disclosure.

[0084] Figure 6A This is a diagram illustrating an exemplary scenario of identifying a code snippet of a first ML model associated with an exemplary first ML pipeline, according to an embodiment of this disclosure. Combined with information from... Figure 1 , Figure 2 , Figure 3 , Figure 4 and Figure 5 To describe the elements Figure 6A . Reference Figure 6AAn exemplary scenario 600A is illustrated. Exemplary scenario 600A may include a first ML pipeline 602. The first ML pipeline 602 may include statement-1 604, statement-2 606, statement-3 608, statement-4 610, statement-5 612, and statement-6 614. As described herein, electronic device 102 or processor 204 may recognize code snippets of a first ML model associated with the first ML pipeline 602.

[0085] Processor 204 can determine one or more input parameters corresponding to the first ML pipeline 602. For example, refer to Figure 6A The determined input parameters can include "X_train", "X_test", "y_train", and "y_test". All assignment expressions can correspond to key-value pairs associated with the hyperparameter set. For example, "family = sm.families.Binomial()" can correspond to hyperparameter key-value pairs that can define the first ML model based on a binomial distribution. API-level information can be used to map existing variables to "X_train", "X_test", "y_train", and "y_test". (See also...) Figure 6A The first ML pipeline 602 can be divided into code segment 616 and code segment 618. Code segment 616 can correspond to an irrelevant portion of the first ML pipeline 602 associated with the first ML model. Code segment 618 can correspond to the first ML model associated with the first ML pipeline 602. The processor 204 can identify code segment 618 corresponding to the first ML model associated with the first ML pipeline 602, and can identify code segment 616 as an irrelevant portion associated with the first ML pipeline 602. Therefore, the processor 204 can extract only the code segment 618 corresponding to the first ML model associated with the first ML pipeline 602 for instantiation. For example, refer to... Figure 6A The first ML model can be identified based on statements -3 608 and -4 610 in code snippet 618. For example, in Figure 8 The document also provides details on the identification of code snippets from the first ML model.

[0086] Figure 6B This is a diagram illustrating an exemplary scenario for instantiating a second ML model selected from a predefined set of models associated with a predefined ML pipeline set, according to an embodiment of this disclosure. Combined with... Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5 and Figure 6ATo describe the elements Figure 6B . Reference Figure 6B An exemplary scenario 600B is illustrated. Exemplary scenario 600B may include code snippets 620 and 622. Code snippet 620 may include statements 624 and 626. Code snippet 620 may include statements 9 and 628, 10 and 630, and 11 and 632. As described herein, electronic device 102 or processor 204 may instantiate a second ML model selected from a predefined set of models associated with a predefined ML pipeline set.

[0087] Reference Figure 6B The selected second ML model can be instantiated by replacing the first ML model with the selected second ML model. Code snippet 620 can correspond to the first ML model associated with the first ML pipeline. Code snippet 622 can correspond to the selected second ML model. The processor 204 can instantiate the selected second ML model by replacing code snippet 620 (including statements -7 624 and -8 626) with code snippet 622 (including statements -9 628, -10 630, and -11 632) in the first ML pipeline. Therefore, the first ML model can be instantiated by replacing the second ML model in the first ML pipeline. It can be noted that one or more input parameters of the first ML model determined for code snippet 620 can be fed as input parameters to the second ML model as defined by code snippet 622. For example, refer to Figure 6B One or more input parameters can be "y_train" and "X_train_sm". Instantiation of the selected second ML model can determine the second ML pipeline set. For example, in... Figure 7 The document also provides details on the instantiation of the selected second ML model.

[0088] Reference Figure 6A and Figure 6B It can be noted that although exemplary scenarios 600A and 600B have been shown in a high-level programming language (e.g., the "Python" programming language) using web-based computational notebooks, the teachings of this disclosure are applicable to other ML pipelines written in different languages ​​and development platforms. It can also be noted that web-based computational notebooks can correspond to computational structures that can be used, particularly during the development phase, to develop / represent ML pipelines. For example, web-based computational notebooks can be used to develop the scripting portion of an ML project (e.g., an ML pipeline).

[0089] It should be noted that Figure 6A and Figure 6BScenarios 600A and 600B are for illustrative purposes and should not be construed as limiting the scope of this disclosure.

[0090] Figure 7 This is a flowchart illustrating an example method for instantiating a second ML model selected from a predefined set of models associated with a predefined ML pipeline set, according to at least one embodiment described in this disclosure. (Combined with...) Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5 , Figure 6A and Figure 6B To describe the elements Figure 7 . Reference Figure 7 The flowchart 700 is shown. The method shown in flowchart 700 can begin at block 702 and can be performed by any suitable system, device, or apparatus, for example by... Figure 1 Example electronic device 102 or Figure 2 The processor 204 executes the process. Although shown in discrete boxes, depending on the specific implementation, the steps and operations associated with one or more boxes of flowchart 700 may be divided into additional boxes, combined into fewer boxes, or removed.

[0091] At box 702, a predefined template associated with the second ML model can be selected. The selected predefined template can be annotated with one or more input parameters of a code snippet from the identified first ML model. Processor 204 can be configured to select the predefined template associated with the second ML model. In this document, the predefined template associated with the second ML model may differ from the template associated with the first ML model. To maintain the functionality of the second ML model identical to that of the first ML model, one or more input parameters of a code snippet from the identified first ML model can be fed to the predefined template associated with the second ML model. For example, refer to... Figure 6BCode snippet 620 can be a code snippet of the identified first ML model. In this document, one or more input parameters may include “y_train” and “X_train_sm”. Code snippet 620 may also include hyperparameters that define the distribution family associated with the first ML model (e.g., a binomial distribution family), as provided in statement -7 624. The template associated with the first ML model can perform regression based on the binomial distribution. The predefined template associated with the second ML model may include a “RandomForestRegressor” function call (according to statements -9 628 and -10 630 of code snippet 622) instead of a function call associated with the “binomial” distribution family to perform regression. Furthermore, at statement -11 632, the input parameters “y_train” and “X_train_sm” can be passed as parameters to the predefined template associated with the second ML model to maintain the functionality of the second ML model identical to that of the first ML model.

[0092] At box 704, a code snippet for constructing a second ML model can be built based on parameterization of one or more function calls in a selected predefined template using one or more annotated input parameters. Processor 204 can be configured to construct a code snippet for constructing a second ML model based on parameterization of one or more function calls in a selected predefined template using one or more annotated input parameters. In this document, parameterization can be used to pass the values ​​of one or more annotated input parameters to one or more function calls in a selected predefined template. In other words, previously collected variable names pointing to appropriate holes in the selected predefined template can be inserted to construct a new model snippet, such as code snippet 622 for the second ML model. For example, refer to... Figure 6B The code snippet 622 for the second ML model can include "RandomForestRegressor" as a function call. In this paper, statement -9 628 calls the "RandomForestRegressor" function, which can be defined in the library "sklearn.ensemble". Statement -10 630 of code snippet 622 assigns the "RandomForestRegressor" function to the variable "logm3". As shown in statement -11 632, one or more annotated input parameters, such as "y_train" and "X_train_sm", can be passed to the "fit" function call associated with the variable "logm3" to construct the code snippet 622 for the second ML model.

[0093] At box 706, the code snippet 622 of the constructed second ML model can be used to replace the code snippet 620 of the identified first ML model to instantiate the second ML model. Processor 204 can be configured to use the code snippet of the constructed second ML model to replace the code snippet 620 of the identified first ML model to instantiate the second ML model. The functionality of the second ML model can be the same as that of the first ML model, since the input parameters can be the same for both. However, one or more functions called in the second ML model can be different from their corresponding functions called in the first ML model. Therefore, the code snippet 620 of the first ML model can be replaced by the code snippet 622 associated with the second ML model to instantiate the second ML model. Control can be passed to the end.

[0094] Although flowchart 700 is shown as discrete operations, such as 702, 704, and 706, in some embodiments, without departing from the essence of the disclosed embodiments, such discrete operations may be divided into other operations, combined into fewer operations, or eliminated, depending on the specific implementation.

[0095] Figure 8 This is a flowchart illustrating an example method for identifying one or more statements associated with a first ML model, according to at least one embodiment described in this disclosure. (Combined with...) Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5 , Figure 6A , Figure 6B and Figure 7 To describe the elements Figure 8 . Reference Figure 8 Flowchart 800 is shown. The method shown in flowchart 800 can begin at 802 and can be performed by any suitable system, device, or apparatus, such as by... Figure 1 Example electronic device 102 or Figure 2 The processor 204 executes the process. Although shown in discrete boxes, depending on the specific implementation, the steps and operations associated with one or more boxes of flowchart 800 may be divided into additional boxes, combined into fewer boxes, or removed.

[0096] At box 802, an Abstract Syntax Tree (AST) associated with the first ML pipeline can be constructed. Processor 204 can be configured to construct the AST associated with the first ML pipeline. It is understood that the AST of code can be a tree representation of the abstract syntactic framework of programming language code in a formal language. The AST may not include every detail of the code or its syntax, but rather may only include abstract syntax elements from the formal language, such as "while" blocks, "if" statements, conditional branches, comparison statements, assignment statements, variable names, etc. Each node of the tree can represent the syntactic structure of the code. Therefore, the AST associated with the first ML pipeline can be constructed based on a tree representation of the abstract syntax of statements in the first ML pipeline in the form of a formal language. The AST helps to easily manipulate and represent statements associated with the first ML model in the first ML pipeline. For example, refer to... Figure 6A This allows us to construct the AST of the first ML pipeline 602.

[0097] At box 804, the last application programming interface (API) call associated with the prediction function in the first ML pipeline can be determined based on the constructed AST. Processor 204 can be configured to determine the last API call associated with the prediction function in the first ML pipeline based on the constructed AST. It is understood that the API can be used to retrieve data based on an API endpoint that can be exposed by the API function. To receive data, a request (also called an API call) can be sent to the address associated with the API endpoint that can be exposed by the API function. The prediction function can predict values ​​based on the training of a given ML model. The prediction function in the first ML pipeline can make predictions based on the training of a first ML model. (See reference...) Figure 6A The prediction function in the first ML pipeline 602 can be predicted using the input parameter "X_test" in statement-6 614. Processor 204 can determine the last API call associated with the prediction function in the first ML pipeline 602 based on the constructed AST. For example, refer to... Figure 6A Statement -6 614 can correspond to the last API call associated with the prediction function (i.e., "predict(X_test)") in the first ML pipeline 602.

[0098] At box 806, the identified last API call can be designated as the target row. Processor 204 can be configured to designate the identified last API call as the target row. The target row can be a row in the first ML pipeline that includes the prediction function. The target row can be a row based on which the prediction output of the first ML model can be retrieved. For example, refer to Figure 6AStatement 6 614 of the first ML pipeline 602 can include a prediction function, and therefore can be set as the target line.

[0099] At box 808, one or more statements associated with the first ML model can be identified based on a specified target line. Processor 204 can be configured to identify one or more statements associated with the first ML model based on a specified target line. As discussed, to instantiate the second ML model, only the code snippet containing the statements associated with the first ML model can be replaced. Therefore, it may be necessary to identify the statements associated with the first ML model in the first ML pipeline. For example, refer to... Figure 6A Based on a specified target line (which could be statement -6 614), statements -3 608, -4 610, and -5 612 of the first ML pipeline 602, along with statement -6 614, can be identified as one or more statements associated with the first ML model. Therefore, code snippet 618 can include one or more statements associated with the first ML model. The remaining statements, such as statements -1 604 and -2 606, may not be associated with the first ML model and can be combined as code snippet 616. Once code snippet 618, which includes one or more statements associated with the first ML model, is identified, the second ML model can be instantiated by replacing one or more statements of the first ML model with equivalent statements of the second ML model using the same input parameters. In this document, the statements in code snippet 616 can remain unchanged.

[0100] In implementation, one or more statements associated with the first ML model can be identified by applying backward program slicing from a specified target line until the model declaration associated with the first ML model is reached. In this document, backward program slicing can be used to obtain a slice of the program by adding related statements one by one, starting from the last statement. In other words, backward program slicing can be used to obtain a portion of the program by traversing backward from the last statement. Processor 204 can identify one or more statements associated with the first ML model by using backward program slicing, which may require traversing backward from the target line of the first ML model's code statements until the model declaration of the first ML model is reached. For example, see... Figure 6AStatement -6 614 can be the target line. The reverse slicer can traverse one or more statements of the first ML model starting from statement -6 614 until it reaches the model declaration of the first ML model in the first ML pipeline 602 to obtain code snippet 618. In this paper, the model declaration of the first ML model can correspond to statement -3 608. For example, the function call "sm.GLM" appearing in statement -3 608 can declare the first ML model.

[0101] In an implementation, processor 204 may also be configured to store the line number of each of one or more statements associated with the first ML model. In this document, one or more statements may correspond to at least one of, but not limited to: model definition, fitting function call, or prediction function call. The model definition statement may define the first ML model associated with the first ML pipeline. It will be understood that in order to create an ML model, it may be necessary to first define the ML model. For example, refer to... Figure 6A The first ML model associated with the first ML pipeline 602 can be defined at statement-3 608. The function "sm.GLM" appearing in statement-3 608 can define the first ML model. The fitting function call can accept input parameters (e.g., a training dataset including features and outputs or labels / scores associated with the training dataset) and can train the first ML model based on the provided input parameters. See reference... Figure 6A The fitting function call associated with the first ML pipeline 602 exists in statement -4 610. In this paper, one or more input parameters corresponding to the training dataset can be provided, such as "y_train" and "X_train". The prediction function call can make predictions based on the training of a given ML model. In this paper, the prediction function call can use the trained first ML model to provide predictions. (See reference...) Figure 6A The prediction function call associated with the first ML pipeline 602 is present in statement -6 614. In this document, the first ML model associated with the first ML pipeline 602 can predict input parameters (e.g., "X_test") associated with a test dataset. Processor 204 can store the line number of each statement in one or more statements associated with the first ML model. For example, processor 204 can store the line numbers of statement -3 608, which includes a model definition call; statement -4 610, which includes a fitting function call; and statement -5 612 and statement -6 614, which includes a prediction function call of the first ML pipeline 602. Storing the line numbers of each statement in one or more statements associated with the first ML model helps identify code snippet 618 and replace code snippet 618 with an equivalent code snippet of the second ML model to instantiate the second ML model. Control can be passed to the end.

[0102] Although flowchart 800 is shown as discrete operations such as 802, 804, 806, and 808, in some embodiments, such discrete operations may be divided into other operations, combined into fewer operations, or eliminated, depending on the specific implementation, without departing from the essence of the disclosed embodiments.

[0103] Figure 9 This is a flowchart illustrating an example method for obtaining model slices to identify code fragments of a first ML model according to at least one embodiment described in this disclosure. (Combined with...) Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5 , Figure 6A , Figure 6B , Figure 7 and Figure 8 To describe the elements Figure 9 . Reference Figure 9 The flowchart 900 is shown. The method shown in flowchart 900 can begin at block 902 and can be performed by any suitable system, device, or apparatus, such as by [unclear]. Figure 1 Example electronic device 102 or Figure 2 The processor 204 executes the process. Although shown in discrete boxes, depending on the specific implementation, the steps and operations associated with one or more boxes of flowchart 900 may be divided into additional boxes, combined into fewer boxes, or removed.

[0104] At box 902, a specified target line can be retrieved from the first ML pipeline. Processor 204 can be configured to retrieve the specified target line from the first ML pipeline. As discussed, the specified target line can be a line in the first ML pipeline that may include a prediction function. (See also...) Figure 6A Statement -6 614 of the first ML pipeline 602 can correspond to the specified target line because statement -6 614 can include a prediction function. The specified target line (i.e., statement -6 614) can be retrieved from the first ML pipeline 602.

[0105] At box 904, the retrieved target line can be added to a queue that includes a set of statements associated with the first ML pipeline. In this document, the set of statements associated with the first ML pipeline can be one or more statements associated with the first ML model, which can be identified based on the specified target line. Processor 204 can be configured to add the retrieved target line to a queue that may include the set of statements associated with the first ML pipeline. (See also...) Figure 6AIn code snippet 618, the set of statements associated with the first ML pipeline 602 may include statement-3608, statement-4610, and statement-5612. A queue may include the set of statements. Retrieved target lines (e.g., statement-6614) may be added to the queue.

[0106] At box 906, the first statement can be popped from the queue. In this document, the first statement can be a specified target line. Processor 204 can be configured to pop the first statement from the queue. For example, refer to... Figure 6A Statement -6 614 can be the first statement popped from the queue.

[0107] At box 908, the execution of the first set of operations 910 can be controlled to obtain model slices associated with the first ML model from the first ML pipeline. Processor 204 can be configured to control the execution of the first set of operations 910 to obtain model slices associated with the first ML model from the first ML pipeline. Model slices can be obtained based on retrieving one or more statements associated with the first ML model from the first ML pipeline. For example, refer to... Figure 6A The model slice can include statement-3 608, statement-4 610, statement-5 612 and statement-6 614.

[0108] The first set of operations 910 may include operations such as first operation 910A, second operation 910B, third operation 910C, fourth operation 910D, fifth operation 910E, and sixth operation 910F. The first set of operations 910 may be executed iteratively by the processor 204 based on checking whether the queue is empty. If the queue is determined to be empty, the execution of the first set of operations 910 may stop, and a model slice may be obtained at operation 912. The first set of operations 910 for obtaining a model slice is described herein.

[0109] At box 910A (i.e., the first operation), one or more variables and objects can be extracted from the first statement. Processor 204 can be configured to extract one or more variables and objects from the first statement. In this document, one or more variables and objects can be extracted from the retrieved target line. For example, refer to... Figure 6A Statement -6 614 can be the first statement that can be popped from the queue. One or more variables and objects extracted from statement -6 614 can include "X_test".

[0110] At box 910B (i.e., the second operation), a set of second statements that appear before the first statement in the first ML pipeline and include at least one of the extracted one or more variables and objects can be identified. Processor 204 can be configured to identify this set of second statements that appear before the first statement in the first ML pipeline and include at least one of the extracted one or more variables and objects. In this document, all statements preceding the first statement in the first ML pipeline that include at least one of the extracted one or more variables and objects can be identified as the second set of statements. For example, refer to... Figure 6A In the first ML pipeline 602, statement -5 612 may include one or more extracted variables and objects, such as "X_test". Therefore, statement -5 612 can be identified as a second set of statements.

[0111] At block 910C (i.e., the third operation), a check can be performed to determine whether a third statement in the identified set of second statements appears before the model definition associated with the first ML model. Processor 204 can be configured to determine whether a third statement in the identified set of second statements appears before the model definition associated with the first ML model. In this document, one of the statements in the identified set of second statements can be designated as the third statement, and processor 204 can determine whether the third statement appears before the model definition associated with the first ML model. If it is determined that the third statement in the identified set of second statements appears before the model definition associated with the first ML model, control can be passed to operation 910D; otherwise, processor 204 can select another statement as the third statement and repeat operation 910C. For example, see reference... Figure 6A Statement -5 612 can be identified as the second statement set. Since the second statement set only includes statement -5 612, statement -5 612 can be designated as the third statement. Furthermore, in the current case, processor 204 can determine that the first ML model can be defined at statement -3 608 in the first ML pipeline 602.

[0112] At box 910D (i.e., the fourth operation), a third statement can be added to a queue based on the determination that it appeared before the model definition. Processor 204 can be configured to add a third statement to the queue based on the determination that it appeared before the model definition. Since statements in the first ML pipeline corresponding only to the first ML model can be identified as one or more statements, if the third statement appears before the model definition, the third statement can be associated with the first ML model; otherwise, the third statement can not be associated with the first ML model. Therefore, if the third statement appears before the model definition, the third statement can be added to the queue. For example, refer to... Figure 6AStatement -5 612, designated as the third statement, appears before the model definition. When processor 204 executes the reverse tracing of each statement in code segment 618 starting from the target line (i.e., statement -6 614), processor 204 can reach statement -5 612. Therefore, statement -5 612 can be added to the queue. The third operation 910C and the fourth operation 910D can be repeated for each statement in the second set of statements.

[0113] At box 910E (i.e., the fifth operation), a first statement can be added to the model slice. Processor 204 can be configured to add the first statement to the model slice. In this document, the identified target statement can be added to the model slice. For example, refer to... Figure 6A You can add statement -6 614 to the model slice.

[0114] At box 910F (i.e., the sixth operation), based on the determination that the queue is not empty, a fourth statement can be popped from the queue as the first statement. Processor 204 can be configured to pop a fourth statement from the queue as the first statement based on the determination that the queue is not empty. Once the first statement can be added to the model slice, the queue can be checked to determine if it contains more statements and is not empty. If the queue is determined to be empty, the model slice can be obtained to identify the code snippet of the first ML model. Furthermore, the model slice can be displayed on display device 208A. However, if the queue is not empty, a fourth statement can be popped from the queue as the first statement.

[0115] At box 912, based on iterative execution of the first set of operations 910, model slices can be obtained to identify code fragments of the first ML model. Processor 204 can be configured to obtain model slices to identify code fragments of the first ML model based on iterative execution of the first set of operations 910 (e.g., Figure 6A (Code snippet 618). Iterative execution may depend on the number of statements in the queue. The first set of operations 910 may involve operations corresponding to boxes 910A through 910F. In the example, the queue may contain three statements. In this paper, the first set of operations may be executed three times to obtain a model slice. The model slice may be a code snippet indicating the first ML model. Control may be passed to the end.

[0116] Although flowchart 900 is shown as discrete operations, such as 902, 904, 906, 908, 910A to 910F and 922, in some embodiments, without departing from the essence of the disclosed embodiments, such discrete operations may be divided into other operations, combined into fewer operations, or eliminated, depending on the specific implementation.

[0117] It can be noted that the purpose of automatically generating ML pipelines could be to learn how to write ML pipelines for a given dataset through the meta-learning model 102A, which could be a completely offline process. In an online setting, when user 112 can provide a new dataset to the meta-learning model 102A, the meta-learning model 102A can automatically generate machine learning pipelines.

[0118] It can be noted that the quality of the meta-learning model 102A can depend on the quality of the ML corpus database, and more specifically, on the quality of the ML models used in the various ML pipelines. However, for many reasons, such as the unavailability of suitable ML models or a lack of awareness among some data scientists about good ML models, the ML pipelines that can be written by data scientists may not be optimal or the best ML models. Low-quality ML models in the ML corpus database can negatively affect the training of the meta-learning model 102A, where the meta-learning model 102A may fail to recognize any learnable patterns regarding what models will be used on what type of dataset. To mitigate the above problems, the meta-learning model 102A of this disclosure (or the disclosed electronic device 102) can be trained on an enhanced ML corpus database that may include a transformed ML pipeline, which may perform better than the original ML pipeline written by the data scientist by hand.

[0119] Figure 10 This is a flowchart illustrating an example method for training a meta-learning model according to at least one embodiment described in this disclosure. (Combined with...) Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5 , Figure 6A , Figure 6B , Figure 7 , Figure 8 and Figure 9 To describe the elements Figure 9 . Reference Figure 10 The flowchart 1000 is shown. The method shown in flowchart 1000 can begin at block 1002 and can be performed by any suitable system, device, or apparatus, for example by... Figure 1 Example electronic device 102 or Figure 2 The processor 204 executes the process. Although shown in discrete boxes, depending on the specific implementation, the steps and operations associated with one or more boxes of flowchart 1000 may be divided into additional boxes, combined into fewer boxes, or removed.

[0120] At box 1002, a set of meta-features can be extracted from a dataset associated with each of multiple ML projects stored in an enhanced ML corpus. Processor 204 can be configured to extract the set of meta-features from the dataset associated with each of the multiple ML projects stored in an enhanced ML corpus. It is understood that features can be independent variables provided to a given ML model that the given ML model may need to learn. Features can include columns of a tabular dataset associated with a given ML model. Meta-features can be used to estimate the performance of a given ML model. Meta-features can be predefined meta-features, which are typically used to learn the relationship between meta-features and ML components. A set of meta-features can be extracted from a dataset associated with each of the multiple ML projects stored in an enhanced ML corpus. The set of meta-features can be extracted based on injecting meta-feature extractor code into the ML pipeline (e.g., a meta-feature method call). For example, a dataset can be passed to a meta-feature method to extract the set of meta-features. For example, in Figure 11 Details of the meta-feature set have already been provided.

[0121] At box 1004, a set of ML pipeline components can be extracted from an ML pipeline set associated with each of multiple ML projects stored in an enhanced ML corpus database. Processor 204 can be configured to extract a set of meta-features from a dataset associated with each of the multiple ML projects stored in an enhanced ML corpus database. It is understood that ML components can include functions used in the ML pipeline of a given ML model. An ML component extractor can be used to extract ML components from an ML pipeline set associated with each of the multiple ML projects that may be stored in an enhanced ML corpus database. For example, in Figure 11 Details on extracting the ML pipeline component set have already been provided.

[0122] At box 1006, a meta-learning model 102A can be trained based on the extracted meta-feature set and the extracted ML pipeline component set. Processor 204 can be configured to train the meta-learning model 102A based on the extracted meta-feature set and the extracted ML pipeline component set. The meta-learning model 102A can use a meta-learning algorithm, which can be trained based on a previously trained learning algorithm. In this paper, the outputs of other learning algorithms for a given dataset, along with learning algorithms applicable to the dataset, can be provided to the meta-learning model 102A. The meta-learning model 102A can learn from the outputs of other learning algorithms. For example, the meta-learning model 102A can learn and predict based on the outputs of other learning algorithms as input. Therefore, the meta-learning model 102A can learn to make predictions based on predictions already made by other learning algorithms.

[0123] It can be noted that the meta-learning model 102A may not be a single ML model internally, but may include multiple ML models. For simplicity, the meta-learning model 102A can be considered a black box that receives a dataset as input and generates an abstract pipeline as output. In this paper, the abstract pipeline can be a sequence of labels that can be converted into code. It can be understood that labels can be names assigned to functions, modules, or sequences of statements to accomplish a specific task. It can be noted that each ML pipeline in the set of ML pipelines associated with each of the multiple ML projects stored in the augmented ML corpus can include several components, which can be in the form of code in the corresponding ML pipeline, allowing developers to avoid writing the functionality of each component. Users may not be able to infer anything from the corresponding ML pipeline unless the ML pipeline can be divided into components that can be assigned unique labels. Explanation augmentation can be used as a technique to provide natural language descriptions of the components used in the corresponding ML pipeline.

[0124] The meta-learning model 102A can be trained based on the extracted meta-feature set and the extracted ML pipeline component set. Since the meta-learning model 102A can be trained based on the meta-feature set extracted from the enhanced ML corpus database and the extracted ML pipeline component set, the quality of the meta-learning model 102A can directly depend on the quality and robustness of the enhanced ML corpus database. For example, in Figure 11 Details of the meta-learning model 102A have already been provided. Control can be passed to the end.

[0125] Although flowchart 1000 is shown as discrete operations, such as 1002, 1004, and 1006, in some embodiments, without departing from the essence of the disclosed embodiments, such discrete operations may be divided into other operations, combined into fewer operations, or eliminated, depending on the specific implementation.

[0126] Figure 11 An exemplary scenario for training a meta-learning model according to at least one embodiment described in this disclosure is shown. Combined with... Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5 , Figure 6A , Figure 6B , Figure 7 , Figure 8 , Figure 9 and Figure 10 To describe the elements Figure 11 . Reference Figure 11An exemplary scenario 1100 is illustrated. Exemplary scenario 1100 may include a database 108, a meta-feature set 1102, an ML pipeline component set 1104, a meta-learning block 1106, a meta-feature set of a topic dataset (represented by 1108), a meta-learning model 102A, and a component set of the topic ML pipeline (represented by 1110). Database 108 may include "n" ML projects, such as ML project-1 114, ML project-2 116, ..., and ML project-n 118. Each of the multiple ML projects may include a dataset and an ML pipeline set applicable to that dataset. For example, ML project-1 114 may include dataset 114A and ML pipeline set 114B, while ML project-2 116 may include dataset 116A and ML pipeline set 116B. Similarly, ML project-n 118 may include dataset 118A and ML pipeline set 118B.

[0127] Figure 11 The "n" ML items shown are for illustrative purposes only. Without departing from the scope of this disclosure, multiple ML items can include only two ML items or more than "n" ML items. For the sake of brevity, ... Figure 11 Only “n” ML items are shown in this document. However, in some implementations, more than “n” ML items may exist without limiting the scope of this disclosure.

[0128] For example, refer to Figure 11 Processor 204 can extract a meta-feature set 1102 from each of the datasets associated with multiple ML projects stored in an enhanced ML corpus database (which may be stored in database 108). The meta-feature set 1102 can be extracted based on injecting a meta-feature extractor (e.g., a meta-feature method call) into the corresponding ML pipeline code. The dataset can be passed to the meta-feature method to extract the meta-feature set 1102. In the example, the meta-feature set 1102 may include rows, columns, missing values, and flags indicating the presence of text.

[0129] Processor 204 can extract ML pipeline component set 1104 from ML pipeline sets associated with each of multiple ML projects stored in an enhanced ML corpus database (e.g., database 108). For example, processor 204 can extract ML pipeline component set 1104 from ML pipeline set 524 associated with the first ML project -1 114, ML pipeline set 526 associated with ML project -2 116, and ML pipeline set 528 corresponding to the nth ML project -n 118. In the example, ML pipeline component set 1104 may include "fillna", "TfidfVectorizer", and "logisticregression". In this paper, "fillna" can be a function used to fill missing values ​​in rows of a dataset. "TfidfVectorizer" can be a word frequency inverse document frequency function that can convert text into meaningful numerical values ​​based on a comparison of the number of times a word appears in a document to the number of documents containing that word. Logisticregression can be a function that predicts values ​​based on logistic regression techniques.

[0130] The meta-learning block 1106 can provide the extracted meta-feature set 1102 and the extracted ML pipeline component set 1104 for training the meta-learning model 102A. Since the ML pipeline component set 1104 can be extracted from an augmented ML corpus database (e.g., database 108) that may contain well-quality transformed ML pipelines, the meta-learning model 102A can be trained well.

[0131] Once the meta-learning model 102A can be trained, the processor 204 can provide the meta-learning model 102A with a set of meta-features (represented by 1108) of the topic dataset, such as rows, columns, missing values, and markers indicating the presence of text. The meta-learning model 102A can generate a set of topic ML pipeline components (represented by 1110) based on the set of meta-features (represented by 1108) of the topic dataset, such as “fillna”, “TfidfVectorizer”, and “logisticregression”. Since the meta-learning model 102A can be trained based on a high-quality transformed ML pipeline from an augmented ML corpus database (e.g., database 108), the generated set of topic ML pipeline components (represented by 1110) can also be of high quality. Therefore, the generated set of topic ML pipeline components (represented by 1110) can perform well for the topic dataset associated with the topic ML pipeline.

[0132] Exemplary experimental setups for this disclosure are presented in Table 1, as follows:

[0133]

[0134] Table 1: Exemplary experimental setup of this disclosure

[0135] It should be noted that the data provided in Table 1 may be considered as experimental data only and should not be construed as limiting the content of this disclosure.

[0136] Exemplary experimental data validating the performance improvement on the training data are presented in Table 2, as follows:

[0137] Percentage of ML pipelines Percentage of performance improvement 17% >5% 11% >3% 13% >1% 21% >0% 38% No improvement

[0138] Table 2: Exemplary experimental data on performance improvement of training data

[0139] As can be observed from Table 2, for 62% of the total 170 ML pipelines, accuracy increased based on the proposed transformation framework. For 17% of the ML pipelines, the performance improvement was significant, exceeding 5%. Furthermore, for 13% of the ML pipelines, the performance improvement was greater than 1%.

[0140] It should be noted that the data provided in Table 2 may be considered as experimental data only and should not be construed as limiting the content of this disclosure.

[0141] Exemplary experimental data on the impact on the test data are shown in Table 3, as follows:

[0142] measure value Average F1 / R2 0.023 >0.5 3 >0.01 5 >0 7

[0143] Table 3: Exemplary experimental data on the impact on test data

[0144] It should be noted that the data provided in Table 3 may be considered as experimental data only and should not be construed as limiting the content of this disclosure.

[0145] Various embodiments of this disclosure may provide one or more non-transitory computer-readable storage media configured to store instructions that, upon execution, cause a system (e.g., example electronic device 102) to perform operations. Operations may include receiving ML items from a plurality of ML items stored in a machine learning (ML) corpus database. Each of the plurality of ML items may include a dataset and a set of ML pipelines applicable to that dataset. Operations may also include, based on a predefined set of ML pipelines, transforming a first ML pipeline in a first set of ML pipelines associated with the received ML item to determine a second set of ML pipelines. Transformation of the first ML pipeline may correspond to replacing a first ML model associated with the first ML pipeline with a second ML model associated with a predefined ML pipeline in the predefined set of ML pipelines. Operations may also include, based on a performance score associated with each ML pipeline in the determined second set of ML pipelines, selecting one or more ML pipelines from the determined second set of ML pipelines. The operation may also include enhancing the ML corpus database to include one or more selected ML pipelines and a first set of ML pipelines associated with the received ML projects.

[0146] As used herein, the terms "module" or "component" can refer to a specific hardware implementation configured to perform actions of software objects or software routines and / or modules or components that can be stored on and / or executed by general-purpose hardware (e.g., computer-readable media, processing devices, etc.) of a computing system. In some embodiments, the different components, modules, engines, and services described herein can be implemented as objects or processes (e.g., as separate threads) that execute on a computing system. While some systems and methods described herein are generally described as being implemented in software (stored on and / or executed by general-purpose hardware), specific hardware implementations or combinations of software with specific hardware implementations are possible and contemplated. In this specification, a "computing entity" can be any computing system as previously defined in this disclosure or any module or combination of modules running on a computing system.

[0147] The terms used in this disclosure and particularly in the appended claims (e.g., the body of the appended claims) are generally intended to be “open-ended” terms (e.g., the term “comprising” should be interpreted as “including but not limited to”, the term “having” should be interpreted as “having at least”, the term “including” should be interpreted as “including but not limited to”, etc.).

[0148] Furthermore, if the intention is to introduce a specific number of claim narratives, such intention will be explicitly stated in the claims, and without such a statement, such intention does not exist. For example, to aid understanding, the appended claims may include the use of the introductory phrases “at least one” and “one or more” to introduce claim narratives. However, even when the same claim includes the introductory phrases “one or more” or “at least one” and indefinite articles such as “a” or “a” (e.g., “a” and / or “a” should be interpreted as meaning “at least one” or “one or more”), the use of such phrases should not be interpreted as implying that the claim narratives introduced by the indefinite article “a” or “a” limit any particular claim containing such introduced claim narratives to an implementation containing only one such narrative; the same applies to the use of definite articles used to introduce claim narratives.

[0149] Furthermore, even when a specific number of introduced claim narratives are explicitly stated, those skilled in the art will recognize that such a statement should be interpreted as at least referring to the stated number (e.g., an unmodified statement "two narratives" without other modifiers means at least two narratives, or two or more narratives). Additionally, in cases where idiomatic expressions such as "at least one of A, B, and C" or "one or more of A, B, and C" are used, such constructions are generally intended to include only A, only B, only C, A and B together, A and C together, B and C together, or A, B, and C together, etc.

[0150] Furthermore, any separator or phrase presenting two or more alternative terms, whether in the specification, claims, or drawings, should be understood to include the possibility of including one, any, or all of the terms. For example, the phrase "A or B" should be understood to include the possibility of including "A" or "B" or "A and B".

[0151] All examples and conditional language described in this disclosure are intended for educational purposes to help the reader understand this disclosure and the concepts contributed by the inventors to promote the art, and should be construed as not being limited to such specific examples and conditions. Although embodiments of this disclosure have been described in detail, various changes, substitutions, and modifications may be made to the embodiments without departing from the spirit and scope of this disclosure.

Claims

1. A method executed by a processor, comprising: Receive ML projects from multiple ML projects stored in a machine learning ML corpus database. Each of the plurality of ML projects includes a dataset and a set of ML pipelines applicable to the dataset; Based on a predefined set of ML pipelines, the first ML pipelines in the first set of ML pipelines associated with the received ML items are transformed to determine the second set of ML pipelines. Wherein, the transformation of the first ML pipeline corresponds to replacing the first ML model associated with the first ML pipeline with a second ML model associated with a predefined ML pipeline in the set of predefined ML pipelines; Based on the performance score associated with each of the determined second ML pipelines in the determined second ML pipeline set, select one or more ML pipelines from the determined second ML pipeline set; Enhance the ML corpus database to include one or more selected ML pipelines and the first ML pipeline set associated with the received ML projects; Identify the code snippet of the first ML model associated with the first ML pipeline; Determine one or more input parameters associated with the identified code segment; Select the second ML model from the set of predefined models associated with the predefined ML pipeline set; The selected second ML model is instantiated by replacing the first ML model with the selected second ML model in the first ML pipeline; Construct an abstract syntax tree (AST) associated with the first ML pipeline; Based on the constructed AST, the last application programming interface (API) call associated with the prediction function in the first ML pipeline is determined; Specify the last identified API call as the target line; and Identify one or more statements associated with the first ML model based on the specified target line.

2. The method according to claim 1, further comprising: Extract a set of meta-features from the dataset associated with each of the plurality of ML projects stored in the enhanced ML corpus database; Extract a set of ML pipeline components from the ML pipeline set associated with each of the plurality of ML projects stored in the enhanced ML corpus database; and The meta-learning model is trained based on the extracted meta-feature set and the extracted ML pipeline component set.

3. The method according to claim 1, wherein, The performance score associated with each of the second ML pipelines in the determined set of second ML pipelines corresponds to at least one of the following: the F1 score or the R2 score associated with the corresponding ML pipeline.

4. The method according to claim 1, wherein, One or more ML pipelines selected from the determined second set of ML pipelines include one of the following: The ML pipelines from the determined second set of ML pipelines that are associated with the maximum performance score. The first set of ML pipelines corresponding to performance scores above the threshold from the determined second set of ML pipelines, or The second set of ML pipelines, which are derived from the determined second set of ML pipelines, corresponds to a predefined number of ML pipelines based on the performance score.

5. The method according to claim 1, wherein, The determined input parameters associated with the identified code fragment include at least one of the following: the training dataset, the test dataset, and the hyperparameter set associated with the first ML model.

6. The method according to claim 1, further comprising: Select a predefined template associated with the second ML model, wherein the selected predefined template is annotated with one or more input parameters of a code snippet of the first ML model that is identified; The code snippet that constructs the second ML model is based on parameterizing one or more function calls in a selected predefined template using one or more annotated input parameters; as well as The second ML model is instantiated by replacing the code fragment of the first ML model with the code fragment of the constructed second ML model.

7. The method according to claim 1, wherein, The one or more statements associated with the first ML model are identified by applying reverse procedural slicing from the specified target line until the model statement associated with the first ML model is reached.

8. The method according to claim 1, further comprising: Store the line number of each of the one or more statements associated with the first ML model. The one or more statements mentioned correspond to at least one of the following: model definition, fitting function call, or prediction function call.

9. The method according to claim 1, further comprising: Retrieve the specified target row from the first ML pipeline; The retrieved target line is added to a queue that includes the set of statements associated with the first ML pipeline; Pop the first statement from the queue; Controlling the execution of a first set of operations to obtain model slices associated with the first ML model from the first ML pipeline, wherein the first set of operations includes: Extract one or more variables and objects from the first statement. Identify a set of second statements that appear in the first ML pipeline before the first statement and include at least one of the extracted variables and objects. Determine whether the third statement in the identified second set of statements appears before the model definition associated with the first ML model. Based on the determination that the third statement appeared before the model definition, the third statement is added to the queue. Add the first statement to the model slice, and Based on the determination that the queue is not empty, a fourth statement is popped from the queue and used as the first statement; and Based on the iterative execution of the first set of operations, the model slices are obtained to identify the code fragments of the first ML model.

10. One or more non-transitory computer-readable storage media configured to store instructions that, in response to being executed, cause an electronic device to perform operations, said operations including: Receive ML projects from multiple ML projects stored in a machine learning ML corpus database. Each of the plurality of ML projects includes a dataset and a set of ML pipelines applicable to the dataset; Based on a predefined set of ML pipelines, the first ML pipelines in the first set of ML pipelines associated with the received ML items are transformed to determine the second set of ML pipelines. Wherein, the transformation of the first ML pipeline corresponds to replacing the first ML model associated with the first ML pipeline with a second ML model associated with a predefined ML pipeline in the set of predefined ML pipelines; Based on the performance score associated with each of the determined second ML pipelines in the determined second ML pipeline set, select one or more ML pipelines from the determined second ML pipeline set; Enhance the ML corpus database to include one or more selected ML pipelines and the first ML pipeline set associated with the received ML projects; Identify the code snippet of the first ML model associated with the first ML pipeline; Determine one or more input parameters associated with the identified code segment; Select the second ML model from the set of predefined models associated with the predefined ML pipeline set; The selected second ML model is instantiated by replacing the first ML model with the selected second ML model in the first ML pipeline; Construct an abstract syntax tree (AST) associated with the first ML pipeline; Based on the constructed AST, the last application programming interface (API) call associated with the prediction function in the first ML pipeline is determined; Specify the last identified API call as the target line; and Identify one or more statements associated with the first ML model based on the specified target line.

11. One or more non-transitory computer-readable storage media according to claim 10, wherein, The operation also includes: Extract a set of meta-features from the dataset associated with each of the plurality of ML projects stored in the enhanced ML corpus database; Extract a set of ML pipeline components from the ML pipeline set associated with each of the plurality of ML projects stored in the enhanced ML corpus database; and The meta-learning model is trained based on the extracted meta-feature set and the extracted ML pipeline component set.

12. The non-transitory computer-readable storage medium according to claim 10, wherein, The performance score associated with each of the second ML pipelines in the determined set of second ML pipelines corresponds to at least one of the following: the F1 score or the R2 score associated with the corresponding ML pipeline.

13. One or more non-transitory computer-readable storage media according to claim 10, wherein, One or more ML pipelines selected from the determined second set of ML pipelines include one of the following: The ML pipelines from the determined second set of ML pipelines that are associated with the maximum performance score. The first set of ML pipelines corresponding to performance scores above the threshold from the determined second set of ML pipelines, or The second set of ML pipelines, which are derived from the determined second set of ML pipelines, corresponds to a predefined number of ML pipelines based on the performance score.

14. One or more non-transitory computer-readable storage media according to claim 10, wherein, The determined input parameters associated with the identified code fragment include at least one of the following: the training dataset, the test dataset, and the hyperparameter set associated with the first ML model.

15. One or more non-transitory computer-readable storage media according to claim 10, wherein, The operation also includes: Select a predefined template associated with the second ML model, wherein the selected predefined template is annotated with one or more input parameters of a code snippet of the first ML model that is identified; The code snippet constructs the second ML model based on parameterizing one or more function calls in a selected predefined template using one or more annotated input parameters; and The second ML model is instantiated by replacing the code fragment of the first ML model with the code fragment of the constructed second ML model.

16. A method executed by a processor, comprising: Receive ML projects from multiple ML projects stored in a machine learning ML corpus database. Each of the plurality of ML projects includes a dataset and a set of ML pipelines applicable to the dataset; Based on a predefined set of ML pipelines, the first ML pipelines in the first set of ML pipelines associated with the received ML items are transformed to determine the second set of ML pipelines. Wherein, the transformation of the first ML pipeline corresponds to replacing the first ML model associated with the first ML pipeline with a second ML model associated with a predefined ML pipeline in the set of predefined ML pipelines; Based on the performance score associated with each of the determined second ML pipelines in the determined second ML pipeline set, select one or more ML pipelines from the determined second ML pipeline set; Enhance the ML corpus database to include one or more selected ML pipelines and the first ML pipeline set associated with the received ML projects; Select a predefined template associated with the second ML model, the selected predefined template being annotated with one or more input parameters of a code snippet of the first ML model that is identified; The code snippet constructs the second ML model based on parameterizing one or more function calls in a selected predefined template using one or more annotated input parameters; and The second ML model is instantiated by replacing the code fragment of the first ML model with the code fragment of the constructed second ML model.

17. One or more non-transitory computer-readable storage media configured to store instructions that, in response to being executed, cause an electronic device to perform operations, said operations including: Receive ML projects from multiple ML projects stored in a machine learning ML corpus database. Each of the plurality of ML projects includes a dataset and a set of ML pipelines applicable to the dataset; Based on a predefined set of ML pipelines, the first ML pipelines in the first set of ML pipelines associated with the received ML items are transformed to determine the second set of ML pipelines. Wherein, the transformation of the first ML pipeline corresponds to replacing the first ML model associated with the first ML pipeline with a second ML model associated with a predefined ML pipeline in the set of predefined ML pipelines; Based on the performance score associated with each of the determined second ML pipelines in the determined second ML pipeline set, select one or more ML pipelines from the determined second ML pipeline set; Enhance the ML corpus database to include one or more selected ML pipelines and the first ML pipeline set associated with the received ML projects; Identify the code snippet of the first ML model associated with the first ML pipeline; Determine one or more input parameters associated with the identified code segment; Select the second ML model from the set of predefined models associated with the predefined ML pipeline set; The selected second ML model is instantiated by replacing the first ML model with the selected second ML model in the first ML pipeline; Select a predefined template associated with the second ML model, wherein the selected predefined template is annotated with one or more input parameters of a code snippet of the first ML model that is identified; The code snippet constructs the second ML model based on parameterizing one or more function calls in a selected predefined template using one or more annotated input parameters; and The second ML model is instantiated by replacing the code fragment of the first ML model with the code fragment of the constructed second ML model.

Citation Information

Patent Citations

  • Systems and methods for operating a data center based on a generated machine learning pipeline

    US20200272909A1

  • Method and apparatus for generating information

    US20200401950A1