Automated Generation of a Pipeline for a New Machine Learning Project from the Pipeline of an Existing Machine Learning Project Stored in a Corpus
The system automatically generates a machine learning pipeline for new projects by merging and adapting existing pipelines, overcoming the shortage of skilled data scientists and enhancing the implementation of new machine learning projects.
Patent Information
- Application Number
- JP2021139558
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-09-02
- Filing Date
- 2021-08-30
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2041-08-30
AI Technical Summary
The shortage of skilled data scientists and the limitations of current Automated Machine Learning (AutoML) solutions make it difficult for non-experts to fully implement new machine learning projects, as existing solutions only provide simple and partial solutions.
A system and method for automatically generating a pipeline for a new machine learning project by storing an existing project in a corpus, generating a query based on a new dataset and task, searching for relevant existing projects, merging their pipelines, and adapting the functional blocks to create an executable pipeline for the new project.
Enables non-expert data scientists to efficiently generate high-quality end-to-end machine learning pipelines for new projects without manual intervention, addressing the shortage of skilled professionals and enhancing the implementation of new machine learning projects.
Smart Images

Figure 0007697320000014 
Figure 0007697320000015 
Figure 0007697320000016
Abstract
Description
Technical Field
[0001] The embodiments discussed in this disclosure relate to the automatic generation of a pipeline for a new machine learning project from the pipeline of an existing machine learning project stored in a corpus.
Background Art
[0002] Machine learning (ML) typically utilizes an ML model trained with training data to make predictions that automatically become more accurate with current training. ML can be used in a wide range of applications including, but not limited to, traffic prediction, web search, online fraud detection, medical diagnosis, speech recognition, email filtering, image recognition, virtual personal assistance, and automatic translation.
[0003] As ML becomes increasingly common, there is often a shortage of ML experts (e.g., skilled data scientists) available to implement new ML projects. For example, according to some estimates, the majority of data scientists currently working on the development of new ML projects are not experts (e.g., relatively inexperienced or beginners), and only about 2 out of 5 with a master's or doctoral degree are qualified to develop increasingly complex ML projects.
[0004] Automated ML (AutoML) is a process that automates the process of applying ML to real-world problems. AutoML can enable non-experts to utilize ML models and techniques without first requiring them to become ML experts. AutoML has been proposed as a solution to the increasingly growing problem of implementing new ML projects despite the shortage of ML experts. However, current AutoML solutions only provide simple and partial solutions that are insufficient to enable non-experts to fully implement new ML projects.
[0005] The subject matter claimed in this disclosure is not limited to embodiments that solve any disadvantages or that operate only in the environments described above. Rather, this background is provided only to explain a technical field in which some embodiments described in this disclosure may be practiced.
Summary of the Invention
[0006] According to aspects of embodiments, an operation may include storing an existing machine learning (ML) project in a corpus. Here, the existing ML project includes an ML pipeline having functional blocks. The operation may further include generating a query for a new ML project based on a new dataset for the new ML project and a new ML task for the new ML project. Further, the operation may include searching a set of existing ML projects based on the query through the existing ML projects stored in the corpus. Further, the operation may include merging the ML pipelines of the set of existing ML projects to generate a new ML pipeline for the new ML project. Here, the new ML pipeline has functional blocks incorporated from the ML pipelines of the set of existing ML projects. Further, the operation may include adapting the functional blocks of the new ML pipeline for the new ML project so that the new ML pipeline is executable to execute the new ML task with the new dataset.
[0007] The objectives and advantages of the embodiments are at least realized and achieved by the elements, features, and combinations particularly pointed out in the claims.
[0008] The foregoing general description and the following detailed description are both given by way of example and for the purpose of explanation and are not limitations of the claimed invention.
Brief Description of the Drawings
[0009] Exemplary embodiments are described and explained with further particularity and detail through the use of the following attached drawings.
[0010]
Figure 1
[0011]
Figure 2
[0012]
Figure 3
[0013]
Figure 4
[0014]
Figure 5
[0015]
Figure 6
[0016]
Figure 7
[0017]
Figure 8A
[0018]
Figure 8B
[0019]
Figure 9
[0020]
Figure 10
[0021]
Figure 11
[0022]
Figure 12
[0023]
Figure 13A
[0024]
Figure 13B
[0025]
Figure 14
[0026]
Figure 15
[0027]
Figure 16
[0028]
Figure 17
[0029]
Figure 18
[0030]
Figure 19
[0031]
Figure 20
DETAILED DESCRIPTION OF THE INVENTION
[0032] Some embodiments described in this disclosure relate to methods and systems for automatically searching existing ML projects and adapting them to new ML projects.
[0033] As ML becomes increasingly common, there is often a shortage of ML experts (e.g., skilled data scientists) available to implement new ML projects. Various AutoML solutions (e.g., Auto-Sklearn, AutoPandas, etc.) have been proposed to address the growing challenge of implementing new ML projects despite the shortage of ML experts, but current AutoML solutions provide only simplified and partial solutions to enable non-experts to fully implement new ML projects. Additionally, open source software (OSS) databases of existing ML projects (e.g., Kaggle, GitHub, etc.) have also been proposed as another solution to the challenge of implementing new ML projects by non-experts, but it can be difficult or impossible for non-experts to find existing ML projects in these databases that may be useful. Further, even if a non-expert succeeds in finding an existing ML project in these databases that may be useful, it can be difficult or impossible for the non-expert to modify the existing ML project to be useful for the new requirements of a new ML project.
[0034] In the present disclosure, the term "ML project" may represent a project that includes a dataset, an ML task defined on the dataset, and an ML pipeline (e.g., a script or program code) configured to perform an operational sequence for training an ML model on the dataset and using the ML model for new predictions. In the present disclosure, the term "computational notebook" may represent a computational structure used to develop and / or represent an ML pipeline, particularly during the development stage (e.g., a Jupyter notebook). The embodiments disclosed herein are illustrated by an ML pipeline in the Python programming language and a computational notebook structured by a Jupyter notebook. It should be understood, however, that other embodiments may include ML pipelines described in different languages and computational notebooks structured on other platforms.
[0035] According to one or more embodiments of the present disclosure, operations may be performed to automatically search for an existing ML project and adapt it for a new ML project. For example, in some embodiments, a computer system may construct a workflow to search for an existing ML project that provides a good starting point for a data scientist to first build a new ML project in a "search-and-adapt" style, and then appropriately adapt the existing ML project to construct an ML pipeline for a new database and new ML tasks for the new ML project, thereby organizationally supporting the natural workflow of the data scientist.
[0036] For example, in some embodiments, a computer system may automatically mine raw ML projects from an OSS database of existing ML projects and may automatically curate new ML projects before storing them in the corpus of existing ML projects. In some embodiments, this mining and curation of existing ML projects from large-scale repositories may result in a corpus of diverse high-quality existing ML projects that can be used in search and adaptation workflows. This curation may also include cleaning the ML pipelines of existing ML projects (e.g., using dynamic program slicing), calculating feature sets to capture the quality and diversity of each ML project, and selecting an optimal number of existing ML projects for these purposes.
[0037] Also, in some embodiments, this curation may involve operations performed to automatically identify and index functional blocks within the ML pipelines of existing ML projects. Unlike traditional software programs, the ML pipelines of ML projects can typically be thought of as a sequence of functional blocks that follow a clearly defined workflow based on dataset properties. Thus, some embodiments include techniques for automatically extracting and labeling functional blocks within the ML pipeline so that they can be correctly indexed within the corpus, such that they can be efficiently searched for use in synthesizing new ML pipelines for new ML tasks. More specifically, this technique may abstract the ML pipeline at an appropriate level and utilize graph-based sequence mining algorithms to extract both custom and idiomatic functional blocks. Finally, each functional block may be semantically labeled.
[0038] In some embodiments, upon receiving a new database and new tasks for a new ML project, such as from a non-expert data scientist, the computer system may first automatically use a hierarchical approach to synthesize a functional block-level pipeline skeleton for the new ML project using an ML model, and then explicitly search through a curated and labeled corpus to identify existing ML projects relevant to instantiating this pipeline skeleton. Next, the computer system may automatically select functional blocks from the ML pipelines of a set of relevant existing ML projects and concretize the pipeline skeleton into a new ML pipeline for the new ML project. Finally, the computer system may adapt the functional blocks of the new ML pipeline so that the new ML pipeline becomes executable to perform the new ML task with the new dataset.
[0039] Accordingly, in some embodiments, a non-expert data scientist may simply formulate a new dataset and new ML tasks for a new ML project. Next, the computer system may perform a tool-assisted interactive search and adaptation workflow to automatically generate a new ML pipeline for the new ML project. This can be immediately executed to perform the new ML task with the new dataset and involves no changes by the non-expert data scientist. Accordingly, some embodiments can empower novice data scientists to efficiently generate new high-quality end-to-end ML pipelines for new ML projects.
[0040] According to one or more embodiments of the present disclosure, the technical field of ML project development can be improved by configuring a computer system to automatically search for existing ML projects and adapt them to new ML projects, as opposed to imposing the task of manually discovering existing ML projects that may be useful to data scientists (who are often non-experts) and modifying the existing ML projects that may be useful for the new requirements of a new ML project. Such a configuration enables the computing system to better search for relevant existing ML projects and adapt them to new ML projects by identifying and extracting functional blocks from existing ML pipelines and automatically adapting them for use in a new ML pipeline.
[0041] Embodiments of the present disclosure are described with reference to the accompanying drawings.
[0042] FIG. 1 is a diagram illustrating an exemplary environment 100 related to automatically searching for existing ML projects and adapting them to new ML projects, configured in accordance with at least one embodiment described in the present disclosure. The environment 100 may include OSS ML project databases 102a - 102n, a curation module 114 configured to curate existing ML projects into an ML project corpus 104, a new dataset 106, and a search module configured to search the ML project corpus 104 for relevant existing ML projects for a new ML project based on the new ML task 108 of the new ML project (which, for example, may be provided by a data scientist 118), and an adaptation module 120 configured to synthesize functional blocks from the ML pipeline 111 of the relevant existing ML project 110 and adapt them to the new ML pipeline 112 of the new ML project.
[0043] The OSS ML project databases 102a - 102n may be large repositories of existing ML projects, where each ML project includes at least a dataset, an ML task defined on the dataset, an ML pipeline (e.g., a script or program code) configured to perform an operation sequence for training an ML model for the ML task and using the ML model for new predictions. Some examples of large repositories of existing ML projects include, but are not limited to, Kaggle and GitHub. In some embodiments, each ML project in the OSS ML project databases 102a - 102n may include a computing notebook. A computing notebook may be a computing structure used to develop and / or represent an ML pipeline, particularly during the development stage. An example of a computing notebook is a Jupyter notebook.
[0044] Each of the curation module 114, the search module 116, and the adaptation module 120 may include code and routines configured to enable a computing device to execute one or more operations. Additionally or alternatively, each of these modules may be implemented using hardware including a processor, a microprocessor (e.g., one or more operations are executed or controlled), an FPGA (field - programmable gate array), or an ASIC (application - specific integrated circuit). In some other examples, each module may be implemented using a combination of hardware and software. In the present disclosure, the operations described as being performed by each of these modules may include operations that can direct each of these modules to execute in a corresponding system.
[0045] The curation module 114 may be configured to execute an operation series on existing ML projects stored in the OSS ML project databases 102a - 102n, either before or after storing the existing ML projects in the ML project corpus 104. For example, the curation module 114 may be configured to automatically mine raw ML projects from the OSS ML project databases 102a - 102n to automatically curate them either before or after storing the raw ML projects in the ML project corpus 104. The ML project corpus 104 may be a repository of existing ML projects curated from the OSS ML project databases 102a - 102n. In some embodiments, the ML project corpus 104 may be a large corpus of cleaned, high-quality, indexed existing ML projects that can be utilized in an automated "search and adapt" style workflow. In this style of workflow, the search may include identifying existing ML projects that should be used as "seeds" for constructing new ML projects related to new ML tasks and new datasets and including new ML pipelines. Further, in this style of workflow, adapting may include using an interactive synthesis approach to adapt relevant existing ML projects and generate new ML pipelines for new ML projects.
[0046] In some embodiments, the curation module 114 may be configured to mine and curate existing ML projects such that only a diverse set of high-quality existing ML projects are stored in the ML project corpus 104. Also, in some embodiments, the curation module 114 may be configured to clean the ML pipelines of existing ML projects (e.g., using dynamic program slicing). Further, in some embodiments, the curation module 114 may be configured to calculate feature sets to capture the quality and diversity of each ML project and select an optimal number of existing ML projects to store in the ML project corpus 104 from the OSS ML project databases 102a-102n. Further, in some embodiments, the curation module 114 may be configured to automatically identify and index functional blocks within the ML pipelines of existing ML projects. Unlike traditional software programs, the ML pipelines of ML projects can typically be considered as sequences of functional blocks following a clearly defined workflow based on dataset properties. Thus, the curation module 114 may be configured to automatically extract and label functional blocks within the ML pipeline (using semantic labels such as "read data") and correctly index them within the ML project corpus 104 such that they can be efficiently searched for new datasets 106 and new ML tasks 108 and new ML pipelines 112 can be efficiently synthesized. More specifically, the curation module 114 may be configured to abstract the ML pipeline at an appropriate level and utilize graph-based sequence mining algorithms to extract both custom and idiomatic functional blocks. Finally, the curation module 114 may be configured to generate semantic labels and assign them to each functional block.
[0047] The search module 116 may be configured to execute a series of operations related to searching through existing ML projects stored in the ML project corpus 104. For example, the search module 116 may be configured to receive a new dataset 106 and a new ML task 108 for a new ML project from, for example, a data scientist 118. Upon receipt, the search module 116 may first be configured to synthesize a functional block-level pipeline skeleton for the new ML project using an ML model, automatically using a hierarchical approach, and then be configured to explicitly search through the ML project corpus 104 to identify relevant existing ML projects 110. From the relevant existing ML projects 110, this pipeline skeleton is instantiated.
[0048] The adaptation module 120 may be configured to execute a series of operations related to synthesizing functional blocks from the ML pipeline 111 of a relevant existing ML project 110 and adapting them to a new ML project 112. For example, the adaptation module 120 may be configured to automatically select functional blocks from the ML pipeline 111 and concretize them into a new ML pipeline 112 for a new ML project (e.g., including a new dataset 106, a new ML task 108, and a new ML pipeline 112). Further, the adaptation module 120 may be configured to adapt the functional blocks of the new ML pipeline 112 so that the new ML pipeline 112 is executable to execute the new ML task 108 with the new dataset 106.
[0049] Accordingly, in some embodiments, a data scientist 118, who may be a non-expert, may only need to formulate a new dataset 106 and a new ML task 108 for a new ML project, and the curation module 114, the search module 116, and the adaptation module 120 may function together (e.g., by performing one or more of the methods disclosed herein) to ultimately generate a new ML pipeline 112 for a new ML project that is immediately executable to perform the new ML task 108 with the new dataset 106, without any changes by the data scientist 118.
[0050] Changes, additions, or omissions may be made to FIG. 1 without departing from the scope of the present disclosure. For example, environment 100 may include more or fewer elements than shown and described in the present disclosure.
[0051] FIG. 2 shows an exemplary environment 200 associated with automatically curating existing ML projects into a corpus, configured in accordance with at least one embodiment described in the present disclosure. Similar to environment 100 of FIG. 1, environment 200 may include OSS ML project databases 102a-102n, a curation module 114, and an ML project corpus 104. Further, as disclosed for environment 200, after data scientists 202a-202n store existing ML projects 204 in the OSS ML project databases 102a-102n, the curation module 114 may be configured to crawl the OSS ML project databases 102a-102n to generate a set of the existing ML projects 204. This set of the existing ML projects 204 may then be further analyzed by the curation module 114.
[0052] While further analyzing the existing ML projects 204, the curation module 114 may be configured to filter 206 the existing ML projects 204 for quality and relevance, and to clean 208 the existing ML projects 204 to identify and / or remove irrelevant content. This filtering 206 and cleaning 208 may be configured to overcome various issues in the existing ML projects 204. For example, some of the computational notebooks in the existing ML projects 204 may not have a high enough quality to build a high-quality ML project corpus. Thus, this filtering 206 may automatically identify higher-quality computational notebooks (e.g., using standard APIs instead of custom code, using appropriate classifiers, and having high precision) for inclusion in the ML project corpus 104. Further, a high-quality ML project corpus should include existing ML projects with diverse computational notebooks. Thus, this filtering 206 may automatically identify computational notebooks with greater diversity for inclusion in the ML project corpus 104. Also, the ML pipelines in the computational notebooks of the existing ML projects 204 may often be noisy, such as Jupyter notebooks that may have a significant amount of irrelevant code (e.g., debugging code, visualization code, and / or experimental code), or deprecated APIs that render good-quality code non-executable. Thus, this cleaning 208 may automatically clean the ML pipelines in the computational notebooks to resolve the noise (e.g., irrelevant code and / or deprecated APIs). Irrelevant code may be resolved by marking the portions of the code that do not programmatically contribute to the overall ML pipeline and thus may add noise to the overall technology. Deprecated APIs may be resolved by automatically replacing the deprecated APIs with new APIs using API adaptation techniques.
[0053] Furthermore, the curation module 114 may be configured to abstract existing ML projects that have been cleaned and filtered 210 and generate project artifacts 212 of the existing ML projects 204 for indexing purposes. This abstraction 210 and the generated project artifacts 212 may be configured to overcome various issues in the existing ML projects 204. For example, it may be difficult to represent ML pipelines in the ML project corpus 104 to enable better search. Thus, this abstraction 210 may automatically identify functional blocks in the ML pipeline code and further identify mappings between specific meta-features in the dataset and the functional blocks. Further, it may be difficult to determine the appropriate level of abstraction to find functional blocks such that they can be identified within any line of code. Thus, this abstraction 210 may automatically identify functional blocks based on the insight that ML pipelines often heavily depend on APIs, similar functional blocks often contain similar API sets, and the structure of computational notebooks (e.g., Jupyter notebooks) can also provide important information regarding functional blocks. Also, it may be difficult to extract the semantic purpose of each functional block and identify alternative implementations of specific functions using semantic labels. Thus, the abstraction 210 utilizes information derived from the markdown cells of computational notebooks (e.g., Jupyter notebooks), and source code comments and library API documentation are provided to automatically generate semantic labels and later use the semantic labels to identify functionally equivalent functional blocks even when they use different syntax. In this way, alternative implementations of functional blocks can be identified and grouped together (e.g., this grouping may be referred to as "clustering").
[0054] Finally, before and / or after filtering 260, cleaning 208, and abstraction 210, the curation module 114 may be configured to store the existing ML projects 204 curated in the ML project corpus 104 for generating project artifacts 212 (e.g., by performing one or more of the methods disclosed herein). Thus, in some embodiments, the environment 200 may be utilized to automatically curate existing ML projects into the ML project corpus 104 to enable the existing ML projects to be later retrieved and adapted for new ML projects.
[0055] Changes, additions, or omissions may be made to FIG. 2 without departing from the scope of the present disclosure. For example, the environment 200 may include more or fewer elements than shown and described in the present disclosure.
[0056] FIG. 3 is a diagram depicting an exemplary environment 300 associated with automatically generating a new ML project pipeline from an existing ML project pipeline stored in a corpus. Similar to the environment 100 of FIG. 1, the environment 300 may include an ML project corpus 104, a new dataset 106, a new ML task 108, an associated existing ML project 110, a new ML pipeline 112, a search module 116, and an adaptation module 120. Further, as disclosed in the environment 300, after an existing ML project is stored in the ML project corpus 104, the search module 116 may be configured to receive, for example, from a data scientist 118, a new ML task 108 for the new dataset 106 and the new ML project 310. Next, the search module 116 may be configured to synthesize a block-level pipeline skeleton 304 for the new ML task project 310 using a pipeline skeleton ML model 302 (which may be pre-trained using training data derived from the ML project corpus 104).
[0057] Next, the search module 116 may be configured to generate a query 306 based on the pipeline skeleton 304 and search for related existing ML projects 110 through the curated and labeled ML project corpus 104. This query 306 may be configured to overcome various challenges. For example, it can be difficult to organize an effective query from a new dataset 106 and a new ML task 108. Thus, the query 306 may be organized based on the insight that there are often mappings between specific meta-features in the new dataset 106 with the new ML task 108 and a set of functional blocks that an ML pipeline solution for this dataset should include. Thus, the set of functional blocks included in the pipeline skeleton 304 can form the basis of the query 306.
[0058] Next, the search module 116 may be configured to search the ML project corpus 104 based on the query 306. This search may be configured to overcome various challenges. For example, it can be difficult to identify the best computational notebook among the existing ML projects of the ML project corpus 104 to be adapted from among many other related computational notebooks. Thus, the search may be organized based on the insight that there may be many computational notebooks with the functional blocks required for the new ML pipeline 112. Thus, a small set of computational notebooks with all the necessary semantic labels can be identified during the search while ensuring quality.
[0059] Next, in some embodiments, the adaptation module 120 may be configured to perform a pipeline merge 308 of functional blocks from the ML pipeline 111 of an associated existing ML project 110 to generate a new ML pipeline 212. This pipeline merge 308 may be configured to overcome various challenges. For example, it may be difficult to merge all the computational notebooks. Thus, the resulting code is syntactically correct and an appropriate solution for the new dataset 106 and the new ML task 108. Accordingly, the pipeline merge 308 may be configured to utilize the pipeline skeleton 304 (as indicated by the arrow from the pipeline skeleton 304 to the pipeline merge 308), and program analysis may be utilized to make the code of the new ML pipeline 112 syntactically correct and executable without further changes.
[0060] Accordingly, in some embodiments, a data scientist 118, who may be a non-expert, may only need to formulate a new dataset 106 and a new ML task 108 for a new ML project, and the search module 116 and the adaptation module 120 may work together (e.g., by performing one or more of the methods disclosed herein) to ultimately generate a new ML pipeline 112 for a new ML project 310 that is immediately executable to perform the new ML task 108 with the new dataset 106, without any further changes, in part, by the data scientist 118.
[0061] Changes, additions, or omissions may be made to FIG. 3 without departing from the scope of the present disclosure. For example, the environment 300 may include more or fewer elements than shown and described in the present disclosure.
[0062] FIG. 4 shows a block diagram of an exemplary computing system 402 according to at least one embodiment of the present disclosure. The computing system 402 may be configured to perform or direct one or more operations associated with one or more modules (e.g., the curation module 114, the search module 116, or the adaptation module 120 of FIGS. 1-3, or any combination thereof). The computing system 402 may include a processor 450, a memory 452, and a data storage device 454. The processor 450, the memory 452, and the data storage device 454 may be communicatively coupled.
[0063] Typically, the processor 450 may include any suitable dedicated or general-purpose computer, computing entity, or processing device, including various computer hardware or software modules, and may be configured to execute instructions stored on any suitable computer-readable storage medium. For example, the processor 450 may include a microprocessor, a microcontroller, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or any other digital or analog circuit configured to interpret and / or execute program instructions and / or process data. Although shown as a single processor in FIG. 4, the processor 450 may include any number of processors configured to perform or direct any number of operations described in the present disclosure, individually or collectively. Further, one or more of the processors may be located on one or more different electronic devices, such as different servers.
[0064] In some embodiments, the processor 450 may be configured to interpret and / or execute program instructions and / or process data stored in the memory 820, the data storage device 454, or both the memory 452 and the data storage device 454. In some embodiments, the processor 450 may fetch program instructions from the data storage device 454 and load the program instructions into the memory 452. After the program instructions are loaded into the memory 452, the processor 450 may execute the program instructions.
[0065] For example, in some embodiments, one or more of the above-described modules (e.g., the curation module 114, the search module 116, or the adaptation module 120, or any combination thereof) may be included in the data storage device 454 as program instructions. The processor 450 may fetch the program instructions of the corresponding module from the data storage device 454 and load the program instructions of the corresponding module into the memory 452. After the program instructions of the corresponding module are loaded into the memory 452, the processor 450 may execute the program instructions, and as a result, the computing system may perform the operations associated with the corresponding module as instructed by the instructions.
[0066] Memory 452 and data storage device 454 may include a computer-readable storage medium carrying or having stored computer-executable instructions or data structures. Such computer-readable storage media may include any commercially available media that can be accessed by a general-purpose or special-purpose computer such as processor 450. By way of example and not limitation, such computer-readable storage media may include random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM), or other optical disk storage, magnetic disk storage, or other magnetic storage devices, flash memory devices (e.g., individual memory devices), or any other storage media tangible or non-transitory that can be used to carry or store particular program code in the form of computer-executable instructions or data structures and that can be accessed by a general-purpose or special-purpose computer. Combinations of the above may also be included within the scope of computer-readable storage media. Computer-executable instructions may include, for example, instructions and data configured to cause processor 450 to perform a particular operation or a group of operations.
[0067] Changes, additions, or omissions may be made to computing system 402 without departing from the scope of the present disclosure. For example, in some embodiments, computing system 402 may include any number of other components that may or may not be explicitly shown or described.
[0068] FIG. 5 is a flowchart of an exemplary method 500 for automatically curating an existing ML project into a corpus adaptable for use in a new ML project, according to at least one embodiment described in the present disclosure. Method 500 may be performed by any suitable system, device, or apparatus. For example, the curation module 114 of FIGS. 1-2 or the computing system 402 of FIG. 4 (when directed by one or more modules) may perform one or more of the operations associated with method 500. Although shown as separate blocks, the steps and operations associated with one or more of the blocks of method 500 may, depending on the particular implementation, be further divided into additional blocks, combined into fewer blocks, or removed.
[0069] In some embodiments, as shown in FIG. 2, method 500 may be utilized by the curation module 114 to filter 206 and clean 208 an existing ML project 204 before storing a subset of the existing ML project 204 in the ML project corpus 104.
[0070] Method 500 may include, at least at block 502, collecting a set of ML projects from a repository of ML projects. In some embodiments, this collecting step may be based on filtering criteria. For example, the curation module 114 may collect a set of existing ML projects 204 from the OSS ML project databases 102a-102n based on filtering criteria. In some embodiments, the set of ML projects may be collected in accordance with one or more operations of method 600, further described below with reference to FIG. 6.
[0071] Method 500 may include, at block 504, the step of ensuring the executability of the ML pipelines in a set of ML projects. For example, the curation module 114 may ensure the executability of a set of existing ML projects 204. In some embodiments, the executability of the ML pipelines may be ensured according to one or more operations of method 700, which will be further described below with reference to FIG. 7. Further, in some embodiments, the executability of the ML pipelines may be ensured as shown in ML pipelines 800 and 850 of FIGS. 8A and 8B, which will be further described below.
[0072] Method 500 may include, at block 506, the step of identifying irrelevant portions of the ML pipelines in a set of ML projects. For example, the curation module 114 may identify and annotate irrelevant portions of a set of existing ML projects 204. In some embodiments, the irrelevant portions of the ML pipelines may be identified as shown in ML pipelines 800 and 850 of FIGS. 8A and 8B, which will be further described below.
[0073] Method 500 may include, at block 508, the step of collecting quality characteristics for a set of ML projects. For example, the curation module 114 may generate quality characteristics for a set of existing ML projects 204. In some embodiments, the quality characteristics may be generated as shown in table 900 of FIG. 9, which will be further described below.
[0074] Method 500 may include, at block 510, the step of generating diversity characteristics for a set of ML projects. For example, the curation module 114 may generate diversity characteristics for a set of existing ML projects 204. In some embodiments, the diversity characteristics may be generated as shown in table 1000 of FIG. 10, which will be further described below.
[0075] Method 500 may include, at block 512, selecting a subset of ML projects from a set of ML projects based on quality features and diversity features. For example, the curation module 114 may select a subset of ML projects from a set of existing ML projects 204 based on quality features and diversity features. In some embodiments, the subset of ML projects may be selected from the set of ML projects according to one or more operations of method 1100, which will be further described later with reference to FIG. 11.
[0076] Method 500 may include, at block 514, storing the subset of ML projects in a corpus of ML projects that can be adapted for use in a new ML project. For example, the curation module 114 may store a subset of existing ML projects 204 in an ML project corpus 104 that can be adapted for a new ML project (such as new ML project 310).
[0077] Changes, additions, or omissions may be made to method 500 without departing from the scope of the present disclosure. For example, some of the operations of method 500 may be performed in a different order. Additionally or alternatively, two or more operations may be performed simultaneously. Further, the steps and operations outlined are provided by way of example, and some of the steps and operations may be optional, combined with fewer steps and operations, or extended with additional steps and operations without detracting from the disclosed embodiments.
[0078] FIG. 6 is a flowchart of an exemplary method 600 for collecting a set of ML projects from one or more repositories of ML projects based on filtering criteria, according to at least one embodiment described in the present disclosure. In some embodiments, the operation of block 502 described above with respect to method 500 of FIG. 5 may be performed according to method 600.
[0079] Method 600 may be performed by any suitable system, device, or apparatus. For example, the curation module 114 of FIGS. 1-2 or the computing system 402 of FIG. 4 (when directed by one or more modules) may perform one or more of the operations associated with method 600. Although shown in separate blocks, the steps and operations associated with one or more of the blocks of method 600 may, depending on the particular implementation, be divided into additional blocks, combined into fewer blocks, or removed.
[0080] Method 600 may include, at block 602, ranking all the datasets of all the ML projects from one or more repositories of an ML project based on the quality of the datasets. For example, the curation module 114 may rank all the datasets of the existing ML projects 204 from the OSS ML project databases 102a-102n based on the quality of the datasets. In some embodiments, the quality of the datasets may be determined based on votes by other users (e.g., votes on Kaggle), on the basis that the dataset is part of a "featured set" of datasets hosted in a project database (e.g., on Kaggle), or on how recent the dataset is.
[0081] Method 600 may include, at block 604, selecting the top first predetermined number of ranked datasets. For example, the curation module 114 may select the top first predetermined number of ranked datasets from the existing ML projects 204.
[0082] Method 600 may include, at block 606, ranking corresponding ML projects for each of the selected datasets based on importance criteria. For example, the curation module 114 may rank corresponding ML projects from existing ML projects 204 for each of the selected datasets based on importance criteria. In some embodiments, the importance of a dataset may be determined based on votes by other users (e.g., votes on Kaggle). In other embodiments, the importance of a dataset may be determined based on the accuracy of the corresponding pipeline.
[0083] Method 600 may include, at block 608, selecting the second predetermined number of top-ranked ML projects for inclusion in a set of ML projects. For example, the curation module 114 may select the second predetermined number of top-ranked ML projects from existing ML projects 204 for inclusion in a set of existing ML projects 20. For example, if the top 15 top-ranked ML projects (based on upvotes by other users on Kaggle) are selected for each of the top 150 datasets on Kaggle, this may result in 2,250 ML projects being generated.
[0084] Changes, additions, or omissions may be made to method 600 without departing from the scope of the present disclosure. For example, the operations of method 600 may be performed in a different order. Further, in some embodiments, method 600 may be repeated or executed concurrently with respect to block 502 of FIG. 5.
[0085] FIG. 7 is a flowchart of an exemplary method 700 for ensuring the executability of ML pipelines in a set of ML projects, according to at least one embodiment described in the present disclosure. In some embodiments, the operation of block 504 described above with respect to method 500 of FIG. 5 may be performed in accordance with method 700.
[0086] Method 700 may be performed by any suitable system, device, or apparatus. For example, the curation module 114 of FIGS. 1-2 or the computing system 402 of FIG. 4 (when directed by one or more modules) may perform one or more of the operations associated with method 600. Although shown as separate blocks, the steps and operations associated with one or more of the blocks of method 700 may, depending on the particular implementation, be further divided into additional blocks, combined into fewer blocks, or removed.
[0087] Method 700 may include, at block 702, determining whether an ML pipeline within an ML project is executable without modification to the ML pipeline. If executable (Yes at block 702), method 700 may proceed to block 704, and if not executable (No at block 702), method 700 may proceed to block 706. For example, the curation module 114 may determine whether an ML pipeline within one of the existing ML projects 204 is executable without modification.
[0088] The method may include, at block 704, maintaining the ML pipeline within a set of ML projects. For example, the curation module 114 may maintain the ML projects within the set of existing ML projects 204 due to the ML pipeline within the ML project being executable, either before or after performing API adaptation with the ML pipeline.
[0089] The method may include, at block 706, determining whether the ML pipeline within the ML project is executable after performing API adaptation with the ML pipeline. If executable (Yes at block 706), method 700 may proceed to block 704, and if not executable (No at block 706), method 700 may proceed to block 708. For example, the curation module 114 may determine whether an ML pipeline within one of the existing ML projects 204 is executable after performing API adaptation to the ML pipeline.
[0090] The method may include, at block 708, removing an ML pipeline from a set of ML projects. For example, the curation module 114 may remove an ML project from an existing set of ML projects 204 because the ML pipeline of the ML project is not executable, either before or after performing API adaptation in the ML pipeline.
[0091] Changes, additions, or omissions may be made to method 700 without departing from the scope of the present disclosure. For example, the operations of method 700 may be performed in a different order. Additionally, in some embodiments, method 700 may be repeated or executed simultaneously with respect to block 504 of FIG. 5.
[0092] FIG. 8A shows exemplary ML pipeline code 800 of an existing ML project, and FIG. 8B shows exemplary cleaned ML pipeline code 850 that results from cleaning the exemplary ML pipeline code 800 of FIG. 8A. The ML pipeline code 800 may be part of a computational notebook (e.g., a Jupyter notebook) of one of the existing ML projects 204. Here, the ML task is to predict whether a person earns more than $50k per year based on a dataset that includes features such as age, workclass, education, occupation, housing price, and the race of multiple workers. FIGS. 8A and 8B are discussed here to provide an example of how blocks 504 and 506 of method 500 may be performed with respect to the ML pipelines in a set of existing ML projects 204.
[0093] In the example shown in FIGS. 8A and 8B, the API adaptation shown may be performed on the ML pipeline code 800, resulting in the ML pipeline code 850 to ensure the executability of the ML pipeline code. Further, the irrelevant portions of the ML pipeline code 800 may be marked in bold and italic in the ML pipeline code 850 and identified as something to be avoided from being executed in a particular environment. In some embodiments, these irrelevant portions may include debugging code, checking code, and graph drawing code.
[0094] In some embodiments, this identification of the irrelevant portions of the ML pipeline code 800 generates the property-preserving sample D of the dataset of the ML project P<D,L> sample (e.g., reducing the time associated with processing the entire dataset but not sacrificing the range of the properties of the dataset by smart sampling the dataset), instrumenting the ML pipeline L to trace the variables and objects changed within each statement of the ML pipeline L, the instrumented ML pipeline L sample in the sample D of the dataset instr to execute and collect the execution trace E, identifying the target statement T in the ML pipeline L (e.g., the target statement shown in the ML pipeline code 850), extracting all variables and objects B from the target statement T, generating the backward slice B of the variables and objects B extracted from the execution trace E, (for generating the annotated ML pipeline L N and) associating and annotating the statements in the ML pipeline from the backward slice B with the ML pipeline L NThis may include annotating all other statements in the as not relevant. In this way, only the statements in the ML pipeline code 800 related to the target statement are annotated as relevant. In some embodiments, smart sampling of the dataset may include retaining all category values at the original ratio of the category column using stratified sampling, sampling uniformly across the indicated range of consecutive (numerical) columns, randomly sampling instances for string columns, and having missing values in that column after reduction of the dataset if the column had missing values in the original data. In some embodiments, this smart sampling of the dataset may result in a significant reduction of the dataset. For example, a 2GB dataset may be reduced to 9MB, resulting in a reduction of the execution time from 10 minutes to 20 seconds.
[0095] In some embodiments, the cleaning of the ML pipeline code 800 applicable to a portion of the Jupyter notebook may result in a cleaned ML pipeline code 850 that is more suitable for feature extraction (for project selection) and subsequent steps of the search and adaptation workflow (e.g., indexing, searching, and adapting the computational notebook).
[0096] Changes, additions, or omissions may be made to the ML pipeline code 800 and the cleaned ML pipeline code 850 without departing from the scope of the present disclosure. For example, some statements of the ML pipeline code 800 may be performed in a different order.
[0097] FIG. 9 is exemplary quality characteristics table 900. FIG. 9 is discussed here to provide an example of how block 508 of method 500 can be executed with respect to a set of existing ML projects 204. In the example shown in FIG. 9, the quality characteristics may include performance characteristics, code quality characteristics, and community metric characteristics. These quality characteristics may be generated for a set of existing ML projects 204. Each of these quality characteristics may include measurements, metrics, and normalized metrics.
[0098] In some embodiments, as disclosed in table 900 of FIG. 9, the step of generating quality characteristics for a set of existing ML projects 204 (e.g., at block 508 of method 500) may include generating a numerical quality score for each ML project in the set of existing ML projects 204 based on the accuracy of the trained ML model of the ML project, the documentation quality of the ML pipeline of the ML project, the modularity quality of the ML pipeline of the ML project, the standard API usage quality of the ML pipeline of the ML project, and the evaluation of the ML project.
[0099] In some embodiments, the purpose of extracting quality characteristics may be to calculate characteristics that can form the basis for the selection of existing ML projects placed within the ML project corpus 104. These quality characteristics may represent the intrinsic value of the ML pipeline in terms of the quality of the trained model, the code structure, and the value by the community. These quality characteristics can be evaluated individually for a given ML pipeline and may be calculated as a single numerical score (e.g., between 0 and 1.0) representing the quality of each ML pipeline. In some embodiments, this single numerical score may be calculated according to the following formula.
Number
[0100] Changes, additions, or omissions may be made to Table 900 without departing from the scope of the present disclosure. For example, additional quality features may be added to the quality features disclosed in Table 900.
[0101] FIG. 10 is an exemplary Table 1000 of diversity features. FIG. 10 is discussed here to provide an example of how block 510 of method 500 may be performed with respect to a set of existing ML projects 204. In the example shown in FIG. 10, the diversity features may include dataset features and computational notebook features. These diversity features may be generated for a set of existing ML projects 204. Each of these diversity features may include intent, metric, and normalization.
[0102] In some embodiments, as disclosed in Table 1000, the step of generating diversity features for a set of existing ML projects 204 (e.g., at block 510 of method 500) may include, for each ML project in the set of existing ML projects 204, the size of the ML project's dataset, the presence or absence of missing values in the ML project's dataset, the type of data in the ML project's dataset, the presence or absence of a given library API in the ML project's ML pipeline, and the relative range of the component data streams in the ML project's dataset, based on which to extract a feature vector for the ML project.
[0103] In some embodiments, the purpose of extracting diversity features may be to calculate features that can form a basis for the selection of existing ML projects placed within the ML project corpus 104. These diversity features may capture the capabilities of the ML pipeline to add a wide range of solutions ultimately available within the ML project corpus 104. These diversity features may be evaluated with respect to the selection of other ML pipelines and may be calculated as a feature vector for each ML project representing its contribution to diversity.
[0104] Changes, additions, or omissions may be made to Table 1000 without departing from the scope of the present disclosure. For example, additional diversity features may be added to the diversity features disclosed in Table 1000.
[0105] FIG. 11 is a flowchart of an exemplary method 1100 for selecting a subset of ML projects from a set of ML projects based on quality features and diversity features, according to at least one embodiment described in the present disclosure. In some embodiments, the operations of block 512 described above with respect to method 500 of FIG. 5 may be performed according to method 1100.
[0106] Method 1100 may be performed by any suitable system, device, or apparatus. For example, the curation module 114 of FIGS. 1-2 or the computing system 402 of FIG. 4 (when directed by one or more modules) may perform one or more of the operations associated with method 600. Although shown as separate blocks, the steps and operations associated with one or more of the blocks of method 1100 may, depending on the particular implementation, be divided into additional blocks, combined into fewer blocks, or removed.
[0107] Method 1100 may include, at block 1102, generating a quality metric for each ML project in a set of ML projects based on the quality characteristics of the ML project. For example, the curation module 114 may generate a quality metric for each ML project in an existing set of ML projects 204 based on the quality characteristics of the ML project. For example, a set of cleaned ML projects is given as follows:
Number
Number
Number
[0108] Method 1100 may include, at block 1104, generating a weight for each ML project in a set of ML projects from the quality metric of the ML project. For example, the curation module 114 may generate a weight for each ML project in an existing set of ML projects 204 from the quality metric of the ML project. For example, for each project, its weight is as follows:
Number
[0109] Method 1100 may include, at block 1106, constructing a Weighted Set Cover (WSC) problem from among the ML projects in a set of ML projects based on weights and diversity features. For example, the curation module 114 may construct a WSC problem from among the ML projects in an existing set of ML projects 204 based on weights and diversity features. Calculating weights from the quality metrics of each ML pipeline enables formulating the original task of selecting a set of pipelines that maximizes the collective quality of the selected pipelines (i.e., the maximization objective) as the problem of selecting a set of pipelines with the minimum weight that can be naturally solved as a WSC problem (i.e., the minimization objective). Further, making weights larger with respect to the quality values (by the choice of W) motivates minimizing the cardinality of the selected set. Since WSC is an NP-Complete problem, well-known approximation algorithms for WSC may be used to solve the above. Two possibilities include the greedy algorithm or an algorithm based on a Linear Programming (LP) relaxation.
[0110] The method may include, at block 1108, solving the WSC problem to select a subset of ML projects. For example, the curation module 114 may solve the WSC problem to select a subset of existing ML projects 204. Solving the WSC problem may be solved for the minimum weight subset, and doing so may indirectly maximize the aggregate quality of the selected pipelines. For example, the goal may be to select the following subset that together exhibits all the features in U:
Number
Number
[0111] Changes, additions, or omissions may be made to method 1100 without departing from the scope of the present disclosure. For example, the operations of method 1100 may be performed in a different order. Further, in some embodiments, method 1100 may be repeated or executed simultaneously with respect to block 512 of FIG. 5.
[0112] FIG. 12 is a flowchart of an exemplary method 1200 for automatically labeling functional blocks in a pipeline of existing ML projects in a corpus adaptable for use in a new ML project, according to at least one embodiment described in the present disclosure. Method 1200 may be performed by any suitable system, device, or apparatus. For example, the curation module 114 of FIGS. 1-2 or the computing system 402 of FIG. 4 (when instructed by one or more modules) may perform one or more of the operations associated with method 1200. Although shown as separate blocks, the steps and operations associated with one or more of the blocks of method 1200 may, depending on the particular implementation, be divided into additional blocks, combined into fewer blocks, or removed.
[0113] In some embodiments, as shown in FIG. 2, method 1200 may be utilized by the curation module 114 to abstract 210 and generate project artifacts 212 on existing ML projects 204 before storing a subset of the existing ML projects 204 in the ML project corpus 104.
[0114] Method 1200 may include, at block 1202, normalizing the ML pipelines of existing ML projects stored in the corpus of existing ML projects. For example, the curation module 114 may normalize the ML pipelines of a subset of existing ML projects 204 stored in the ML project corpus 104 (possibly after the existing ML projects 204 have been filtered 206 and cleaned 208). In some embodiments, the ML pipelines may be normalized as shown in the original ML pipeline code 1300 and the normalized ML pipeline code 1350 of FIGS. 13A and 13B, described further below.
[0115] Method 1200 may include, at block 1204, extracting functional blocks from the normalized ML pipeline. For example, the curation module 114 may extract functional blocks from the normalized ML pipeline. In some embodiments, the functional blocks may be extracted according to one or more operations of method 1400, described further below with reference to FIG. 14.
[0116] Method 1200 may include, at block 1206, assigning a label to each of the functional blocks in the normalized ML pipeline. For example, the curation module 114 may assign a label to each of the functional blocks in the normalized ML pipeline. In some embodiments, the labels may be assigned according to one or more operations of method 1600, described further below with reference to FIG. 16.
[0117] Method 1200 may include, at block 1208, indexing each of the ML pipelines in the corpus based on the labels assigned to the functional blocks. For example, the curation module 114 may index each of the ML pipelines in the ML project corpus based on the labels assigned to the functional blocks.
[0118] Method 1200 may include, at block 1210, generating a new ML pipeline to execute a new ML task on a new dataset for a new ML project, using labels assigned to functional blocks within a corpus. For example, search module 116 and adaptation module 120 may generate a new ML pipeline 112 to execute a new ML task 108 on a new dataset 106 for a new ML project 310, using labels assigned to functional blocks within ML project corpus 104.
[0119] Changes, additions, or omissions may be made to method 1200 without departing from the scope of the present disclosure. For example, some of the operations of method 1200 may be performed in a different order. Additionally or alternatively, two or more operations may be performed simultaneously. Further, the steps and operations outlined are provided as examples, and some of the steps and operations may be optional, combined with fewer steps and operations, or extended with additional steps and operations, without detracting from the disclosed embodiments.
[0120] FIG. 13A shows exemplary original ML pipeline code 1300 of an existing ML project prior to normalization, and FIG. 13B shows exemplary normalized ML pipeline code 1350 after normalization of the original ML pipeline code 1300. The original ML pipeline code 1300 may be part of a computational notebook (e.g., a Jupyter notebook) of one of the existing ML projects 204. Here, the ML task is to predict whether a person makes more than $50k per year, based on a dataset that includes features such as age, workclass, education, occupation, household income, and race of multiple workers. FIGS. 13A and 13B are discussed here to provide an example of how block 1202 of method 1200 may be executed with respect to ML pipelines within a set of existing ML projects 204.
[0121] In the example shown in FIGS. 13A and 13B, the original ML pipeline code 1300 may be normalized in various ways including normalizing variable names, replacing column names with column data types, removing parameters from API statements, and collapsing repeated instances of API statements into a single instance of the API statement. For example, the variable names "dataset", "array", "X", "Y", "X_train", "X_validation", "Y_train", "Y_validation", "random_forest", "predictions" in the original ML pipeline code 1300 may each be normalized to "_var_" in the normalized ML pipeline code 1350. Also, the columns "workclass", "occupation", "native.country", "sex" in the original ML pipeline code 1300 may each be normalized to "_str_category_" in the normalized ML pipeline code 1350. Further, each of the parameters "filna", "map", "train_test_split", "RandomForestClassifier", "fit", "predict" (e.g., parameters such as "X", "United-States", "Male", "0", "Female", "1", etc.) of the API statements in the original ML pipeline code 1300 may each be normalized by removing the parameters in the normalized ML pipeline code 1350. Also, the API statement repeated three times in the original ML pipeline code 1300:
Number
[0122] Changes, additions, or omissions may be made to the original ML pipeline code 1300 and the normalized ML pipeline code 1350 without departing from the scope of the present disclosure. For example, some statements of the original ML pipeline code 1300 may be executed in a different order, and other normalizations may be performed on the original ML pipeline code 1300.
[0123] FIG. 14 shows a flowchart of an exemplary method 1400 for identifying idiom function blocks and custom function blocks according to at least one embodiment described in the present disclosure. In some embodiments, the operations of block 1204 described above with respect to method 1200 of FIG. 12 may be performed according to method 1400.
[0124] Method 1400 may be performed by any suitable system, device, or apparatus. For example, the curation module 114 of FIGS. 1-2 or the computing system 402 of FIG. 4 (when instructed by one or more modules) may perform one or more of the operations associated with method 1400. Although shown as separate blocks, the steps and operations associated with one or more of the blocks of method 1400 may be divided into additional blocks, combined into fewer blocks, or removed depending on the particular implementation.
[0125] In some embodiments, method 1400 may be utilized to split each ML pipeline of an existing ML project stored in the ML project corpus 104 into code cells. In some embodiments, a computational notebook (e.g., a Jupyter notebook) is naturally structured such that all the code within the computational notebook is organized into a set of code cells, each of which may initially be assumed to be a different functional block, but this assumption may be invalidated after additional analysis. Method 1400 may then be utilized to identify groups of statements that are repeated across code cells as idiom function blocks and all other groups of statements within the code cells as custom function blocks.
[0126] Method 1400 may include, at block 1402, generating a directed graph. For example, the curation module 114 may generate a directed graph (see, e.g., the directed graph shown in FIG. 15). In some embodiments, each node in the directed graph may represent a statement normalized by their occurrence within the ML project corpus 104, and each edge may represent the average probability score of the co-occurrence of the statements corresponding to the two nodes that the edge connects. There may be no connection between the last statement of one cell and the first statement of the next cell. For example, given two nodes A and B, the value of their edge may be expressed as follows:
Number
[0127] Method 1400 may include, at block 1404, for each statement in each code cell, adding the statement as a node in a directed graph or incrementing a count among existing nodes in the directed graph of the statement, and at block 1404b, calculating a co-occurrence score between a statement in a code cell and the statement immediately preceding the statement, and at block 1404c, generating an edge between the node of the statement and the node of the statement immediately preceding the statement when the co-occurrence score is greater than a co-occurrence score threshold. For example, for each statement in each code cell (e.g., each code cell in a computational notebook such as a Jupyter notebook), the curation module 114 adds the statement as a node in a directed graph or increments a count among existing nodes in the directed graph of the statement (see, e.g., the nodes of the directed graph shown in FIG. 15 that have a count inside each node), calculates a co-occurrence score between a statement in a code cell and the statement immediately preceding the statement (see, e.g., the co-occurrence scores in the directed graph of multi-line idioms shown in FIG. 15), and may generate an edge between the node of the statement and the node of the statement immediately preceding the statement when the co-occurrence score is greater than a co-occurrence score threshold (see, e.g., the edges between the nodes in the directed graph of multi-line idioms shown in FIG. 15). In some embodiments, this generation of edges may alternatively be performed by first generating all edges regardless of the co-occurrence score, but then removing all edges having a co-occurrence score less than a particular threshold. The threshold may be determined empirically. After the removal of edges, a set of connected components may remain. Here, each component represents an idiom (e.g., a set of statements / APIs that occur together to implement a function).
[0128] Method 1400 may include, at block 1406, as an idiom feature block, marking all nodes connected by an edge. For example, the curation module 114 may mark all nodes connected by an edge as an idiom feature block (e.g., refer to the nodes connected by an edge in the directed graph of the multi-line idiom shown in FIG. 15).
[0129] Method 1400 may include, at block 1408, for any of the normalization ML pipelines, marking each set of nodes corresponding to a consecutive statement, such as a custom feature block, that is not marked as one of the idiom feature blocks. For example, the curation module 114 may mark each set of nodes corresponding to a consecutive statement in any of the normalization ML pipelines, such as a custom feature block, that is not marked as one of the idiom feature blocks (e.g., refer to the single-line feature block and the multi-line feature blocks shown in FIG. 15).
[0130] Changes, additions, or omissions may be made to method 1400 without departing from the scope of the present disclosure. For example, the operations of method 1400 may be performed in a different order. Further, in some embodiments, method 1400 may be repeated or executed simultaneously with respect to block 1204 of FIG. 12.
[0131] FIG. 15 shows the extraction of functional blocks from a normalized ML pipeline using a directed graph. FIG. 15 is discussed here to provide an example of how block 1204 of method 1200, and blocks 1402 - 1408 of method 1400 may be executed with respect to normalized ML pipeline code 1500. In the example shown in FIG. 15, the normalized ML pipeline code 1500 may be used to generate one or more directed graphs. Here, each node represents a statement and each edge represents a co-occurrence between statements having a score higher than a threshold. As disclosed in the single-line functional blocks, statement 1 occurs 94 times, statement 2 occurs 6 times, and statement 3 occurs 6 times across all normalized ML pipelines. Further, as disclosed in the multi-line functional blocks, statements 4, 5, 6, 7 occur only once. This indicates that these statements appear only in the normalized ML pipeline code 1500 and not in any other normalized ML pipeline. Also, as disclosed in the multi-line idioms, statement 8 occurs 10 times, statement 9 occurs 10 times, and statement 11 occurs 2 times. Edges with corresponding co-occurrence scores higher than a threshold score (e.g., a threshold of 0.5) indicate that the statement sequences 8, 9, 10, and 8, 9, 11 appear together within multiple normalized ML pipelines and thus should be marked together as an idiom functional block within each normalized ML pipeline in which they appear. Further, after marking the idiom 8, 9, 10 as an idiom functional block within the normalized ML pipeline code 1500, the remaining statements within the normalized ML pipeline code 1500 can be decomposed into single-line functional blocks 1, 2, 3, and multi-line custom functional blocks including statements 4, 5, 6, 7 by the boundaries of the code cells in which they exist.
[0132] Changes, additions, or omissions may be made to the normalized ML pipeline code 1500 and the directed graph of FIG. 15 without departing from the scope of the present disclosure. For example, additional directed graphs for additional normalized ML pipeline code may be added.
[0133] Figure 16 is a flowchart of an exemplary method 1600 for assigning labels to each of the functional blocks in a normalized ML pipeline. In some embodiments, the operations of block 1206 described above with respect to method 1200 of FIG. 12 may be performed in accordance with method 1600.
[0134] Method 1600 may be performed by any suitable system, device, or apparatus. For example, the curation module 114 of FIGS. 1-2 or the computing system 402 of FIG. 4 (when directed by one or more modules) may perform one or more of the operations associated with method 1600. Although shown as separate blocks, the steps and operations associated with one or more of the blocks of method 1600 may, depending on the particular implementation, be divided into additional blocks, combined into fewer blocks, or removed.
[0135] Method 1600 may include, at block 1602, extracting text from any comments within the functional block, or, when no comments exist within the functional block, mining text from the documentation of any API statements within the functional block. For example, the curation module 114 may extract text from any comments (e.g., statements beginning with "#" preceding each functional block within the normalized ML pipeline code 1500) within the functional block, or, when no comments exist within the functional block, mine text from the documentation of any API statements (e.g., documentation of API statements obtainable in a repository of API code such as on a website hosting the API code) within the functional block. This extraction or mining may, additionally or alternatively, include preprocessing the extracted or mined text by applying standard preprocessing techniques (e.g., stop word removal, stemming, tokenization, etc.).
[0136] Method 1600 may include, at block 1604, identifying a first common verb and a noun or noun phrase in the extracted or mined text. For example, the curation module 114 may identify common verbs (e.g., "scale" or "apply") and nouns or noun phrases (e.g., "data" or "standard" or "unit variance") in the extracted or mined text. In the context of method 1600, the term "common" may refer to a verb, noun, or noun phrase that is common across multiple instances of the idiom functional block or common across labels. In some embodiments, block 1604 may include, for all instances of the idiom block, extracting noun / verb candidates for each instance of the idiom functional block from the extracted or mined text that may form a label.
[0137] Method 1600 may include, at block 1606, generating a label for a functional block based on a first common verb and a noun or noun phrase. For example, the curation module 114 may generate a label "scale date" from a comment "scale the data to be between -1 and 1". In some embodiments, generating this label may include incorporating the most important verb and noun or noun phrase and assigning these words as semantic labels for the functional block and other instances of the same idiom functional block. In these embodiments, the most important words may be determined as the top N frequently used words, or may be determined through topic modeling, or may be determined by some other method. In some embodiments, block 1606 may include performing a consensus operation among the label candidates contributed by each instance of the idiom functional block to find the most important common noun / verb instances across these different candidates. This may form an initial label for all instances of the idiom functional block. For example, four labels "scale data", "apply standard", "scale numerical column data", "standard feature remove mean scale unit variance" from the idiom functional block may be relabeled with a single common label "scale standard data".
[0138] In some embodiments, blocks 1604 and 1606 may be performed for idiom functional blocks, but may be varied for custom functional blocks. In these embodiments, blocks 1604 and 1606 may be varied for each custom functional block by incorporating the most important nouns and verbs from the extracted or mined text of the custom functional block, and rather than performing a consensus operation, provide a starting point for block 1608.
[0139] Method 1600 may include, at block 1608, generating a similarity score for each pair of functional blocks in the normalized ML pipeline. For example, the curation module 114 may generate a similarity score (e.g., a score between 0 and 1.0) for each pair of functional blocks within the normalized ML pipeline. In some embodiments, the similarity score can be calculated through simple word matching. For example, for two labels having word sets A and B, the similarity score may be calculated as follows: [Number]
[0140] Method 1600 may include, at block 1610, generating a group of functional blocks if the similarity score is greater than a similarity score threshold. For example, the curation module 114 may generate a group of functional blocks if the similarity score is greater than a similarity score threshold (e.g., a threshold of 0.5). In some embodiments, for a given functional block, other functional blocks may be sorted based on the similarity score based on semantic labels, and the top K may be marked as different implementations of the same function. In some embodiments, the similarity score threshold may be adjusted empirically.
[0141] Method 1600 may include, at block 1612, identifying a second common verb and noun or noun phrase within each label of the functional blocks in the functional block group. For example, the curation module 114 may identify a second common verb (e.g., "scale") and a noun or noun phrase (e.g., "data") for each of the functional blocks in the functional block group. This second identification may allow a second iteration after the first label is generated to further strengthen the functional block group with labels that are similar enough to be considered functionally equivalent.
[0142] Method 1600 may include, at block 1614, generating a common label for a functional block group based on a second common verb and a noun or noun phrase. For example, the curation module 114 may generate a common label from the second common verb and the noun or noun phrase. In some embodiments, common or frequently occurring words of the semantic labels may be assigned as the common semantic label for the entire group. For example, the curation module 114 may update the label assigned to each of the functional blocks within each functional block group to a common label. For example, two labels, "scale standard data" and "scale data feature", from functional blocks determined to be functionally equivalent may be relabeled with a single common label, "scale data".
[0143] Method 1600 may include, at block 1616, updating the label assigned to each of the functional blocks within each functional block group to a common label.
[0144] Changes, additions, or omissions may be made to method 1600 without departing from the scope of the present disclosure. For example, the operations of method 1600 may be performed in a different order. Further, in some embodiments, method 1600 may be repeated or executed simultaneously with respect to block 1206 of FIG. 12.
[0145] Figure 17 shows the automatic labeling of functional blocks within an ML pipeline. Figure 17 is discussed here to provide an example of how the various blocks of method 1600 can be executed. In the example shown in Figure 17, the functional block Block-1 may include two normalized statements, namely, "_var_=StandardScaler()" and "_var1_=_var_.fit_transform()". Since this functional block appears in four separate computational notebooks, namely notebook-1, notebook-2, notebook-3, and notebook-4, it may be an idiomatic functional block. Further, the second functional block Block-2 may include two normalized statements, namely, "_var_=MinMaxScaler()" and "_var1_=_var_.fit_transform()". Although these two functional blocks are not the same, they may be determined to be functionally equivalent based on the similarity between their assigned labels, as will be described later.
[0146] Regarding Block-1, in block 1602 of method 1600, text may be extracted from the comments for notebook-1, notebook-2, and notebook-3, and there may be no available comments for notebook-4. Thus, the text may be mined from an alternative source (e.g., API documentation) for notebook-4. Next, in block 1604 of method 1600, common verbs (e.g., "scale" or "apply") and nouns or noun phrases (e.g., "data" or "standard" or "unit variance") may be identified from the extracted or mined text. Next, in block 1606 of method 1600, based on the first common verb and noun or noun phrase, the label "scale standard data" may be generated for Block-1. Similarly for Block-2, in block 1602 and (as described above) modified versions of blocks 1604 and 1606, the label "scale data feature" may be generated.
[0147] In block 1608 of method 1600, a similarity score of 0.67 may be generated for a pair of Block-1 and Block-2. In block 1610 of method 1600, Block-1 and Block-2 may be grouped together because their similarity score (0.67) is higher than a similarity score threshold (e.g., a threshold of 0.60). In block 1612 of method 1600, common verbs (e.g., "scale") and nouns (e.g., "data") may be identified among the labels of Block-1 and Block-2. Method 1600 may include, in block 1614, the step of generating a common label ("scale data") for Block-1 and Block-2 based on the common verbs (e.g., "scale") and nouns (e.g., "data").
[0148] Changes, additions, or omissions may be made to the functional blocks, the extracted or mined text, the similarity scores, and the automatically assigned labels without departing from the scope of the present disclosure.
[0149] FIG. 18 is a flowchart of an exemplary method 1800 for automatically generating a pipeline for a new ML project from a pipeline of an existing ML project stored in a corpus according to at least one embodiment described in the present disclosure. Method 1800 may be executed by any suitable system, device, or apparatus. For example, the curation module 114, the search module 116, and the adaptation module 120 of FIGS. 1-3, or the computing system 402 of FIG. 4 (when directed by one or more modules) may perform one or more of the operations associated with method 1800. Although shown in separate blocks, the steps and operations associated with one or more of the blocks of method 1800 may be divided into additional blocks, combined into fewer blocks, or removed depending on a particular implementation.
[0150] In some embodiments, method 1800 may be utilized by curation module 114, search module 116, and adaptation module 120 to perform the operations disclosed in FIGS. 1 and 2.
[0151] Method 1800 may include, at block 1802, storing an existing ML project in a corpus. Here, the existing ML project includes an ML pipeline having functional blocks. For example, curation module 114 may store existing ML project 204 in ML project corpus 104. In some embodiments, existing ML project 204 may include an ML pipeline having functional blocks. In some embodiments, these functional blocks may be identified according to the operations of block 1204 of method 1200.
[0152] Method 1800 may include, at block 1804, generating a search query for a new ML project based on a new dataset for the new ML project and a new ML task for the new ML project. For example, search module 116 may generate query 306 from new ML project 310 based on new dataset 106 for new ML project 310 and new ML task 108 for new ML project 310.
[0153] Method 1800 may include, at block 1806, searching for a set of related existing ML projects based on the search query through the existing ML projects stored in the corpus. For example, search module 116 may search for related existing ML project 110 based on query 306 through the existing ML projects stored in ML project corpus 104.
[0154] Method 1800 may include, at block 1808, merging the ML pipelines of a set of related existing ML projects to generate a new ML pipeline for a new ML project. Here, the new ML pipeline has functional blocks incorporated from the ML pipelines of the set of related existing ML projects. For example, adaptation module 120 may perform pipeline merge 308 of the ML pipeline 111 of related existing ML project 110 to generate a new ML pipeline 112 for new ML project 310. In this example, the new ML pipeline 112 may have functional blocks incorporated from the ML pipeline 111 of related existing ML project 110.
[0155] Method 1800 may include, at block 1810, adapting the functional blocks of the new ML pipeline for the new ML project so that the new ML pipeline is executable to perform a new ML task with a new dataset. For example, adaptation module 120 may adapt the functional blocks of the new ML pipeline 112 for new ML project 310 so that the new ML pipeline 112 is executable to perform a new ML task 108 with new dataset 106.
[0156] Changes, additions, or omissions may be made to method 1800 without departing from the scope of the present disclosure. For example, some of the operations of method 1800 may be performed in a different order. Additionally or alternatively, two or more operations may be performed simultaneously. Further, the steps and operations outlined are provided by way of example, and some of the steps and operations may be optional, combined into fewer steps and operations, or expanded into additional steps and operations without detracting from the disclosed embodiments.
[0157] FIG. 19 shows a sequence graph 1900 and a pipeline skeleton 2002 for a new ML project (e.g., new ML project 1310). FIG. 20 shows the pipeline skeleton 2002 and a table 2050 of ML pipelines that can be searched for functional blocks that match the pipeline skeleton 2002. FIGS. 19 and 20 are discussed here to provide an example of how blocks 1804, 1806, 1808, 1810 of method 1800 can be performed with respect to the ML pipeline corpus 104.
[0158] As shown in FIGS. 19 and 20, the pipeline skeleton 2002 may be an ordered set of functional blocks for a new ML pipeline 112 of a new ML project 310, and may correspond to labels assigned to functional blocks of ML pipelines of existing ML projects stored in the ML project corpus 104. In some embodiments, the pipeline skeleton 2002 may be generated by a pipeline skeleton ML model 302. The pipeline skeleton ML model 302 (or a set of ML models) may be trained to learn a mapping between dataset meta-features and semantic labels. For example, given the meta-features of a new dataset 1056, the pipeline skeleton ML model 302 may be trained to synthesize a pipeline skeleton 2002 that includes the required semantic labels having those sequences.
[0159] In some embodiments, the pipeline skeleton ML model 302 may include a multivariate multi-class classifier that is trained before generating the pipeline skeleton 2002. The multivariate multi-class classifier may be configured to map dataset meta-features to an unordered set of functional blocks (indicated by corresponding semantic labels) that the pipeline skeleton 304 or 2002 should include. This training includes extracting dataset features from the datasets of existing ML projects in the ML project corpus 104 associated with a particular label, identifying the set of all labels from the functional blocks of the existing ML projects, preparing training data including an input vector having the dataset features and a binary output tuple indicating the presence or absence of each of the set of all labels, and training the pipeline skeleton ML model 302 to learn the mapping between the dataset features and the corresponding labels of the set of all labels. In some embodiments, training the pipeline skeleton ML model 302 may enable the pipeline skeleton ML model 302 to predict an ordered set of functional blocks (e.g., within the pipeline skeleton 304 or 2002) that can be used to construct the ML pipeline of the new ML pipeline 112 using the prominent characteristics of the new dataset 106 and the new ML task 108 (meta-features). The meta-features of the dataset may include, but are not limited to, the number of rows, the number of features, the presence of numerical values, the presence of missing values, the presence of counts, the presence of numerical categories, the presence of string categories, the presence of text, and the type of target.
[0160] In some embodiments, the pipeline skeleton ML model 302 may further include a sequence graph (similar to the sequence graph 1900) representing a partial order among functional blocks learned from training data. The sequence graph may be configured to map an unordered set of blocks to an ordered set (e.g., as shown in the pipeline skeleton 2002) based on a partial order among the blocks learned from a training project corpus. The sequence graph may include nodes for each label of the set of all labels from the functional blocks of an existing ML project. The sequence graph may also include directed edges between each pair of a first node and a second node, where the first node precedes the second node in one of the existing ML projects.
[0161] Once the pipeline skeleton ML model 302 is trained, the pipeline skeleton ML model 302 may be utilized to generate queries 306 for a new ML project 310. In some embodiments, this generation of the query 306 may include mapping dataset features to an unordered set of labels of a new ML pipeline 112 of the new ML project 310, and may further include mapping the unordered set of labels to an ordered set of labels using a partial order represented in a sequence graph (e.g., sequence graph 1900). The query 306 may include such an ordered sequence of labels as a pipeline skeleton 2002. For example, FIG. 19 shows an example of the step of mapping an unordered set of labels generated by a pipeline skeleton ML model to an ordered sequence of labels using the sequence graph 1900. The unordered set of labels may first be mapped to the corresponding nodes in the sequence graph 1900 represented by the set of bold nodes, namely, "Read Data", "Fill Missing Values", "Convert String to Int", "Split Train Test", "Random Forest". Next, a subgraph of the sequence graph 1900 represented by these nodes may be extracted, and the topological order of the nodes may be calculated based on this subgraph to provide the ordered sequence of these labels represented in the pipeline skeleton 2002.
[0162] In some embodiments, query 306 may be utilized to search through existing ML projects stored in ML project corpus 104. This search may include generating a label vector and generating weights from the quality metrics of the existing ML projects for each of the existing ML projects stored in ML project corpus 104. Next, this search may include a Weighted Set Cover (WSC) problem based on those weights and label vectors from the existing ML projects stored in ML project corpus 104, including solving the WCS problem to select a set of existing ML projects that together contain all of the labels in the ordered label set. For example, given a set of cleaned candidate compute notebooks: J = {J1, J2,..., J n} that collectively includes semantic labels derived from region: U = {s1, s2,..., s i , s j ,..., s k} and a set of required semantic labels: R = {s m}, the search may be formulated to select the following subset that together contains all of the semantic labels in R:
Number
Number
[0163] After the search is complete, search results such as related existing ML projects 110 (e.g., in pipeline merge 308) may be merged to generate a new ML pipeline 112 for a new ML project 310. This pipeline merge 308 may include the step of incorporating all the functional blocks of the new ML project (e.g., corresponding to an ordered set of labels) from the set of ML pipelines 111 of the related existing ML project 110. For example, as disclosed in FIG. 20, if the related existing ML projects 110 labeled "Mushroom Classification", "WorldHappinessReport2019", and "Cardio" in Table 2050 are represented by three ML projects, each of the functional blocks in the pipeline skeleton 2002 may be incorporated from the functional blocks of these three ML projects. Since the ML project labeled "Mushroom Classification" has most of the required functional blocks, it may be treated as the main ML project, while the remaining functional blocks may be incorporated from the ML project "World Happiness Report 2019", which may be treated as an auxiliary ML project. In some embodiments, if the same label exists in multiple auxiliary compute notebooks, one of the compute notebooks may be selected (e.g., based on quality, randomly, etc.). For example, FIG. 20 shows a case where the ML project corpus 103 includes a total of three ML projects, and the search (e.g., solved through the WSC problem disclosed herein) reads the first two ML projects to fit the pipeline skeleton 2002 in ten bn.
[0164] Pipeline merge 308 may further include steps of adapting functional blocks of a new ML pipeline for a new ML project. This adaptation may include resolving conflicts of various names or object names (e.g., adapting names based on program analysis) and enabling the new ML pipeline 112 to execute a new ML task 108 with a new dataset 106.
[0165] Changes, additions, or omissions may be made to the sequence graph 1900, pipeline skeleton 2002, and table 2050 without departing from the scope of the present disclosure. For example, each of the sequence graph 1900, pipeline skeleton 2002, and table 2050 may include fewer or more components than those shown in FIGS. 19 and 20.
[0166] As described above, the embodiments described herein may include the use of special-purpose or general-purpose computers that include various computer hardware or software modules, as will be discussed in more detail below. Further, as described above, the embodiments described in the present disclosure may be implemented using a computer-readable medium having stored computer-executable instructions or data structures.
[0167] As used in this disclosure, the terms "module" or "component" may refer to a specific hardware implementation configured to perform the operations of the module or component, and / or a software object or software routine that can be stored and / or executed by general-purpose hardware of a computing system (e.g., a computer-readable medium, a processing device, etc.). In some embodiments, components, modules, engines, and services different from those described in this disclosure may be implemented as objects or processes (e.g., separate threads) running on a computing system. Although some of the systems and methods described in this disclosure are described as being implemented generally in software (stored in and / or executed by general-purpose hardware), dedicated hardware implementations or combinations of software and dedicated hardware implementations are also possible and contemplated. In this description, a "computing entity" may be any computing system described above in this disclosure, or any module or combination of modules running on a computing system.
[0168] The terms used in this disclosure and particularly in the appended claims (e.g., the body of the appended claims) are generally intended to be terms of a "broad" nature (e.g., the term "comprising" should be interpreted as "comprising, but not limited to," the term "having" should be interpreted as "having, but not limited to," etc.).
[0169] Furthermore, if an enumeration of a specific number of introduced claims is intended, such intention shall be explicitly indicated in the claims, and if there is no such enumeration, such intention does not exist. For example, for the purpose of assisting understanding, the following appended claims may include the use of introductory phrases "at least one" and "one or more" to introduce an enumeration of claims. However, the use of such phrases shall not be construed to limit any particular claim that includes such an introduced enumeration of claims to an embodiment that includes only one such enumeration even when the same claim includes the introductory phrase "one or more" or "at least one" and the indefinite article "a" or "an" (e.g., "a" and / or "an" should be construed to mean "at least one" or "one or more"). That is, the same applies to the use of the definite article used to introduce an enumeration of claims.
[0170] Furthermore, when an enumeration of a specific number of introduced claims is explicitly recited, one of ordinary skill in the art should understand that such enumeration should be construed to mean at least the recited number (e.g., a recitation of "two enumerations" without other modifiers means at least two enumerations, or two or more enumerations). Further, in examples where a recitation such as "at least one of A, B, and C, etc." or "one or more of A, B, and C, etc." is used, typically such a configuration is intended to include A alone, B alone, C alone, A and B together, A and C together, B and C together, or A, B, and C together, etc. This interpretation of the phrase "A or B" remains applicable even when the term "A and / or B" is used to sometimes include the possibility of "A" or "B" or "A and B".
[0171] Furthermore, any disjunctive word or phrase representing two or more alternative terms should be understood to contemplate the possibility of including one of the terms, any of the terms, or both terms, regardless of whether in the description, claims, or drawings. For example, the phrase "A or B" should be understood to include the possibility of "A" or "B" or "A and B".
[0172] All examples and conditional language recited herein are intended for the teaching purpose of helping the reader understand the present disclosure and the concepts that the present disclosure contributes to the further development of the technology, and should be construed as not being limited to such specifically recited examples and conditions. Although the embodiments of the present disclosure have been described in detail, various changes, alternatives, and selections can be made without departing from the spirit and scope of the present disclosure.
[0173] In addition to the above embodiments, the following appendices are further disclosed. (Appendix 1) A method comprising: storing an existing machine learning (ML) project in a corpus, wherein the existing ML project includes an ML pipeline having functional blocks; generating a search query for a new ML project based on a new dataset for the new ML project and a new ML task for the new ML project; searching through the existing ML projects stored in the corpus based on the search query to obtain a set of existing ML projects; merging the ML pipelines of the set of existing ML projects to generate a new ML pipeline for the new ML project, wherein the new ML pipeline has functional blocks incorporated from the ML pipelines of the set of existing ML projects; adapting the functional blocks of the new ML pipeline for the new ML project so that the new ML pipeline is executable to execute the new ML task with the new dataset; and a method comprising the above steps. (Appendix 2) The method according to Appendix 1, wherein the search query includes a pipeline skeleton including an ordered set of functional blocks for the new ML project. (Appendix 3) The functional blocks of the ML pipeline of the existing ML project include labels, The method further includes a step of training a pipeline skeleton ML model to generate the pipeline skeleton, the pipeline skeleton ML model including a multivariate multi-valued classifier, The step of training the pipeline skeleton ML model, comprises extracting dataset features from a dataset of the existing ML project correlated with a specific label, identifying a set of all labels from the functional blocks of the existing ML project, preparing training data including an input vector having the dataset features and a binary output tuple indicating the presence or absence of each of the set of all labels, training the pipeline skeleton ML model to learn a mapping between the dataset features and the corresponding labels among the set of all labels, The method according to Appendix 2. (Appendix 4) The pipeline skeleton ML model further includes a sequence graph representing a partial order between the functional blocks learned from the training, the sequence graph including, nodes for each label among the set of all labels from the functional blocks of the existing ML project, edges between each pair of a first node and a second node, the first node preceding or being preceded by the second node in one of the existing ML projects, The method according to Appendix 3. (Appendix 5) The step of generating the search query for the new ML project, Mapping the dataset features to an unordered set of labels for the new ML pipeline of the new ML project; Using the partial order represented in the sequence graph to map the unordered set of labels to a query for the search that includes an ordered set of labels as parts of the pipeline skeleton; The method according to appendix 4, comprising the above. (Appendix 6) The step of searching through the existing ML projects stored in the corpus: Generating a label vector for each existing ML project stored in the corpus; Generating weights from the quality metrics of the existing ML projects for each existing ML project stored in the corpus; Constructing a weighted set cover (WSC) problem from the existing ML projects stored in the corpus based on their weights and label vectors; Solving the WSC problem to select a set of existing ML projects that together contain all the labels in the ordered set of labels; The method according to appendix 5, comprising the above. (Appendix 7) The step of merging the set of existing ML projects to generate the new ML pipeline for the new ML project includes incorporating all functional blocks corresponding to the ordered set of labels from the set of existing ML projects. The method according to appendix 6. (Appendix 8) The step of adapting the functional blocks of the new ML pipeline for the new ML project includes resolving conflicts in variable names or object names so that the new ML pipeline is executable for performing the new ML task with the new dataset. The method according to appendix 7. (Appendix 9) One or more non-transitory computer-readable storage media configured to store instructions that, in response to being executed, cause a system to perform operations, the operations comprising: Storing an existing machine learning (ML) project in a corpus, the existing ML project including an ML pipeline having functional blocks; Generating a search query for a new ML project based on a new dataset for the new ML project and a new ML task for the new ML project; Searching, through the existing ML projects stored in the corpus, for a set of existing ML projects based on the search query; Merging the ML pipelines of the set of existing ML projects to generate a new ML pipeline for the new ML project, the new ML pipeline having functional blocks incorporated from the ML pipelines of the set of existing ML projects; Adapting the functional blocks of the new ML pipeline for the new ML project so that the new ML pipeline is executable to perform the new ML task with the new dataset. One or more non-transitory computer-readable storage media comprising the foregoing. (Appendix 10) The one or more non-transitory computer-readable storage media of Appendix 9, wherein the search query includes a pipeline skeleton including an ordered set of functional blocks for the new ML project. (Appendix 11) The functional blocks of the ML pipeline of the existing ML project include labels, The operations further comprising training a pipeline skeleton ML model to generate the pipeline skeleton, the pipeline skeleton ML model including a multivariate multi-valued classifier. The step of training the pipeline skeleton ML model comprises: extracting dataset features from a dataset of an existing ML project correlated with a specific label; identifying a set of all labels from the functional blocks of the existing ML project; preparing training data including an input vector having a binary output tuple indicating the presence or absence of each of the dataset features and the set of all labels; training the pipeline skeleton ML model to learn a mapping between the dataset features and the corresponding labels of the set of all labels; one or more non-transitory computer-readable storage media according to Appendix 10. (Appendix 12) The pipeline skeleton ML model further includes a sequence graph representing a partial order between the functional blocks learned in the training, the sequence graph comprising: nodes for each label of the set of all labels from the functional blocks of the existing ML project; edges between each pair of a first node and a second node, wherein the first node precedes or is preceded by the second node in one of the existing ML projects; one or more non-transitory computer-readable storage media according to Appendix 11. (Appendix 13) The step of generating the search query for the new ML project comprises: mapping the dataset features to an unordered set of labels for the new ML pipeline of the new ML project; using the partial order represented in the sequence graph to map the unordered set of labels to the search query including an ordered set of labels as a part of the pipeline skeleton; One or more non-transitory computer-readable storage media according to appendix 12, including (Appendix 14) The step of searching through the existing ML projects stored in the corpus includes For each existing ML project stored in the corpus, generating a label vector; For each existing ML project stored in the corpus, generating a weight from the quality metrics of the existing ML project; From the existing ML projects stored in the corpus, constructing a weighted set cover (WSC) problem based on their weights and label vectors; Solving the WSC problem to select a set of existing ML projects that together contain all of the labels in the ordered set of labels; One or more non-transitory computer-readable storage media according to appendix 13, including (Appendix 15) The step of merging the set of existing ML projects to generate the new ML pipeline for the new ML project includes taking in all functional blocks corresponding to the ordered set of labels from the set of existing ML projects. One or more non-transitory computer-readable storage media according to appendix 14. (Appendix 16) The step of adapting the functional blocks of the new ML pipeline for the new ML project includes resolving conflicts in variable names or object names so that the new ML pipeline is executable to perform the new ML task with the new dataset. One or more non-transitory computer-readable storage media according to appendix 15. (Appendix 17) A system comprising One or more processors; One or more non-transitory computer-readable storage media configured to store instructions that, responsive to being executed by the one or more processors, cause the system to perform operations, the operations comprising: Storing an existing machine learning (ML) project in a corpus, the existing ML project including an ML pipeline having functional blocks; Generating a search query for a new ML project based on a new dataset for the new ML project and a new ML task for the new ML project; Searching, through the existing ML projects stored in the corpus, for a set of existing ML projects based on the search query; Merging the ML pipelines of the set of existing ML projects to generate a new ML pipeline for the new ML project, the new ML pipeline having functional blocks incorporated from the ML pipelines of the set of existing ML projects; Adapting the functional blocks of the new ML pipeline for the new ML project such that the new ML pipeline is executable to perform the new ML task with the new dataset; A system comprising. (Appendix 18) The search query includes a pipeline skeleton including an ordered set of functional blocks for the new ML project, and the functional blocks of the ML pipeline of the existing ML project include labels; The operations further include training a pipeline skeleton ML model to generate the pipeline skeleton, the pipeline skeleton ML model including a multivariate multi-valued classifier; The step of training the pipeline skeleton ML model Steps of extracting dataset features from a dataset of an existing ML project correlated with a specific label, Steps of identifying a set of all labels from the functional blocks of the existing ML project, Steps of preparing training data including an input vector having a binary output tuple indicating the presence or absence of each of the dataset features and the set of all labels, Steps of training the pipeline skeleton ML model to learn a mapping between the dataset features and the corresponding labels among the set of all labels, The system according to Appendix 17, including (Appendix 19) The pipeline skeleton ML model includes a sequence graph representing a partial order between functional blocks learned from the training data, and the sequence graph Nodes for each label among the set of all labels from the functional blocks of the existing ML project, Edges between each pair of a first node and a second node, where the first node precedes or is preceded by the second node in one of the existing ML projects, including The steps of generating the search query for a bona fide new ML project Steps of mapping the dataset features to an unordered set of labels for the new ML pipeline of the new ML project, Steps of mapping to the search query including an ordered set of labels as a part of the pipeline skeleton using the partial order represented in the sequence graph, The system according to Appendix 18, including (Appendix 20) The steps of searching through the existing ML projects stored in the corpus For each existing ML project stored in the corpus, generating a label vector; For each existing ML project stored in the corpus, generating a weight from the quality metrics of the existing ML project; From the existing ML projects stored in the corpus, constructing a weighted set cover (WSC) problem based on their weights and label vectors; Solving the WSC problem to select a set of existing ML projects that together contain all the labels in the ordered set of labels; including The step of merging the set of existing ML projects to generate the new ML pipeline for the new ML project includes incorporating all the functional blocks corresponding to the ordered set of labels from the set of existing ML projects. The step of adapting the functional blocks of the new ML pipeline for the new ML project includes resolving contradictions in variable names or object names so that the new ML pipeline is executable to perform the new ML task with the new dataset. The system according to Appendix 19.
Explanation of Symbols
[0174] 102 OSS ML Project Database 104 ML Project Corpus 106 New Dataset 108 New ML Task 109 Dataset 110 Related Existing ML Projects 111 ML Pipeline 112 New ML Pipeline 114 Curation Module 116 Search Module 120 Adaptation Module
Claims
A method executed by a system, wherein the system stores an existing machine learning (ML) project in a corpus, the existing ML project including an ML pipeline having functional blocks; generates a search query for a new ML project based on a new dataset for the new ML project and a new ML task for the new ML project; searches a set of existing ML projects through the existing ML projects stored in the corpus based on the search query; merges the ML pipelines of the set of existing ML projects to generate a new ML pipeline for the new ML project, the new ML pipeline having functional blocks incorporated from the ML pipelines of the set of existing ML projects; adapts the functional blocks of the new ML pipeline for the new ML project by resolving conflicts in variable names or object names so that the new ML pipeline is executable to execute the new ML task with the new dataset; A method comprising the steps of. **Claim 2** The method according to claim 1, wherein the search query includes a pipeline skeleton including an ordered set of functional blocks for the new ML project. **Claim 3** The functional blocks of the ML pipeline of the existing ML project include labels, The method further includes the step of the system training a pipeline skeleton ML model to generate the pipeline skeleton, the pipeline skeleton ML model including a multivariate multi-valued classifier, The step of training the pipeline skeleton ML model includes extracting dataset features from a dataset of existing ML projects correlated with specific labels; identifying a set of all labels from the functional blocks of the existing ML projects; preparing training data including an input vector having the dataset features and a binary output tuple indicating the presence or absence of each of the set of all labels; training the pipeline skeleton ML model to learn a mapping between the dataset features and corresponding labels among the set of all labels; The method according to claim 2, comprising:
4. The pipeline skeleton ML model further includes a sequence graph representing a partial order among the functional blocks learned in the training, and the sequence graph nodes for each label among the set of all labels from the functional blocks of the existing ML project; an edge between each pair of a first node and a second node, where the first node precedes or is preceded by the second node in one of the existing ML projects; The method according to claim 3, comprising:
5. The step of generating the search query for the new ML project mapping the dataset features to an unordered set of labels for the new ML pipeline of the new ML project; using the partial order represented in the sequence graph to map the unordered set of labels to the search query including an ordered set of labels as a part of the pipeline skeleton; The method according to claim 4, comprising:
6. The step of searching through the existing ML projects stored in the corpus generating a label vector for each existing ML project stored in the corpus; generating a weight from the quality metric of the existing ML project for each existing ML project stored in the corpus; constructing a weighted set cover (WSC) problem from the existing ML projects stored in the corpus based on their weights and label vectors; solving the WSC problem to select a set of existing ML projects that together include all the labels in the ordered set of labels; The method according to claim 5, comprising:
7. The step of merging the set of the existing ML projects to generate the new ML pipeline for the new ML project includes the step of incorporating all functional blocks corresponding to the ordered set of the labels from the set of the existing ML projects, the method according to claim 6.
8. One or more non-transitory computer-readable storage media configured to store instructions, which, in response to being executed, cause a system to perform operations, the operations being Storing an existing machine learning (ML) project in a corpus, the existing ML project including an ML pipeline having functional blocks, Generating a search query for a new ML project based on a new dataset for the new ML project and a new ML task for the new ML project, Searching a set of existing ML projects through the existing ML projects stored in the corpus based on the search query, Merging the ML pipelines of the set of the existing ML projects to generate a new ML pipeline for the new ML project, the new ML pipeline having functional blocks incorporated from the ML pipelines of the set of the existing ML projects, Adapting the functional blocks of the new ML pipeline for the new ML project by resolving conflicts in variable names or object names so that the new ML pipeline is executable to execute the new ML task with the new dataset, One or more non-transitory computer-readable storage media including.
9. The search query includes a pipeline skeleton including an ordered set of functional blocks for the new ML project, the one or more non-transitory computer-readable storage media according to claim 8.
10. The functional blocks of the ML pipeline of the existing ML project include labels, The operations further include training a pipeline skeleton ML model to generate the pipeline skeleton, the pipeline skeleton ML model including a multivariate multi-valued classifier. The step of training the pipeline skeleton ML model comprises: extracting dataset features from a dataset of an existing ML project correlated with a specific label; identifying a set of all labels from the functional blocks of the existing ML project; preparing training data including an input vector having a binary output tuple indicating the presence or absence of each of the dataset features and the set of all labels; training the pipeline skeleton ML model to learn a mapping between the dataset features and the corresponding labels of the set of all labels; One or more non-transitory computer-readable storage media according to claim 9. **Claim 11** The pipeline skeleton ML model further includes a sequence graph representing a partial order between the functional blocks learned in the training, the sequence graph comprising: nodes for each label of the set of all labels from the functional blocks of the existing ML project; edges between each pair of a first node and a second node, where the first node precedes or is preceded by the second node in one of the existing ML projects; One or more non-transitory computer-readable storage media according to claim 10. **Claim 12** The step of generating the search query for the new ML project comprises: mapping the dataset features to an unordered set of labels for the new ML pipeline of the new ML project; using the partial order represented in the sequence graph to map the unordered set of labels to the search query including an ordered set of labels as parts of the pipeline skeleton; One or more non-transitory computer-readable storage media according to claim 11. **Claim 13** The step of searching through the existing ML projects stored in the corpus comprises: generating a label vector for each existing ML project stored in the corpus; generating a weight from the quality metric of the existing ML project for each existing ML project stored in the corpus; From the existing ML projects stored in the corpus, steps of constructing a weighted set cover (WSC) problem based on their weights and label vectors; Steps of solving the WSC problem to select a set of existing ML projects that together contain all of the labels in the ordered set of labels; The one or more non-transitory computer-readable storage media according to claim 12, comprising:
14. The step of generating the new ML pipeline for the new ML project by merging the set of existing ML projects includes the step of incorporating all functional blocks corresponding to the ordered set of labels from the set of existing ML projects. The one or more non-transitory computer-readable storage media according to claim 13.
15. A system, comprising: One or more processors; One or more non-transitory computer-readable storage media configured to store instructions that, in response to being executed by the one or more processors, cause the system to perform operations, the operations including: Storing existing machine learning (ML) projects in a corpus, the existing ML projects including an ML pipeline having functional blocks; Generating a search query for the new ML project based on a new dataset for the new ML project and a new ML task for the new ML project; Searching through the existing ML projects stored in the corpus based on the search query to obtain a set of existing ML projects; Merging the ML pipelines of the set of existing ML projects to generate a new ML pipeline for the new ML project, the new ML pipeline having functional blocks incorporated from the ML pipelines of the set of existing ML projects; Adapting the functional blocks of the new ML pipeline for the new ML project by resolving conflicts in variable names or object names so that the new ML pipeline is executable to perform the new ML task with the new dataset; A system, comprising:
16. The search query includes a pipeline skeleton including an ordered set of functional blocks for the new ML project, and the functional blocks of the ML pipeline of the existing ML project include labels, The operation further includes the step of training a pipeline skeleton ML model to generate the pipeline skeleton, the pipeline skeleton ML model including a multivariate multi-valued classifier, The step of training the pipeline skeleton ML model, extracting dataset features from a dataset of the existing ML project correlated with a specific label, identifying a set of all labels from the functional blocks of the existing ML project, preparing training data including an input vector having the dataset features and a binary output tuple indicating the presence or absence of each of the set of all labels, training the pipeline skeleton ML model to learn a mapping between the dataset features and the corresponding label among the set of all labels, The system according to claim 15, comprising:
17. The pipeline skeleton ML model includes a sequence graph representing a partial order between functional blocks learned from the training data, the sequence graph including: nodes for each label among the set of all labels from the functional blocks of the existing ML project, edges between each pair of a first node and a second node, the first node preceding or being preceded by the second node in one of the existing ML projects, including, The step of generating the search query for the new ML project, mapping the dataset features to an unordered set of labels for the new ML pipeline of the new ML project, mapping to the search query including an ordered set of labels as a part of the pipeline skeleton using the partial order represented in the sequence graph, The system according to claim 16, comprising:
18. The step of searching through the existing ML projects stored in the corpus, For each existing ML project stored in the corpus, generating a label vector; For each existing ML project stored in the corpus, generating weights from the quality metrics of the existing ML project; Constructing a weighted set cover (WSC) problem from the existing ML projects stored in the corpus based on their weights and label vectors; Solving the WSC problem to select a set of the existing ML projects that together contain all of the labels in the ordered set of labels; comprising; The step of merging the set of existing ML projects to generate the new ML pipeline for the new ML project includes the step of incorporating all of the functional blocks corresponding to the ordered set of labels from the set of existing ML projects. The system according to claim 17.
Citation Information
Patent Citations
Data orchestration platform management
JP2019133610A
Recipe generation for improved modeling
US20190018821A1
Method and system for flexible pipeline generation
WO2019144240A1
Method for identifying project component, and reusability detection system therefor
WO2020026680A1