Method and apparatus, and device for generating machine learning flowcharts

By configuring modules, the user can be guided to generate machine learning flowcharts, which solves the problem of the lack of professional teams in the enterprise, and realizes low-threshold, efficient machine learning model construction and flexible customization capabilities, which are suitable for various information processing equipment.

CN112884166BActive Publication Date: 2025-06-24LENOVO (BEIJING) LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110351770.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-03-31
Publication Date
2025-06-24
Estimated Expiration
2041-03-31

AI Technical Summary

Technical Problem

The lack of professional artificial intelligence technology teams in enterprises makes it difficult to efficiently build machine learning flowcharts, and the existing platforms are not flexible or lack intuitive, which increases learning costs and construction difficulties.

Method used

Through the configuration module, the user can be guided to configure parameters and automatically generate a machine learning flowchart, including the first configuration module specifying target tasks and the second configuration module configuration related parameters, and gradually guide the user to complete the generation and editing of the flowchart.

Benefits of technology

It lowers the threshold for using machine learning, broadens the population, improves construction efficiency and accuracy, provides flexible customization capabilities, and meets the needs of different users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112884166B_ABST
    Figure CN112884166B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a method, apparatus, and device for generating a machine learning flow chart. The method includes: determining a first parameter configured based on a first configuration module, where the first parameter is used to specify a target task, and the first configuration module includes at least one candidate task; determining a second configuration module associated with the target task; determining a second parameter configured based on the second configuration module; and generating a machine learning flow chart according to the first parameter and the second parameter, where the machine learning model generated by the machine learning flow chart is used to execute the target task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to artificial intelligence technology, including but not limited to methods, devices, and equipment for generating machine learning flowcharts. Background Art

[0002] In recent years, with the rapid development of artificial intelligence technology, a large number of enterprises hope to apply artificial intelligence technology to their industries to analyze the production data accumulated over a long time, promote the informatization and intelligent transformation of the enterprises, and improve production efficiency and benefits. However, most enterprises do not have a professional artificial intelligence technology team, and forming such a team takes a long time and is costly.

[0003] Therefore, there is an urgent need in the current market for a machine learning platform with low-threshold and highly customizable modeling capabilities, enabling artificial intelligence technology experts (senior users) or industry experts without artificial intelligence background knowledge (beginner users) to efficiently construct machine learning flowcharts, reducing the usage threshold and cost of machine learning, and making artificial intelligence technology truly benefit all industries. Summary of the Invention

[0004] In view of this, the method, device, and equipment for generating machine learning flowcharts provided by embodiments of the present application use a configuration module to guide users to perform simple parameter configuration and automatically generate machine learning flowcharts, thereby reducing the usage threshold and cost of machine learning. The method, device, and equipment for generating machine learning flowcharts provided by embodiments of the present application are implemented as follows:

[0005] The method for generating a machine learning flowchart provided by an embodiment of the present application includes: determining a first parameter guided by a first configuration module for specifying a target task, where the first configuration module includes at least one candidate task; determining a second configuration module associated with the target task; determining a second parameter guided by the second configuration module; and generating a machine learning flowchart based on the first parameter and the second parameter, where the machine learning model generated by the machine learning flowchart is used to perform the target task.

[0006] The device for generating a machine learning flowchart provided by an embodiment of the present application includes: a determination unit for determining a first parameter guided by a first configuration module for specifying a target task, where the first configuration module includes at least one candidate task; the determination unit is further configured to determine a second configuration module associated with the target task; the determination unit is further configured to determine a second parameter guided by the second configuration module; and a generation unit for generating a machine learning flowchart based on the first parameter and the second parameter, where the machine learning model generated by the machine learning flowchart is used to perform the target task.

[0007] The electronic device provided by the embodiment of the present application includes a memory and a processor. The memory stores a computer program that can run on the processor, and when the processor executes the program, the method described in the embodiment of the present application is implemented.

[0008] In the embodiment of the present application, the user is guided to configure the first parameter based on the first configuration module, and the first parameter is used to specify the target task; wherein, the first configuration module includes at least one candidate task; the second configuration module associated with the target task is determined; the user is guided to configure the second parameter based on the second configuration module; according to the first parameter and the second parameter, a machine learning flow chart is generated, and the machine learning model generated by the machine learning flow chart is used to execute the target task. In this way, on the one hand, guiding the user to configure the parameters for generating the machine learning flow chart step by step can greatly reduce the threshold for constructing and using the machine learning flow chart, effectively broadening the user group; on the other hand, since the user can obtain the machine learning flow chart through simple configuration, there is no need for the user to drag and connect one node by one node, effectively improving the efficiency of building the machine learning model. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The drawings here are incorporated into the specification and form a part of this specification. These drawings show embodiments consistent with the present application and are used together with the specification to illustrate the technical solutions of the present application.

[0010] Figure 1 It is a schematic implementation flow chart of the method for generating the machine learning flow chart provided by the embodiment of the present application;

[0011] Figure 2 It is a schematic diagram of the first configuration interface provided by the embodiment of the present application;

[0012] Figure 3 It is a schematic diagram of the configuration window provided by the embodiment of the present application;

[0013] Figure 4 It is a schematic diagram of the second configuration interface provided by the embodiment of the present application;

[0014] Figure 5 It is the editing interface of the machine learning flow chart provided by the embodiment of the present application;

[0015] Figure 6 It is another schematic implementation flow chart of the method for generating the machine learning flow chart provided by the embodiment of the present application;

[0016] Figure 7 It is yet another schematic implementation flow chart of the method for generating the machine learning flow chart provided by the embodiment of the present application;

[0017] Figure 8It is another schematic implementation flowchart of the method for generating a machine learning flowchart provided by an embodiment of the present application;

[0018] Figure 9 It is a schematic diagram of another second configuration interface provided by an embodiment of the present application;

[0019] Figure 10 It is a machine learning flowchart automatically generated according to user configuration provided by an embodiment of the present application;

[0020] Figure 11 It is a schematic structural diagram of a device for generating a machine learning flowchart according to an embodiment of the present application;

[0021] Figure 12 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0022] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will further describe the specific technical solutions of the present application in detail with reference to the accompanying drawings in the embodiments of the present application. The following embodiments are used to illustrate the present application but are not intended to limit the scope of the present application.

[0023] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0024] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0025] It should be noted that the terms "first / second / third" involved in the embodiments of the present application are used to distinguish similar or different objects, and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein.

[0026] The embodiments of the present application provide a method for generating a machine learning flowchart. This method is applied to an electronic device, which can be various types of devices with information processing capabilities during implementation. For example, the electronic device may include a personal computer, a laptop computer, a server, a cluster server, a mobile phone, or a tablet computer, etc. The functions implemented by this method can be realized by a processor in the electronic device calling program code. Of course, the program code can be stored in a computer storage medium. It can be seen that the electronic device includes at least a processor and a storage medium.

[0027] Figure 1 This is a schematic diagram of the implementation process of the method for generating a machine learning flowchart provided by an embodiment of the present application. As Figure 1 shown, the method may include the following steps 101 to 104:

[0028] Step 101, determine a first parameter configured based on a first configuration module, where the first parameter is used to specify a target task; wherein, the first configuration module includes at least one candidate task.

[0029] A so-called machine learning flowchart refers to a flowchart that includes relevant nodes for generating a machine learning model, that is, an abstract structure of the code of the machine learning model. After constructing the machine learning flowchart, the electronic device runs the code corresponding to the machine learning flowchart to obtain the machine learning model.

[0030] The target task may be a task that the user wants to complete through the machine learning model, and the candidate task may be a task that the machine learning model can complete. For example, classification, clustering, regression, or other industry-related tasks.

[0031] In some embodiments, the first configuration module may include a first configuration interface. When the user wants to generate a machine learning flowchart, the electronic device displays the first configuration interface including at least one candidate task. Figure 2 This is a schematic diagram of the first configuration interface provided by an embodiment of the present application. As Figure 2 shown, the first configuration interface 21 includes candidate tasks "clustering", "regression", and / or "classification". Each candidate task is correspondingly provided with a selection button 211. After the electronic device detects the user's click on the selection button 211, the first parameter is determined, that is, the candidate task corresponding to the selection button that accepts the click operation.

[0032] In some embodiments, the first configuration interface 21 further includes information corresponding to each candidate task. For example, by the user clicking Figure 2 the "Learn More" button 213 in, a note description of the candidate task corresponding to the clicked "Learn More" button 213 pops up. The note description at least includes: the applicable scenario of the candidate task, a scenario example, relevant parameters of the scenario example, and an example of the corresponding machine learning flowchart. The user can determine which selection button 211 to click based on this information.

[0033] In some embodiments, the first configuration interface 21 further includes an icon 212 corresponding to each candidate task. The user can quickly understand the usage scenario of the candidate task through the icon 212.

[0034] In some embodiments, the first configuration interface 21 further includes a brief introduction to each candidate task. For example, the brief introduction to clustering includes various clustering algorithms, and different clustering algorithms are applicable to data with different densities and distributions; the brief introduction to regression includes mainstream regression algorithms from which the most suitable algorithm for the data can be selected; the brief introduction to classification includes the mainstream classification process, and, preferably, one of them will be displayed in the results.

[0035] Of course, the first configuration module can be various types of modules, and the way of interaction between the first configuration module and the user is not limited. In some embodiments, the first configuration module may include a configuration window through which the user can input the target task to be completed. For example, Figure 3 is a schematic diagram of the configuration window provided by an embodiment of the present application, as Figure 3 shown, the configuration window 31 includes an input box 311. The user inputs the word "clustering" in the input window 311, and at this time, the electronic device determines that the target task is "clustering".

[0036] Step 102, determine the second configuration module associated with the target task.

[0037] After detecting the user's click on the selection button and determining the target task, the electronic device jumps to the second configuration module associated with the target task. For example, after determining that the target task is "clustering", it jumps to the second configuration module associated with "clustering", and the second configuration module is used to guide the user to configure the second parameters for generating the machine learning flow chart of "clustering". Similarly, the second configuration module can be various types of modules, and the way of interaction between the second configuration module and the user is also not limited. In some embodiments, the second configuration module may include a second configuration interface, and the second configuration interface is used to guide the user to configure the second parameters for generating the machine learning flow chart.

[0038] Step 103, determine the second parameters configured based on the guidance of the second configuration module.

[0039] For example, Figure 4 is a schematic diagram of the second configuration interface provided by an embodiment of the present application, as Figure 4 shown, the second configuration interface 41 includes a data upload sub-module 411, a data analysis sub-module 412, and a parameter configuration sub-module 413. Among them, the data analysis module 412 is used to analyze the distribution of the feature values of the selected target attributes in the imported sample dataset, so that the user can configure the first information and the second information according to the distribution of the feature values of the target attributes; among them, the first information is used to indicate the processing method of the feature values of the specified attributes, and the second information is used to indicate the prediction target. In the embodiments of the present application, steps 603 to 605 of the following embodiments can be executed to implement step 103, and for the sake of avoiding repetition, it will not be elaborated here.

[0040] Step 104: Generate a machine learning flow chart according to the first parameter and the second parameter. The machine learning model generated by the machine learning flow chart is used to execute the target task.

[0041] After the user completes the configuration, an automatically generated machine learning flow chart can be obtained. The generation mechanism of the machine learning flow chart is summarized from the modeling best practices of a large number of machine learning solutions. This application embodiment does not limit it here.

[0042] In some embodiments, the machine learning flow chart includes nodes and directed connections between the nodes. Among them, the nodes represent the code modules of the machine learning model and are executable units that can complete independent tasks (such as data import, data preprocessing, feature engineering, and prediction, etc.). The directed connections between the nodes represent the data flow when the machine learning model generated according to the machine learning flow chart executes tasks.

[0043] It can be understood that in the embodiments of this application, by gradually guiding the user to configure the parameters for generating the machine learning flow chart, the threshold for constructing and using the machine learning model can be greatly reduced, effectively broadening the user group. And because the user can obtain the machine learning flow chart through simple configuration, the user does not need to drag and connect nodes one by one, thereby effectively improving the efficiency of constructing the machine learning flow chart.

[0044] For step 104, when the electronic device is implemented, it can generate a machine learning flow chart in non-editable mode, or it can also generate a machine learning flow chart in editable mode. Regarding whether to generate a machine learning flow chart in editable mode, the user can pre-configure it, for example, configure the parameter value for whether to generate an editable mode based on the second configuration module. Of course, in some other embodiments, the second configuration module may not provide the configuration function for the parameter value of whether to generate an editable mode, that is, the electronic device defaults to generate a machine learning flow chart in non-editable mode or editable mode. In the case of defaulting to generate a machine learning flow chart in non-editable mode, in some embodiments, after step 104, the method further includes: receiving an editable instruction; where the editable instruction is used to indicate editing the machine learning flow chart; in response to the editable instruction, generating the editable machine learning flow chart.

[0045] Thus, in the editable mode, the user can edit the machine learning flowchart according to their own needs. For example, modify the node parameters of the machine learning flowchart, add or delete nodes, or modify the connection relationships between nodes. There is no limitation on the way for the user to trigger the entry into the editable mode. For example, the user can click or double-click on any blank area of the interface where the machine learning flowchart is located, and at this time the electronic device determines that it has received an editable instruction; or, the user can input a specific gesture or click a specific button to trigger the electronic device to generate an editable machine learning flowchart.

[0046] Figure 5 This is the editing interface of the machine learning flowchart provided by the embodiment of the present application. As Figure 5 shown, when a node in the machine learning flowchart is selected, the parameters that can be edited for the node will be displayed on the interface of the editable flowchart, which are used to edit the node parameters of the selected node. For example, when the node "Principal Component Analysis" is selected, the right sidebar will display parameters that can be edited such as output mode, processing type, number of components to retain, singular value decomposition, allowable deviation, iteration power, random state, data items, etc. Thus, when the generated machine learning flowchart does not meet the user's requirements, the user can select the machine learning flowchart for secondary editing to flexibly achieve custom requirements.

[0047] Figure 6 This is another schematic diagram of the implementation process of the method for generating a machine learning flowchart provided by the embodiment of the present application. As Figure 6 shown, the method may include the following steps 601 to step 606:

[0048] Step 601, determine the first parameter guided by the first configuration module for configuration, where the first parameter is used to specify the target task; wherein, the first configuration module includes at least one candidate task.

[0049] In some embodiments, information corresponding to each candidate task is set on the first configuration module; correspondingly, in response to a user operation, further display the annotation description of the candidate task corresponding to the target information specified by the user operation; wherein, the annotation description is used to explain the candidate task.

[0050] For example Figure 2 shown, the information corresponding to each candidate task may be the relevant configuration button for each candidate task (for example Figure 2 the "Learn More" button 213 in ), when the electronic device detects a click operation on this configuration button by the user, in response to this operation, determine the target information, and pop up the annotation description of the candidate task corresponding to the target information. The annotation description at least includes: the applicable scenario of the candidate task, a sample scenario, the relevant parameters of the sample scenario, and an example of the corresponding machine learning flowchart.

[0051] For example, when the electronic device detects a click operation on the configuration button 213 corresponding to the candidate task of "classification", the popped-up scenario can be the definition and explanation of classification. The scenario example can be classifying pictures into different animals. The scenario example of the parameters to be configured for completing this picture classification task, and the corresponding example of the machine learning flow chart can be the machine learning flow chart for classifying pictures. Thus, through a complete description guidance, a configuration example and result are given to the user, guiding low-level users to finally obtain a satisfactory machine learning model.

[0052] Step 602: Jump to the second configuration module associated with the target task;

[0053] Step 603: Determine at least one of the sample data sets configured based on the second configuration module.

[0054] In some embodiments, the user is guided to configure the data through the data upload sub-module 411 on the second configuration module. Hint information is set on the data upload sub-module 411. In some embodiments, the sample data set may include a training data set and / or a test data set. Correspondingly, the hint information may be "Please upload the training data and test data according to the selected problem."

[0055] Step 604: Analyze the eigenvalue distribution of the selected target attribute in the sample data set; wherein, the eigenvalue distribution is used to guide the configuration of the processing method for the eigenvalues of the target attribute.

[0056] The eigenvalue distribution is diverse. For example: the average value, median, standard deviation, unique value, maximum value, minimum value, and / or missing value of the eigenvalues of the target attribute, etc. Among them, the unique value represents the degree of repetition of the eigenvalues under the target attribute, that is, how many values are not the same as other values, and the missing value represents no data. For example, in the data column corresponding to the target attribute, that is, in the eigenvalue column, how many positions have no data values. The electronic device can analyze the eigenvalue distribution of the selected target attribute through the data analysis sub-module 412 and display it to the user.

[0057] Furthermore, the user can configure the processing method for the feature values of the target attribute according to the distribution of the feature values. For example, if there are many missing feature values for the target attribute, that is, when the missing value is large, it indicates that the data column of the target attribute is not sufficient to train a machine learning model. Then the user can configure this target attribute as a prohibited attribute, that is, the processing method for the feature values of the target attribute is to prohibit the use of the feature values of the target attribute. Another example is that the user configures the feature values of the target attribute as the average value, that is, the processing method for the feature values of the target attribute is to use the average value to replace these original feature values. In this way, when generating the machine learning flow chart, the feature values of the target attribute used are the average values, rather than the original feature values.

[0058] In some other embodiments, the user can perform more advanced configurations according to the distribution of the feature values, such as some configurations related to the model algorithm. For example, according to the distribution of the feature values of a target attribute, determine what kind of preprocessing to perform on the feature values of the target attribute and configure it in the subsequent configuration items.

[0059] Step 605: Determine the first information and the second information configured based on the second configuration module; wherein, the first information is used to indicate the processing method for the feature values of the specified attribute, and the second information is used to indicate the prediction target.

[0060] In the embodiments of the present application, the processing method for the feature values of the specified attribute can be various. For example, prohibit the use of this attribute, that is, do not use the feature values under this attribute when generating the machine learning flow chart; another example is to replace the original feature values of this attribute with the median of the feature values of this attribute, that is, use the median under this attribute to train relevant parameters when generating the machine learning flow chart. The prediction target is what the machine learning flow chart ultimately aims to achieve. For example, predicting income in a census. The user can configure a certain attribute in the sample dataset as the prediction target.

[0061] The user can configure the first information and the second information through the parameter configuration sub-module 413. It should be understood that steps 604 and 605 are optional steps. That is to say, in the actual configuration process, the user can only configure the sample dataset; or after determining the first information and the second information according to experience, directly configure them on the parameter configuration sub-module 413, that is, only configure the sample dataset, the first information and the second information; or only configure the sample dataset. After the electronic device analyzes the distribution of the feature values of the selected target attribute, do not configure the first information and the second information, but perform more advanced configurations according to the distribution of the feature values.

[0062] Step 606: Generate a machine learning flow chart according to the first parameter and the second parameter. The machine learning model generated by the machine learning flow chart is used to execute the target task.

[0063] Figure 7 This is another schematic diagram of the implementation process for the method of generating a machine learning flow chart provided by the embodiments of the present application. As Figure 7 shown, the method may include the following steps 701 to step 706:

[0064] Step 701, determine a first parameter configured based on the guidance of a first configuration module, where the first parameter is used to specify a target task; wherein, the first configuration module includes at least one candidate task; information corresponding to each candidate task is set on the first configuration module;

[0065] Step 702, jump to a second configuration module associated with the target task;

[0066] Step 703, determine at least one of the sample data sets configured based on the second configuration module;

[0067] Step 704, analyze the eigenvalue distribution of the selected target attribute; wherein, the eigenvalue distribution is used to guide the configuration of the processing method for the eigenvalues of the target attribute;

[0068] Step 705, determine a first piece of information and a second piece of information configured based on the second configuration module; wherein, the first piece of information is used to indicate the processing method for the eigenvalues of the specified attribute, and the second piece of information is used to indicate the prediction target.

[0069] Step 706, generate a machine learning flow chart according to the first parameter, the second parameter, and a preset default parameter, and the machine learning model generated by the machine learning flow chart is used to execute the target task.

[0070] The preset default parameter refers to the default parameter value determined according to the best practice of current machine learning model building. For example, some professional parameters related to the model require users who are relatively familiar with the model or algorithm to be able to set them. For example, random parameters, the number of training rounds (i.e., the number of times the machine learning model is trained), the image preprocessing method, etc. These parameters are preset according to the best practice of machine learning model building and do not require users to set them. In this way, users only need to complete a small number of required configurations to obtain results, reducing the usage threshold of the embodiments of the present application.

[0071] In some embodiments, the second parameter further includes a parameter used to indicate the generation method of the machine learning flow chart; in the case where the indicated generation method is the dynamic generation method, generate multiple different versions of the machine learning flow chart according to the first parameter, the second parameter, and the preset default parameter.

[0072] That is to say, when the indicated generation method is dynamic generation, multiple machine learning flowcharts with different processing complexities or different computational complexities can be generated. For example, the nodes and the directed connections between the nodes in the generated machine learning flowcharts are the same, but the complexity of the algorithms used in the nodes is different, or the latter machine learning flowchart adds nodes on the basis of the previous machine learning flowchart to perform more refined processing on the data.

[0073] In this way, on the one hand, for ordinary users, the user can select the one with better effect and matching the current data characteristics as the machine learning flowchart for constructing the machine learning model; on the other hand, for senior users, the user can select the flowchart that meets their requirements as the machine learning flowchart for constructing the machine learning model; in this way, the editing operations of senior users can be simplified; it can be seen that a machine learning flowchart satisfactory to the user can be obtained through one configuration.

[0074] In some embodiments, after generating multiple different versions of the machine learning flowchart, the identification keys of the multiple different versions of the machine learning flowchart are presented in the first window; the target identification key for receiving the selection operation is determined; in response to the selection operation, the target machine learning flowchart corresponding to the target identification key and the performance parameters of the target machine learning flowchart are presented in the second window.

[0075] The performance parameters of the target machine learning flowchart refer to the performance parameters achieved according to the preset performance evaluation indicators when the electronic device runs the target machine learning model generated by the target machine learning flowchart. For example, the performance parameters of a machine learning model for image recognition can be the recognition accuracy, etc. In this way, a selection reference is provided for the user, the problem of the user's decision-making difficulty due to lack of professional knowledge is solved, the construction and use thresholds of the machine learning model are further reduced, and the user population is effectively broadened.

[0076] In some embodiments, after step 706, the method further includes: after receiving an editable instruction, the electronic device jumps to the editable interface of the target machine learning flowchart according to the selected target machine learning flowchart indicated by the editable instruction, so that the user can edit the machine learning flowchart. In this way, when none of the multiple machine learning flowcharts can satisfy the user, the user can select the target machine learning flowchart closest to the expected goal for editing, so as to obtain a machine learning flowchart satisfactory to the user.

[0077] In recent years, with the rapid development of artificial intelligence technology, a large number of enterprises hope to apply artificial intelligence technology to the industry, analyze the production data accumulated over the long term for the enterprise, promote the informatization and intelligent transformation of the enterprise, and improve production efficiency and revenue. However, most enterprises do not have a professional artificial intelligence technology team, and forming such a team takes a long time and is costly.

[0078] Therefore, the market urgently needs a machine learning platform with low-threshold and high-customization modeling capabilities, enabling artificial intelligence experts (senior users) or industry experts without artificial intelligence background knowledge (beginner users) to efficiently build machine learning models, reduce the usage threshold and cost of machine learning, and make artificial intelligence technology truly benefit all industries, allowing each enterprise to enjoy the dividends brought by artificial intelligence technology.

[0079] Some related machine learning and artificial intelligence platforms only use the form of dragging algorithm nodes to construct a machine learning flow chart for modeling. A machine learning flow chart is a flow chart that contains the relevant nodes for generating a machine learning model and is an abstract structure for generating the code of the machine learning model. After the machine learning flow chart is constructed, running the code corresponding to the machine learning flow chart can obtain the machine learning model. Users can build a machine learning model by independently selecting the required nodes, configuring the parameters of the nodes one by one, and then connecting them according to the data relationship between the nodes to form a flow chart.

[0080] The disadvantages of this solution are as follows: On the one hand, if users do not have a strong relevant knowledge background, they may not know which modules to use and how to connect the modules into an effective machine learning process. Beginner users are difficult to successfully construct the required machine learning process, and the learning cost is too high. On the other hand, for senior users, although this method provides high flexibility, the process of dragging and connecting lines takes a long time and is prone to errors and omissions, making it difficult to quickly and accurately build a machine learning model.

[0081] Another part of machine learning and artificial intelligence platforms do not require users to draw a machine learning flow chart, but use a form-filling method to configure the machine learning process and directly display the modeling results after model training.

[0082] The disadvantages of this solution are as follows: On the one hand, since such products do not display the specific machine learning flow chart, the model construction process is like a black box, and it is difficult for users to intuitively understand the modeling process. On the other hand, although this modeling method is convenient, it lacks flexibility, and users cannot adjust the parameters of the model or the structure of the flow chart according to specific requirements.

[0083] Based on this, the exemplary application of the embodiments of the present application in an actual application scenario will be described below.

[0084] In the embodiments of the present application, through step-by-step guided configuration, an editable machine learning flow chart is generated and displayed.

[0085] In the embodiments of the present application, according to simple configurations input by the user, such as the type of problem to be solved, training / test data, features to be predicted, and other parameter information, an automatically generated machine learning flow chart can be obtained. Optionally, the user can edit and adjust the flow chart as needed. The obtained flow chart can be used for model training and use.

[0086] In the embodiments of the present application, on the one hand, it enables users (regardless of whether they have a machine learning technical background) to easily apply machine learning technology with only simple configurations, thus greatly reducing the threshold for building and using machine learning models and effectively expanding the user population. On the other hand, since users can obtain a machine learning flow chart through simple configurations, they do not need to drag and connect modules one by one, thus effectively improving the efficiency of building machine learning models. Also, it avoids the problem that users cannot use the machine learning model properly due to selecting incorrect modules or connections, thereby improving the accuracy of building machine learning models. On yet another hand, users can edit the flow chart based on the flow chart obtained through configuration to flexibly achieve customized requirements. Therefore, the technical solution provided by the embodiments of the present application has high customization capabilities.

[0087] In the embodiments of the present application, user configurations are obtained through step-by-step guidance, and an editable machine learning flow chart is generated and displayed;

[0088] Figure 8 is another schematic implementation flow chart of the method for generating a machine learning flow chart provided by the embodiments of the present application. As Figure 8 shown, the method may include the following steps 801 to step 803:

[0089] Step 801, display a graphical interface and prompt the user for input;

[0090] Step 802, detect the user input and obtain the machine learning process configuration;

[0091] Step 803, generate and display an editable flow chart according to the input configuration parameters.

[0092] The following will specifically introduce the implementation process of this method from several key points:

[0093] Configure through step-by-step guidance;

[0094] Guide the user to configure the machine learning process through a graphical interface. The configuration content includes but is not limited to:

[0095] (1) Type of problem to be solved (such as classification, clustering, regression, or other industry-related task types);

[0096] (2) Training / test data;

[0097] (3) Feature to be predicted;

[0098] Among them, the feature to be predicted refers to the data item or label where the predicted answer is located. For example, when predicting whether an image is a cat, there is a column of data indicating "yes" or "no", and the name of this column of data is the feature to be predicted; or when predicting which category a certain data belongs to, this category is the feature to be predicted.

[0099] (4) Other parameter information, etc.

[0100] Among them, some configurations may be displayed or hidden according to the previous configurations, so as to play a role in guiding the user step by step. For example, the user first selects the type of problem to be solved, and then in the interface for displaying configuration data according to the problem type. Suppose the task type he selects is classification, then the data he needs to configure later is different from that of clustering. That is to say, first configure the type of problem to be solved, and then automatically determine the configuration data according to the task type; in this way, from the interaction level, gradually present the configuration data that the user needs to fill in, and gradually guide the user to fill in the configuration data, improving the user experience and avoiding problems such as the user losing patience and having a poor experience due to too many required information or too long a form during the process of configuring information.

[0101] Some configurations will set default values in advance according to the best practices of machine learning modeling. The user only needs to complete a small amount of required configurations to obtain the result, which reduces the usage threshold of this method. For example, some configurations are professional parameters related to the model, and only users who are relatively familiar with the model or algorithm have the ability to set them. These professional parameters are set according to empirical values and do not require the user to set them. For example, random parameters, number of training rounds, image preprocessing methods, etc.

[0102] As Figure 2 shown, the schematic diagram of the first configuration interface shows multiple candidate tasks, a brief introduction to each candidate task, and related configuration buttons, Select buttons, and icons corresponding to each candidate task.

[0103] For example, candidate tasks include Clustering, Regression, and Classification. The explanation for Clustering is that it includes various clustering algorithms, and different clustering algorithms are applicable to data with different densities and distributions; the explanation for Regression is that it includes mainstream regression algorithms, and one that fits the data best can be selected from them; the explanation for Classification is that it includes mainstream classification pipelines, and the best one will be demonstrated in the results. Additionally, users can determine the usage scenarios of candidate tasks by observing the icons and then click the corresponding selection buttons.

[0104] The relevant configuration button is "Learn more". When the electronic device detects a click operation on this configuration button, further explanations and more detailed descriptions of the candidate task corresponding to this configuration button will pop up, including: descriptions of the applicable scenarios of the candidate task, examples and configuration examples of the examples, and result examples. For example, when the electronic device detects a click operation on the "Learn more" button under Classification, the example that pops up can be an example of classifying pictures into different animals, the parameters to be configured and the corresponding results. In this way, through a complete instruction guide, a configuration example and results are given to the user to guide low-level users to ultimately obtain a satisfactory machine learning flow chart and thus a satisfactory machine learning flow chart.

[0105] After the electronic device detects a click on the Select button, it switches from the first configuration interface to the second configuration interface (i.e., part of the content of the second configuration module). Figure 9 This is another schematic diagram of the second configuration interface provided by the embodiment of the present application, as Figure 9 shown. This interface shows the window, buttons for configuring test data, and corresponding explanations; users can click the corresponding buttons according to the interface prompts of this interface to complete the configuration of the test data.

[0106] Configuring the test data window may include an Upload data window, a Check data type window, and a Setting window. The explanatory note corresponding to the Upload data window is "Please upload training data and test data (i.e., an example of a sample data set) based on the problem selected". Among them, the Upload data window includes upload windows for Training data and Test data, and there are local file and data source buttons on this window; the explanatory notes corresponding to the local file and data source buttons are "Drop the data file here or import form".

[0107] The Upload data window also has a data insights window for presenting the eigenvalue distribution of the selected data column in response to the user's selection of the data column. For example Figure 9A list for displaying the content of a sample data set. Each column in the table presents the feature values of the corresponding attribute. For example, the column of user ID presents the feature values of the attribute "user ID"; another example is that the column of gender presents the feature values of the attribute "gender". Understandably, the distribution of feature values helps users understand the distribution of feature values of the corresponding attribute. Users can configure the processing method for the feature values of this attribute according to these feature value distributions, such as prohibiting the use of this attribute or using a certain value (such as the average value, median, etc.) to replace it, or performing more advanced configurations, such as configurations related to model algorithms. Advanced users can make some decisions, for example, configure what preprocessing to perform on the feature values under a certain attribute according to the distribution of feature values of this attribute, and configure it in subsequent configuration items. The data insight window can include the mean, median, standard deviation, unique value, max, min, and / or missing value of the selected data column, etc. Among them, the unique value represents the degree of repetition of this column of data, that is, how many values are not the same as other values, and the missing value represents no data. For example, the blank positions (i.e., no data values) in the selected data column in the table above the data insight. In addition, the distribution of feature values of this data column can also be displayed through a data histogram, enabling users to more intuitively understand the distribution of feature values of this attribute.

[0108] The configuration window is for configuring some other parameters. For example, configure ID and target. Among them, ID is used to configure the data column, and target is used to configure the prediction target; for example, the user configures ID (i.e., a certain attribute) according to the data insight, and in subsequent model training, the data column corresponding to this ID will not be used.

[0109] After the user completes the configuration, an automatically generated flow chart can be obtained. The generation mechanism of the flow chart is summarized from the modeling best practices of a large number of machine learning solutions.

[0110] Figure 10 The machine learning flow chart automatically generated according to the user configuration provided by the embodiment of the present application, as Figure 10 shown, the nodes in the flow chart are all executable units that can complete independent tasks (such as data import, data preprocessing, feature engineering, prediction, etc.). The connection lines between the nodes represent the relevant relationships of the data, and the connected flow chart can be used for model training and use.

[0111] In some embodiments, the configured machine learning flowchart may include the following nodes: a data import node (table reader) for importing tabular data, an industry solution feature engineering process node (industrysolution feature engineering pipeline Evaluation), an industry solution hyperparameter node (industry solution hyper parameter), a hyperparameter query node (hyper parameter search), a model selection node, a visualization output node, a prediction node, an evaluation node, a merge result node, a data output node, and a statistics node.

[0112] Optionally, the user can jump to the flowchart editing interface with one click and perform secondary editing on the machine learning flowchart as needed. For example, modify parameters, add or delete nodes, or the connection relationships between nodes; as Figure 5 shown, on the left is the directory of optional nodes, including Input, Preprocess, Analysis, FeatureEngineering, Algorithms (regression algorithms, classification algorithms, clustering algorithms), Output, and Script. Clicking on the corresponding directory will expand the modules under the directory for the user to select. For example, clicking on the Input directory will expand the directory of input nodes, such as data import (table reader) for importing data; the nodes under the Analysis directory are used to analyze data, such as data insights. The algorithm nodes under the Algorithms directory are used to set different algorithms. Clicking and dragging the nodes can add them to the flowchart; the Script node indicates that the user can write their own script code. Additionally, there is a directory of some non-selectable nodes. For example, the prediction node is used to obtain prediction results; the evaluation node is used to display the evaluation results of the machine learning model to the user.

[0113] When a node in the flowchart is selected, each column on the right will display the parameters that can be edited for editing the corresponding node parameters of the selected node. As Figure 5As shown in the figure, when the node of principal component analysis is selected, the right sidebar will display the parameters that can be edited, such as Output Mode, Process Types, Number of Components to keep, SvdSolver, Tolerance, Iterated Power, Random State, Columns, etc. After the configuration is completed, when the node runs, a specific identifier will appear in the corresponding node to indicate the running state of this node. For example, a "tick" indicates that the node has run successfully.

[0114] This method does not limit the implementation mechanism of generating a flow chart after obtaining the configuration, the implementation form of editing the generated flow chart, and the utilization form of the flow chart after generating the flow chart. Any interactive form that gradually guides to obtain the user configuration and generates an editable machine learning flow chart is within the scope of this method and is easy to detect.

[0115] Based on the foregoing embodiments, an embodiment of the present application provides a device for generating a machine learning flow chart. The device includes each unit included, which can be implemented by a processor; of course, it can also be implemented by specific logic circuits; during the implementation process, the processor can be a central processing unit (CPU), a microprocessor (MPU), a digital signal processor (DSP), or a field programmable gate array (FPGA), etc.

[0116] Figure 11 It is a schematic structural diagram of the device for generating a machine learning flow chart according to an embodiment of the present application. As Figure 11 shown, the device 110 includes a determination unit 111 and a generation unit 112, where:

[0117] The determination unit 111 is configured to determine a first parameter guided by a first configuration module, where the first parameter is used to specify a target task; where the first configuration module includes at least one candidate task;

[0118] The determination unit 111 is configured to associate with a second configuration module of the target task;

[0119] The determination unit 111 is further configured to determine a second parameter guided by the second configuration module;

[0120] In some embodiments, the generation unit 112 is configured to generate a machine learning flow chart according to the first parameter and the second parameter, and the machine learning model generated by the machine learning flow chart is used to execute the target task.

[0121] In some embodiments, the determining unit 111 is configured to determine at least one of the configured sample data sets; and / or analyze the eigenvalue distribution of the selected target attribute in the sample data set; wherein the eigenvalue distribution is used to guide the processing mode of the eigenvalue of the target attribute; and / or determine the first information and the second information of the configuration; wherein the first information is used to indicate the processing mode of the eigenvalue of the specified attribute, and the second information is used to indicate the prediction target.

[0122] In some embodiments, information corresponding to each candidate task is set on the first configuration module; correspondingly, the determining unit 111 is further configured to, in response to a user operation, further display an annotation description of the candidate task corresponding to the target information specified by the user operation; wherein the annotation description is used to explain the candidate task. The annotation description at least includes: the applicable scenario of the candidate task, a scenario example, the relevant parameters of the scenario example, and an example of the corresponding machine learning flow chart.

[0123] In some embodiments, the generating unit 112 is further configured to receive an editable instruction; wherein the editable instruction is used to indicate editing the machine learning flow chart; in response to the editable instruction, generate the editable machine learning flow chart.

[0124] In some embodiments, the generating unit 112 is further configured to generate a machine learning flow chart according to the first parameter, the second parameter, and a preset default parameter.

[0125] In some embodiments, the generating unit 112 is further configured to, when the indicated generating mode is a dynamic generating mode, generate multiple different versions of the machine learning flow chart according to the first parameter, the second parameter, and a preset default parameter.

[0126] In some embodiments, the generating unit 112 is further configured to present an identification key of the multiple different versions of the machine learning flow chart in a first window; determine the target identification key for receiving a selection operation; in response to the selection operation, present the target machine learning flow chart corresponding to the target identification key and the performance parameters of the target machine learning flow chart in a second window.

[0127] The description of the above device embodiments is similar to the description of the above method embodiments and has similar beneficial effects to the method embodiments. For the technical details not disclosed in the device embodiments of the present application, please refer to the description of the method embodiments of the present application for understanding.

[0128] It should be noted that in the embodiments of the present application Figure 12The division of units in the generating device of the machine learning flowchart shown is schematic, merely a logical function division, and there may be other division methods in actual implementation. In addition, each functional unit in various embodiments of the present application may be integrated in a processing module, may exist alone physically, or two or more modules may be integrated in one module. The above integrated modules may be implemented in the form of hardware, may also be implemented in the form of software functional modules, or may be implemented in the form of a combination of software and hardware.

[0129] It should be noted that in the embodiments of the present application, if the above method is implemented in the form of software functional units and sold or used as an independent product, it may also be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the related technology, may be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable an electronic device to execute all or part of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), magnetic disks, or optical discs that can store program codes. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0130] Embodiments of the present application provide an electronic device Figure 12 which is a schematic diagram of the hardware entity of the electronic device in the embodiments of the present application. As Figure 12 shown, the electronic device 120 includes a memory 121 and a processor 122. The memory 121 stores a computer program that can run on the processor 122. When the processor 122 executes the program, it implements the steps in the method provided in the above embodiments.

[0131] It should be noted that the memory 121 is configured to store instructions and applications executable by the processor 122, and can also cache data to be processed or already processed by the processor 122 and each unit in the electronic device 120 (for example, image data, audio data, voice communication data, and video communication data), and can be implemented by flash memory (FLASH) or random access memory (RAM).

[0132] Embodiments of the present application provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps in the method provided in the above embodiments.

[0133] Embodiments of the present application provide a computer program product containing instructions. When it runs on a computer, it causes the computer to execute the steps in the method provided in the above method embodiments.

[0134] It should be noted that the descriptions of the above storage medium and device embodiments are similar to those of the above method embodiments and have similar beneficial effects to the method embodiments. For the technical details not disclosed in the storage medium, storage medium and device embodiments of the present application, please refer to the descriptions of the method embodiments of the present application for understanding.

[0135] It should be understood that the "one embodiment" or "an embodiment" or "some embodiments" mentioned throughout the specification means that the specific features, structures or characteristics related to the embodiment are included in at least one embodiment of the present application. Therefore, the appearances of "in one embodiment" or "in an embodiment" or "in some embodiments" throughout the specification do not necessarily refer to the same embodiment. In addition, these specific features, structures or characteristics can be combined in one or more embodiments in any suitable manner. It should be understood that in various embodiments of the present application, the sequence numbers of the above processes do not mean the order of execution is prior or subsequent. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application. The sequence numbers of the embodiments of the present application above are only for description and do not represent the advantages or disadvantages of the embodiments. The above descriptions of each embodiment tend to emphasize the differences between the embodiments, and their similarities or similarities can be referred to each other. For the sake of brevity, they will not be repeated herein.

[0136] The term "and / or" in this article is only a description of the association relationship of the associated objects, indicating that there can be three relationships. For example, object A and / or object B can represent: object A exists alone, object A and object B exist simultaneously, and object B exists alone.

[0137] It should be noted that in this article, the term "comprise", "include" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0138] In several embodiments provided by this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The embodiments described above are merely illustrative. For example, the division of the modules is only a logical functional division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the couplings, direct couplings, or communication connections between the components shown or discussed with each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be electrical, mechanical, or other forms.

[0139] The modules described above as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules. They can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0140] In addition, in each embodiment of this application, each functional unit can be all integrated in a processing module, or each unit can be separately used as a module, or two or more units can be integrated in a module. The above integrated units can be implemented in the form of hardware or in the form of a combination of hardware and software functional modules.

[0141] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments. The foregoing storage medium includes various media that can store program codes, such as removable storage devices, read-only memory (ROM), magnetic disks, or optical discs.

[0142] Alternatively, if the above integrated modules of this application are implemented in the form of software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of this application essentially or the part that contributes to the related technology can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable an electronic device to execute all or part of the methods described in each embodiment of this application. The foregoing storage medium includes various media that can store program codes, such as removable storage devices, ROM, magnetic disks, or optical discs.

[0143] The methods disclosed in several method embodiments provided by this application can be arbitrarily combined without conflict to obtain new method embodiments.

[0144] The features disclosed in several product embodiments provided by this application can be arbitrarily combined without conflict to obtain new product embodiments.

[0145] The features disclosed in several method or device embodiments provided by this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0146] As described above, it is only the implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed by this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claimed rights.

Claims

1. A method for generating a machine learning flow chart, characterized in that, The method includes: Determine a first parameter based on the boot configuration of the first configuration module, where the first parameter is used to specify a target task; wherein, the first configuration module includes at least one candidate task; information corresponding to each candidate task is set on the first configuration module; In response to a user operation, further display an annotation description of the candidate task corresponding to the target information specified by the user operation; wherein, the annotation description is used to explain the candidate task; Determine a second configuration module associated with the target task; Determine at least one sample data set configured based on the second configuration module; Analyze the eigenvalue distribution of the selected target attribute in the sample data set; wherein, the eigenvalue distribution is used to guide the configuration of the processing method for the eigenvalues of the target attribute; Determine first information and second information configured based on the eigenvalue distribution of the target attribute; wherein, the first information is used to indicate the processing method for the eigenvalues of the specified attribute, and the second information is used to indicate the prediction target; Generate a machine learning flow chart according to the first parameter, the first information, and the second information, and the machine learning model generated by the machine learning flow chart is used to execute the target task; The method further includes: Receive an editable instruction; wherein, the editable instruction is used to indicate editing the machine learning flow chart; In response to the editable instruction, generate an editable machine learning flow chart; In response to a node in the displayed editable machine learning flow chart being selected, display the parameters that can be edited for the node to edit the node parameters of the selected node.

2. The method according to claim 1, characterized in that The annotation description at least includes: The applicable scenario of the candidate task, a scenario example, the relevant parameters of the scenario example, and an example of the corresponding machine learning flow chart.

3. The method according to any one of claims 1 to 2, characterized in that, The generating a machine learning flow chart according to the first parameter, the first information, and the second information includes: Generate a machine learning flow chart according to the first parameter, the second parameter, and a preset default parameter; wherein, the second parameter is guided and configured based on the second configuration module, and the preset default parameter is determined based on machine learning model modeling practice.

4. The method according to claim 3, characterized in that, The second parameter further includes a parameter used to indicate the generation method of the machine learning flow chart; The generating a machine learning flow chart according to the first parameter, the second parameter, and a preset default parameter includes: In the case where the indicated generation method is a dynamic generation method, generate multiple different versions of the machine learning flow chart according to the first parameter, the second parameter, and a preset default parameter.

5. The method according to claim 4, wherein The method further includes: Present the identification keys of the multiple different versions of the machine learning flow chart in a first window; Determine the target identification key for which a selection operation is received; In response to the selection operation, present the target machine learning flow chart corresponding to the target identification key and the performance parameters of the target machine learning flow chart in a second window.

6. A generating device for a machine learning flow chart, characterized in that, Includes: A determining unit, configured to determine a first parameter based on a boot configuration of a first configuration module, where the first parameter is used to specify a target task; wherein, the first configuration module includes at least one candidate task; information corresponding to each candidate task is set on the first configuration module; The determining unit is further configured to, in response to a user operation, further display an explanatory note of a candidate task corresponding to target information specified by the user operation; wherein, the explanatory note is used to explain the candidate task; The determining unit is further configured to determine a second configuration module associated with the target task; The determining unit is further configured to determine at least one sample data set configured based on the second configuration module; analyze the eigenvalue distribution of a selected target attribute in the sample data set; wherein, the eigenvalue distribution is used to guide the configuration of the processing method for the eigenvalues of the target attribute; determine first information and second information configured based on the eigenvalue distribution of the target attribute; wherein, the first information is used to indicate the processing method for the eigenvalues of a specified attribute, and the second information is used to indicate a prediction target; A generating unit, configured to generate a machine learning flowchart according to the first parameter, the first information, and the second information, where the machine learning model generated by the machine learning flowchart is used to execute the target task; The generating unit is further configured to receive an editable instruction, where the editable instruction is used to indicate editing the machine learning flowchart, and in response to the editable instruction, generate the editable machine learning flowchart; in response to a node displayed in the editable machine learning flowchart being selected, display parameters that can be edited for the node, for editing the node parameters of the selected node.

7. An electronic device, comprising a memory and a processor, the memory storing a computer program that can run on the processor, characterized in that, When the processor executes the program, the method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Method and system for building machine learning modeling template

    CN108710949A