Apparatus for automating machine learning and method for automating machine learning
A semi-automated white-box approach using semantic technology allows non-experts to build and configure machine learning pipelines, addressing complexity and enabling efficient, explainable pipeline development.
Patent Information
- Application Number
- JP2021166890
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-10-12
- Filing Date
- 2021-10-11
- Publication Date
- 2026-01-29
- Estimated Expiration
- 2041-10-11
AI Technical Summary
Developing a machine learning pipeline requires specialized training and deep understanding of data and domain knowledge, making it complex for non-experts to build and configure.
A semi-automated white-box approach using semantic technology to enable non-machine learning experts to build and configure machine learning pipelines, leveraging semantic technologies to code domain and machine learning knowledge, allowing explainable and extensible pipeline development.
Enables non-experts to efficiently develop and maintain machine learning pipelines across various datasets, ensuring explainability and adaptability.
Smart Images

Figure 0007808315000001 
Figure 0007808315000002 
Figure 0007808315000003
Abstract
Description
[Technical Field]
[0001] background The present invention relates to the automation of machine learning, and in particular to the automation of pipeline construction and pipeline configuration. [Background technology]
[0002] A machine learning pipeline transforms raw data into conclusions and functioning machine learning models.
[0003] Developing a machine learning pipeline is a complex process that requires a deep understanding of the data, as well as an understanding of the domain and the problem being addressed, which requires specialized training in information processing and data analysis, especially machine learning. Summary of the Invention [Means for solving the problem]
[0004] Disclosure of the Invention The apparatus for automating machine learning and the computer-implemented machine learning automation method provide a semi-automated white-box approach to machine learning development supported by semantic technology, allowing even non-machine learning experts to build and configure machine learning pipelines in a convenient and adaptable manner.
[0005] The apparatus and method enable explainable, extensible, and configurable machine learning pipeline development that can be used by non-machine learning experts with minimal machine learning training. This is achieved by using semantic technologies in the machine learning pipeline to code a formal representation of domain knowledge and machine learning knowledge.
[0006] Machine learning pipelines can be developed on multiple different datasets to solve similar tasks, and these pipelines can be efficiently maintained and scaled for future situations.
[0007] The computer-implemented method includes the steps of determining, in a representation of relationships between elements, an element representing a first characteristic of the machine learning pipeline; determining, in the representation, an element representing a second characteristic of the machine learning pipeline in dependence on the element representing the first characteristic; outputting an output for the element representing the second characteristic; detecting an input, in particular a user input; and determining parameters of the machine learning pipeline in dependence on the element representing the second characteristic if the input satisfies predetermined requirements, or not determining parameters of the machine learning pipeline in dependence on the element representing the second characteristic if the input does not satisfy the predetermined requirements.
[0008] The parameters configure the machine learning pipeline. The second characteristic makes the machine learning pipeline explainable to a non-machine learning expert user. The non-machine learning expert user can configure the machine learning pipeline via the inputs. Knowledge of the first characteristic of the machine learning pipeline is not required.
[0009] According to one aspect, the method includes a step of determining whether an element representing a first property and an element representing a second property have a relationship that satisfies a predetermined condition, in particular whether the relationship satisfies the condition that the element representing the first property is semantically reachable from the element representing the second property in the representation according to semantics coded in the representation.
[0010] The step of outputting the output may include, inter alia, prompting the user to respond by indicating a like or dislike of an element representing the second characteristic, or by selecting an element representing the second characteristic.
[0011] The method preferably comprises the steps of determining a relationship by a function for assessing semantic reachability, detecting a response, and modifying at least one parameter of the function depending on the response, such that the reachability function is updated based on the input.
[0012] The method may include determining a link between the two elements depending on the response, and constructing a semantic reachability graph representing the relationship by storing machine learning templates corresponding to the two elements and the link between those elements in the semantic reachability graph.
[0013] The method may comprise determining a plurality of elements in the representation that are semantically reachable from an element representing the first property, and determining an element representing the second property from the plurality of elements in dependence on an input, wherein a user may select, for example from a list, the element representing the second property that is used to determine the parameter.
[0014] The method may include determining other elements representing the second property that are semantically reachable in the representation from the element representing the first property if the input does not satisfy the requirement, e.g., if the user does not like the second element, the calculation of the elements used for the parameters may be repeated.
[0015] According to one embodiment, the method includes determining the condition depending on the input, thus updating the concept of semantic reachability according to the user's preferences.
[0016] According to one aspect, the method includes determining parameters for a component of a machine learning pipeline and determining a condition dependent on a function of the component within the machine learning pipeline.
[0017] The method may include determining a group of elements depending on the input, and determining a plurality of parameters for representing the machine learning pipeline depending on the group of elements, such that a plurality of parameters related to a first element that represents a characteristic of the machine learning pipeline are determined.
[0018] The method may include determining a machine learning model relying on a representation of the machine learning pipeline.
[0019] The method may include providing raw data having an image, determining parameters of a machine learning pipeline for an image classifier model, determining a machine learning model depending on the parameters, and training the image classifier model using at least one image of the raw data, wherein the machine learning model trained in this manner is or includes an image classifier model.
[0020] The method may include providing an image and classifying the image using a machine learning model.
[0021] An apparatus for determining a machine learning model is configured to perform this method.
[0022] Further advantageous embodiments can be derived from the following description and drawings. [Brief explanation of the drawings]
[0023] [Figure 1] FIG. 1 is a diagram illustrating a schematic diagram of an apparatus for determining a machine learning model. [Figure 2] FIG. 1 shows a schematic diagram of steps in a method for determining a machine learning model. [Figure 3] FIG. 2 is a diagram illustrating an example of user interaction. DETAILED DESCRIPTION OF THE INVENTION
[0024] FIG. 1 shows an apparatus 100 for determining a machine learning model 102 .
[0025] The apparatus 100 comprises a user interface layer 104, a semantic layer 106, a data and mapping layer 108, and a machine learning layer 110. The apparatus 100 is configured to perform the computer-implemented methods described below.
[0026] The user interface layer 104 provides the user with an interface for using the semantic components of the semantic layer 106, i.e., the complete system.
[0027] The user interface layer 104 has a first function 112 configured to dynamically visualize information related to data annotations 114 and machine learning models 116.
[0028] The user interface layer 104 may have graphical user interface functionality for displaying information about current data annotations 114 and machine learning models 116.
[0029] The first function 112 is configured to elicit information from a user regarding links L between data annotations 114 and / or machine learning models 116 and / or machine learning ontology templates T.
[0030] In this embodiment, a representation of a machine learning pipeline can be selected by a user by selecting the representation from a plurality of machine learning pipelines 118 displayed on a graphical user interface.
[0031] In this embodiment, the link L between the machine learning ontology templates T is selectable by the user by selecting the machine learning ontology template T to be linked from multiple machine learning templates displayed on the graphical user interface.
[0032] The information about the machine learning model 116 has a visualization section 120 .
[0033] The visualization unit 120 has a feature group 122. In this embodiment, the feature group 122 has icons for a first feature group 122-1, a second feature group 122-2, and a third feature group 122-3. In this embodiment, the first feature group 122-1 is a group for a single feature. In this embodiment, the second feature group 122-2 is a group for a time series. In this embodiment, the third feature group 122-3 is a group for a quality indicator.
[0034] The visualization unit 120 includes a graphical representation of the processing algorithms 124. The graphical representation of the processing algorithms 124 includes an icon for a first algorithm 124-1, an icon for a second algorithm 124-2, and a graph 124-3. In this example, the first algorithm 124-1 is an algorithm for a first feature group 122-1, e.g., a group for single features. In this example, the second algorithm 124-2 is an algorithm for a second group 122-2, e.g., a group for time series features. The graph 124-3, in this example, displays the progression of resistance over time starting from the origin of a Cartesian coordinate system, including the length, peak (maximum value), and drop from the peak to a final value.
[0035] The visualizer 120 includes a graphical representation of the machine learning algorithm 126. The graphical representation of the machine learning algorithm 126 in this example includes an icon 126-1 that depicts one aspect of the machine learning algorithm 126.
[0036] The information about the current data annotation 114 includes feature groups 128. In this example, feature groups 128 include icons for a first feature group 128-1, a second feature group 128-2, and a third feature group 128-3.
[0037] The information about the current data annotation 114 includes domain feature names 130. In this example, the domain feature names 130 include icons for a first domain feature name 130-1, a second domain feature name 130-2, a third domain feature name 130-3, a fourth domain feature name 130-4, and a fifth domain feature name 130-5.
[0038] In this example, the first domain feature name 130-1 represents the status of the data. In this example, the second domain feature name 130-2 represents a characteristic of the data. In this example, the third domain feature name 130-3 represents one type of data. In this example, the fourth domain feature name 130-4 represents another type of data. In this example, the fifth domain feature name 130-5 represents the quality of the data.
[0039] In this example, the first domain feature name 130-1 and the second domain feature name 130-2 are mapped to the first feature group 128-1. In this example, the third domain feature name 130-3 and the fourth domain feature name 130-4 are mapped to the second feature group 128-2. In this example, the fifth domain feature name 130-5 is mapped to the third feature group 128-3.
[0040] The information about the current data annotation 114 includes raw feature names 132. In this example, raw feature names 132 include icons for a first raw feature name 132-1, a second raw feature name 132-2, a third raw feature name 132-3, a fourth raw feature name 132-4, and a fifth raw feature name 132-5.
[0041] In this example, the first raw feature name 132-1 represents a status code for the data. In this example, the second raw feature name 132-2 represents a characteristic of the data. In this example, the third raw feature name 132-3 represents one type of data. In this example, the fourth raw feature name 132-4 represents another type of data. In this example, the fifth raw feature name 132-5 represents the quality of the data.
[0042] In this example, the first raw feature name 132-1 is mapped to the first domain feature name 130-1. In this example, the second raw feature name 132-2 is mapped to the second domain feature name 130-2. In this example, the third raw feature name 132-3 is mapped to the third domain feature name 130-3. In this example, the fourth raw feature name 132-4 is mapped to the fourth domain feature name 130-4. In this example, the fifth raw feature name 132-5 is mapped to the fifth domain feature name 130-5.
[0043] The visualizer 120 has a representation T' of a plurality of linkable machine learning ontology templates T. A machine learning ontology template T can reference one of the feature groups, one of the machine learning algorithms, one of the domain feature names, or one of the raw feature names, and is represented by one of a plurality of icons.
[0044] The arrows connecting these items in FIG. 1 represent the links L between these elements provided by the user.
[0045] The semantic layer 106 includes a representation O of relationships for a plurality of elements. More specifically, the representation O includes a first ontology 134, a second ontology 136, and a third ontology 138.
[0046] The first ontology 134 in this embodiment is a domain ontology, which encodes domain knowledge using a formal representation including the classes and attributes of the domain and their relationships.
[0047] The second ontology 136 in this example is a feature group ontology, which stores links between terms in a domain ontology and feature groups in a catalog of pre-designed machine learning pipelines.
[0048] The third ontology 138 in this example is a machine learning pipeline ontology, which codes machine learning knowledge in a formal representation, including allowed and default feature groups, appropriate feature processing algorithms for those feature groups, corresponding feature-processed groups, appropriate machine learning modeling algorithms for each feature-processed group, and their relationships. The third ontology 138, in this example, codes a catalog. This catalog stores several successful and practically common machine learning pipelines pre-designed in a formal representation by machine learning experts. A machine learning pipeline is a pre-designed mapping from feature groups to their feature processing algorithms for the feature groups, to their corresponding feature-processed groups, and to their specified machine learning modeling algorithms for the feature-processed groups.
[0049] The semantic layer 106 comprises a machine learning ontology MLO and a machine learning ontology template T.
[0050] The machine learning ontology MLO codes machine learning knowledge through a formal expression. For example, the machine learning ontology MLO codes which feature groups or combinations thereof are allowed or not allowed for processing input data, and / or default feature groups for processing input data. For example, the machine learning ontology MLO codes which feature processing algorithms are suitable for which feature groups. For example, the machine learning ontology MLO codes feature processed groups corresponding to feature processing algorithms, appropriate machine learning modeling algorithms for the feature processed groups and / or their relationships.
[0051] In one embodiment, the machine learning ontology MLO defines the permitted search space of the semantic reachability graph.
[0052] The machine learning ontology template T, in one embodiment, is a fragment of an ontology that includes variables used to generate an instance of the third ontology 138, the machine learning pipeline ontology.
[0053] The semantic layer 106 includes a dynamic extender, for example a dynamic smart reasoner DSR, which takes as input the machine learning ontology MLO, machine learning templates T, and user-provided links L, and dynamically extends, configures, and builds a third ontology 138, i.e., the machine learning pipeline ontology.
[0054] The third ontology 138, i.e., the machine learning pipeline ontology constructed using the machine learning ontology MLO and the template T, encodes a semantic reachability graph G for a set of elements such as feature groups, feature processing algorithms, feature-processed groups, and their specified machine learning modeling algorithms for each feature-processed group.
[0055] According to one embodiment, a semantic reachability graph G is constructed by linking machine learning templates T using links L.
[0056] The Dynamic Smart Reasoner DSR in this embodiment has two layers of functionality.
[0057] First, the dynamic start reasoner DSR is configured to receive links L as input and link machine learning templates T to dynamically configure, extend, and build a semantic reachability graph G.
[0058] In this embodiment, the permitted search space of the semantic reachability graph G is defined based on the machine learning pipeline ontology MLO.
[0059] Second, the dynamic smart reasoner DSR is configured to receive as input, e.g., annotations A of an input dataset D, and to compute a set of elements S1,...,Sn in the representation O that are semantically reachable from links L and possibly have semantic reachability relations between them. Elements S1,...,Sn that have semantic reachability relations to other elements Sj are semantically reachable from said other elements Sj.
[0060] The dynamic smart reasoner DSR is configured to dynamically update the semantic reachability relations in the semantic reachability graph G.
[0061] Semantic layer 106 includes a second function SR, which in this example includes a first reasoner 140, a second reasoner 142, and a third reasoner 144. The second function SR may also include an annotator 146.
[0062] Annotator 146 allows a user to annotate the raw data with terms from the first ontology 134, in this example a domain ontology.
[0063] The first function 112 can be configured to determine an annotation A for the data from a user input. The second function SR can be configured to receive the annotation A. The first function 112 can be configured to determine a representation M of the machine learning pipeline from the user input. The second function SR can be configured to receive the representation M of the machine learning pipeline.
[0064] The first function 112 is configured for user interaction. More specifically, the first function 112 is configured to determine and output an output prompting the user to respond. According to one embodiment, the first function 112 is configured to output a prompt indicating an element of the representation O for which the user initiated the user interaction. The first function 112 can be configured to request the user to indicate whether they like or dislike the element. The first function 112 can be configured to request the user to select one or more elements from a list of elements. The first function 112 can be configured to detect a user input in response to the output. The first function 112 can be configured to determine a result of the user interaction depending on the input. The first function 112 can be configured to determine a result indicating whether the element for which the user initiated the user interaction should be used to determine a machine learning pipeline. The first function 112 can be configured to determine a result indicating one or more elements selected by the user from the list according to the input. The first function 112 can be configured to determine a result that includes the element for which the user initiated the user interaction.
[0065] According to one embodiment, the first function 112 is configured to determine at least one condition for evaluating the semantic reachability of elements in the representation O.
[0066] According to this embodiment, the annotator 146 is configured to process annotations A and elements from the first ontology 134 to determine a first dataset 148, which comprises raw data 150 from the raw data lake 152 and a first mapping 154 of raw feature names from the raw data 150 to domain feature names according to the first ontology 134.
[0067] According to this embodiment, first reasoner 140 is configured to process first mapping 154 and determine a second mapping 156 of domain feature names to feature groups according to second ontology 136. First reasoner 140 can be configured to auto-generate and user-configure second mapping 156 based on second ontology 136 and user configuration, as well as results of user interaction, e.g., input from a user.
[0068] According to this embodiment, second reasoner 142 is configured to process second mapping 156 to determine a third mapping 158 of feature groups to processing algorithms for the feature-processed groups according to second ontology 136 and third ontology 138. Second reasoner 142 can be configured to automatically generate third mapping 158 based on the catalog that is second ontology 136, e.g., third ontology 138, second mapping 156, a user-selected representation M of a machine learning pipeline, and results of user interaction, e.g., input from a user.
[0069] According to this embodiment, the second reasoner 142 is configured to determine a fourth mapping 160 of domain names to feature groups for a processing algorithm depending on the second mapping 156 and the third mapping 158 and the results of the user interaction, e.g., input from the user.
[0070] According to this embodiment, the second reasoner 142 is configured to determine a fifth mapping 162 of feature groups to feature processed groups depending on the third mapping 158 and the results of the user interaction, e.g., input from the user.
[0071] According to this embodiment, the third reasoner 144 is configured to determine a sixth mapping 164 of feature-processed groups to machine learning algorithms depending on the third mapping 158. The third reasoner 144 can be configured to automatically generate the sixth mapping based on the catalog, e.g., based on the third ontology 138, a user-selected representation M of the machine learning pipeline, and the results of a user interaction, e.g., input from a user.
[0072] The machine learning layer 110 includes a first data integrator 166 configured to determine raw data 150 from the raw data lake 152. The machine learning layer 110 includes a second data integrator 168 configured to determine integrated data 170 depending on the first dataset 148. The raw data stored in the raw data lake 152 is transformed by the second data integrator 168 into integrated data 170 suitable for machine learning. The integrated data 170 is prepared for a second dataset 172 having the integrated data 170 and a fourth mapping 160.
[0073] The machine learning layer 110 includes a first processor 174 configured to determine features 176 for a third data set 178 dependent on the second data set 172. The third data set 178 includes the features 176 and a fifth mapping 162.
[0074] The machine learning layer 110 includes a second processor 180 configured to determine selected features 182 for a fourth dataset 184 depending on the third dataset 178. The fourth dataset 184 includes the selected features 182 and a sixth mapping 164.
[0075] The machine learning layer 110 includes a third processor 186 configured to determine the machine learning model 102 depending on a fourth dataset 184 .
[0076] The data and mapping layer 108 includes a raw data lake 152 , a first dataset 148 , a second dataset 172 , a third dataset 178 , a fourth dataset 184 and a machine learning model 102 .
[0077] In FIG. 1, double arrows indicate data information flow, solid arrows indicate semantic information flow, and dashed lines indicate connections by class of elements in representation O.
[0078] Below, we describe an exemplary method for automated machine learning pipeline construction and configuration that uses semantics that allow for user interaction during the automated construction of the machine learning pipeline.
[0079] According to this method, users need to annotate their data with domain ontology terms and select a pre-designed machine learning pipeline from a catalog. The user can interact with intermediate steps in the sequence to build a representation of the machine learning pipeline M. Everything else is automated. This is achieved by coding a formal representation of domain and machine learning knowledge using semantic technologies in the catalog of pre-designed machine learning pipelines.
[0080] According to one embodiment, the representation of the machine learning pipeline M has multiple parameters. The representation O can include semantics that code domain knowledge and machine learning knowledge. The semantics have multiple elements that represent the domain knowledge or the machine learning knowledge and the relationships between them.
[0081] According to one embodiment, once annotation A is prepared, a first element in representation O is determined depending on annotation A. Then, a second element in representation O can be determined depending on the first element. According to this embodiment, the second element is determined such that the first element and the second element have a relationship that satisfies a first condition. The first condition can be that the first element is semantically reachable from the first element to the second element in representation O according to the semantics coded in representation O. If the first element and the second element satisfy the first condition, the parameter is determined depending on the second element. If the first element and the second element do not satisfy the first condition, the parameter is not determined depending on the second element.
[0082] According to another aspect, a third element in the representation O can be determined depending on a second element. According to this embodiment, the second element and the third element have a relationship that satisfies a second condition. The second condition can be that the third element in the representation O is semantically reachable from the second element. According to this aspect, if the second condition is satisfied, the parameter can be determined depending on the second element and the third element. If the second condition is not satisfied, the parameter is not determined depending on the second element and the third element.
[0083] For example, element S1 may instantiate a first parameter P1 with one feature group appearing in representation O.
[0084] For example, element S2 may instantiate a second parameter P2 by one or more processing algorithms appearing in representation O that are semantically reachable from the feature group of element S1.
[0085] For example, element S3 may instantiate a third parameter P3 by one or more processed feature groups appearing in representation O that are semantically reachable from one or more processing algorithms of element S2.
[0086] For example, element S4 may instantiate a fourth parameter P4 with processed features appearing in representation O that are semantically reachable from one or more processed feature groups of element S3.
[0087] For example, element S5 may instantiate a fifth parameter P5 by one or more machine learning algorithms appearing in representation O that are semantically reachable from one or more processed features of element S4.
[0088] Since ontologies are not graphs, the notion of semantic reachability can vary and depend on the use case and application.
[0089] According to one aspect, semantic reachability can be determined based on projecting one or more of the above ontologies onto a graph structure and computing graph reachability, which can also take into account the path tightness between ontology elements, thereby taking into account explicit and implicit relationships between elements of representation O.
[0090] This method is described below with reference to Figure 2. The method is for a machine learning pipeline with parameters P1, ..., Pk as an example. A machine learning pipeline may have multiple sets of parameters, each represented by one element in the representation O.
[0091] The method comprises the step 200 of providing a representation M of a machine learning pipeline having parameters P1, . . . , Pk.
[0092] The step of providing a representation M includes the steps of detecting a user input that identifies one of a plurality of representations 118 of the machine learning pipeline, and selecting the representation M of the machine learning pipeline identified in the user input.
[0093] The method comprises the step 202 of providing a representation O. The representation O comprises a number of elements and their relationships.
[0094] The representation O in this example comprises a first ontology 134, a second ontology 136 and a third ontology 138. Graphs representing these can be used in a similar manner.
[0095] The method includes a step 204 of providing input data.
[0096] The method comprises a step 206 of preparing an annotation A of the features of the input data, in particular the names of the variables in the input data.
[0097] The method comprises a step 208 of providing a first element and a second element of a representation O. The step of providing a first element, in this embodiment, comprises determining one element of the representation O that represents a feature of the input data in the representation O with respect to the annotation A. According to one embodiment, names of variables of the input data according to the first ontology are determined.
[0098] The step of providing the second element in this embodiment comprises determining, with respect to the first element, an element of the representation O that represents a feature name in the representation O. According to one embodiment, names of variables in the domain according to the first ontology are determined.
[0099] A second element can be determined in representation O depending on a first element, such that the first element and the second element have a relationship that satisfies a first condition. This means, according to this embodiment, that the second element is semantically reachable from the first element in representation O according to the semantics coded in representation O. A plurality of second elements that satisfy the first condition can be determined. The method is described with respect to one of these second elements and applies equally to any number of these second elements.
[0100] The method further includes determining 210 a first data set 148 having input data 150, a first element, and a second element.
[0101] The method further comprises a step 212 of determining a third element in the representation O representing one feature group. A first parameter P1 can be determined depending on the third element.
[0102] Step 212 in this embodiment comprises determining an element of representation O that represents a group of features in representation O that satisfy a second condition with respect to the second element, in particular according to the second ontology 36. The method is described with respect to one of these third elements and applies similarly to any number of these third elements.
[0103] A third element in representation O can be determined depending on a second element such that the third element is semantically reachable from the second element in representation O.
[0104] The third element may be determined according to the second ontology 136 .
[0105] The method may include determining a number of third elements that are semantically reachable from the second element in a representation O of the relationships between the elements.
[0106] The method optionally comprises a step 213 of performing a user interaction for one or more third elements.
[0107] This user interaction will be explained below with reference to Figure 3. The first parameter P1 can be determined depending on the third element or independently of the result of the user interaction for the third element.
[0108] If the result of the user interaction indicates that the third element satisfies the first requirement, step 214 is performed. If not, step 212 is performed. If the user likes or selects the third element, the first requirement can be satisfied. If not, the first requirement cannot be satisfied.
[0109] To perform step 212 with an updated concept of semantic reachability, the second condition may be updated, modified or determined depending on the outcome of the user interaction.
[0110] The second condition can be updated, changed or determined depending on the function of the component in the machine learning pipeline for which the first parameter P1 is used.
[0111] The result of the user interaction may indicate a group of elements having a third element determined depending on input received in the user interaction, as described below. The method may include determining parameters for the representation of the machine learning pipeline M, including parameters for the third element, depending on the group of elements.
[0112] The method further includes a step 214 of determining a fourth element in the representation O. In this embodiment, determining the fourth element includes determining an element of the representation O that represents a processing algorithm that satisfies a third condition with respect to the third element. The fourth element in the representation O can be determined depending on the third element such that the fourth element is semantically reachable from the third element. The fourth element represents a processing algorithm for the feature group represented by the third element. The second parameter P2 can be determined depending on the fourth element. The method is described with respect to one of these fourth elements and applies similarly to any number of these fourth elements.
[0113] The fourth element may be determined according to the third ontology 138 .
[0114] The method may include determining a number of fourth elements that are semantically reachable from a third element in a representation O of the relationships between the elements.
[0115] The method optionally comprises a step 215 of performing a user interaction for one or more fourth elements.
[0116] This user interaction is explained below with reference to FIG.
[0117] The second parameter P2 can be determined depending on the fourth factor or independently of the result of a user interaction for the fourth factor.
[0118] If the result of the user interaction indicates that the fourth element satisfies the second requirement, then step 216 is performed. If not, then step 214 is performed.
[0119] If the user likes or selects the fourth element, the second requirement can be met; otherwise, the second requirement cannot be met.
[0120] To perform step 214 with an updated concept of semantic reachability, the third condition may be updated, modified or determined depending on the outcome of the user interaction.
[0121] The third condition can be updated, changed or determined depending on the function of the component in the machine learning pipeline for which the second parameter P2 is used.
[0122] The result of the user interaction may indicate a group of elements having a fourth element determined depending on input received in the user interaction, as described below. The method may include determining parameters for the representation of the machine learning pipeline M, including a parameter for the fourth element, depending on the group of elements.
[0123] The method further includes determining 216 a second data set 172 having the set of integrated data 170, the second element, the third element, and the fourth element.
[0124] The method further includes determining 218 a set of features 176 by processing the second data set 172 with a processing algorithm represented by a fourth element.
[0125] The method further includes determining 220 a fifth element. In this embodiment, determining the fifth element includes determining an element of representation O that represents feature set 172 that satisfies the fourth condition with respect to the third element. In this embodiment, the fifth element is determined such that the fifth element is semantically reachable from the third element in representation O and such that the fifth element is semantically reachable from the fourth element in representation O. In this embodiment, the fifth element represents one or more processed feature groups.
[0126] The method may include determining a plurality of fifth elements that are semantically reachable from the third element and the fourth element in a representation O of the relationships between the elements.
[0127] The method optionally comprises a step 221 of performing a user interaction for one or more fifth elements.
[0128] This user interaction is explained below with reference to FIG.
[0129] The third parameter P3 can be determined depending on the fifth element or independently of the result of a user interaction for the fifth element. The fourth parameter P4 can be instantiated by the processed features, i.e., the set of features 176, or independently of the result of a user interaction for the fifth element.
[0130] If the results of the user interaction indicate that the fifth element satisfies the third requirement, then step 222 is performed. If not, then step 220 is performed.
[0131] If the user likes or selects the fifth element, the third requirement can be met; otherwise, the third requirement cannot be met.
[0132] To perform step 220 with an updated concept of semantic reachability, the fourth condition may be updated, modified or determined depending on the outcome of the user interaction.
[0133] The fourth condition can be updated, changed or determined depending on the function of the component in the machine learning pipeline for which the third parameter P3 is used.
[0134] The fourth condition can be updated, changed or determined depending on the function of the component in the machine learning pipeline for which the fourth parameter P4 is used.
[0135] The result of the user interaction may indicate a group of elements having a fifth element determined depending on input received in the user interaction, as described below. The method may include determining parameters for the representation of the machine learning pipeline M, including a parameter for the fifth element depending on the group of elements.
[0136] The method further includes determining 222 a third data set 178 having the set of features 176, a third element, and a fifth element.
[0137] The method further includes determining 224 a sixth element. Determining the sixth element includes determining an element of representation O that represents a machine learning modeling algorithm that satisfies a fifth condition with respect to the fifth element. In this example, the sixth element is determined such that the sixth element is semantically reachable from the fifth element in representation O.
[0138] The method may include determining a number of sixth elements that are semantically reachable from the fifth element in a representation O of the relationships between the elements.
[0139] The method optionally comprises a step 225 of performing a user interaction for one or more sixth elements.
[0140] This user interaction is explained below with reference to FIG.
[0141] The fifth parameter P5 may be determined depending on one or more machine learning algorithms represented by the sixth element, or independently of the result of a user interaction of the sixth element.
[0142] If the result of the user interaction indicates that the sixth element is to be used, then step 226 is performed; if not, then step 224 is performed.
[0143] To perform step 224 with an updated concept of semantic reachability, the fifth condition may be updated, modified or determined depending on the outcome of the user interaction.
[0144] The fifth condition can be updated, changed or determined depending on the function of the component within the machine learning pipeline for which the fifth parameter P5 is used.
[0145] The result of the user interaction may indicate a group of elements having a sixth element determined depending on input received in the user interaction, as described below. The method may include determining parameters for the representation of the machine learning pipeline M, including a parameter for the sixth element depending on the group of elements.
[0146] The method further includes determining 226 a fourth data set 184 having the selected feature from the set of features 176 and the fifth and sixth elements.
[0147] The method may further include instructions for processing the data and / or examining 228 the data defined for the fourth data set 184 according to the sixth element.
[0148] In this way, a machine learning pipeline transforms raw input data into conclusions and functioning machine learning models.
[0149] The machine learning pipeline can be applied, for example, to provide an image classifier model that can be used for monitoring and process control based on images acquired during a process. The process can be resistance welding. The images can depict the welded point and / or weld seam.
[0150] For purposes of training an image classifier model, a user can annotate training images from a process or a simulation thereof. During user interaction, a user can select features to use for image classification from domain feature names. During user interaction, a user can select a specific processing algorithm to process or classify the images. The machine learning modeling algorithm can be automatically determined from the user selection. To automatically form a trained machine learning model from the training images, the machine learning modeling algorithm can be executed with parameters according to a representation of the machine learning pipeline M.
[0151] For training, the raw data 150 comprises a plurality of images. During training, at least one parameter of the machine learning pipeline M is determined, which is a parameter for an image classifier model. The parameter for the image classifier model may be a hyperparameter or weight of an artificial neural network. A machine learning model is determined depending on the parameter and trained using at least one image of the raw data 150.
[0152] After training, the machine learning model trained in this way can be used to classify images.
[0153] In the above-described method, a user interacts with intermediate steps to construct a representation of the machine learning pipeline M. One or more parameters for a step correspond to components of the machine learning pipeline and are determined semi-automatically. In this context, semi-automatically means that a series of steps are repeated during the interaction until the results of the user interaction indicate that the user has selected or preferred one or more elements of the corresponding components of the machine learning pipeline.
[0154] Upon user interaction, an element representing a first characteristic of the machine learning pipeline M is determined in the representation O of the relationships between elements. Depending on which step in the sequence is being processed, this element may be the second, third, fourth, or fifth element.
[0155] In a user interaction, an element representing a second characteristic of the machine learning pipeline M is determined in the representation O depending on the element representing the first characteristic. Which element is determined depends on the sequence of steps being processed. In a corresponding user interaction, a third element is determined depending on the second element, a fourth element is determined depending on the second and third elements, a fifth element is determined depending on the fourth element, or a sixth element is determined depending on the fifth element.
[0156] The element representing the first characteristic and the element representing the second characteristic have a relationship that satisfies a condition for evaluating semantic reachability applicable to a series of steps.
[0157] When a graph is used to evaluate semantic reachability, nodes in the graph can represent those elements. The condition can be determined from attributes of a first node in the graph for an element that represents a first characteristic. Whether the condition is satisfied can be determined depending on attributes of a second node in the graph for an element that represents a second characteristic. According to one embodiment, the condition is defined by attributes of the first node, and if the attributes of the second node satisfy the condition, then the second node is semantically reachable from the first node. Changes to the condition are stored in the graph, for example, as updated or new attributes of the first node or the second node.
[0158] FIG. 3 shows the steps in user interaction for an iteration.
[0159] In step 300, user interaction includes determining an output that prompts the user for a response.
[0160] This output can be, for example, a prompt that displays the element with which the user interaction was initiated, the prompt can ask the user to like or dislike the element, or the prompt can ask the user to select an element.
[0161] When user interaction is initiated with an element group, the prompt may include a list of elements in the group and a request to select one or more of the elements.
[0162] User interaction includes outputting step 302. In this embodiment, a prompt is displayed to the user.
[0163] User interaction includes detecting 304 an input, which can be a like / dislike attribute for an element displayed to the user, or a selection of multiple elements, e.g., an element from a list of elements displayed to the user.
[0164] This input can be used to determine parameters of the machine learning pipeline M depending on the element if the element is selected, or can be used to not determine parameters of the machine learning pipeline M depending on the element if the element is not selected.
[0165] User interactions are a feature for assessing semantic reachability By the above In this example, the function is a function of a component within the machine learning pipeline that uses the parameters that are subject to user interaction. In this case, the user interaction may include detecting a response and modifying at least one parameter of the function depending on the response.
[0166] In this embodiment, the functionality is implemented as follows.
[0167] The dynamic start reasoner DSR receives links L as input and links machine learning templates T to dynamically configure, extend and / or build a semantic reachability graph G.
[0168] In this embodiment, the permitted search space of the semantic reachability graph G is defined based on the machine learning pipeline ontology MLO.
[0169] The Dynamic Smart Reasoner DSR takes as input an annotation A and computes sets of elements S1,...,Sn in the representation O that are semantically reachable due to links L and possibly have semantic reachability relations between them.
[0170] In this embodiment, the dynamic smart reasoner DSR dynamically updates a third ontology 138, the machine learning pipeline ontology.
[0171] This function allows determining whether a condition is satisfied, e.g., it can say that if the condition is satisfied, then an element in the expression that exhibits a first property is semantically reachable to an element that exhibits a second property, according to the semantics coded in the expression, and it cannot say this if the condition is not satisfied.
[0172] The user interaction includes a step 306 of determining the outcome of the user interaction depending on the input.
[0173] The result can indicate whether the element that initiated the user interaction should be used to determine the machine learning pipeline.
[0174] The result may indicate one or more of a plurality of elements, particularly elements selected by the user according to the input.
[0175] The result can indicate other element groups that contain the element with which the user interaction was initiated.
Claims
1. A computer-implemented method comprising: in a representation (O) of relationships between elements, determining elements representing a first characteristic of a machine learning pipeline (M) (steps 212, 214, 220); in the representation (O), determining elements representing a second characteristic of the machine learning pipeline (M) depending on the elements representing the first characteristic (steps 214, 220, 224); outputting an output for the elements representing the second characteristic (step 302); detecting an input, particularly a user input (step 304); when the input meets a predetermined requirement, determining parameters of the machine learning pipeline (M) depending on the elements representing the second characteristic, or when the input does not meet the predetermined requirement, not determining the parameters of the machine learning pipeline (M) depending on the elements representing the second characteristic (steps 214, 220, 224); A method characterized by including the above steps.
2. Determining whether the elements representing the first characteristic and the elements representing the second characteristic have a relationship satisfying a predetermined condition, particularly, according to the semantics coded in the representation (O), whether they have a relationship satisfying the condition that they are semantically reachable from the elements representing the first characteristic to the elements representing the second characteristic in the representation (O) (steps 214, 220, 224), the method according to claim 1. The method according to claim 1.
3. The step of outputting the output (step 302) particularly includes steps of prompting a user to respond so as to represent a preference for the elements representing the second characteristic or to select the elements representing the second characteristic, the method according to claim 1 or 2. The method according to claim 1 or 2.
4. Determining the relationship by a function for evaluating semantic reachability, detecting the response, and changing at least one parameter of the function depending on the response, the method according to claim 3. The method according to claim 3.
5. Determining a link (L) between two elements depending on the response, and constructing a semantic reachability graph representing the relationship by storing in the semantic reachability graph the machine learning template (T) corresponding to the two elements and the link (L) between the two elements, the method according to claim 4. The method according to claim 4.
6. Determining a plurality of elements semantically reachable in the representation (O) from the element representing the first characteristic (steps 214, 220, 224); and determining the element representing the second characteristic from the plurality of elements depending on the input (step 306). The method according to any one of claims 1 to 5.
7. When the input fails to meet the requirement, determining other elements representing the second characteristic that are semantically reachable in the representation (O) from the element representing the first characteristic (steps 214, 220, 224). The method according to any one of claims 1 to 6.
8. Determining the condition depending on the input (steps 215, 221, 225). The method according to claim 2.
9. Determining parameters for one component of the machine learning pipeline; and determining the condition depending on the function of the component within the machine learning pipeline (steps 215, 221, 225). The method according to claim 2.
10. Determining an element group depending on the input (step 306); and determining a plurality of parameters for the representation of the machine learning pipeline (M) depending on the element group (steps 216, 222, 226). The method according to any one of claims 1 to 9.
11. Determining a machine learning model depending on the representation of the machine learning pipeline (M) (step 228). The method according to any one of claims 1 to 10.
12. Providing unprocessed data (150) having an image; determining the parameters of the machine learning pipeline (M) for an image classifier model; determining the machine learning model depending on the parameters; and training the image classifier model using at least one image of the unprocessed data (150). The method according to claim 11.
13. Providing an image; and classifying the image using the machine learning model. The method according to claim 12.
14. An apparatus (100) for determining a machine learning model (102). An apparatus (100), characterized in that the apparatus (100) is configured to perform the method according to any one of claims 1 to 13. **Claim 15** A computer program, wherein the computer program includes computer-readable instructions for causing the computer to perform the steps in the method according to any one of claims 1 to 13 when the computer program is executed on the computer.
Citation Information
Patent Citations
Learning apparatus, analyzing system, learning method, and learning program
JP2019079392A
Nearline updates to personalized models and features
US20190188591A1
Machine learning engineering through hybrid knowledge representation
US20200265324A1