METHOD FOR QUALIFYING A MACHINE LEARNING MODEL
The method for qualifying machine learning models through defined criteria and test dataset evaluation addresses the lack of standardized trustworthiness assessment, enhancing safety and efficiency in ML model deployment and certification.
Patent Information
- Application Number
- DE102024201221
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-09
- Publication Date
- 2025-08-14
AI Technical Summary
Existing methods for evaluating machine learning models lack reliable and standardized procedures to ensure trustworthiness and quality, particularly in safety-critical applications, relying heavily on expert judgment and incomplete specifications.
A method for qualifying trained machine learning models by determining model behavior features and evaluating test datasets using defined criteria, enabling automated and standardized qualification processes.
This approach minimizes risks and uncertainties in ML model applications, enhances development efficiency, and ensures compliance with safety and legal requirements, facilitating self-certification and improving reproducibility across different systems.
Smart Images

Figure 00000011_0000 
Figure 00000011_0001 
Figure 00000012_0000
Abstract
Description
State of the art
[0001] The development and use of machine learning methods and associated models is increasingly gaining attention in various fields of technology. Their use extends across a wide range of applications, including autonomous driving, robotics, medical diagnostics, speech recognition, tools, household products, and many more. The trustworthiness and quality of such models are crucial, especially when used in safety-critical areas and where errors in the model's functionality can have serious consequences.
[0002] The evaluation of machine learning (ML) models and the assessment of their performance are central aspects in the development and application of such systems. The qualitative evaluation of ML models is carried out based on a validation and a test data set. These data sets are typically obtained by randomly dividing a total data set into a training, a validation, and a test data set, often according to a heuristic split ratio. A suitability assessment of the resulting test data set is usually performed by experts.
[0003] After the model has been trained on the training data, it is typically tested on the validation data to see how well it generalizes. This can, for example, allow for the detection of overfitting or the optimal combination of hyperparameters (e.g., learning rate, number of layers). Test data is a separate and independent dataset used for the final evaluation of model performance after the model has been optimized with the training and validation data. The test data was not presented to the model during the training process and is used to evaluate the model's ability to generalize to new, unknown data.
[0004] In addition to the quantitative performance criteria used in evaluating ML models, experts also play an important role. Their experience and knowledge provide in-depth insight into the trustworthiness and quality of models. Experts use checklists to evaluate qualitative criteria that are crucial for the trustworthiness of ML models.
[0005] For a comprehensive evaluation of ML models, statements must be made about the desired behavior of the ML model and the expected operating conditions. For ML models, specifications for this are usually incomplete. Assumptions about the desired behavior and operating conditions are usually indirectly confirmed through visual inspection of samples and expert approval.
[0006] Given the growing importance of ML models and their diverse applications, there is an urgent need for reliable evaluation methods to ensure that these models meet the required quality and trustworthiness standards. Disclosure of the invention
[0007] A first general aspect of the present disclosure relates to a method for qualifying a trained machine learning model. The method includes receiving a trained machine learning model, determining one or more model behavioral characteristics, performing an evaluation of a test data set based on one or more test data criteria, and determining a qualification result based on the one or more model behavioral characteristics and the evaluation of the test data set.
[0008] A second general aspect of the present disclosure relates to a computer system configured to perform the method for qualifying a trained machine learning model according to the first general aspect (or an embodiment thereof).
[0009] A third general aspect of the present disclosure relates to a computer program configured to perform the method for qualifying a trained machine learning model according to the first general aspect (or an embodiment thereof).
[0010] A fourth general aspect of the present disclosure relates to a computer-readable medium or signal storing and / or containing the computer program according to the third general aspect (or an embodiment thereof).
[0011] The method proposed in this disclosure according to the first general aspect (or an embodiment thereof) can serve to provide a method for qualifying a trained machine learning (ML) model. These methods can help minimize potential risks and uncertainties in the application of ML models while simultaneously making development and implementation more efficient in shortened development cycles. In examples, the qualification method can help implement the release of ML models with regard to security goals, compliance with legal requirements, and / or operational risk.A further advantage may be that the method can be used in the context of various technical functions and systems, such as autonomous driving functions, control unit functions in vehicles, such as control unit functions that replace physical sensors, calculate correction values or calculate abstract system states, (cloud-based) monitoring of vehicle fleets, and / or ML-based systems in the field of tools and / or household products.
[0012] The techniques of the present disclosure can contribute to standardization in the qualification of ML models. Various ML model architectures can be qualified using the disclosed techniques. A further advantage can be that the standardized approach can enable the automation of ML model qualification, for example, computer-implemented. In examples, industrial self-certification can be enabled. This can include the ability of ML models that are integrated into complex and specific system environments and that could only be replicated by external certification bodies with considerable technical and time-consuming effort to certify themselves within the existing (target) system environment. Furthermore, the comparability and reproducibility of ML models in different products and / or product generations can be improved.A uniform definition of test data criteria allows for the consistent use of reference models to qualify the test data set. The techniques of the present disclosure can enable the modular implementation of the method and can enable the addition of additional model behavioral characteristics and / or test data criteria regarding new developments and domain properties. The method can contribute to reducing (safety-)technical risks when using ML models in technical systems.
[0013] Some terms are used in this disclosure as follows: A "machine learning model" can include any dataset that is trained to learn specific patterns in datasets and extract patterns, insights, and predictions from them. A machine learning model can be trained using, for example, supervised learning, unsupervised learning, reinforcement learning, semi-supervised learning, and / or transfer learning. A machine learning model may include, for example, a linear regression model, a logistic regression model, a support vector machine (SVM), a k-nearest neighbors classifier, decision trees and random forests, a k-means clustering model, a neural network such as a recurrent neural network (RNN), a long short-term memory network (LSTM), and / or a transformer model.Furthermore, a machine learning model can include algorithms for Q-learning, linear interpolation, nearest neighbor interpolation, and / or principal component analysis.
[0014] A "vehicle" can be any device that transports passengers and / or cargo. A vehicle can be a motor vehicle (e.g., a car or a truck), but also a rail vehicle. A vehicle can also be a motorized two- or three-wheeler. However, floating and flying devices can also be vehicles. Vehicles can be at least partially autonomous, autonomously operating, or assisted. Short description of the characters Fig. 1-A to 1-B schematically illustrates a procedure for qualifying a trained machine learning model. Fig. 2 schematically illustrates an exemplary qualification matrix according to one or more embodiments of the present disclosure. Detailed description
[0015] Fig. 1-A and Fig. 1-B are flowcharts that include possible method steps of the method 100 for qualifying a trained machine learning model. The method 100 for qualifying a trained machine learning model includes receiving 110 a trained machine learning model. The method further includes determining 120 one or more model behavioral characteristics 13. The method further includes performing 130 an evaluation of a test data set based on one or more test data criteria 14. The method further includes determining 150 a qualification result based on the one or more model behavioral characteristics 13 and the evaluation of the test data set.
[0016] In examples, determining 120 one or more model behavioral characteristics 13 may include specifying one or more model behavioral characteristics 13 that are examined as part of the qualification of the ML model. For this purpose, one or more test data criteria 14 may subsequently be specified, with the aid of which the test data set is evaluated. As in Fig. 2, in examples, the one or more test data criteria 14 may be based on the one or more model behavioral characteristics 13. When qualifying the machine learning model, it may be advantageous not only to evaluate the output result of the machine learning model, but also to evaluate the test data set with respect to the model behavioral characteristic 13 under investigation.
[0017] In some examples, the method 100 for qualifying a trained machine learning model can be understood as a chain of reasoning. Fig.2 shows an exemplary qualification matrix 10. In examples, the quality objective 11 of an overall system is to be achieved. The overall system can comprise one or more system components. The machine learning model can be a component of the one or more system components. A quality objective 12 can be specifically defined for the machine learning model. To qualify the machine learning model, the following method 100 of an embodiment of the present disclosure can be carried out. In a first level, the one or more model behavior characteristics 13 can be used to determine whether the machine learning model achieves relevant (desired) functional objectives. This can be examined based on a qualified test data set. In a second level, the evaluation of the test data set is carried out using the one or more test data criteria 14.For this purpose, one or more reference models can be defined that can be applied to the test data set to obtain a quantification of one or more test data criteria 14. In a third level, the inaccuracies in the reference models are analyzed by comparing them with operational data. If necessary, the test data set can be recreated or adapted. In examples, the new test data set can be generated with adapted reference models. This approach can be advantageous for achieving a continuous chain of reasoning for qualifying the machine learning model.
[0018] In examples, one or more test data criteria 14 may be determined for each of the one or more model behavioral characteristics 13. In examples, the method may include determining 120 a plurality of model behavioral characteristics 13. In this case, different test data criteria 14 may be determined for each of the plurality of model behavioral characteristics 13. In each case, a different number and type of the one or more test data criteria 14 may be determined for the one or more model behavioral characteristics 13. For example, for a first model behavioral characteristic 13, 1, 2, ..., 10, or more test data criteria 14 may be determined, and for a second model behavioral characteristic 13, 1, 2, ..., 10, or more test data criteria 14 may be determined.In examples, the one or more model behavior characteristics may comprise properties of the machine learning model that are directed toward the (desired) function of the machine learning model. In examples, the one or more model behavior characteristics may comprise at least one of an objective function, domain robustness, generalization behavior of the network, input context, or output context. For example, the objective function may be described by a previously defined function of the machine learning model. For example, the objective function may be described by physical model knowledge, by simulation models, and / or semantic functional relationships, for example, in relation to the overall system. For example, domain robustness may be described by the relationship between the variances of input variables of the machine learning model and the invariances of output variables of the machine learning model.For example, the generalization behavior of the network can be described by the acceptable maximum deviation from the (test) dataset and / or by the maximum permissible interpolation behavior of the machine learning model. In examples, the input context can be described by previously defined operating ranges of the machine learning model, by safety-critical boundaries, and / or by distances from known datasets. In examples, the output context can be described by physical boundaries, by previously defined operating ranges of the machine learning model, by safety-critical boundaries, and / or by distances from known datasets.
[0019] In examples, the one or more test data criteria 14 may include at least one of coverage, input context validity, input distribution, output context validity, output distribution, functional accuracy, concept drift, or dataset dependency.
[0020] In examples, the one or more test data criteria may be evaluated by comparing a quantification of the one or more test data criteria on the test data set with the quantification of the same or more test data criteria on a data set of a model serving as a reference / pattern. For example, based on the data distribution of the reference / pattern data set, the data distribution of the test data set may be evaluated as a test data criterion.
[0021] In examples, the degree of coverage can include a criterion for describing the content completeness of the test data set with regard to relevant properties. For example, the degree of coverage can be evaluated based on a classification, e.g., with a maximum permissible variance of the nearest neighbor points of the output values in the test data set. In examples, the degree of coverage can be evaluated based on a percentage. For example, the degree of coverage can be evaluated based on a percentage of the contents of the test data set, e.g., compared to a data set of a model serving as a pattern / reference. An example of this could be traffic scenarios (e.g., weather, roads, lighting, pedestrians) in autonomous driving.
[0022] In examples, the validity of the input context can include a criterion for describing the validity of test data input data in the respective system context. For example, the validity of the input data of the test data set can be evaluated using known properties in the (input) data sets, e.g., from resampling or unlabeled data. For example, the validity of the input data of the test data set can be evaluated by accepting exemplary sample / reference data sets with permitted and / or prohibited data points, and / or by interpolation and extrapolation rules with respect to existing data points. In examples, physical model knowledge, e.g., from simulations related to the overall system, can also be used to evaluate the validity of the input data in the test data set in the respective system context.In another example, the validity of the input data in the test data set can be evaluated using a semantic rule set that can define possible or impossible input data combinations (scenarios). In some examples, the validity of the input context can be evaluated based on a percentage. For example, the validity of the input context can be evaluated based on a context distribution of the test data set, for example, compared to a data set of a context model serving as a pattern / reference. An example of this could be the operating range of an internal combustion engine.
[0023] In examples, the input distribution can include a criterion for describing the distribution of the test data input. For example, the input distribution can be evaluated based on the acceptance of an exemplary distributed data set (e.g., from measurements with a known data generation process), by means of a statistical analysis (e.g., outside temperature distribution), by means of an established domain standard, such as operating points of a standardized driving cycle, and / or legal requirements. In examples, the input distribution can be evaluated based on a kernel-Stein discrepancy. For example, the input distribution can be evaluated based on a data distribution of the test data set compared to a data set of a model serving as a pattern / reference. An example of this can be the frequency distribution of wall materials in measurements with a power tool locating device, a so-called “wall scanner.”
[0024] In examples, the validity of the initial context can include a criterion for describing the admissibility of output data from the machine learning function in the respective system context. In examples, the validity of the initial context can include a criterion for describing the admissibility of test data output data (target values) in the respective context of the overall system. For example, the validity of the initial context can be evaluated using admissible / prohibited thresholds from safety requirements (e.g., based on physical models). For example, the validity of the initial context can be evaluated using the acceptance of exemplary data sets from models serving as reference / patterns with permitted / prohibited output values, interpolation rules, and / or extrapolation rules with respect to existing data points.In another example, the validity of input data can be evaluated using a semantic set of rules that can define possible or impossible input data combinations (scenarios).
[0025] In examples, the output distribution can include a criterion for describing the distribution of output data. In examples, the output data can include the target values in the test data set to be evaluated. For example, the output distribution can be evaluated using the acceptance of an exemplary distributed data set (for example, from measurements with a known data generation process). In examples, the output distribution can be evaluated based on a kernel-Stein discrepancy. For example, the output distribution can be evaluated based on a data distribution of the test data set compared to a data set of a model serving as a pattern / reference. An example of this can be the frequency distribution of wall materials in measurements with a power tool locating device, a so-called "wall scanner."
[0026] In examples, functional accuracy may include a criterion for the respective model behavioral characteristic to describe the content correctness of the input data with respect to an assumed ground truth. For example, functional accuracy may be evaluated based on the objective function, domain robustness, the generalization behavior of the network, the input context, and / or the output context.
[0027] In examples, concept drift may include a criterion for describing the transferability of a test data domain and / or a training data domain to an operating data domain. For example, concept drift may be evaluated by interpolating the test data set and comparing the interpolation with labeled operating data. In examples, concept drift may be evaluated based on a kernel-Stein discrepancy. For example, concept drift may be evaluated based on a data distribution of the test data set compared to a data set of a model serving as a pattern / reference. An example of this may be the data distribution of the in-vehicle test data set compared to a test bench test data set of a temperature estimation model for the stator temperature of an electric vehicle.
[0028] In examples, the dataset dependency may include a criterion for describing the dependencies of the test dataset on other datasets, in particular the training dataset.
[0029] In examples, the method may include determining 140 one or more evaluation metrics based on the one or more model behavioral characteristics 13. In examples, determining 150 the qualification result may be further based on the one or more evaluation metrics. For example, the one or more evaluation metrics may include a misclassification rate. In examples, the one or more evaluation metrics may include a confidence interval, for example, based on an application-specific metric. In examples, the one or more evaluation metrics may include a deviation frequency. For example, the one or more evaluation metrics may include a deviation frequency against a robustness requirement of a model serving as a reference, for example, robustness against image rotation of an optical pass / fail component analysis in manufacturing.In examples, the one or more evaluation metrics may include a deviation. For example, the one or more evaluation metrics may include a deviation from functional requirements from a reference model. An example would be the cumulative fuel consumption deviation for an ECU function used to estimate fuel consumption in a vehicle. In examples, the one or more evaluation metrics may include a mean squared error (MSE).
[0030] In examples, the method may include using the machine learning model if the qualification result falls within a clearance range. In examples, the method may include using the machine learning model if the one or more evaluation metrics fall within a clearance range. In examples, the qualification result may include a separate result for the one or more model behavioral characteristics 13 and a separate result for the evaluation of the test data set. In examples, a clearance range may be specified for each of the one or more model behavioral characteristics 13 and for the evaluation of the test data set. In examples, using the machine learning model may include switching from a conventional method to a method based on the machine learning model. In examples, the switching may be performed in a computer-implemented manner in an operational environment.In examples, the approval range may include one or more first thresholds for the one or more model behavioral features 13 and / or one or more second thresholds for the one or more test data criteria 14. For example, the first threshold for the one or more model behavioral features 13 may be a threshold for the aforementioned misclassification rate.
[0031] In examples, the evaluation 130 of the test data set may include defining 131 one or more reference models for each test data criterion of the one or more test data criteria 14, applying 132 the one or more reference models to the test data set to obtain a test data quantification of at least one test criterion of the one or more test data criteria 14. In examples, the one or more reference models may include one or more of a data density estimation and / or data density function, semantic models, principal component analysis, linear interpolation, frequency distribution modeling, k-nearest neighbors classification, and / or nearest neighbor interpolation. For example, the data density estimation may be performed based on given data, based on kernel-based methods, based on histograms, and / or with respect to neighboring points using k-nearest neighbors.For example, semantic models can be represented in the form of ontologies and / or rule sets.
[0032] For example, for the test data criterion “coverage”, the one or more reference models may include the (test) data density estimation and / or semantic models.
[0033] For example, for the test data criterion "validity of the input context," the one or more reference models may comprise linear interpolation. In examples, the linear interpolation may be applied to a subspace generated by principal component analysis. For example, for the test data criterion "validity of the input context," the one or more reference models may comprise semantic models. For example, for the test data criterion "input distribution," the one or more reference models may comprise modeling of the frequency distribution. In examples, this may comprise partitioning the input data of the machine learning model. For example, for the test data criterion "validity of the output context," the one or more reference models may comprise semantic models.
[0034] For example, for the test data criterion “initial distribution,” the one or more reference models may include a data density function.
[0035] For example, for the test data criterion “concept drift,” the one or more reference models may include a nearest neighbor interpolation and / or a local linear interpolation based on the test data set and a comparison of the interpolation with labeled operational data.
[0036] In examples, the one or more test data criteria can be quantified using the one or more reference models. In examples, the quantifications of the one or more test data criteria can be combined into the test data quantification.
[0037] In some embodiments, the test data set may comprise a first original test data set and / or a generated test data set. In examples, the generated test data set may be generated by applying at least one reference model of the one or more reference models to the first original test data set. In examples, the generated test data set may be generated by selecting a set from a plurality of second original test data sets. For example, a reference model may comprise generating the test data set using a manifold model and / or an approximate k-nearest neighbors classifier. For example, a reference model for the test criterion "validity of the input context" may comprise generating the test data set using a manifold model and / or an approximate k-nearest neighbors classifier.
[0038] In some embodiments, the method 100 may further comprise applying 160 the one or more reference models to an operational data set to obtain an operational data quantification 15 of at least one test criterion of the one or more test data criteria 14. In some examples, the method may further comprise comparing 170 the test data quantification with the operational data quantification of the at least one test criterion of the one or more test data criteria 14 to obtain a comparison result 15. In examples, if the comparison result 15 lies outside a defined acceptance range, the method 100 may comprise generating 180 a new test data set and / or modifying the original test data set. The method 100 may further comprise performing 190 an evaluation of the new test data set based on the one or more test data criteria 14.In examples, the method 100 may further comprise determining 200 a new qualification result based on the one or more evaluation metrics and the evaluation of the new test data set. For example, if the comparison result 15 indicates an unacceptable deviation of the operating data from the test data set, determining 200 the new qualification result may be necessary to ensure that the machine learning model also delivers correct results during operation. In examples, the method 100 may comprise switching from a method based on the machine learning model to a conventional method if the comparison result 15 lies outside a defined acceptance range. This may, for example, represent a reverse processing of the method step already mentioned above of switching from a conventional method to a method based on the machine learning model.In examples, the application 160 of the one or more reference models to the operational data set and / or the comparison 170 of the test data quantification with the operational data quantification can be performed in live operation. For example, this can be performed continuously, at specific time intervals, multiple times a day, and / or by means of a trigger. In examples, the application 160 of the one or more reference models to the operational data set can be part of a surveillance and / or monitoring of the operational data and / or the model behavior in operational mode.
[0039] In examples, at least a subset of the operational dataset may include labeled data. For example, for the test criteria functional accuracy and / or concept drift, the operational dataset may include labeled data. For example, for the test criteria coverage, input context validity, and / or input distribution, labeled data may be superfluous.
[0040] In examples, the method 100 for qualifying the trained machine learning model and / or the machine learning model (qualified and / or unqualified) may be configured for execution in a vehicle, a robot, a building, a power tool, and / or a household appliance, and / or for controlling and / or monitoring a vehicle function, a robot function, a building automation function, a power tool automation function, and / or a household appliance automation function. In examples, the method 100 may include loading the qualified machine learning model into a computer system in a vehicle, a robot, a building, a power tool, and / or a household appliance.
[0041] In examples, a robot may comprise a vehicle (autonomously or assisted driving). For example, the vehicle function may be a function for autonomous and / or assisted driving. In some examples, the method 100 for qualifying the trained machine learning model and / or the machine learning model (qualified and / or unqualified) may be designed to execute on a computer system of a vehicle (e.g., an autonomous, highly automated, or assisted driving vehicle). For example, the computer system may be implemented locally in the vehicle or (at least partially) implemented in a backend that is communicatively connected to the vehicle. For example, the computer system may comprise a control unit on which the method 100 for qualifying the trained machine learning model and / or the machine learning model may be executed.In some examples, the vehicle may include a computer system with a communication interface that enables communication with a backend. For example, the method 100 for qualifying the trained machine learning model and / or the machine learning model may be executed in this backend. In examples, control unit functions in vehicles may include or access the method 100 for qualifying the machine learning model and / or the machine learning model, for example, when they are executed remotely in a cloud. For example, control unit functions that replace physical sensors, calculate correction values, and / or abstract system states (such as aging) may include or access the method 100 for qualifying the machine learning model and / or the machine learning model, for example, when they are executed remotely in a cloud.
[0042] In other examples and as indicated above, the method 100 for qualifying the trained machine learning model and / or the machine learning model (qualified and / or unqualified) may be configured for execution in a robot, and / or configured for controlling and / or monitoring a robot function (in particular for controlling and / or monitoring a movement function of a robot). In some examples, the method 100 for qualifying the trained machine learning model and / or the machine learning model may be executed on a computer system of a robot. For example, the computer system may be implemented locally in the robot or (at least partially) implemented in a backend communicatively connected to the robot.
[0043] In one example, the method 100 for qualifying the trained machine learning model and / or the machine learning model (qualified and / or unqualified) may be designed for execution in a building and / or for controlling building functions (in particular for controlling building automation functions). For example, the building function may be a function for regulating room temperature, lighting, and / or security devices. In some examples, the method 100 for qualifying the trained machine learning model and / or the machine learning model may be designed for execution on a computer system within the building. For example, the computer system may be implemented locally in the building or (at least partially) implemented in a backend that is communicatively connected to the building.For example, the computer system may include a control system or a building automation control device on which the method 100 for qualifying the trained machine learning model and / or the machine learning model may be executed. In examples, the building may have a computer system with a communication interface that enables communication with an external backend. For example, the method 100 for qualifying the trained machine learning model and / or the machine learning model may be executed in this backend. In examples, input data of the machine learning model and / or operational data may be based on information such as room temperature, brightness, or the presence of people.In some cases, input data from the machine learning model and / or operational data may include a relative temperature difference, illuminance, or distance to a specific location or object within the building. Information may, in some examples, originate from a network, such as sensor data or settings from other buildings or building components. This information may be provided through communication between buildings or building components or via an external backend.
[0044] In other examples, the method 100 for qualifying the trained machine learning model and / or the machine learning model (qualified and / or unqualified) may be configured for execution in a power tool, and / or configured for controlling and / or monitoring a power tool function (in particular for controlling and / or monitoring a work function of the power tool). In some examples, the method 100 for qualifying the trained machine learning model and / or the machine learning model may be executed on a computer system of the power tool. For example, the computer system may be implemented locally in the power tool or (at least partially) implemented in a backend communicatively connected to the power tool.
[0045] In other examples, the method 100 for qualifying the trained machine learning model and / or the machine learning model (qualified and / or unqualified) may be configured for execution in a household appliance, and / or configured for controlling and / or monitoring a household appliance function (in particular, for controlling and / or monitoring a working function of the household appliance). In some examples, the method 100 for qualifying the trained machine learning model and / or the machine learning model may be executed on a computer system of the household appliance. For example, the computer system may be implemented locally in the household appliance or (at least partially) implemented in a backend communicatively connected to the household appliance.
[0046] In further examples, the method 100 for qualifying the trained machine learning model and / or the machine learning model (qualified and / or unqualified) may be configured for execution in a machine tool, a personal assistant, an access control system, and / or a medical device, for example, for medical imaging, and / or may be accessible via a network. In examples, the method 100 may include loading the qualified machine learning model into a computer system of a machine tool, a personal assistant, an access control system, and / or a medical device.
[0047] In examples, the machine learning model can be used to analyze audio data and / or video data. In examples, the machine learning model can be used to classify sensor data, detect objects, or semantically segment sensor data, for example, with respect to traffic signs, road surfaces, pedestrians, vehicles, electrical wires, water pipes, gas pipes, biochemical reactions, or physical blockages, for example, in path planning. In examples, the machine learning model can be used to determine continuous values, for example, using regression with respect to distance, speed, acceleration, gas concentration, azimuth, altitude, fuel consumption, emissions, temperature, current, voltage, aging, yaw rate, humidity, pressure, vibration, and / or cloud cover. In examples, the machine learning model can be used for object tracking, for example, based on pixel attributes.In further examples, the machine learning model can be used for anomaly detection, optionally using an autoencoder. In examples, the machine learning model can be used to estimate sensor signals, in particular sensor signals from a sensor measuring speed, angular rate, current, voltage, temperature, pressure, air pressure, weight, deformation, flow of a liquid or gas, or composition of a gas.
[0048] For example, the labeled operating data mentioned above can be obtained using sensors. This could mean, for example, that the output signals of a controller are recorded using sensors, so that a labeled operating data set can be obtained from the input data and target data.
[0049] Also disclosed is a computer system configured to execute method 100 for qualifying a machine learning model. The computer system may include at least one processor and / or at least one main memory. The computer system may further include a (non-volatile) memory.
[0050] Also disclosed is a computer program designed to execute the method 100 for qualifying a machine learning model. The computer program can be in interpretable or compiled form, for example. It can be loaded (even in parts) into the RAM of a computer for execution, for example, as a bit or byte sequence.
[0051] Further disclosed is a computer-readable medium or signal that stores and / or contains the computer program or at least a portion thereof. The medium may, for example, comprise one of RAM, ROM, EPROM, HDD, SDD, etc., on / in which the signal is stored.
Claims
[1] A method (100) for qualifying a trained machine learning model, comprising: - receiving (110) a trained machine learning model, - determining (120) one or more model behavior characteristics (13), - carrying out (130) an evaluation of a test data set based on one or more test data criteria (14), and - Determining (150) a qualification result based on the one or more model behavior characteristics (13) and the evaluation of the test data set. [2] A method according to claim 1, wherein the method comprises - determining (140) one or more evaluation metrics based on the one or more model behavioral characteristics (13), and wherein determining (150) the qualification result is further based on the one or more evaluation metrics. [3] A method according to claim 1 or 2, wherein the method further comprises: - Use the machine learning model when the qualification result falls within a clearance range. [4] Method according to claim 3, wherein the release range comprises one or more first threshold values for the one or more model behavioral characteristics (13) and / or one or more second threshold values for the one or more test data criteria (14). [5] The method according to claim 3 or 4, wherein using the machine learning model comprises switching from a conventional method to a method based on the machine learning model. [6] Method according to one of the preceding claims, wherein the evaluation (130) of the test data set comprises - defining (131) one or more reference models for each test data criterion of the one or more test data criteria (14), - applying (132) the one or more reference models to the test data set to obtain a test data quantification of at least one test criterion of the one or more test data criteria (14). [7] The method according to claim 6, wherein the test data set comprises a first original test data set and / or a generated test data set, wherein the generated test data set is generated by applying at least one reference model of the one or more reference models to the first original test data set and / or wherein the generated test data set is generated by a selection set from a plurality of second original test data sets. [8] Method according to claim 6 or 7, wherein the method (100) further comprises - applying (160) the one or more reference models to an operational data set to obtain an operational data quantification (15) of at least one test criterion of the one or more test data criteria (14), - comparing (170) the test data quantification with the operating data quantification of the at least one test criterion of the one or more test data criteria (14) to obtain a comparison result (15) and, if the comparison result (15) lies outside a defined acceptance range, - generating (180) a new test data set and / or modifying the original test data set, - carrying out (190) an evaluation of the new test data set on the basis of the one or more test data criteria (14), - Determining (200) a new qualification result based on the one or more evaluation metrics and the evaluation of the new test data set. [9] The method of claim 8, wherein at least a subset of the operational data set comprises labeled data. [10] The method of any preceding claim, wherein the one or more model behavioral features comprise at least one of an objective function, domain robustness, generalization behavior of the network, input context, or output context. [11] Method according to one of the preceding claims, wherein the one or more test data criteria (14) comprise at least one of coverage level, validity of the input context, input distribution, validity of the output context, output distribution, functional accuracy, concept drift, or data set dependency. [12] Method (100) according to one of the preceding claims, wherein the method 100 for qualifying the trained machine learning model and / or the machine learning model is designed for execution in a vehicle, a robot, a building, a power tool and / or a household appliance, and / or for controlling and / or monitoring a vehicle function, a robot function, a building automation function, a power tool automation function, and / or a household appliance automation function. [13] A computer system adapted to carry out the method (100) for qualifying a machine learning model according to any one of the preceding claims 1 to 12. [14] A computer program comprising instructions which, when executed by a computer system, cause the computer program to execute the method (100) for qualifying a machine learning model according to any one of the preceding claims 1 to 12. [15] A computer-readable medium or signal storing and / or containing the computer program according to claim 14.