Methods, systems, and computer-readable media for providing a testing and validation platform for testing and validating a machine learning lifecycle
Patent Information
- Application Number
- DE102025106915
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-22
- Filing Date
- 2025-02-24
- Publication Date
- 2025-09-04
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
PRIORITY CLAIM
[0001] This application claims priority to U.S. Provisional Patent Application Serial No. 63 / 559,569, filed February 29, 2024, the disclosure of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0002] The subject matter described herein relates to testing a machine learning model. More specifically, the subject matter relates to methods, systems, and computer-readable media for providing a testing and validation platform for testing and validating a machine learning lifecycle. BACKGROUND
[0003] As the use of artificial intelligence (AI), and especially machine learning, increases, there is a need for ongoing quality assurance through testing and documentation. During the lifecycle of a product or solution containing a machine learning model (MLM), the MLM will be updated several times to adapt to changes in the solution domain. All different models must be tested and validated to ensure that the required quality standards are met for each stage of the development cycle.
[0004] MLM testing and validation procedures do not provide a uniform set of tests for different stages of an MLM and thus cannot ensure that every stage and every aspect of the lifecycle is addressed accordingly. Each test is independent of the others and separate from different stages and iterations. Without a proper process and test sequencing, the development time is extended and the susceptibility to errors increases. Furthermore, standardized and complete documentation cannot be ensured.
[0005] Comprehensive knowledge of the domain and machine learning (ML) is required to use existing methodologies, challenging teams of domain experts and ML experts to find common ground to address problems in both fields. This increases the likelihood of errors and the time required to repeat the development of customized models. There is a need for continuous testing and validation of MLMs across all stages and multiple iterations, from the initial planning of an MLM to the end of the solution's life cycle, by providing an iterative validation and testing process that connects all stages of the development cycle while remaining compatible with established software development processes such as the V-model or Waterfall, regardless of the domain. SUMMARY
[0006] Methods, systems, and computer-readable media for providing a testing and validation platform for testing and validating a machine learning lifecycle are disclosed. An exemplary method for providing a testing and validation platform for testing and validating a machine learning lifecycle includes analyzing, at the testing and validation platform, a received data set by performing at least one statistical test on the data set, the data set comprising data representative of samples each having features and a corresponding target. The method further includes displaying, by the testing and validation platform, results of the analysis. The method further includes receiving, at the testing and validation platform, a selected at least one feature of the features for training a machine learning model.The method further comprises validating, by the testing and validation platform, subsets of the dataset, wherein the subsets comprise an expert dataset, a training dataset, and a validation dataset and / or a test dataset, wherein each subset comprises data representative of different sample values. The method further comprises testing, by the testing and validation platform, and after the machine learning model has been trained, the machine learning model by using the expert dataset to determine at least one metric for each sample value of the expert dataset, and comparing each metric of the at least one metric to a corresponding defined confidence interval.
[0007] According to another aspect of the subject matter described herein, if the metric determined from the expert dataset is within the corresponding defined confidence interval, the method further comprises testing, by the testing and validation platform, the machine learning model by determining at least one metric for each sample of the validation dataset and / or the test dataset.
[0008] According to another aspect of the subject matter described herein, the method further comprises comparing the metric determined from the expert data set with the metric determined from the validation data set and / or the test data set.
[0009] According to another aspect of the subject matter described herein, the method further comprises defining clusters of the metric determined from the expert data, determining the cluster boundaries, defining clusters of the metric determined from the validation data set and / or the test data set, and comparing the clusters of the metric determined from the expert data set to the cluster boundaries.
[0010] According to another aspect of the subject matter described herein, the method further comprises determining that the machine learning model is secure if the clusters of the metric determined from the validation dataset and / or the test dataset are within the cluster boundaries.
[0011] According to another aspect of the subject matter described herein, the method further comprises determining a value for the machine learning model based on the percentage of the metric determined from the validation dataset and / or the test dataset that lies within the cluster boundaries.
[0012] According to another aspect of the subject matter described herein, the method further comprises iteratively testing the machine learning model by determining metrics from the expert dataset, determining metrics from the validation dataset and / or the test dataset, and comparing the metric from the expert dataset with the metric from the validation dataset and / or the test dataset.
[0013] According to another aspect of the method described herein, results from each test of each iteration of the machine model are stored for comparison.
[0014] According to another aspect of the method described herein, the at least one statistical test determines correlations between the features of the data set and the displayed results include the determined correlations.
[0015] According to a further aspect of the method described herein, the at least one statistical test identifies a degree of influence that each of the features has on the at least one target, and the displayed results include the features and the corresponding identified degree of influence.
[0016] An exemplary system for providing a testing and validation platform for testing and validating a machine learning lifecycle includes a testing and validation platform configured to analyze a received data set by performing at least one statistical test on the data set, the data set comprising data representative of samples each having features and a corresponding target. The system is further configured to display results of the analysis. The system is further configured to receive a selected at least one feature of the features for training a machine learning model.The system is further configured to validate subsets of the dataset, wherein the subsets comprise an expert dataset, a training dataset, and a validation dataset and / or a test dataset, wherein each subset comprises data representative of different sample values. The system is further configured to test the machine learning model, and after the machine learning model has been trained, by using the expert dataset to determine at least one metric for each sample value of the expert dataset and comparing each metric of the at least one metric to a corresponding defined confidence interval.
[0017] According to another aspect of the system described herein, if the metric determined from the expert dataset is within the corresponding defined confidence interval, the testing and validation platform is configured to test the machine learning model by determining at least one metric for each sample of the validation dataset and / or the test dataset.
[0018] According to another aspect of the system described herein, the testing and validation platform is configured to compare the metric determined from the expert dataset with the metric determined from the validation dataset and / or the test dataset.
[0019] According to another aspect of the system described herein, the testing and validation platform is configured to define clusters of the metric determined from the expert dataset, determine the cluster boundaries, define clusters of the metric determined from the validation dataset and / or the test dataset, and compare the clusters of the metric determined from the expert dataset to the cluster boundaries.
[0020] According to another aspect of the system described herein, the testing and validation platform is configured to determine that the machine learning model is secure if the clusters of the metric determined from the validation dataset and / or the test dataset are within the cluster boundaries.
[0021] According to another aspect of the system described herein, the testing and validation platform is configured to determine a value for the machine learning model based on the percentage of the metric determined by the validation dataset and / or the test dataset that lies within the cluster boundaries.
[0022] According to another aspect of the system described herein, the testing and validation platform is configured to iteratively test the machine learning model by determining a metric from the expert dataset, determining a metric from the validation dataset and / or the test dataset, and comparing the metric from the expert dataset with the metric from the validation dataset and / or the test dataset.
[0023] According to another aspect of the system described herein, the at least one statistical test identifies a degree of influence that each of the features has on the at least one target, and the displayed results include the features and the corresponding identified degrees of influence.
[0024] An exemplary non-transitory computer-readable medium has executable instructions stored thereon that, when executed by at least one processor of at least one computer, cause the at least one computer to perform steps comprising analyzing a received data set by performing at least one statistical test on the data set, the data set comprising data representative of samples each having features and a corresponding target. The non-transitory computer-readable medium is further configured to display results of the analysis. The non-transitory computer-readable medium is further configured to receive a selected at least one feature of the feature for training a machine learning model.The non-transitory computer-readable medium is further configured to validate, through the testing and validation platform, subsets of the dataset, wherein the subsets comprise an expert dataset, a training dataset, and a validation dataset and / or a test dataset, each subset comprising data representative of different sample values. The non-transitory computer-readable medium is further configured to test the machine learning model, and after the machine learning model has been trained, by using the expert dataset to determine at least one metric for each sample value of the expert dataset, and comparing each metric of the at least one metric to a corresponding defined confidence interval.
[0025] According to another aspect of the non-transitory computer-readable medium described herein, if the metric determined from the expert data set is within the corresponding defined confidence interval, the non-transitory computer-readable medium is further configured to test the machine learning model by determining at least one metric for each sample of the validation data set and / or the test data set.
[0026] The subject matter described herein may be implemented in software in combination with hardware and / or firmware. For example, the subject matter described herein may be implemented in software executed by a processor (e.g., a hardware-based or physical processor). In an example implementation, the subject matter described herein may be implemented using a non-transitory computer-readable medium having stored thereon computer-executable instructions that, when executed by the processor of a computer, direct the computer to perform the steps. Example computer-readable media suitable for implementing the subject matter described herein include non-transitory devices, such as disk storage devices, chip memory devices, programmable logic devices, such asfield-programmable gate arrays and application-specific integrated circuits. Furthermore, a computer-readable medium implementing the subject matter described herein may be located on a single device or computing platform or may be distributed across multiple devices or computing platforms. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The subject matter described herein will now be explained in more detail with reference to the accompanying drawings. They show: Fig. 1 is a block diagram illustrating an example system for providing a testing and validation platform for testing and validating a machine learning lifecycle; Fig. 2 a schematic diagram of five stages of testing and validation; Fig. 3A is an exemplary graphical user interface display of a configuration for a YData profile report generated by an open source YData profile analyzer; Fig. 3B is an exemplary graphical user interface (GUI) display of an overview of a YData profile report generated by an open source YData profile analyzer; Fig. 4A shows an example GUI display of a report summary for a mutual information test; Fig. 4B is an example GUI display of a mutual information value diagram; Fig. 5 an example GUI display of a value analysis using metrics; Fig. 6A-6G together show an example JSON test result of a mutual information metric; and Fig. 7 is a flowchart illustrating an example method for deploying a testing and validation platform for testing and validating a machine learning lifecycle. DETAILED DESCRIPTION
[0028] The subject matter described herein includes methods, systems, and computer-readable media for providing a testing and validation platform for testing and validating a machine learning lifecycle. The testing and validation platform provides iterative testing throughout the lifecycle of the MLM, also referred to herein as the model. The lifecycle is divided into five stages, with each stage containing a set of tests to evaluate the datasets and the model. The stages use metrics for analysis and interpretability by comparing metric results to an expert dataset to determine the quality of the MLM. Test results, as well as recommendations for improving the performance of a particular dataset or model, are retained along with the test configuration, enabling automation of these tests for later iterations of the five-stage process.The methodology allows for easy adaptation of further tests if the area or state of the art requires additional testing, enabling the coordination of testing and validation across a wide range of domains.
[0029] This standardized result schema enables automated comparison between different models, thus significantly accelerating the development process. For example, if the same system is deployed in a continuous integration / continuous delivery (CI / CD) pipeline, several different model configurations could be trained, and based on the automatically generated test results, only the best models could be used for the next stage, reducing the amount of manual testing and validation. For example, it is possible for an experienced ML expert to set up the test configuration for an ML project, thus enabling domain experts to analyze the results provided by the tests with the help of easy-to-understand visualizations and explanatory texts.
[0030] The testing and validation platform creates a standardized way to test and validate MLMs that AI and domain experts can follow, establishing a common foundation for domain and ML experts. Relying on statistics and easy-to-understand visualizations, it facilitates communication between the two groups, leading to faster and higher-quality results. The testing and validation platform creates a common foundation for domain experts experienced in the industry of the problem the MLM is designed to solve, and AI experts experienced in AI development and implementing MLMs.
[0031] The test and validation platform documents the test results, allowing users to repeat tests with the same configuration for a different model iteration. For example, in an MLM deployed to detect pedestrians at intersections, a possible special case might include a pedestrian pushing a bicycle. If the deployed MLM is unable to capture this special case, the model must be updated to capture such special cases. It is essential that this updated model passes all quality tests. With the test and validation platform, the engineer can automatically repeat tests for all stages using the adapted datasets and models, enabling qualified comparison of MLMs and export of results for documentation purposes.
[0032] The testing and validation platform can provide a checklist indicating which lifecycle stages and corresponding testing have been completed, confirming that all security aspects and necessary tests have been performed and documented. Using this methodology, it is possible to align the model's development and lifecycle with published regulatory guidelines, such as the EU AI Act and the US AI Act, and standards such as DIN SPEC 13266.
[0033] Fig. 1 is a block diagram illustrating an example system 100 for providing a test and validation platform for testing and validating a machine learning lifecycle. The system 100 includes a test and validation platform 102 having at least one processor 104 and a memory 106. The test and validation platform 102 may include, without limitation, a microcontroller, microprocessor, digital signal processor (DSP), and / or system on a chip (SoC) as described herein. The test and validation platform 102 may include a single computing device operating independently or may include two or more computing devices operating collaboratively, in parallel, sequentially, or the like; two or more computing devices may be included together in a single computing device or in two or more computing devices.The test and validation platform 102, using the processor 104 and the memory 106, may be configured to perform any of the steps described herein. The test and validation platform 102 may include a database 108 from which the test and validation platform 102 may store, access, manipulate, and retrieve information, such as datasets and MLMs. The database 108 may include a cloud drive. The test and validation platform 102 may communicate with a machine learning model 110, which may be stored locally, such as in the memory 106 or the database 108, or stored remotely, such as on one or more other computing devices with which the test and validation platform 102 is configured to communicate.
[0034] Fig. Figure 2 shows a schematic diagram 200 of the five stages of testing and validation that the test and validation platform 102 shown in Fig. 1, iteratively over the course of the MLM's life cycle. The testing and validation platform 102 may provide one or more tests for each stage. In one aspect of the described subject matter, the testing and validation platform 102 may provide a preselected set of one or more tests for each stage and automatically execute the tests for the corresponding stage. In another aspect of the described subject matter, a user may select one or more tests at each stage, and the testing and validation platform 102 may automatically execute the selected tests. The testing and validation platform 102 may also upload specific tests for the area in which the MLM operates. The testing and validation platform 102 may further provide recommendations based on the test results as described herein.The testing and validation platform 102 can be configured to automatically advance to the next stage once the MLM passes one or more tests in a current stage. A passing score can be based on user instructions, industry standards, security requirements, regulatory guidelines, and the like. If an issue arises in the subsequent stage that cannot be resolved in the current stage, the user can easily return to the previous stage, adjust the dataset and model accordingly, and repeat the stage's tests to ensure that all required quality checks are still met.
[0035] Stage 1 includes data and problem analysis. Users of the testing and validation platform 102 may include AI experts and experts in the relevant field for which the MLM is being implemented, referred to herein as domain experts. The testing and validation platform 102 receives a dataset used to train, test, evaluate, and infer the MLM. The dataset includes data representing samples, and each sample includes variables, namely features, and at least one corresponding target.For example, each sample in a dataset for an MLM designed to predict the number of bicycles rented from a vehicle rental company at a given hour, which is the goal, can include features such as the time of day, the weather, the day of the week, and whether the day is a holiday. Each sample can also identify a value for the corresponding goal, which in this example is the number of vehicles rented within the corresponding hour. At Stage 1, the testing and validation platform 102 supports AI experts and domain experts in performing data analysis and problem analysis, respectively. Understanding the domain and the actual current problem is an essential part of machine learning.Domain experts are familiar with the domain and can provide insight to identify the task for the MLM, the goal and potential characteristics that affect the goal, special cases the MLM must pass, and applicable safety standards. Without a thorough understanding of the data at hand, biases and false assumptions can detrimentally impact the further development of the MLM. The 102 testing and validation platform supports developers by automating the data analysis process and detecting biases and hidden correlations in the data, as well as by providing a common ground for domain and ML experts, relying on statistics and easy-to-understand visualizations.
[0036] At Stage 1, users evaluate the dataset, which may include determining whether it contains sufficient data, how the data is scaled or what ranges of values are available, whether the dataset contains constant values, how the data is distributed, and the level of data quality. The testing and validation platform 102 provides one or more tests that it can automatically run on the received dataset and a corresponding report to assist users in their analysis.
[0037] The one or more tests may include third-party or open source tests, such as the YData profile analyzer that generates a YData profile report as described in Fig. 3A to 3B. The YData Profile report can highlight important information, allowing users to quickly understand data characteristics. For example, the YData Profile report can identify duplicate rows of data in the dataset, the degree of correlation between variables, such as between different features and between a feature and the target, and uneven representation for a variable.
[0038] Stage 2 includes feature engineering. In stage 2, the features that will be used for training are selected. The testing and validation platform 102 supports this decision by providing metrics to measure the influence of features on the target, measuring correlations, and recommending which features should be used. The testing and validation platform 102 analyzes the received data set by performing at least one statistical test on the data set. The statistical tests may determine correlations between the variables of the data set, such as between features and / or between features and the at least one target. The statistical tests may determine mutual information of the features, which includes identifying a degree of influence that each of the features has on the at least one target, as described in Fig. 4B. For example, the testing and validation platform 102 may determine whether each feature has a high, medium, low, or no / negligible correlation with the at least one target. Based on the results, the testing and validation platform 102 may provide a general recommendation to the user to remove the features with little or no information gain, such as those features that have little or no correlation with the target.
[0039] The testing and validation platform 102 displays results of the analysis, which may include the determined correlations between features and / or the at least one target. As an example, the results may include mutual information metrics. The displayed results may include the features and the corresponding identified influence levels. The results may be displayed in a diagram showing the mutual information for each feature. Depending on the domain, there may be additional knowledge leading to a different meaning of a feature that only domain experts can know. The resulting representation is a basis for discussing these issues and selecting the features that best fit the problem. The testing and validation platform 102 assists users in their selection of one or more features with the highest meaning, namely the highest correlation with the target.The testing and validation platform 102 receives a user selection of at least one feature.
[0040] To evaluate the MLM, data unknown to the model is required. The testing and validation platform 102 supports users in strictly separating datasets. A user can separate the datasets into different subsets, such as three, four, or more subsets of data. The subsets may include an expert dataset, a training dataset, a validation dataset, and / or a test dataset, where each subset has data representing different samples. The testing and validation platform 102 validates the subsets of the dataset, where the subsets include an expert dataset, a training dataset, a validation dataset, and / or a test dataset, where each subset has data representing different samples.Validating subsets may involve identifying and flagging duplicate samples between subsets of data, thereby preventing erroneous evaluation results caused by known data in the test set. Validating subsets may also involve ensuring that subsets of data, such as the training dataset, the validation dataset, and / or the test dataset, have similar distributions and no overlap. In this stage, users determine whether the subsets of data meet operational design domain (ODD) requirements, adequately cover outliers, and identify biases and hidden relationships in the data.
[0041] The expert dataset is a subset of the received dataset selected and refined by domain experts for domain representativeness and by AI experts for meeting dataset requirements, such as high accuracy and no bias. The expert dataset is a trusted interval in the results for each metric, such as a confidence interval, within which results are considered certain / accurate. The confidence interval is a necessary condition for calculating a confidence value. The testing and validation platform 102 can compare distributions and coverage of the dataset to assist users in selecting the expert dataset, and the testing and validation platform 102 receives the user selection for the expert dataset.For the example of the car rental dataset, a user can define that the Monte Carlo dropout (MCD) error must not differ by more than -0.15 and +0.6, while the fast gradient sign method (FGSM) must not differ by more than -0.3 and +0.4 compared to the original data.
[0042] Stage 3 includes model training. In the third stage, the testing and validation platform 102 supports the use of methods to analyze how well the model is adapted to the problem, such as determining whether there is overfitting or underfitting and how vulnerable the model is to attacks. Based on the results, a recommendation is provided to continue, modify, or stop training, as well as options for improvement. A user can either use the tests in this stage after a model has been trained or use the tests during the validation phase of model training (e.g., after a specified number of epochs) by using the API of the testing and validation platform 102 to launch tests directly within the training setup.When the MLM is trained to perform a safety-critical task, general metrics such as accuracy, recall, and precision may not be sufficient to determine whether the model has achieved the required level of quality and reliability. The testing and validation platform 102 uses local (instance-based) and global (general behavior of the model) interpretability metrics, as well as calibration checks, to determine the maturity of the model, thus helping users save costs by stopping training at the right time without sacrificing quality, as well as recommending steps to improve predictions.
[0043] Stage 4 includes model evaluation. In this stage, the testing and validation platform 102 performs a thorough analysis of the model using the tests from the previous stage and more complex tests, such as local and global interpretability. The testing and validation platform 102 can provide standardized tests to determine the quality of the model. After the model has been trained, the testing and validation platform 102 tests the MLM by using the expert dataset to determine at least one metric for each sample of the expert dataset and comparing each metric of the at least one metric with a corresponding defined confidence interval. The results are automatically retained and can be used to compare the different model iterations, as well as as proof that the evaluation was performed with the highest standards and without negligence.A given configuration can be reused for a new testing iteration, and the results can be automatically compared. When integrated into an automated CI / CD pipeline, this leads to significantly increased development speed.
[0044] The testing and validation platform 102 calculates the metric using the MLM and the expert dataset. If the results are within the confidence intervals for the corresponding metric, as determined by the user, the testing and validation platform 102 can automatically continue testing. If the results are outside the confidence interval, the testing and validation platform 102 can pause further testing for tuning the model. The performance of the model with the expert dataset is the sufficient condition. If the results are within the expected range, tests can be performed with the validation dataset and / or test dataset in the later stages. The testing and validation platform 102 can define clusters of the metric determined by the expert data, which determines cluster boundaries, as described in Fig. 5. Using a density-based clustering algorithm such as dbscan, the test and validation platform 102 can define the cluster automatically.
[0045] If the metric determined from the expert dataset is within the corresponding defined confidence interval, the testing and validation platform 102 may test the MLM by determining at least one metric for each sample of the validation dataset and / or the test dataset.
[0046] The test and validation platform 102 may compare the metric determined from the expert data set with the metric determined from the validation data set and / or the test data set, as described in Fig. 5. The testing and validation platform 102 may define clusters of the metric determined by the validation dataset and / or the test dataset and compare these clusters to the clusters of the metric from the expert dataset and / or the cluster boundaries defined by the clusters of the metric from the expert dataset. The testing and validation platform 102 may determine that the machine learning model is secure if the clusters of the metric determined by the validation dataset and / or the test dataset fall within the cluster boundaries. The testing and validation platform 102 may determine a score for the machine learning model based on the percentage of the metric determined by the validation dataset and / or the test dataset that falls within the cluster boundaries.In the case of individual outliers, an investigation of these data points is recommended, and the model score is reduced by the percentage of data points outside the cluster boundaries. If the number of clusters found in the validation dataset and / or the test dataset differs from the number of clusters in the expert dataset, the testing and validation platform 102 determines that the model is not secure. Users can use the metric from the validation dataset to tune hyperparameters and the metric from the test dataset to ensure minimal model overfitting.
[0047] Based on all information from this and previous stages, the testing and validation platform 102 can generate a recommendation that evaluates the model and provides insight into its maturity level and production readiness. In this stage, the testing and validation platform 102 can automatically compare models based on a set of different metrics that can be preselected. This allows the user to train multiple models simultaneously, for example, in different configurations or network sizes, and the testing and validation platform 102 can automatically select the model with the highest scores for further evaluation or deployment.
[0048] Stage 5 includes model inference, where the testing and validation platform 102 runs tests on collected real-world data to examine data and range deviations and recommend a retraining and adaptation process if a deviation is detected. The real-world data may include live data and / or recorded live data, which may be added to the testing and validation platform 102 in one or more batches. During this stage, the testing and validation platform 102 may test the MLM in different scenarios to ensure that the model is not overfit to certain scenarios.For example, if the MLM is trained to determine a safe stopping distance using a dataset with features including weather, temperature, road conditions, and the estimated size of the vehicle ahead, but the future automobile manufacturing trend leads to smaller and lighter compact vehicles with shorter stopping distances, the trained MLM may not be equipped to accurately predict a safe stopping distance in this new scenario. After a model is deployed, ongoing observation of the model's behavior is necessary to ensure that performance is not degraded by data or range drift. At this stage, the test and validation platform 102 serves as predictive maintenance to detect drift before the drift has a significant impact on model performance.The testing and validation platform 102 can automate iterative testing of the model and provide users with valuable insight into the inner workings of the model by recommending steps to further improve predictions in the next iteration. Based on knowledge gained in the inference stage, the testing and validation platform 102 can begin an iterative process of testing and validating the MLM, with each iteration of the five stages continuously improving the quality of the model. Therefore, the testing and validation platform 102 is configured to validate the model from the design phase of an ML project through the end of the product's life cycle. Fig. Figure 2 also shows the continuous process by which the user re-evaluates the problem the MLM is intended to address in Stage A, re-evaluates the features, such as all identified features and / or the selected features and their correlation with the goal in Stage B, and re-evaluates the model in Stage C, all of which the testing and validation platform 102 supports as described herein. The testing and validation platform 102 can perform this evaluation in a flow process using live data from production systems sent to a backend. Constantly processing and analyzing a stream of live data allows engineers to have near-real-time insight into the health of the model, using a dashboard with the most important information, as well as alerts regarding unusual data and / or decisions.This insight will help engineers accelerate the continuous improvement process of their models and thus help meet the requirement of updatable AI systems.
[0049] The five stages are iterative, meaning that after publishing a model, improvement begins by re-evaluating the dataset, closing the loop between steps five and one. This iterative process is necessary for a current model. After the first initial iteration of the stages, the testing and validation platform 102 can automatically execute the tests using the configurations saved along with the test results from the initial iteration, thus enabling automatic comparison and, if implemented on the development side, automatic hyperparameter tuning. The proposed process reuses the metric configuration performed by default in previous iterations, thereby accelerating the development lifecycle and making different models comparable.The expert dataset can be updated between iterations as new insights and corner cases are discovered in the inference stage. The testing and validation platform 102 can iteratively test the MLM by determining metrics from the expert dataset, determining metrics from the validation dataset and / or the test dataset, and comparing the metrics from the expert dataset with the metrics from the validation dataset and / or the test dataset. The testing and validation platform 102 can store results from each test of each machine model iteration for comparison between different iterations of the same model and / or between iterations of different models if multiple models have been trained.
[0050] In the event of serious issues, the process allows for reverting to the previous stage or iteration to correct issues found in later stages, while maintaining comprehensive documentation of changes by default. This allows the user to provide complete documentation to authorities to ensure that there was no negligence and that all risk mitigation steps have been implemented. Each report can be exported and includes comprehensive information on the datasets used, the model used, execution time, metric configuration, and results to comply with regulatory requirements.
[0051] The reports provided by the testing and validation platform 102 include a text description of the results and a visual representation of the results, generally in the form of a graph. Fig. 3A shows an exemplary graphical user interface display 300 of a configuration of a YData profile report generated by an open-source YData profile analyzer. Display 300 includes a textual description of the results in the YData profile report. Fig. 3B shows a display 320 of an exemplary graphical user interface (GUI) of an overview of a YData profile report. The overview includes a visual representation of the YData profile report, allowing users to quickly understand characteristics of the variables in the YData profile report, such as the level or extent to which two variables are correlated and identified duplicates in the dataset.
[0052] Fig. 4A shows an example GUI display 400 of a report summary for a mutual information test, including an analysis of the results and recommendations, which in this example is to remove the features that do not affect the target from the model. Fig. Figure 4B shows an example GUI display of a mutual information value chart, including a bar chart with features plotted on the x-axis and a mutual information value plotted on the y-axis. Ranges of mutual information values are identified as high influence, medium influence, and no influence, allowing users to readily understand the levels or degrees of influence the feature has on the at least one target.
[0053] Fig. 5 shows an example GUI display 500 of a value analysis using metrics and generated by the test and validation platform. Fig. 5 includes metrics from the expert dataset and the test dataset, but it is clear that the display 500 may include metrics from the validation dataset or another dataset. Metrics from the MCD test are plotted on the y-axis and metrics from the FGSM test are plotted on the x-axis. It is clear that various metrics from other statistical tests may be plotted in addition to or instead of the MCD and FGSM metrics. The metrics 502 of sample values from the expert dataset are all close to each other and form a cluster 504, which defines the cluster boundaries 506, approximated by the Fig. 5. Thus, the sufficient condition for calculating the value is met, since the necessary condition that the metric from the expert dataset lies within selected confidence intervals of [-0.15, +0.6] for MCD and within [-0.3, +0.4] for FGSM is met. Not all metric 510 of samples from the dataset that were not included in every training of the model and are thus unknown to it lie within the defined cluster boundaries 506 of the rectangle. The testing and validation platform can determine a percentage of the metric 510 of samples from the test dataset that lies within cluster boundaries 506, which in the example shown is 93.47357065803668%. Since only about 93% of the metric 510 of samples from the test dataset lies within cluster boundaries 506, the test and validation platform determines a value of about 93% for the combination of these two metrics.This system for evaluating a model's value can be used for a varying number of metrics, such as all metrics relevant to the model and the problem at hand. The confidence intervals defined by the experts for the expert dataset should be directly linked to the level of safety criticality and the "socially acceptable risk."
[0054] Fig. Figures 6A-6G collectively show an example JSON test result of a mutual information metric in displays 600, 610, 620, 630, 640, 650, and 660. This example test result includes a machine-readable and comparable version of the test result. For automatic comparison, the value of the "resultMetricString" identifier, which stores the current test results, is used. Using the underlying data structure, these automatic comparisons can be used to select the better-fitting model and / or dataset, while ensuring that the same dataset was used for the tests.
[0055] Fig.7 is a flowchart illustrating an example method 700 for providing a testing and validation platform for testing and validating a machine learning lifecycle. At step 702, the testing and validation platform analyzes a received data set by performing at least one statistical test on the data set, wherein the data set includes data representative of samples each including features and a corresponding target.
[0056] At step 704, the testing and validation platform displays results of the analysis. The at least one statistical test may determine correlations between the features of the data set, and the displayed results may include the determined correlations. The at least one statistical test may identify a degree of influence that each of the features has on the at least one target, and the displayed results may include the features and the corresponding identified degree of influence.
[0057] At step 706, the testing and validation platform receives at least one selected feature of the features for training a machine learning model.
[0058] At step 708, the testing and validation platform validates subsets of the dataset, the subsets comprising an expert dataset, a training dataset, and a validation dataset and / or a test dataset, each subset comprising data representing different sample values.
[0059] At step 710, after the machine learning model has been trained, the testing and validation platform tests the machine learning model by using the expert dataset to determine at least one metric for each sample of the expert dataset, and by comparing each metric of the at least one metric to a corresponding defined confidence interval.
[0060] If the metric determined from the expert dataset lies within the corresponding defined confidence interval, the testing and validation platform may test the machine learning model by determining at least one metric for each sample of the validation dataset and / or the test dataset. The testing and validation platform may compare the metric determined from the expert dataset with the metric determined from the validation dataset and / or the test dataset. The testing and validation platform may define clusters of the metric determined from the expert dataset, determine the cluster boundaries, and define clusters of the metric determined from the validation dataset and / or the test dataset, and compare the clusters of the metric determined from the expert dataset with the cluster boundaries.
[0061] The testing and validation platform may determine that the machine learning model is secure if the clusters of the metric determined from the validation dataset and / or the dataset fall within the cluster boundaries. The testing and validation platform may determine a score for the machine learning model based on the percentage of the metric determined from the validation dataset and / or the test dataset that falls within the cluster boundaries. The testing and validation platform may iteratively test the machine learning model by determining metrics from the expert dataset, determining metrics from the validation dataset and / or the test dataset, and comparing the metric from the expert dataset with the metric from the validation dataset and / or the test dataset. Results from each test of each machine learning model iteration may be saved for comparison.
[0062] It is understood that method 700 is for illustrative purposes only, and that different and / or additional steps may be used. It is also understood that various steps described herein may occur in different orders or sequences. It is understood that various details of the subject matter described herein may be changed without departing from the scope of the subject matter described herein. Furthermore, the foregoing description is for illustrative purposes only and not for limiting purposes, as the subject matter described herein is defined by the claims set forth below. QUOTES CONTAINED IN THE DESCRIPTION
[0000] This list of documents submitted by the applicant was generated automatically and is included solely for the convenience of the reader. This list is not part of the German patent or utility model application. The DPMA assumes no liability for any errors or omissions. Cited patent literature
[0000] US 63 / 559,569
[0001]
Claims
[1] A method for providing a testing and validation platform for testing and validating a machine learning lifecycle, the method comprising the steps of: Analyzing, at the testing and validation platform, a received data set by performing at least one statistical test on the data set, the data set comprising data representative of samples each having features and a corresponding target; Displaying, through the testing and validation platform, results of the analysis; Receiving, at the testing and validation platform, a selected at least one feature of the features for training a machine learning model; Validating, by the test and validation platform, subsets of the dataset, wherein the subsets comprise an expert dataset, a training dataset, and a validation dataset and / or a test dataset, each subset comprising data representative of different sample values; and Testing, by the testing and validation platform and after the machine learning model has been trained, the machine learning model by using the expert dataset to determine at least one metric for each sample of the expert dataset, and comparing each metric of the at least one metric to a corresponding defined confidence interval. [2] The method of claim 1, comprising, if the metric determined from the expert dataset is within the corresponding defined confidence interval, testing, by the testing and validation platform, the machine learning model by determining at least one metric for each sample of the validation dataset and / or the test dataset. [3] The method of claim 2, comprising comparing the metric determined from the expert data set with the metric determined from the validation data set and / or the test data set. [4] The method of claim 3, comprising defining clusters of the metric determined from the expert data, determining the cluster boundaries, defining clusters of the metric determined from the validation data set and / or test data set, and comparing the clusters of the metric determined from the expert data set with the cluster boundaries. [5] The method of claim 4, comprising determining that the machine learning model is secure if the clusters of the metric determined from the validation dataset and / or the test dataset are within the cluster boundaries. [6] The method of claim 4, comprising determining a value for the machine learning model based on the percentage of the metric determined from the validation dataset and / or the test dataset that lies within the cluster boundaries. [7] The method of claim 3, comprising iteratively testing the machine learning model by determining metrics from the expert dataset, determining metrics from the validation dataset and / or the test dataset, and comparing the metric from the expert dataset with the metric from the validation dataset and / or the test dataset. [8] The method of claim 1, wherein the at least one statistical test determines correlations between the features of the data set and the displayed results comprise the determined correlations. [9] The method of claim 1, wherein the at least one statistical test identifies a degree of influence that each of the features has on the at least one target, and the displayed results comprise the features and the corresponding identified degrees of influence. [10] A system for implementing the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
63/559,569