Systems and methods for evaluating and deploying data models with improved performance measures - Patents.com
Patent Information
- Application Number
- JP2024540015
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-11
- Filing Date
- 2022-12-14
- Publication Date
- 2025-11-26
AI Technical Summary
Existing systems lack a proactive method to evaluate and compare multiple versions of machine learning data models for technical performance and predictive accuracy before deployment, often leading to suboptimal performance post-deployment.
A system and method for autonomously evaluating and comparing multiple versions of machine learning data models by generating performance reports based on configuration and mapping files, using baseline files to validate performance, and deploying the leading version with improved accuracy.
Ensures the deployment of data models with superior technical performance and predictive accuracy by comparing versions pre-deployment, enhancing model performance and accuracy proactively.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Technical field]
[0001] FIELD OF THEINVENTION This application relates to machine learning data models and, more particularly, to evaluating the predictive accuracy and technical performance measures of data models. Summary of the Invention
[0002] Summary of the Invention This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used as an aid in determining the scope of the claimed subject matter. The present invention is defined by the claims, as supported by this specification including the Detailed Description. [Means for solving the problem]
[0003] Briefly and in general terms, the present disclosure describes, among other things, methods, systems, and computer-readable media for comparatively evaluating distinct versions of a machine learning data model (hereinafter, "model") for technical performance and / or predictive accuracy, and deploying versions that demonstrate improved technical performance and / or predictive accuracy over other versions. As described below, aspects of the invention discussed below monitor and comparatively evaluate technical performance and / or predictive accuracy by monitoring multiple diverse versions of a model in order to select and deploy (e.g., manually or autonomously via a processor without user input) a leading version that is indicated as having the greatest predictive accuracy and / or other indicia of superior performance (e.g., metrics, bias, data drift). Prior to deployment, one or more versions of the model can be evaluated autonomously (e.g., without user selection, input, and / or intervention) against one or more other (e.g., in-use, currently deployed, previously deployed) versions of the model.
[0004] In one aspect of the invention, a computerized method for evaluating and improving the performance and accuracy of a version of a model is provided. According to the method, a plurality of datasets are received for a plurality of versions of the model. Each of the plurality of datasets includes a plurality of predictions for a corresponding version of the model. In an aspect, a configuration file and a mapping file are received. In various aspects, a plurality of version-performance reports are generated from the plurality of datasets based on the received configuration file and mapping file. Each of the plurality of version-performance reports includes a plurality of performance measures determined for a corresponding version of the model. In an aspect, a baseline file is received. The plurality of version-performance reports are validated based on the baseline file. According to the method, a leading version in the plurality of versions is determined based on the corresponding plurality of performance measures for the plurality of version-performance reports of the plurality of versions. In some aspects, the leading version of the model is deployed.
[0005] Another aspect provides one or more non-transitory computer-readable media having computer-executable instructions embodied thereon that, when executed, perform a method for evaluating and improving the performance and accuracy of a version of a model. According to the media, a plurality of datasets are received for a plurality of versions of the model. In various aspects, each of the plurality of datasets includes a plurality of predictions of a corresponding version of the model. A configuration file and a mapping file are received. In aspects, a plurality of version-performance reports are generated from the plurality of datasets based on the configuration file and the mapping file. In some aspects, each of the plurality of version-performance reports includes a plurality of performance measures determined for the corresponding version of the model. A baseline file is received. In aspects, the plurality of version-performance reports are validated based on the baseline file. In some aspects, a leading version in the plurality of versions is determined based on a corresponding version-performance report of the leading version that indicates that the leading version has improved performance relative to at least one other version in the plurality of versions of the model. In some aspects, the leading version of the model is deployed.
[0006] In another embodiment, a system for evaluating and improving the performance and accuracy of a model version is provided. The system includes a data model performance monitoring system that executes a script via one or more processors. The data model performance monitoring system receives a plurality of datasets for a plurality of versions of the model, each of the plurality of datasets including a plurality of predictions of a corresponding version of the model. In some embodiments, the data model performance monitoring system receives a configuration file and a mapping file. In an embodiment, via the script, the data model performance monitoring system generates a plurality of version-performance reports from the plurality of datasets based on the configuration file and the mapping file. In some embodiments, each of the plurality of version-performance reports includes a plurality of performance measures determined for a corresponding version of the model. The data model performance monitoring system receives a baseline file and validates the plurality of version-performance reports based on the baseline file. The system includes a monitoring dashboard module that, in some embodiments, determines a leading version in the plurality of versions based on a corresponding plurality of performance measures in the version-performance report of the leading version for the plurality of version-performance reports of the plurality of versions. In such an embodiment, the corresponding version-performance report of the leading version indicates that the leading version has improved performance relative to other versions in the plurality of versions of the model. In some embodiments, the monitoring dashboard module communicates to deploy the leading version of the model.
[0007] BRIEF DESCRIPTION OF THE DRAWINGS Aspects will now be described in detail below with reference to the accompanying drawing figures. [Brief description of the drawings]
[0008] [Figure 1] FIG. 1 is a block diagram of an example system according to an aspect discussed herein. [Diagram 2]2 illustrates example computer-executable instructions suitable for implementation via the example system of FIG. 1 according to aspects discussed herein. [Diagram 3] 2 illustrates example computer-executable instructions suitable for implementation via the example system of FIG. 1 according to aspects discussed herein. [Figure 4] 2 illustrates example computer-executable instructions suitable for implementation via the example system of FIG. 1 according to aspects discussed herein. [Diagram 5] 2 illustrates example computer-executable instructions suitable for implementation via the example system of FIG. 1 according to aspects discussed herein. [Figure 6] 2 illustrates example computer-executable instructions suitable for implementation via the example system of FIG. 1 according to aspects discussed herein. [Figure 7] 2 illustrates example computer-executable instructions suitable for implementation via the example system of FIG. 1 according to aspects discussed herein. [Figure 8] 2 illustrates example computer-executable instructions for a configuration file suitable for implementation via the example system of FIG. 1, according to aspects discussed herein. [Figure 9] 2 illustrates an example report generated by the example system of FIG. 1 according to aspects discussed herein. [Figure 10] 2 illustrates an example performance report generated by the example system of FIG. 1 according to aspects discussed herein. [Figure 11] 2 illustrates an example alert generated by the example system of FIG. 1 according to aspects discussed herein. [Figure 12] 2 illustrates an example graphical user interface corresponding to the dashboard monitoring module of FIG. 1 according to aspects discussed herein. [Figure 13] 1 is a flowchart of a method for evaluating and improving the performance and accuracy of a model version according to aspects discussed herein. [Figure 14]FIG. 1 is a block diagram of an example environment suitable for implementing aspects of the invention discussed herein. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0009] Detailed Description of the Invention The subject matter of the present invention is specifically described herein to meet statutory requirements. However, the description itself is not intended to limit the scope of this patent. Rather, the inventors intend that the claimed subject matter may be embodied in other ways, including different steps or combinations of steps similar to those described herein, together with other current or future technologies. Furthermore, although the terms "step" and / or "block" may be used herein to refer to different elements of the method employed, this term should not be interpreted as implying any particular order between the various steps disclosed herein, unless the order of the individual steps is explicitly described.
[0010] overview Aspects of the invention herein provide a pre-deployment environment that simultaneously tests and measures technical performance and output accuracy of different versions of a particular data model. Testing of multiple versions can be performed autonomously in parallel. Aspects herein enable direct comparison of technical performance and output accuracy across different versions of a particular data model. They also facilitate and / or enable manual (e.g., semi-autonomous with user interaction) and / or autonomous (e.g., no user input required) selection of one or more versions of a data model that are superior compared to other versions of the data model that have reduced, impaired, or degraded technical performance and / or output accuracy for deployment. Technical performance measures such as metrics and / or forecast accuracy can be compared to baselines, minimums, thresholds, and / or ranges (e.g., with or without upper and lower / adjacent buffers) that serve to "validate" a version of the data model. Technical performance measures can be defined by user customization, industry-based standards, or a combination thereof.
[0011] The selected data model version is deployed, for example, to "upgrade", update, and / or replace another data model version already deployed and / or currently in use. In such an embodiment, the selected data model version is selected for deployment, in particular because it has demonstrated improved or superior technical performance measures and / or predictive accuracy relative to the technical performance measures and / or predictive accuracy of the currently in use data model version, based on autonomous testing and evaluation, as discussed below. In this manner, the selected data model version can then be subsequently deployed based on its demonstrated stability and improvement compared to the existing in use data model version. The newly deployed data model version thus replaces the current data model version.
[0012] More specifically, the systems, methods, and media herein obtain, acquire, and / or receive one or more datasets for one or more distinct machine learning / artificial intelligence data model versions, which in various embodiments may include different data models and / or different data model types. In some embodiments, one or more datasets are received for each of a plurality of distinct machine learning / artificial intelligence data model versions for one or more distinct or different models. The datasets include observational data and predicted data stored in one or more databases. In embodiments, the datasets are ingested and consumed by a computer programming script. The script, in various embodiments, uses a version mapping file and a configuration file to convert and reconstruct each version-specific dataset into a corresponding report. Each version-specific report is evaluated to identify, locate, and extract technical performance measures, such as metrics, prediction accuracy, bias, data drift, etc., from the corresponding dataset. From this evaluation, a version-performance report is generated and / or compiled so that each data model version can be directly compared to the others. Based on the comparison of the version-performance reports, an alert can be autonomously issued if a performance measure is determined by the computer-implemented methods, systems, and media herein to violate baselines, minimums, thresholds, and / or ranges (e.g., with or without upper and lower / adjacent buffers) that serve to "validate" the performance measure in the version-performance report. The version-performance reports, violations, alerts, etc. can also be stored in one or more databases. Also, based on the comparison, one or more data model versions are manually or autonomously selected for deployment as having demonstrated improved technical performance measures and / or superior predictive accuracy relative to the technical performance measures and / or predictive accuracy of another currently in use version of the data model.
[0013] Thus, the methods, systems, and media discussed herein provide technical improvements in the technical field of data model testing and evaluation. For example, the methods, systems, and media discussed herein provide a technical solution to a technical shortcoming, namely, the absence of a concurrent comparison tool to evaluate multiple versions of a data model in a pre-deployment environment to ensure that any subsequently deployed version performs better (e.g., better technical performance and / or output accuracy) than the currently deployed version of the data model. Other systems are reactive in nature and cannot accurately and simultaneously evaluate the performance of a new data model version until after deployment and output capture for the new data model. In contrast, aspects herein represent a paradigm shift to a proactive approach that provides a concurrent comparison tool to evaluate multiple versions of a data model in a pre-deployment environment to ensure that a subsequently deployed version performs better than the currently deployed version of the data model.
[0014] definition As used herein, the terms "observed data," "ground truth," "actual values," and "target values" are used interchangeably to refer to empirical data and / or observed real-world information coded as data. For example, observed data includes measured, captured, or recorded values that represent and / or quantify variables for an event or outcome that occurred. In one example, observed data includes values for a particular medical institution's total patient volume accrued over a defined six-month period as recorded in the institution's historical reporting data.
[0015] The term "prediction data" as used herein refers to any and all data input into and output from a version of a data model. For example, prediction data can include inputs such as training data sets that are ingested to generate and trigger outputs. Additionally or alternatively, prediction data can include outputs generated or created from a data model version, e.g., predictions made by that version of the data model using the inputs. Prediction data can also include metadata associated with the data model, metadata associated with a data version of the data model, metadata associated with inputs of the data model version, and / or metadata associated with outputs of the data model version. Prediction data can reference other outputs of the data model version.
[0016] As used herein, the terms "model" and "data model" are used interchangeably to refer to a machine learning / artificial intelligence type data model defined by algorithmic decision logic. A data model (and any version thereof) can include features such as decision logic, computational layers, neural networks, Markov chains, weighting algorithms (specific or non-specific to variables, values, layers, sub-models), and / or random forests. Although referred to in the singular, it will be understood that a data model (and any version thereof) can include multiple specific sub-models operating together in a specific sequence or in parallel that contribute to an output, such as a prediction.
[0017] As used herein, "version" and "data model version" are used interchangeably to refer to a particular iteration of a data model having a defined configuration for inputs, (e.g., decision tree) operations, and / or outputs that are specific or unique to that particular iteration.
[0018] As used herein, the terms "script" and "computer programming script" are used interchangeably to refer to computer readable and executable instructions / programming code that are expressions of instructions that cause, manage and facilitate the performance of a sequence of operational steps by a computer in an automated or semi-automated manner.
[0019] "Performance Measure" as used herein refers to a captured measurement that represents and quantifies an aspect of the technical performance and predictive accuracy (or inaccuracy) of a model version and / or other behavior. Performance measures may include, for example, metrics, predictive accuracy, bias, data drift, noise, variance, etc. Examples of metrics include Measured Absolute Percentage Error (MAPE), Mean Absolute Error (MAE), and / or Root Mean Squared Error (RMSE), although other metrics and corresponding algorithms are contemplated and within the scope of the present invention.
[0020] Embodiment Beginning with FIG. 1, an example of a system environment 100 is presented. The system environment 100 includes a data model performance monitoring system 102, hereinafter referred to as the "system." The system 102 receives, obtains, acquires, and / or imports multiple different versions of the same data model. In some embodiments, the multiple different versions of the data model include at least a currently in use (i.e., deployed) version of the data model and an updated, undeployed version of the same data model. Alternatively, the system 102 receives, obtains, acquires, and / or imports multiple different versions for each of the multiple different data models. In such embodiments, the multiple different versions of the different data models include at least a currently in use (i.e., deployed) version of the data model and an updated, undeployed version of the same data model for each of the multiple different data models. Thus, in various embodiments, multiple versions of multiple data models can be evaluated simultaneously using the system 102 and methods discussed below, although the discussion generally refers to evaluating multiple versions of the same data model for simplicity.
[0021] 1, model versions 104A, 104B, 104n are stored in a database (not shown). Each of the model versions 104A, 104B, 104n is associated with observed data 106A, 106B, 106n and predicted data 108A, 108B, 108n corresponding to the respective version, e.g., model version 104A includes observed data 106A and predicted data 108A, model version 104B stores observed data 106B and predicted data 108B, etc.
[0022] The system 102, in an embodiment, receives, obtains, retrieves, and / or imports a version mapping file 110 and a configuration file 112. The system 102 also includes a script 114. The script 114 operates to receive, obtain, retrieve, and / or import the model versions 104A, 104B, 104n including the observed data 106A, 106B, 106n and the predicted data 108A, 108B, 108n. Additionally, the script 114 operates to receive, obtain, retrieve, and / or import the version mapping file 110 and the configuration file 112. The script 114 can receive and / or extract data / files in any order, in parallel, and / or simultaneously.
[0023] Generally, the script 114 operates to "pre-process" the model versions 104A, 104B, 104n (including the observed data 106A, 106B, 106n and the predicted data 108A, 108B, 108n) based on information in the version mapping file 110 and the configuration file 112. In some embodiments, the model versions 104A, 104B, 104n are each transformed by extracting certain information and reconstructing the data into a coherent format for use in performing performance evaluations. More specifically, the script 114 captures the observed data 106A, 106B, 106n and the predicted data 108A, 108B, 108n from each or corresponding model versions 104A, 104B, 104n.
[0024] In an embodiment, the script 114 utilizes the version mapping file 110 to identify and locate specific data points in each model version to be evaluated. For example, the version mapping file 110 specifies a number or set of data points in the model version 104A to be extracted and another number or set of data points in the model version 104B to be extracted, where the data points to be extracted in the different versions correspond to the same or similar variables in the model. More simply, the version mapping file 110 is retrieved and utilized by the script 114 so that the script can know which data points in the different versions correspond to the same variables, events, etc. for subsequent comparison, i.e., the version mapping file 110 allows the script to map between the model versions 104A, 104B, 104n. The data points can correspond to the observed data 106A, 106B, 106n and / or the predicted data 108A, 108B, 108n. In some embodiments, the version mapping file 110 is a .json file.
[0025] The script 114, in an embodiment, utilizes the configuration file 112 to identify a particular configuration of the model version 104A, 104B, 104n for evaluation using the data points identified from the mapping. The script 114 utilizes the configuration file 112, where the configuration file 112, in an embodiment, can specify, define, and / or indicate a particular type of monitoring for the performance measures for subsequent operation. The configuration file 112 includes and / or defines one or more parameters and / or identifies one or more particular performance measures to be captured for all versions of a particular model and / or for a particular model version being evaluated. The configuration file 112 can also include metadata information that the script 114 utilizes to obtain and / or query the observed data 106A, 106B, 106n and / or predicted data 108A, 108B, 108n associated with the model version 104A, 104B, 104n. In one example, the configuration file 112 may include certain computer operations / functions, such as query_params and data_params, that the script 114 may utilize to retrieve and separate the observed data 106A, 106B, 106n and / or predicted data 108A, 108B, 108n from the model versions 104A, 104B, 104n, for example, sorting and / or aggregating by various subcategories and / or by version. FIG. 8 illustrates an example of computer-executable instructions for a configuration file 800.
[0026] Referring to Figure 2, an example of computer executable instructions 200 implemented via system 102 is shown. For example, Figure 3 shows box 300 enclosing an example portion of computer executable instructions for a script to call and receive config_file and mapping_versions files. In Figure 4, box 400 encloses another example portion of computer executable instructions for extracting predicted data (e.g., insight_df) and / or observed data (e.g., actuals_df) using parameters defined in config_file. Figure 5 shows box 500 enclosing yet another example portion of computer executable instructions for utilizing version_map in mapping_versions to recognize which data points in the extracted predicted data and / or extracted observed data correspond to the same variables, events, etc. for subsequent comparison across separate versions of the data model. Using computer executable instructions, for example, the script 114 can query the observed data 106A, 106B, 106n and predicted data 108A, 108B, 108n of the model versions 104A, 104B, 104n using parameters of the configuration file 112. Then, using computer executable instructions, for example, the script can filter only the necessary data points from the version based on the version mapping file 110. In one such example, for each version, the filtered data points of the predicted data and observed data of that model are merged and communicated to a computational dictionary with a data frame. In such an embodiment, the merged and filtered data and the corresponding data frame form a final_df file, and the computational dictionary acts to provide metadata about the final_df file.
[0027] Returning to FIG. 1, the script 114 generates the version data files 116A, 116B, 116n. FIG. 9 shows an example of the version data files 900, such as the version data files 116A, 116B, 116n. In an embodiment, each data model version is converted to a corresponding version data file. The system 102 performs the report generation 118 from the version data files 116A, 116B, 116n. The system 102 uses the version data files 116A, 116B, 116n to determine a set of performance measures to be calculated. For example, the set of performance measures to be calculated by the system 102 for the report generation 118 is defined in the configuration file 112 as Measured Absolute Percentage Error (MAPE), Mean Absolute Error (MAE), Root Mean Squared Error (RMSE), model bias, data drift, and / or any combination thereof. The system 102 generates reports 120A, 120B, 120n from the version data files 116A, 116B, 116n.
[0028] The system 102 makes a performance determination 122 from the reports 120A, 120B, 120n. The system 102 receives, obtains, acquires, and / or imports a baseline file 124 that includes one or more baselines and one or more thresholds 126. Using the baseline file 124 and the thresholds 126, the information in the reports 120A, 120B, 120n is verified and a version-performance report 130A, 130B, 130n is generated based on the performance determination 122. With reference to FIG. 6, for example, box 600 encloses a portion of computer executable instructions for generating a report having a performance measure, such as reports 120A, 120B, 120n. As discussed further below, FIG. 7 illustrates a box 700 that encloses a portion of computer executable instructions for outputting a report having a performance measure that can be used to make a performance determination.
[0029] To make the performance determination 122 and validate the data for inclusion in the validated version-performance report 130A, 130B, 130n, one or more baseline values and / or one or more thresholds defined in the baseline file 124 are applied to the data in the report 120A, 120B, 120n. Generally, the one or more baseline values define the expected values predicted for a variable or event by a model version based on specified known inputs. Thus, the baseline "expected" value can be used to determine whether the model version generated the same or similar values in predictions for that variable. Thus, the baseline value can be used to evaluate the prediction accuracy of each of the model versions 104A, 104B, 104n in comparison to such expected values. The baseline value can further include a margin, for example, to determine whether the model version generated a predicted value that is within a predefined buffer range of the baseline expected value for the corresponding variable or a predicted value that is not within the buffer range. The threshold 126 can be, for example, a target value that is a customized and / or leading minimum or maximum value for measuring and evaluating a performance measure, such as a metric (e.g., MAPE, MAE, and / or RMSE). Additionally or alternatively, the threshold 126 can define a value for evaluating other performance measures, such as, for example, data drift, model bias, and / or noise. For each model, the comparison and determination of each performance measure evaluated considering the baseline and / or threshold is included in the version-performance report 130A, 130B, 130n generated by the system 102. FIG. 10 shows an example of a version-specific performance report 1000. The version-performance report 130A, 130B, 130n is thus verified using the baseline and threshold that act as a quality control guideline or evaluation yardstick.
[0030] If one or more baseline values and / or one or more thresholds for one or more performance measures are not met, one or more alerts 128 can be automatically generated and communicated to a database, another system, and / or a user. When a threshold and / or baseline is violated, the system 102 generates a corresponding alert that includes, for example, an identifier of the baseline or threshold that was violated, an identifier of the model version and the particular performance measure in which the violation occurred, the expected and / or target values of the baseline and / or threshold that was violated, the value of the performance at which the violation was determined, etc. FIG. 11 illustrates an example alert 1100. Additionally, features in the model are analyzed with the baseline. In some embodiments, the baseline is created from the training data. During the analysis of the model, the occurrence of any new or missing features in the model data is captured. For any valid features in the model data, the data is analyzed with the baseline to capture any data type mismatches, positive, negative, and non-zero variations or violations.
[0031] Subsequently, the verified version-performance reports 130A, 130B, 130n are communicated to the monitoring dashboard module 132. FIG. 12 illustrates an example graphical user interface 1200 corresponding to the monitoring dashboard module 132. Although not illustrated in FIG. 1, additional verifications may be performed in the system 102, for example, by the monitoring dashboard module 132. In general, the monitoring dashboard module 132 determines one leading version 134 based on the verified performance measure data in the verified version-performance reports 130A, 130B, 130n. Alternatively, multiple leading versions may be selected as candidates for deployment. The leading version 134 may be identified and selected autonomously by the system 102, for example, if the leading version 134 demonstrates improved performance measures and / or superior predictive accuracy relative to at least one other version in the multiple versions of the data model in light of a comparison of system behavior of various performance measures in validated version-performance reports 130A, 130B, 130n for model versions 104A, 104B, 104n. The leading version 134 may alternatively be selected semi-autonomously by responding to user input, such as manual user selection of a leading version 134 from a list and / or user confirmation of a system recommended leading version 134. The system 102 may then execute and / or trigger a deployment 136 of the leading version 134 (or multiple leading versions of the data model) to other downstream applications, such as a data science workflow 138.
[0032] Although the system environment 100 and its components have been described, those skilled in the art will appreciate that the system environment 100 is merely one example of a suitable system and is not intended to limit the scope of use or functionality of the present invention. Similarly, the system environment 100 should not be construed as imposing any dependencies and / or any requirements with respect to each component and combination of components depicted in FIG. 1. Those skilled in the art will appreciate that the locations of components depicted in FIG. 1 are exemplary, as other methods, hardware, software, components, and devices for establishing communication links between the components depicted in FIG. 1 may be utilized in embodiments of the present invention. Those skilled in the art will appreciate that the components may be connected in a variety of ways, hardwired or wireless, and may use intermediate components that are omitted or not included in FIG. 1 for simplicity. Thus, the absence of components from FIG. 1 should not be construed as limiting the present invention to exclude additional components and combinations of components. Additionally, while components are depicted in FIG. 1 as singular components, it will be appreciated that some embodiments may include multiple devices and / or components, such that FIG. 1 should not be considered as limiting the number of devices or components.
[0033] With reference to FIG. 13, a flowchart of a method 1300 for evaluating and improving the performance and accuracy of a data model version is provided. In some embodiments, the method 1300 can be computer-executed. For example, the method 1300 can be implemented and / or performed autonomously or semi-autonomously using one or more non-transitory computer-readable storage media having computer-readable instructions embodied therein for execution by one or more processors. In embodiments, the computer-readable and executable instructions can include one or more scripts, such as the scripts and script portions discussed with respect to FIGS. 2-7, that specify the execution of the method 1300. The method 1300 can be implemented and / or performed in some embodiments using software and / or hardware components. For example, the method 1300 can be performed using the software, hardware, components, and / or devices illustrated in the system environment 100 of FIG. 1. The computer-readable and executable instructions can correspond to one or more applications, in one embodiment, that can implement and / or perform all or a portion of the method 1300 autonomously or semi-autonomously.
[0034] At block 1302, multiple datasets are received for multiple versions of the model, where each of the multiple datasets includes multiple predictions for a corresponding version of the model. In one embodiment, the multiple datasets include observed data and predicted data for versions 104A, 104B, 104n of the model of FIG. 1. At block 1304, a configuration file and a mapping file are received. In one embodiment, the configuration file 112 and the version mapping file 110 of FIG. 1 are received and / or imported by the script 114. In the configuration file, multiple performance measures can be specified, where the configuration file includes information customized for the model. The multiple performance measures can include data drift and bias, metrics such as MAPE, MAE, and / or RMSE, and the like.
[0035] At block 1306, multiple version-performance reports are generated from the multiple data sets based on the configuration file and the mapping file, where each of the multiple version-performance reports includes the multiple performance measures determined for the corresponding version of the model. In one embodiment, the multiple version-performance reports are generated using script 114 of FIG. 1. In one such embodiment, the multiple version-performance reports correspond to version data files 116A, 116B, 116n shown in FIG. 1. In various embodiments, multiple data subsets are identified from the mapping file and the subsets are extracted from the multiple data sets to generate the multiple version-performance reports. In one such embodiment, each of the multiple data subsets is extracted from one of the multiple data sets for a corresponding version in the multiple versions of the model. Multiple performance measures to be calculated can be identified from the configuration file. For each of the multiple versions, a script (e.g., script 114 of FIG. 1) is executed that calculates the multiple performance measures for the corresponding version of the model based on the corresponding data subset. The script also generates a version-performance report for the corresponding version of the model in such an embodiment of method 1300.
[0036] At block 1308, a baseline file is received. In one embodiment, the baseline file is the baseline file 124 of FIG. 1, which may include the threshold value 126, as described above. At block 1310, the multiple version-performance reports are validated based on the baseline file. In one embodiment, the multiple version-performance reports are validated using the baseline file 124 of FIG. 1 and the threshold value 126. The validated multiple version-performance reports may correspond to the version-performance reports 130A, 130B, 130n of FIG. 1, which are generated based on the performance determination 122 of the system 102. Validating the multiple version-performance reports may include determining, for each of the multiple versions, whether each of the multiple performance measures in the version-performance report of the corresponding version at least meets a corresponding baseline and / or threshold value defined in the baseline file for evaluation of a particular event, variable, prediction, or other performance measure. The corresponding baseline and / or threshold value may define a measure of model prediction accuracy. Additionally, the method 1300 may, in various embodiments, compare each of the multiple performance measures across the multiple version-performance reports of the multiple versions of the model.
[0037] In block 1312, a leading version is determined for the multiple versions based on corresponding performance measures against multiple version-performance reports of the other versions. Additionally or alternatively, a leading version is determined for the multiple versions based on corresponding performance measures against a baseline and / or threshold. In one embodiment, the leading version corresponds to leading version 134 of FIG. 1, which may be determined manually or autonomously by system 102.
[0038] If the multiple performance measures include data drift, in such embodiments, the leading version can be determined by determining whether a value quantifying the data drift of the leading version in the corresponding version-performance report indicates improved performance compared to values quantifying the data drift of the multiple versions. In general, improvement is indicated when the data drift value demonstrates that one or more mathematical measures quantifying the behavior of the data model version are stable and not fluctuating (e.g., fluctuating by changing over time such that a mathematical mean "moves" in a certain direction).
[0039] In an example drift calculation, the features of a model are analyzed using a baseline created from training data. In an exemplary drift calculation, statistics of the model data with respect to the baseline data are obtained. The drift calculation can then identify any drift that exists for features in the model. Furthermore, the drift calculation can operate to analyze multiple models with respective baselines for each model. In the case of multiple models, a baseline file can exist for each model, and the features of each model can be mapped to the respective model baselines for each model.
[0040] When the multiple performance measures include model bias, in such an embodiment, the leading version can be determined by determining whether the value quantifying the model bias of the leading version shows an improvement in performance compared to the values quantifying the model bias of the multiple versions in the corresponding version-performance report. Bias helps to understand the progress of the model with respect to a particular feature. In general, improvement is indicated by a model bias value that demonstrates that the predicted output of the data model version is the same or similar to the expected output based on the training data. Data model bias refers to a quantified value that represents the accuracy of the prediction of the data model version to match or deviate from the training set. Thus, data model bias is quantified for a variable as the difference between the prediction of the variable's value from the data model version and the expected value of the variable obtained through the training data. When calculating the bias of a model version, insight features such as predicted values from the model data and predicted values from the training data of the model are analyzed with a bias baseline created with the training data. This helps to generate a result. The result can be a comprehensive report that describes the progress of the feature-level bias of the data against the baseline over time.
[0041] In some embodiments, pre-training bias (evaluating features with actual labels) and post-training bias (evaluating features with actual and predicted label values) are supported. For example, when model pipeline data (actual values) are evaluated, the model monitoring system loads the pre-processed and baseline files into the system and runs the pre-processing algorithm to obtain model insight features and actual values. The bias and configured metrics can then be used to analyze the data in the baseline file to calculate any pre-training bias. In another example, when model pipeline data (actual values and predicted values) are evaluated, the model monitoring system loads the pre-processed and baseline files into the system and runs the pre-processing algorithm to obtain model insight features and actual values. The bias and configured metrics can then be used to analyze the data in the baseline file to calculate any post-training bias.
[0042] Where the multiple performance measures include metrics such as Measured Absolute Percentage Error (MAPE), Mean Absolute Error (MAE), or Root Mean Squared Error (RMSE), or a combination thereof, the leading version may, in such embodiments, be determined by determining whether a value quantifying the metric for the leading version in the corresponding version-performance report indicates improved performance as compared to values quantifying the performance metrics of the multiple versions.
[0043] MAPE can generally be expressed in the following examples, although other expressions of MAPE are contemplated to be within the scope of the embodiments discussed herein.
[0044]
number
[0045] In general, the MAE may be expressed in the following examples, although other expressions of the MAE are intended to be within the scope of the embodiments discussed herein.
[0046]
number
[0047] RMSE is generally the standard deviation of the prediction errors. Thus, RMSE may be expressed in the following examples, although other expressions of RMSE are intended to be within the scope of the aspects discussed herein.
[0048]
number
[0049] The reading version, in some embodiments, may be displayed via a graphical user interface corresponding to the monitoring dashboard module 132 of FIG.
[0050] At block 1314, the leading version of the model may be deployed. The leading version is deployed because it has demonstrated improved technical performance measures and / or superior predictive accuracy over another version (e.g., the data model version currently in use). Thus, the newly deployed leading version replaces another version that does not perform as well. Additionally or alternatively, the leading version may be used as input to retrain the corresponding data model and generate additional updated versions of the data model.
[0051] Referring to FIG. 14, an example of a computing environment according to an embodiment of the present invention is illustrated. Those skilled in the art will appreciate that the example computing environment 1400 is merely one example of a suitable computing environment and is not intended to limit the scope of use or functionality of the present invention. Likewise, the computing environment 1400 should not be interpreted as imposing any dependency and / or any requirement with respect to each component and combination of components illustrated in FIG. 14. Those skilled in the art will appreciate that the connections illustrated in FIG. 14 are also exemplary, as other methods, hardware, software, and devices for establishing communication links between components, devices, systems, and entities as illustrated in FIG. 14 may be utilized in embodiments of the present invention. Although the connections are illustrated using one or more solid lines, those skilled in the art will appreciate that the example connections in FIG. 14 may be hardwired or wireless and may use intermediate components that are omitted or not included in FIG. 14 for simplicity. Thus, the absence of components from FIG. 14 should not be interpreted as limiting the present invention to exclude additional components and combinations of components. Additionally, although devices and components are depicted in FIG. 14 as singular devices and components, it will be understood that some embodiments may include multiple devices and components, such that FIG. 14 should not be considered as limiting the number of devices or components.
[0052] Continuing, the computing environment 1400 of FIG. 14 is illustrated as being a distributed environment in which components and devices may be remote from one another and may perform separate tasks. The components and devices may communicate with one another and may be linked to one another using a network 1402. The network 1402 may include wireless and / or physical (e.g., hardwired) connections. Exemplary networks include a service provider or carrier telecommunications network, a wide area network (WAN), a local area network (LAN), a wireless local area network (WLAN), a cellular telecommunications network, a Wi-Fi network, a short-range wireless network, a wireless metropolitan area network (WMAN), a Bluetooth®-enabled network, an optical fiber network, or a combination thereof. The network 1402 generally provides the components and devices with access to the Internet and web-based applications.
[0053] The computing environment 1400 includes a computing device 1404 in the form of a server. Although shown as a single component in FIG. 14, the present invention may employ multiple local and / or remote servers in the computing environment 1400. The computing device 1404 may include components such as a processing unit, internal system memory, and a suitable system bus for coupling to various components including a database or database cluster. In some embodiments, the database cluster takes the form of a cloud-based data store, which in some embodiments is accessible by a cloud-based computing platform. The system bus may be any of several types of bus structures, including a memory bus or memory controller, a peripheral bus, and a local bus using any of a variety of bus architectures. By way of example and not limitation, such architectures include the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MCA) bus, the Enhanced ISA (EISA) bus, the Video Electronics Standards Association (VESA)® local bus, and the Peripheral Component Interconnect (PCI) bus, also known as the Mezzanine bus.
[0054] The computing device 1404 may include or have access to computer-readable media. Computer-readable media may be any available media that can be accessed by the computing device 1404, including volatile and nonvolatile media, and removable and non-removable media. By way of example, and not limitation, computer-readable media may include computer storage media and communication media. Computer storage media may include, but is not limited to, volatile and nonvolatile media, and removable and non-removable media implemented in any method or technology for storage of information, such as computer-readable instructions, data structures, program modules, or other data. In this regard, computer storage media may include, but is not limited to, random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, CD-ROM, digital versatile disk (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store the desired information and which can be accessed by the computing device 1404. Computer storage media do not include transitory signals.
[0055] Communication media typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave or other transport mechanism, and may include any information delivery media. As used herein, the term "modulated data signal" refers to a signal that has one or more of its attributes set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, communication media include wired media, such as a wired network or direct-wired connection, and wireless media, such as acoustic, radio frequency (RF), infrared and other wireless media. Combinations of any of the above may also be included within the scope of computer-readable media.
[0056] In an embodiment, the computing device 1404 communicates with one or more remote computers 1406 in the computing environment 1400 using logical connections. In an embodiment in which the network 1402 includes a wireless network, the computing device 1404 can employ a modem to establish communication with the Internet, the computing device 1404 can connect to the Internet using a Wi-Fi or wireless access point, or the server can use a wireless network adapter to access the Internet. The computing device 1404 uses the network 1402 to perform bidirectional communication with any or all of the components and devices depicted in FIG. 14 . Thus, the computing device 1404 can transmit data to and receive data from the remote computer 1406 via the network 1402.
[0057] Although shown as a single device, the remote computer 1406 can include multiple computing devices. In an embodiment having a distributed network, the remote computer 1406 may be located in one or more different geographic locations. In an embodiment in which the remote computer 1406 is multiple computing devices, each of the multiple computing devices may be located across various locations, such as buildings on a campus, medical and research facilities in a medical complex, offices or "branch offices" of a bank / credit company, or may be a mobile device, for example, wearable or carried by a person or attached to a trackable item in a vehicle or warehouse.
[0058] In some aspects, the remote computer 1406 is physically located in a medical environment, such as, for example, a laboratory, an inpatient room, an outpatient room, a hospital, a medical vehicle, a veterinary environment, an outpatient environment, a medical billing office, a financial or administrative office, a hospital administration environment, a home medical environment, and / or a medical professional's office. By way of example, the medical personnel can include medical professionals such as physicians, surgeons, radiologists, cardiologists, and oncologists, paramedics, physician assistants, clinical nurse practitioners, nurses, nursing assistants, pharmacists, nutritionists, microbiologists, laboratory professionals, genetic counselors, researchers, veterinarians, students, etc. In other aspects, the remote computer 1406 may be physically located in a non-medical environment, such as a packaging and shipping facility, or deployed within a fleet of delivery or courier vehicles.
[0059] Continuing, the computing environment 1400 includes a data store 1408. Although the data store 1408 is shown as a single component, it may be implemented using multiple data stores communicatively coupled to one another regardless of the geographic or physical location of the memory devices. The exemplary data store may store data in the form of artifacts, server lists, properties associated with servers, environments, properties associated with environments, computer instructions coded in multiple different computer programming languages, deployment scripts, applications, properties associated with applications, release packages, version information for release packages, build levels associated with applications, identifiers of applications, identifiers of release packages, users, roles associated with users, permissions associated with roles, workflows and steps within workflows, clients, servers associated with clients, attributes associated with properties, audit information, and / or audit trails of workflows. The exemplary data store may also store data in the form of electronic records, such as, for example, electronic medical records of patients, transaction records, billing records, task and workflow records, time series event records, and the like.
[0060] Generally, the data store 1408 includes physical memory configured to store information encoded with data. For example, the data store 1408 can provide storage for computer-readable instructions, computer-executable instructions, data structures, data arrays, computer programs, applications, and other data that support functions and operations undertaken using the exemplary computing environment 1400 and components shown in FIG.
[0061] In a computing environment having distributed components communicatively coupled via network 1402, program modules may be located in local and / or remote computer storage media, including, by way of example only, memory storage devices. Aspects of the invention may be described in the context of computer-executable instructions, such as program modules, being executed by a computing device. Program modules may include, but are not limited to, routines, programs, objects, components, and data structures that perform particular tasks or implement particular abstract data types. In aspects, computing device 1404 may access, retrieve, communicate, receive, and update information stored in data store 1408, including program modules. Thus, computing device 1404 may use a processor to execute computer instructions stored in data store 1408 to perform aspects described herein.
[0062] Although the internal components of the devices of Figure 14, such as the computing device 1404, are not shown, one skilled in the art would understand that the internal components and their interconnections are present in the devices of Figure 14. Accordingly, further details regarding the internal structural devices will not be further disclosed herein.
[0063] The present invention has been described in relation to particular embodiments, which are intended in all respects to be illustrative and not restrictive, and the invention is not limited to these embodiments, and variations and modifications can be made without departing from the scope of the invention.
Claims
1. 1. A computer-implemented method for evaluating and improving performance and accuracy of versions of a model, comprising: receiving a plurality of datasets for a plurality of versions of a model, each of the plurality of datasets comprising a plurality of predictions for a corresponding version of the model; the method further comprising: receiving a configuration file and a mapping file; and generating a plurality of version-performance reports from the plurality of datasets based on the configuration file and the mapping file, each of the plurality of version-performance reports comprising a plurality of performance measures determined for the corresponding version of the model; the method further comprising: receiving a baseline file; validating the plurality of version-performance reports based on the baseline file; determining a leading version in the plurality of versions based on corresponding performance measures for the plurality of version-performance reports for the plurality of versions; and deploying the leading version of the model.
2. Determining a leading version in the plurality of versions includes:
2. The method of claim 1, comprising determining, based on a corresponding version-performance report of the leading version, that the leading version indicates improved performance relative to at least one other version in the plurality of versions of the model.
3. 2. The method of claim 1 , wherein generating the multiple version-performance reports from the multiple datasets based on the configuration file and the mapping file includes identifying multiple data subsets to extract from the multiple datasets from the mapping file; and extracting the multiple data subsets from the multiple datasets, each of the multiple data subsets being extracted from one of the multiple datasets for the corresponding version in the multiple versions of the model.
4. 4. The method of claim 3, wherein generating the multiple version-performance reports from the multiple data sets based on the configuration file and the mapping file includes identifying the multiple performance measures to be calculated from the configuration file.
5. 5. The method of claim 4, wherein generating the plurality of version-performance reports from the plurality of data sets based on the configuration file and the mapping file comprises: executing a computer script that, for each of the plurality of versions, calculates the plurality of performance measures for the corresponding version of the model based on the corresponding data subset, and generates the version-performance report for the corresponding version of the model.
6. 6. The method of claim 5, wherein validating the plurality of version-performance reports based on the baseline file includes at least determining, for each of the plurality of versions, whether each of the plurality of performance measures in the version-performance report of the corresponding version satisfies a corresponding threshold defined in the baseline file, the corresponding threshold defining a measure of model predictive accuracy.
7. The method of claim 6 , further comprising comparing each of the plurality of performance measures across the plurality of version-performance reports for the plurality of versions of the model.
8. 2. The method of claim 1 , wherein the plurality of performance measures includes data drift, and determining the leading version includes determining that a value quantifying data drift for the leading version in the corresponding version-performance report indicates improved performance relative to a value quantifying data drift for the plurality of versions.
9. 2. The method of claim 1 , wherein the plurality of performance measures includes model bias, and wherein determining the leading version includes determining that a value quantifying model bias of the leading version in the corresponding version-performance report exhibits improved performance relative to a value quantifying model bias of the plurality of versions.
10. 2. The method of claim 1 , wherein the plurality of performance measures include at least one metric of a Measured Absolute Percentage Error (MAPE), a Mean Absolute Error (MAE), or a Root Mean Squared Error (RMSE), and wherein determining the leading version includes determining that a value quantifying the at least one metric of the leading version in the corresponding Version-Performance Report indicates improved performance relative to values quantifying performance metrics of the plurality of versions.
11. The method of claim 1 , further comprising retraining the model using the leading version.
12. The method of claim 1 , wherein the plurality of performance measures are specified in the configuration file customized for the model, and the plurality of performance measures include data drift and bias.
13. 3. The method of claim 2, wherein the plurality of performance measures include at least one metric of Measured Absolute Percentage Error (MAPE), Mean Absolute Error (MAE), or Root Mean Squared Error (RMSE), and wherein determining the leading version includes two or more of: determining that a value quantifying data drift for the leading version in the corresponding version-performance report indicates improved performance relative to values quantifying data drift for the plurality of versions; determining that a value quantifying model bias for the leading version in the corresponding version-performance report indicates improved performance relative to values quantifying model bias for the plurality of versions; and determining that a value quantifying the at least one metric of MAPE, MAE, or RMSE for the leading version in the corresponding version-performance report indicates improved performance relative to values quantifying model bias for the plurality of versions.
14. 10. A system for evaluating and improving the performance and accuracy of model versions, the system comprising: a data model performance monitoring system that receives, via one or more processors executing a script, a plurality of datasets for a plurality of versions of a model, each of the plurality of datasets including a plurality of predictions for a corresponding version of the model; receives a configuration file and a mapping file; generates a plurality of version-performance reports from the plurality of datasets based on the configuration file and the mapping file, each of the plurality of version-performance reports including a plurality of performance measures determined for the corresponding version of the model; receives a baseline file and validates the plurality of version-performance reports based on the baseline file; and a monitoring dashboard module that communicates to deploy the leading version of the model by determining a leading version among the plurality of versions based on a corresponding plurality of performance measures in the version-performance report of the leading version against the plurality of version-performance reports of the plurality of versions, the corresponding version-performance report of the leading version indicating that the leading version has improved performance relative to other versions in the plurality of versions of the model.
15. A program for causing one or more processors to execute the method according to any one of claims 1 to 13.