Confidence metrics for deployed machine learning models
By modifying the input data and analyzing multiple results from the machine learning model, it provides input-data-specific confidence assessments, addressing the unreliability issue of client-side machine learning models and improving the accuracy and efficiency of data analysis and decision-making, particularly in the medical field.
Patent Information
- Application Number
- CN202080027009.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-02-01
- Filing Date
- 2020-01-21
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2040-01-21
AI Technical Summary
Existing machine learning models cannot be retrained on the client side due to limited computing resources, licensing issues, or FDA constraints, resulting in their lack of security and reliability. The lack of effective confidence assessment methods also affects the accuracy of data analysis and decision-making.
By modifying or expanding the input data, multiple results are generated using machine learning models, and the changes in the results are analyzed to determine confidence metrics, providing confidence assessments specific to the input data. By combining traditional machine learning and image processing techniques, confidence estimates and uncertainty indicators are generated.
It enables reliability assessment of machine learning model outputs, improving the accuracy and efficiency of data analysis, especially in the medical field, helping medical professionals quickly identify and adjust model results, and improving the efficiency and effectiveness of clinical decision support systems.
Smart Images

Figure CN113661498B_ABST
Abstract
Description
Technical Field
[0001] This invention generally relates to machine learning, and more specifically to obtaining a confidence metric for a deployed machine learning model. Background Technology
[0002] Recent technological advancements have led to the use of machine learning (ML) models designed to assist in data analysis (e.g., for identifying medical features and / or making clinical decisions). Typical data analysis applications include identification, characterization (e.g., semantic segmentation, voxel labeling), and grading (e.g., classification).
[0003] ML models are typically trained using training datasets of limited size and / or variability. For example, in the medical field, the variability represented by all training data is limited due to the lack of large databases. Therefore, so-called 'amplification' methods are often used to increase the size and / or variability of the training dataset in order to improve the performance, reliability, and / or robustness of the ML model.
[0004] After training and deployment, customers use the final (fixed) ML model to evaluate input data (e.g., new medical cases).
[0005] For the client (i.e., on the client side), the ML component / system is typically a closed / fixed 'black box' configured to receive input data and generate / output results or decisions based on that input data. Therefore, in typical use cases, the ML component / system is 'sealed' (or fixed), and it is not possible to retrain the ML model on the client side. This sealing (or fixing) of the ML component / system can be due to a number of different reasons, including, for example, limited computational resources; licensing issues; the infeasibility of on-site label correction; or FDA constraints. Summary of the Invention
[0006] A method is provided for obtaining a confidence metric for a machine learning (ML) model, the method comprising: processing input data using the ML model to generate preliminary results; generating multiple modified instances of the input data; processing the multiple modified instances of the input data using the ML model to generate corresponding multiple secondary results; and determining a confidence metric related to the preliminary results based on the secondary results.
[0007] A concept for determining confidence (i.e., confidence metric) associated with a deployed ML model is proposed. Specifically, it is suggested that the confidence metric can be determined by modifying (or expanding) the input data and analyzing the results provided by the ML model for the modified (or expanded) input data. This proposal may rely on the concept that an acceptable ML model should have 'good performance,' for example, such that small perturbations to the input data should have a relatively small impact on the model output.
[0008] For example, input data may be automatically modified (or expanded) several times before being processed by an ML model. Based on the results associated with the modified (or expanded) data, variations in the results can be analyzed to assess their robustness or variability. This allows for the determination of a confidence metric for the ML model. For instance, a confidence metric specific to certain input data can be determined, and this can be based on the variance of the results provided by the modified (or expanded) version of the ML model that processes the specific input data.
[0009] For example, the proposed embodiments can be used to identify whether the results of a client-side ML model (e.g., provided by processing specific input data) are reliable.
[0010] Furthermore, the embodiments can facilitate the provision of additional information related to uncertainties associated with the output or results of the ML model.
[0011] Therefore, the proposed embodiments may be particularly advantageous for applications that preferably indicate the perceived accuracy or reliability of the output of a deployed (e.g., client-side) ML model. This could be especially important in the healthcare field, where healthcare practitioners need to understand and evaluate ML model results (and accordingly accept or adjust ML model decisions).
[0012] Unlike traditional ML models (which may have indications of global / general confidence levels provided by the model provider), the proposed implementation can provide confidence metrics specific to the input data (e.g., a single medical case).
[0013] Therefore, concepts that go beyond traditional methods of simply highlighting general or global confidence levels can be provided. For example, the proposed embodiments can correlate confidence metrics with ML model results / outputs based on specific input data of the ML model. This allows results to be provided with supplementary information (such as image overlays and related text descriptions), which can enable experts (e.g., clinicians, technicians, data analysts, engineers, medical practitioners, radiologists, etc.) to quickly evaluate the model's results by focusing on results / outputs associated with higher confidence metrics.
[0014] The ML model used in the embodiments can be constructed using traditional machine learning and / or image processing techniques, thereby leveraging historical data and / or established knowledge to improve the accuracy of the determinations / results provided by the proposed embodiments.
[0015] The embodiments can provide confidence estimates (e.g., measures of uncertainty) of outcomes associated with specific input data (e.g., image features or regions). In this way, the embodiments can help identify input data and / or output outcomes with high ML model uncertainty.
[0016] Therefore, the proposed embodiments can identify input data (e.g., medical image regions) that are important to the ML model output, and also associate such input data with visual features that may be useful to users (e.g., medical practitioners) (e.g., image overlays with relevant textual descriptions). This allows users to quickly and easily verify the model's results and identify situations where the model has not made correct or trustworthy decisions. Furthermore, the embodiments can identify the uncertainty (i.e., confidence measure) associated with each input data (e.g., each medical case). For example, this allows users (such as medical practitioners) to start viewing the model output from the most uncertain output (i.e., with the lowest confidence measure).
[0017] Therefore, the proposed embodiments can facilitate improvements in data analysis and case diagnosis (e.g., making them more accurate and / or easier). The embodiments can also be used to improve the efficiency and / or effectiveness of clinical decision support (CDS) systems. Thus, the proposed embodiments can provide an improved CDS concept.
[0018] Therefore, the proposed embodiments may be particularly relevant to medical data analysis and medical image analysis. For example, it may help identify the input / output data of an ML model (e.g., medical case or medical image features) and the uncertainty (i.e., confidence measure) associated with the ML model output of the input data. Thus, the proposed concept can also facilitate the accurate assessment or diagnosis of an individual's health using medical analytics. Consequently, the input data may include, for example, medical data, medical images, or medical features. Furthermore, the results generated by processing the input data using an ML model may include inference, medical decisions, diagnoses, judgments, or recommendations.
[0019] In some of the proposed embodiments, determining the confidence metric may include: determining a measure of the distribution or variance of the quadratic outcome; and determining the confidence metric based on the determined measure of distribution or variance. For example, determining a measure of the distribution or variance of the quadratic outcome may include determining at least one of the following: the inverse variance of the quadratic outcome; the Shannon entropy of the quadratic outcome; the Gini coefficient of the quadratic outcome; the Kullback-Liebler divergence of the quadratic outcome; and a measure of the concentration of the quadratic outcome. Therefore, the confidence metric of a machine learning model can be determined using simple mathematical methods or formulas. Thus, accurate and / or informed data analysis can be easily achieved using ML models with reduced complexity.
[0020] It should be understood that various approaches, methods, or functionalities can be used to provide confidence measures based on quadratic results. For example, some embodiments may use the inverse variance of the quadratic result, while others may use methods for measuring histograms of categorical data, such as Shannon entropy, Gini coefficient, Kullback-Liebler (KL) divergence, etc. Alternatively or additionally, measures of concentration of empirical distributions may be used.
[0021] In some embodiments, generating multiple modified instances of the input data may include applying a first spatial warping transformation to the input data to generate a first modified instance of the input data. Therefore, a simple modification / expansion method can be employed to generate modified instances of the input data. This also allows for control over the modifications made. Thus, generating modified instances of the input data can be easily implemented with reduced complexity.
[0022] Furthermore, the embodiment may also include applying a first inverse space curvature transformation to a quadratic result generated for a first modified instance of the input data. This allows the result to be transformed back for comparison with the result on the unmodified input data, thereby making the evaluation easier and / or more accurate.
[0023] In some embodiments, generating multiple modified instances of the input data may include adding noise to the input data to generate a second modified instance of the input data. Adding noise makes it easy to add random modifications, thereby ensuring small random perturbations. Therefore, the proposed embodiments enable easy implementation of small or minor modifications to the input data with reduced complexity.
[0024] Furthermore, generating multiple modified instances of the input data may include applying a local deformation transformation (e.g., a bending function) to the input data to generate a third modified instance of the input data. Further, such an embodiment may also include applying a first inverse local deformation transformation to a quadratic result generated for the third modified instance of the input data.
[0025] As an example, machine learning models can include artificial neural networks, generative adversarial networks (GANs), Bayesian networks, or combinations thereof.
[0026] The embodiment may also include the step of correlating the determined confidence metric with the preliminary results.
[0027] The embodiments may further include the step of generating an output signal based on the determined confidence metric. The embodiments may be adapted to provide such an output signal to at least one of the following: a subject, a medical practitioner, a medical imaging device operator, and a radiologist. Therefore, the output signal can be provided to a user or medical device to indicate the calculation result / decision and its associated confidence metric.
[0028] Some embodiments may also include the step of generating control signals for modifying graphical elements based on a determined confidence metric. The graphical elements can then be displayed according to the control signals. In this way, a user (such as a radiologist) can have a suitably configured display system that can receive and display information about the results provided by the machine learning model. Therefore, embodiments can enable users to remotely analyze results (e.g., outputs, decisions, inferences, etc.) from a deployed (e.g., client-side) machine learning model.
[0029] According to another aspect of the invention, a computer program product for obtaining a confidence metric of a machine learning model is provided, wherein the computer program product includes a computer-readable storage medium having computer-readable program code embedded therein, the computer-readable program code being configured to perform all the steps of the embodiments when executed on at least one processor.
[0030] A computer system may be provided, including a computer program product according to an embodiment; and one or more processors adapted to perform the method according to an embodiment by executing computer-readable program code of the computer program product.
[0031] In another aspect, the present invention relates to a computer-readable non-transitory storage medium comprising instructions which, when executed by a processing device, perform steps of a method for identifying features in a medical imaging of a subject according to an embodiment.
[0032] According to another aspect of the present invention, a system for obtaining a confidence metric of a machine learning model is provided. The system includes an input interface configured to acquire input data; a data modification component configured to generate multiple modified instances of the input data; a machine learning model interface configured to transmit the input data and the multiple modified instances of the input data to a machine learning model, and further configured to receive preliminary results generated by the machine learning model processing the input data and to receive multiple secondary results generated by the machine learning model processing the corresponding multiple modified instances of the input data; and an analysis component configured to determine a confidence metric related to the preliminary results based on the secondary results.
[0033] It should be understood that all or part of the proposed system may include one or more data processors. For example, the system may be implemented using a single processor adapted to perform data processing to determine a confidence metric for a deployed machine learning model.
[0034] The system used to obtain confidence metrics for a machine learning model can be located far from the machine learning model, and data can be communicated between the machine learning model and system units via communication links.
[0035] The system may include server equipment with input interfaces, data modification components, and machine learning model interfaces; and client equipment with analysis components. Therefore, dedicated data processing devices can be used to determine confidence metrics, thereby reducing the processing requirements or capabilities of other components or devices in the system.
[0036] The system includes client devices, which comprise input interfaces, data modification components, client-side machine learning models, and analysis components. In other words, users (such as medical professionals) may have appropriately positioned client devices (such as laptops, tablets, mobile phones, PDAs, etc.) that process received input data (e.g., medical data) to generate preliminary results and relevant confidence metrics.
[0037] Therefore, processing can be hosted at a location different from where the input data is generated and / or processed. For example, it may be advantageous to perform only part of the processing at a specific location for computational efficiency reasons, thereby reducing associated costs, processing power, transmission requirements, etc.
[0038] Therefore, it should be understood that processing capacity can be distributed throughout the system in different ways depending on predetermined constraints and / or the availability of processing resources.
[0039] The embodiments can also allow some of the processing load to be distributed throughout the system. For example, preprocessing can be performed at a data acquisition system (e.g., a medical imaging / sensing system). Alternatively or additionally, processing can be performed at a communication gateway. In some embodiments, processing can be performed at a remote gateway or server, thereby abandoning processing requests from end users or output devices. This distribution of processing and / or hardware can allow for improved maintainability (e.g., by centralizing complex or expensive hardware in a preferred location). It can also allow for the design or positioning of computational loads and / or traffic within a networked system based on available processing power. A preferred approach might be to process initial / source data locally and transfer the extracted data for full processing at a remote server.
[0040] Implementations may incorporate pre-existing, pre-installed, or otherwise separately supplied machine learning models. Other implementations may include (e.g., integrated into) new devices that incorporate machine learning models.
[0041] These and other aspects of the invention will become apparent and be elucidated with reference to the embodiments described below. Attached Figure Description
[0042] Now, examples of various aspects of the invention will be described in detail with reference to the accompanying drawings, wherein...
[0043] Figure 1 This is a simplified block diagram of a system for obtaining a confidence metric for a machine learning model, according to an embodiment.
[0044] Figure 2 This is a flowchart of a method for obtaining a confidence metric for a machine learning model according to an embodiment; and
[0045] Figure 3 This is a simplified block diagram of a system for obtaining a confidence metric for a machine learning model, according to another embodiment. Detailed Implementation
[0046] A concept for obtaining a confidence metric for a machine learning model is proposed. Furthermore, the confidence metric can be associated with specific input data provided to the machine learning model. Therefore, embodiments can enable the provision of information that may be useful for evaluating the model's output.
[0047] Specifically, confidence metrics can be determined by modifying (or expanding) the input data and analyzing the results provided by the ML model for the modified (or expanded) input data. For example, the input data might be automatically modified (or expanded) several times, which is then processed by the ML model. The ML model output (i.e., the results) of the modified (or expanded) data can then be analyzed to assess the robustness or variability of the results. This allows for the determination of a confidence metric for the ML model.
[0048] For example, a confidence metric can be determined that is associated with the input data of an ML model, and this can be based on the variance of the ML model output (i.e., the result) of the modified (or expanded) input data.
[0049] Associating such a confidence metric with the input data of an ML model allows indicators (e.g., a graphical overlay of textual descriptions with the relevant confidence metric) to be correlated with the ML model's output of the input data. This can facilitate a simple and quick evaluation of ML model results (e.g., by identifying outputs that the model is less confident about).
[0050] The embodiments can provide estimates of the uncertainty of the ML model output. Therefore, the proposed embodiments can, for example, be used to identify whether the output of a client-side ML model is reliable (e.g., by processing specific input data).
[0051] For example, the embodiments can be used to improve the analysis of medical data of subjects. Therefore, the illustrative embodiments can be used in many different types of medical assessment devices and / or medical assessment facilities, such as hospitals, wards, research facilities, etc.
[0052] Through examples, the confidence assessment of ML model outputs can be used to understand and / or evaluate the decisions made by the ML model. Using the proposed embodiments, users can, for example, identify less reliable or more reliable model outputs / results.
[0053] Furthermore, the implementation can be integrated into a data analytics system or ML decision-making system to provide users (e.g., technicians, data analysts) with real-time information about the results. Using this information, technicians can examine the model outputs and / or decisions, and, if necessary, adjust or modify the outputs and / or decisions.
[0054] The proposed implementation can identify uncertain decisions or outputs from an ML model. Such decisions / outputs can then be improved by focusing on and / or (e.g., by learning from more information sources).
[0055] To provide context for the description of the elements and functions of the illustrative embodiments, the accompanying drawings are provided as examples of how aspects of the illustrative embodiments can be implemented. Therefore, it should be understood that the drawings are merely examples and are not intended to assert or imply any limitation regarding the environment, system, or method in which aspects or embodiments of the invention may be implemented.
[0056] Embodiments of the present invention can be designed to enable potential grading of ML model results. This can be useful, for example, for evaluating deployed (e.g., client-side) ML models by identifying uncertain decisions or outputs. This can help reduce the impact of erroneous or inaccurate decisions, thereby providing improved data analysis. Therefore, embodiments can be used for real-time data evaluation purposes, such as evaluating whether a medical image analysis model is suitable for a specific object and / or medical scanning procedure.
[0057] Figure 1 An embodiment of a system 100 for obtaining a confidence metric for an ML model 105 is shown. In this document, the ML model 105 is deployed to a client and has therefore already been trained and completed. Therefore, retraining the ML model 105 on the client side is neither feasible nor possible.
[0058] System 100 includes an interface component 110 adapted to acquire input data 10. Herein, interface component 110 is adapted to receive input data 10 in the form of a medical image 10 from a medical imaging device 115 (such as, for example, an MRI device).
[0059] Medical image 10 is transmitted to interface component 110 via a wired or wireless connection. By way of example, the wireless connection may include a short-to-medium range communication link. For the avoidance of ambiguity, a short-to-medium range communication link may be considered to mean a short- or medium-range communication link with a range of up to about one hundred (100) meters. In a short-range communication link designed for very short communication distances, the signal travels from a few centimeters to a few meters, while in a medium-range communication link designed for short-to-medium range communication, the signal travels from a distance of up to one hundred (100) meters. Examples of short-range wireless communication links include ANT+, Bluetooth, Bluetooth Low Energy, IEEE 802.15.4, ISA 100a, Infrared (IrDA), Near Field Communication (NFC), RFID, 6LoWPAN, UWB, Wireless HART, Wireless HD, Wireless USB, and ZigBee. Examples of medium-range communication links include Wi-Fi, ISM band, and Z-Wave. In this document, the output signal is not encrypted for secure communication via wired or wireless connection. However, it should be understood that in other embodiments, one or more encryption technologies and / or one or more secure communication links may be used for signal / data communication in the system.
[0060] System 100 also includes a data modification component 120 configured to generate multiple modified instances of the input data. Specifically, in this embodiment, the data modification component 120 is configured to apply multiple different spatial curvature transformations to the medical image 10 to generate corresponding multiple modified instances of the medical image 10. Thus, the data modification component 120 makes small / minor modifications to the medical image 10 to generate multiple modified (or expanded) versions of the medical image.
[0061] System 100 also includes an ML model interface 122 configured to transmit medical image 10 and multiple modified instances of the medical image to ML model 105. For this purpose, the machine learning model interface 122 of system 100 can transmit machine learning model 105 via the Internet or a “cloud” 50.
[0062] In response to receiving medical image 10 and multiple modified instances of the medical image, the ML model processes the received data to generate corresponding results. More specifically, medical image 10 is processed by ML model 105 to generate preliminary results, and modified instances of the input data are processed by ML model 105 to generate corresponding secondary results.
[0063] These results generated by ML model 105 are then communicated back to system 100. Therefore, machine learning model interface 122 is also configured to receive preliminary results (generated by ML model 105 processing medical image 10) and multiple secondary results (generated by ML model 105 processing corresponding multiple modified instances of medical image 10).
[0064] In this embodiment, because the data modification unit 120 applies multiple different spatial curvature transformations to the medical image 10 (to generate corresponding multiple modified instances of the medical image 10), the data modification unit 120 is also configured to apply corresponding inverse spatial curvature transformations to the received multiple secondary results. Thus, the secondary results are transformed back or normalized for reference to the preliminary results.
[0065] It should be noted in this document that, for this purpose, by applying transformations and / or inverse transformations, the data modification component 120 can communicate with one or more data processing resources available on the Internet or "cloud" 50. Such data processing resources can undertake some or all of the processing required to implement the transformation. Therefore, it should be understood that this embodiment can employ distributed processing principles.
[0066] System 100 also includes an analysis component 124 configured to determine a confidence metric based on the quadratic result. In this document, analysis component 124 is configured to determine the confidence metric based on the inverse variance of the quadratic result. Therefore, a lower variance in the quadratic result corresponds to a higher confidence metric value. Conversely, a higher variance in the quadratic result corresponds to a lower confidence metric value. The confidence metric is then correlated with the input medical image 10 and its preliminary results processed by the ML model 105.
[0067] Furthermore, it should be noted that in order to determine the confidence metric based on the quadratic result, the analysis component 124 can communicate with one or more data processing resources available on the Internet or "cloud" 50. Such data processing resources can perform some or all of the processing required to determine the confidence metric. Therefore, it should be understood that this embodiment can employ distributed processing principles.
[0068] The analysis unit 124 is also adapted to generate an output signal 130 representing the preliminary results and the determined confidence level. In other words, after determining the reliability level of the preliminary results, an output signal 130 representing the reliability level of the preliminary results is generated.
[0069] The system also includes a graphical user interface (GUI) 160 for providing information to one or more users. Output signals 130 are provided to the GUI 160 via a wired or wireless connection. By way of example, a wireless connection may include a short- to medium-range communication link. Figure 1 As indicated, output signal 130 is provided from data processing unit 110 to GUI 160. However, if the system is already using data processing resources via the Internet or cloud 50, output signal GUI 160 can be used by GUI 160 via the Internet or cloud 50.
[0070] Based on the output signal 130, the GUI 160 is adapted to convey information by displaying one or more graphical elements in its display area. In this way, the system can use an ML model 105, which can be used to indicate a level of determinism or confidence associated with the processing result, to convey information about the outcome of processing the medical image 10. For example, the GUI 160 can be used to display graphical elements to medical practitioners, data analysts, engineers, medical imaging device operators, technicians, etc.
[0071] Despite the details described above Figure 1 The example embodiments described relate to medical imaging, but it should be understood that the proposed concepts can be extended to other forms of input data, such as medical case notes, engineering images, etc.
[0072] Furthermore, as described above, it should be understood that ML models can take any suitable form. For example, ML models can include artificial neural networks, generative adversarial networks (GANs), Bayesian networks, or combinations thereof.
[0073] according to Figure 1 From the above description, it should be understood that the embodiments can provide an estimate of the confidence level of the results returned from the ML model. The confidence level estimate can be derived by augmenting / modifying the input data and is obtained for a deployed ML model on the client side.
[0074] The implementation may be based on the premise that an accurate or reliable / trustworthy ML model should have "good performance," meaning that small perturbations to the input data should have a small corresponding effect on the output data. To analyze the ML model, the input data to be processed (e.g., a new case) undergoes multiple augmentations / modifications, all of which are then processed by the ML model. The variance from the results of processing the augmented / modified input can then be used to determine a measure of confidence (which may represent the accuracy, robustness, or reliability of the ML model). Thus, the implementation can provide input-specific (e.g., case-specific) confidence measures, and this may be independent of any claims or recommendations made by the provider of the ML model.
[0075] Therefore, the proposed implementation can be summarized as a proposal based on interpreting the generalization ability of ML models as a continuous property. In other words, if an ML model (such as a neural network) generalizes well to unseen input data (e.g., case C0), then a small perturbation to the input data (C0) should only have a small corresponding effect on the ML model output.
[0076] The implementation example can be applied to the client-side deployment phase of the deep learning module (rather than the training phase).
[0077] Furthermore, by way of example, the method for obtaining a confidence metric for an ML model according to the embodiment can be summarized as follows:
[0078] (i) A new case C0 is input into the ML model, resulting in a specific preliminary result R0; in the case of semantic segmentation via voxel labeling, C0 corresponds to the input image volume, while R0 corresponds to a volume of the same size, where each voxel contains label information, for example, specifying the level as a single integer value for the corresponding image voxel in C0. Alternatively, for each voxel from C0, R0 may contain a probability vector whose entries correspond to the probabilities of the relevant level. For example, for an ML model that can decide between n different levels, the vector associated with a certain image voxel consists of n entries totaling 1.
[0079] (ii) Case C0 is modified ('expanded') in various ways, such as by spatial registration with other patients, by arbitrary local deformation (bending), by adding noise, or a combination of these methods. In the case of semantic segmentation, when applying spatial transformation, for each voxel in the curved image, the coordinates in the original image are stored.
[0080] (iii) Each extended version Ci of C0 is also passed to the ML model, and each yields a corresponding result Ri; for semantic segmentation, the commonly used ML model is the U-Net architecture.
[0081] (iv) A preliminary result R0 is provided to the user, supplemented by the variability of the set {Ri}, which indicates the confidence level associated with the preliminary result R0 (where the confidence level is inversely proportional to the variability of the set {Ri}). For semantic segmentation, the confidence level is calculated separately for each voxel.
[0082] For each voxel in C0, we look up the corresponding position in each Ci and collect the output of the ML algorithm from Ri. All outputs associated with the voxels under consideration are summarized using a histogram (level frequency summation). Based on the analysis of the histogram and related empirical distributions, such as Shannon entropy (i.e., Where pi is the probability of a given sign) or various deterministic measures such as the Gini coefficient (which is equal to the area under the absolute parity line) (defined as 0.5) are subtracted from the area under the Lorenz curve and divided by the area under the absolute parity line. In other words, twice the area between the Lorenz curve and the absolute parity line can be calculated.
[0083] For voxel labeling (e.g., semantic segmentation) and localization tasks, augmentation using differential homeomorphic space transformations (e.g., rigid, affine, differential homeomorphic bending) allows embodiments to uniquely transform the resulting labeled image or localization coordinates back to the original voxel grid (by applying an inverse transformation). Thus, in the labeling task, for each voxel in the original voxel grid, an entire label ensemble is generated. In addition to the Shannon entropy or Gini coefficient mentioned above, the sample variance of this label ensemble can then be used to derive a quantitative voxel-by-voxel confidence score. More precisely, the confidence score is inversely proportional to the sample variance or Shannon entropy.
[0084] For visualization, confidence can be encoded in the display as color saturation or opacity using the color hue of the level labels. In localization tasks, a complete set of point locations can be generated, and their distributions can also be superimposed as individual points or via a fitted point density function (e.g., using a normal distribution).
[0085] For classification tasks, the level assignment {R} deviates from level R0 (which is assigned to C0 by the network). iThe proportion of} can provide a robustness measure (i.e., a confidence measure) of the transformations used for the expansion.
[0086] Now, for reference Figure 2 The diagram depicts a flowchart of a method 200 for obtaining a confidence metric for an ML model. For the purposes of this example, the ML model includes at least one of the following: an artificial neural network, a generative adversarial network (GAN), and a Bayesian network.
[0087] exist Figure 2 In step 210, the input data d0 is processed using an ML model to generate a preliminary result R0.
[0088] Step 220 includes: generating multiple modified instances (d0, ... i d ii d iii More specifically, step 220, which generates multiple modified instances of the input data, includes several steps, namely steps 222, 224, and 226. Step 222 includes applying a first spatial curvature transformation (e.g., a rigid or affine transformation) to the input data d0 to generate a first modified instance d of the input data. i Step 224 includes: adding noise to the input data d0 to generate a second modified instance d of the input data. ii Step 226 includes: applying a local deformation transformation to the input data d0 to generate a third modified instance d of the input data. iii .
[0089] Step 230 includes: processing multiple modified instances of the input data using an ML model (d i d ii d iii To generate multiple corresponding quadratic results (R) i R ii R iii It should be noted in this paper that, where appropriate, the corresponding inverse transform can also be used to transform the quadratic result back (i.e., normalize) for comparison with the initial result R0. For example, the inverse space curvature transform is applied to the first modified instance d of the input data. i The generated quadratic result R i Furthermore, the inverse local deformation transformation is applied to the third modified instance d of the input data. iii The generated quadratic result R iii .
[0090] Step 240 includes: based on the (normalized) quadratic result (R i R ii R iiiThis paper determines the confidence measure associated with the ML model. In this paper, determining the confidence measure includes: based on (normalized) quadratic results (R²). i R ii R iii The confidence metric is determined by the inverse variance of the equation. As mentioned above, in other embodiments, determining the confidence metric may include determining the Shannon entropy or the Gini coefficient.
[0091] Figure 2 An exemplary embodiment further includes step 250: associating the determined confidence measure with the preliminary result R0. This allows for the generation and association of a simple representation of the reliability of the preliminary result R0 with the preliminary result R0. This enables users to quickly and easily identify and evaluate the importance and / or relevance of the preliminary result R0 relative to the analysis ML model's decision or output on the input data d0.
[0092] Now, for reference Figure 3 This describes another embodiment of the system according to the invention, which includes a suitable ML model 410. In this document, the ML model 410 includes a conventional neural network 410, which can be used, for example, for medical / clinical decision services.
[0093] The neural network 410 transmits an output signal representing the result or decision from processing the input data to a remotely located data processing system 430 via the Internet 420 (e.g., using a wired or wireless connection) to obtain a confidence metric of the ML model (such as a server).
[0094] The data processing system 430 is adapted to acquire and process input data according to the method of the proposed embodiment to identify the confidence metric of the ML model 410.
[0095] More specifically, the data processing system 430 acquires input data and generates multiple modified instances of the input data. Then, the data processing system 430 transmits the input data and the multiple modified instances of the input data to the ML model 410. The ML model 410 processes the data received from the data processing system 430 to generate corresponding results. More specifically, the input data is processed by the ML model 410 to generate preliminary results, and the modified instances of the input data are processed by the ML model 410 to generate corresponding multiple secondary results. Then, the data processing system 430 acquires the preliminary results and the multiple secondary results from the ML model 410. Based on the variability of the secondary results, the data processing system 430 determines a confidence measure for the preliminary results created by the ML model 410.
[0096] The data processing system 430 is also adapted to generate an output signal representing a confidence metric. Therefore, the data processing system 430 provides a centrally accessible processing resource that can receive input data and run one or more algorithms to identify and assess the reliability of preliminary results (e.g., decisions) output by the ML model 410 for the input data. Information related to the acquired confidence metrics can be stored by the data processing system (e.g., in a database) and made available to other components of the system. This provision of information about the ML model and preliminary results can be made in response to a request (e.g., via the Internet 420) and / or can be made without a request (i.e., 'push').
[0097] In order to receive information about the reliability of the ML model or preliminary results, and thus to perform model / data analysis or evaluation, the system also includes a first mobile computing device 440 and a second mobile computing device 450.
[0098] In this document, the first mobile computing device 440 is a mobile phone device (such as a smartphone) having a display for showing graphical elements representing confidence metrics. The second mobile computing device 450 is a mobile computer, such as a laptop computer or tablet computer, having a display for showing graphical elements representing ML model results and related confidence metrics.
[0099] The data processing system 430 is adapted to transmit output signals to a first mobile computing device 440 and a second mobile computing device 450 via the Internet 420 (e.g., using a wired or wireless connection). As mentioned above, this can be done in response to receiving a request from the first mobile computing device 440 or the second mobile computing device 450.
[0100] Based on the received output signals, the first mobile computing device 440 and the second mobile computing device 450 are adapted to display one or more graphic elements in a display area provided by their respective displays. To this end, both the first mobile computing device 440 and the second mobile computing device 450 include software applications for processing, decrypting, and / or interpreting the received output signals to determine how to display the graphic elements. Therefore, each of the first mobile computing device 440 and the second mobile computing device 450 includes a processing arrangement adapted to determine one or more values representing a confidence metric and to generate graphic elements for modifying their size, shape, position, orientation, pulsation, or color based on the confidence metric.
[0101] Therefore, the system can convey information about the features in the ML model results to the users of the first mobile computing device 440 and the second mobile computing device 450. For example, each of the first mobile computing device 440 and the second mobile computing device 450 can be used to display graphical elements to medical practitioners, data analysts, or technicians.
[0102] Figure 3 The implementation of the system can vary between the following scenarios: (i) where the data processing system 430 conveys display-ready data, which may include, for example, display data containing graphic elements (e.g., in JPEG or other image formats) simply displayed to a user of the mobile computing device using conventional images or web page displays (which may be web-based browsers, etc.); and (ii) where the data processing system 430 conveys raw data information, which the receiving mobile computing device then processes to generate preliminary results and relevant confidence metrics (e.g., using local software running on the mobile computing device). Of course, in other implementations, processing can be shared between the data processing system 430 and the receiving mobile computing device, such that a portion of the data generated by the data processing system 430 is sent to the mobile computing device for further processing by the mobile computing device's local dedicated software. Therefore, embodiments can employ server-side processing, client-side processing, or any combination thereof.
[0103] Furthermore, in cases where the data processing system 430 does not 'push' information (e.g., output a signal) but instead conveys information in response to a request to receive it, the user of the device making such a request may need to confirm or verify their identity and / or security credentials in order to convey the information.
[0104] This invention can be a system, method, and / or computer program product for obtaining a confidence metric of a machine learning model. The computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to perform aspects of this invention.
[0105] Computer-readable storage media can be tangible devices capable of retaining and storing instructions for use by an instruction execution device. Computer-readable storage media can be, for example, but not limited to, electronic storage devices, magnetic storage devices, optical storage devices, electromagnetic storage devices, semiconductor storage devices, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of computer-readable storage media includes: portable computer floppy disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable optical disc read-only memory (CD-ROM), digital versatile disk (DVD), memory sticks, floppy disks, mechanical encoding devices such as punched cards or recessed structures in which instructions are recorded, and any suitable combination of the foregoing. The computer-readable storage media used herein should not be construed as transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides or other transmission media (e.g., light pulses through optical fibers), or electrical signals transmitted through wires.
[0106] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to a suitable computing / processing device or via a network (e.g., the Internet, a local area network, a wide area network, and / or a wireless network) to an external computer or external storage device. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards them to a computer-readable storage medium within the suitable computing / processing device.
[0107] Computer-readable program instructions can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., via the Internet using an Internet service provider). In some embodiments, an electronic circuit system including, for example, a programmable logic circuit system, a field-programmable gate array (FPGA), or a programmable logic array (PLA) can execute computer-readable program instructions by utilizing state information of the computer-readable program instructions to personalize the electronic circuit system in order to perform aspects of the present invention.
[0108] This document describes aspects of the invention with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.
[0109] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / actions specified in one or more blocks of a flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that can direct a computer, programmable data processing apparatus, and / or other apparatus to operate in a particular manner, such that the computer-readable storage medium in which the instructions are stored includes an article of writing comprising instructions for implementing aspects of the functions / actions specified in one or more blocks of a flowchart and / or block diagram.
[0110] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer-implemented process, such that the instructions, which execute on the computer, other programmable apparatus or other device, implement the functions / actions specified in one or more boxes of a flowchart and / or block diagram.
[0111] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each box in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which includes one or more executable instructions for implementing one or more specified logical functions. In some alternative implementations, the functions marked in the boxes may occur in a non-diagramal order. For example, depending on the functions involved, two consecutively shown boxes may actually be executed substantially simultaneously, or these boxes may sometimes be executed in reverse order. It should also be noted that each box in the block diagram and / or flowchart, and combinations of boxes in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified function or action or executes a combination of dedicated hardware and computer instructions.
[0112] Therefore, based on the above description, it should be understood that the embodiments can be used to determine a confidence metric independent of the ML model provider's claims. This confidence metric can be specific to the input data (e.g., specific to the input case). Thus, the embodiments can provide information that can be used as a robustness metric for grading and detection tasks. Furthermore, data modifications / expansions can be specifically designed for particular use cases.
[0113] Therefore, the proposed embodiments are applicable to a wide range of data analysis concepts / domains, including medical data analysis and clinical decision support applications. For example, the embodiments can be used for medical image screening, where medical images of objects are used to investigate and / or evaluate objects. In this case, pixel-level information about gradational uncertainty (i.e., decision confidence) can be provided, which can interpret or supplement the ML model output of medical professionals (e.g., radiologists).
[0114] This description is presented for illustrative and descriptive purposes and is not intended to be exhaustive or limiting of the invention in the disclosed forms. Many modifications and variations will be apparent to those skilled in the art. Embodiments have been selected and described to best explain the principles of the proposed embodiments, one or more practical applications, and to enable others skilled in the art to understand various embodiments with various modifications.
Claims
1. A method for obtaining a confidence metric for a machine learning model, the method comprising: The machine learning model is used to process the input data to generate preliminary results, wherein the input data includes images or text descriptions; Generate multiple modified instances of the input data; The machine learning model is used to process the multiple modified instances of the input data to generate multiple corresponding secondary results; as well as Based on the secondary results, a confidence measure related to the preliminary results is determined, wherein determining the confidence measure includes: Determine the distribution or measure of the variance of the quadratic result; and Determining a confidence measure based on the determined distribution or variance measure, wherein determining the distribution or variance measure of the quadratic result includes determining at least one of the following: The inverse variance of the quadratic result; The Shannon entropy of the quadratic result; The Gini coefficient of the second result; The Kullback-Liebler divergence of the quadratic result; and The concentration measure of the secondary result, wherein the multiple modification instances that generated the input data include: A first spatial curvature transformation is applied to the input data to generate a first modified instance of the input data; The first inverse space curvature transformation is applied to the quadratic result generated for the first modified instance of the input data.
2. The method according to claim 1, wherein generating multiple modified instances of the input data includes: Noise is added to the input data to generate a second modified instance of the input data.
3. The method according to claim 1, wherein generating multiple modified instances of the input data includes: The local deformation transformation is applied to the input data to generate a third modified instance of the input data.
4. The method according to claim 3, further comprising: The first inverse local deformation transformation is applied to the second result generated by the third modification instance of the input data.
5. The method according to any one of claims 1 to 4, wherein the machine learning model comprises at least one of the following: Artificial neural networks; Generative Adversarial Networks (GANs); and Bayesian networks.
6. The method according to any one of claims 1 to 4, further comprising: The determined confidence metric is correlated with the preliminary results.
7. A computer program product for obtaining a confidence metric of a machine learning model, wherein the computer program product includes a computer-readable storage medium having computer-readable program code embodied therein, the computer-readable program code being configured to perform the steps of any one of claims 1 to 6 when executed on at least one processor.
8. A system comprising at least one processor and a computer program product according to claim 7.
9. A system for obtaining a confidence metric for a machine learning model, the system comprising: An input interface (110) is configured to acquire input data, wherein the input data includes an image or text description; The data modification component (120) is configured to generate multiple modified instances of the input data; The machine learning model interface (122) is configured to transmit the input data and the plurality of modified instances of the input data to the machine learning model (105), and is also configured to receive preliminary results generated by the machine learning model processing the input data, and is configured to receive a plurality of secondary results generated by the machine learning model processing the corresponding plurality of modified instances of the input data; as well as An analysis component (124) is configured to determine a confidence measure related to the preliminary result based on the quadratic result, wherein the data modification component is configured to apply a first spatial curvature transformation to the input data to generate a first modified instance of the input data, wherein the analysis component is configured to determine a measure of the distribution or variance of the quadratic result, and is configured to determine a confidence measure based on the determined measure of the distribution or variance, wherein determining the measure of the distribution or variance of the quadratic result includes determining at least one of the following: The inverse variance of the quadratic result; The Shannon entropy of the quadratic result; The Gini coefficient of the second result; The Kullback-Liebler divergence of the quadratic result; as well as The concentration measure of the quadratic result, wherein the data modification component is further configured to apply a first inverse space curvature transformation to the quadratic result generated for the first modification instance of the input data.
Citation Information
Patent Citations
Systems and methods for automated diagnosis and decision support for breast imaging
WO2005001740A2
Registration apparatus for registering images
WO2017001210A1