Self-surveillance monitoring system for ai solutions
The AI monitoring system addresses reliability issues in medical imaging diagnostics by generating 'out-of-distribution' clusters to detect and alert errors, enhancing the confidence and adoption of AI models in clinical practice.
Patent Information
- Application Number
- PCT/EP2025/059529
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-19
- Filing Date
- 2025-04-08
- Publication Date
- 2025-10-23
AI Technical Summary
Existing artificial intelligence (AI) models in medical imaging diagnostics face challenges in clinical practice adoption due to unreliable performance and the inability to detect errors, particularly when applied to data outside their training domain, leading to reduced clinician trust and hindered widespread adoption.
A monitoring system for AI models that uses a self-surveillance mechanism to continuously check performance by generating 'out-of-distribution' clusters, providing confidence indicators and alerts for potential errors, allowing for post-market surveillance and improved reliability.
Enhances confidence in AI-assisted diagnostics by flagging potential errors and providing insights into cases not covered by the current model, thereby improving the reliability and adoption of AI solutions in medical imaging.
Smart Images

Figure EP2025059529_23102025_PF_FP_ABST
Abstract
Description
SELF-SURVEILLANCE MONITORING SYSTEM FOR Al SOLUTIONSFIELD OF THE INVENTION
[0001] The present invention is generally related to machine learning, and more particularly, to artificial intelligence-assisted, medical image processing. In particular, the invention relates to a monitoring system and a method for a medical diagnostic Al model trained on a training data set.BACKGROUND OF THE INVENTION
[0002] Machine learning models, neural networks or similar forms of artificial intelligence (Al) are becoming more and more popular to support various medical tasks, including helping with worklist prioritization, serving as a second opinion observer, preprocessing images, and / or automating workflows.
[0003] Al has been extensively investigated as a tool for computer-aided diagnosis (CAD) in radiology. Thousands of Al solutions have been developed in research labs, and, to date, some of these perform on par, or even better than clinicians in some circumstances. For example, Al systems can accurately diagnose pneumonia on chest X-ray images, and detect breast cancer on mammographies and pulmonary nodules on chest computed tomography (CT). Despite such performance, Al solutions face challenges in clinical practice adoption. For instance, for every new approved Al solution for a clinical problem, it is in its regular (e.g., daily) clinical application that one observes where the Al solution works as expected, and where it does not work. In theory, a human observer or user should catch all issues or deficiencies of the Al-tool, and, if the Al solution indicates poor performance (e.g., the Al-tool provides false or hallucinated results), the user may choose not to follow the recommendation of the Al-tool. However, in practice, questions are raised with such human monitoring of the Al-tool. For instance, what if less experienced clinical personnel are not able to catch Al errors? Or, what if clinicians do catch Al errors, but as such errors become more and more frequent, the clinicians lose trust in the Al-model, and consequently, lose the time and / or speed benefit an Al-tool can provide. In both cases, the benefit of Al are reduced significantly, and its widespread adoption in clinical practice is hindered.SUMMARY OF THE INVENTION
[0004] One object of the present invention is to provide confidence in the use of Al-assisted medical imaging diagnostics. To better address such concerns, in a first aspect of the invention, a monitoring system for a medical diagnostic Al model trained on a training data set is disclosed that includes a training data set and a library comprising one or more clusters, and a processor configured to receive input data, receive a medical diagnostic result from an Al model based on the input data, determine a similarity score between the input data and the training data set and the input data and the one or more clusters, and provide an indication of confidence of applicability of the Al model to the input data based on the determined similarity scores. Though Al models for medical data promise to revolutionize clinical practice, at the same time, Al models should be very accurate to avoid errors. Training data is generally expected to not represent all possible cases, and it is likely that in real clinical applications, the Al model may be applied to cases it cannot handle well, possibly resulting in errors (e.g., false results). Certain embodiments of the monitoring system provide a post-market, self-surveillance monitoring system that continuously checks the performance of Al models in application through the use of similarity scores and (failure) clusters, which improves confidence in such Al-solutions and helps in more widespread adoption of Al- solutions in medical imaging diagnostics, which may provide more consistent results when compared to the variability in user abilities in medical imaging diagnostics.
[0005] In one embodiment, the monitoring system is configured to output to one of the one or more clusters of the library the similarity score and the medical diagnostic result based on a similarity to the one of the one or more clusters; or instantiate (e.g., create and enter) a new cluster to the library and associate the medical diagnostic result, the similarity score, and the input data to the new cluster based on the similarity score. In some embodiments, the monitoring system collects the Al-failure cases to generate “out-of-distribution” (failure) clusters of data within the clinical application (e.g., due to clinical and / or demographic differences), which enables the monitoring system to flag cases where the Al model likely will not work as expected, and to provide insightinto cases not covered by the current model. In some embodiments, cluster generation is performed live during operation of the model, without manual intervention and without any intervention required by the model manufacturer.
[0006] These and other aspects of the invention will be apparent from and elucidated with reference to the embodiment(s) described hereinafter.BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Many aspects of the invention can be better understood with reference to the following drawings, which are diagrammatic. The components in the drawings are not necessarily to scale, emphasis instead being placed upon clearly illustrating the principles of the present invention. Moreover, in the drawings, like reference numerals designate corresponding parts throughout the several views.
[0008] FIG. 1 is a schematic diagram of an example medical imaging and processing system, in accordance with an embodiment of the invention.
[0009] FIG. 2 is a schematic diagram of an example processing / control system that comprises Al-assisted medical imaging diagnostics and Al-model result monitoring, in accordance with an embodiment of the invention.
[0010] FIGS. 3A-3D are schematic diagrams that illustrate an example library of one or more clusters, and a training data set, in a similarity space of failure cases for an example processing / control system, in accordance with an embodiment of the invention.
[0011] FIG. 4 is a schematic diagram of report generation in a processing / control system comprising a library of one or more clusters, in accordance with an embodiment of the invention.
[0012] FIG. 5 is a block diagram of an example processing / control system, in accordance with an embodiment of the invention.
[0013] FIG. 6 is a flow diagram of an example method of monitoring a medical diagnostic result of an Al model, in accordance with an embodiment of the invention.DETAILED DESCRIPTION OF EMBODIMENTS
[0014] Disclosed herein are certain embodiments of a (self-surveillance) monitoring system for a medical diagnostic Al model trained on a training data set. Inone embodiment, the monitoring system continuously checks the performance of an Al model in applications using different input data by collecting Al-failure cases to generate “out-of-distribution” clusters of data within the clinical application (e.g., failed because of clinical or demographic differences when compared to training data set). Through operation of the monitoring system, cases where the monitored Al model is not likely to work as expected may be flagged to a user, which avoids reliance on false results while also providing insight into cases not covered by the current model.
[0015] As explained above, there are limits to the training data set used by a given Al-model, and hence false results may occur when the input data is not similar to the training data set. There is an expectation that human intervention in assessing the Al model results helps to catch these errors and prevent application of the Al model to data that may give rise to such errors, though differences in technical ability in assessment and / or the potential for recurring errors in the application of the Al model to different input data may hinder the adoption of Al solutions in clinical practice.
[0016] Further, training data is not always locally available (e.g., due to privacy, regulatory, storage, and / or processing time constraints) and can encompass vast amounts of data with high dimensionality, especially moving towards multiomics approaches (with joint input of, for instance, images and meta data including demographics, anamnesis, blood test data, etc.). With a trained Al model in clinical operation, comparing new input to the model with previous training data “on the fly” may not be practically feasible without a suitable strategy for the compression, representation and availability of training data metrics to judge whether an inference query is covered by the training domain or not. As many Al models today interpolate well between training data points, an extrapolation to input data outside the training domain often produces hallucinated or false results - a no-go in radiology and medical practice. One problem is that Al models existing today do not intrinsically know whether a new data point is “in-distribution" or “out-of-distribution" compared to the training data. In contrast, certain embodiments of a monitoring system provides for feasible data compression and either a cloud-based, or local availability, of outlier assessment metrics for new input data, based on the training data of the model.
[0017] Having summarized certain features of a monitoring system of the present disclosure, reference will now be made in detail to the description of a monitoring system as illustrated in the drawings. While a monitoring system will be described in connection with these drawings, with emphasis on disease classification and localization in (e.g., chest x-ray) images, there is no intent to limit it to the embodiment or embodiments disclosed herein. For instance, the monitoring system may be used in other Al-based imaging detection and diagnosis applications in medical and / or research industries, including digital pathology or for abnormal growth detection / diagnosis in other (non-lung, such as kidney stones or uterine or ovary cysts) anatomical structures of a subject. Further, although the description identifies or describes specifics of one or more embodiments, such specifics are not necessarily part of every embodiment, nor are all of any various stated advantages necessarily associated with a single embodiment. On the contrary, the intent is to cover alternatives, modifications and equivalents included within the principles and scope of the disclosure as defined by the appended claims. For instance, two or more embodiments may be interchanged or combined in any combination. Further, it should be appreciated in the context of the present disclosure that the claims are not necessarily limited to the particular embodiments set out in the description.
[0018] Turning attention now to the figures, FIG. 1 is a schematic diagram of an example medical imaging and processing system 10. The medical imaging and processing system 10 depicted in FIG. 1 comprises a medical imaging system 12. The medical imaging system 12 may include any one of a variety of medical imaging systems, including a magnetic resonance imaging system, a computed tomography system, a positron emission tomography system, a single photon emission tomography system, an ultrasound system, a digital X-ray system, and a digital fluoroscope. In some embodiments, other types of imaging systems may be used for other purposes, such as histological tissue scanning, scanning electron microscope imaging, etc.
[0019] In this example a subject (e.g., patient) 14 is shown reposing on a support 16, which supports at least a portion of the subject 14 within an imaging zone. Control of the medical imaging system 12 is via a processing / control system 18. The processing / control system 18 executes (machine) executable instructions to cause theprocessing / control system 18 to control the medical imaging system 12 via, for instance, a hardware interface. The processing / control system 18 controls the medical imaging system 12 to acquire and process (e.g., reconstruct) a medical image. In one embodiment, the processing / control system 18 uses artificial intelligence to perform various processing operations on the medical image, including localization and delineation of a region of interest (e.g., a tumor and / or anatomical structure or organ), and / or disease classification.
[0020] The medical imaging and processing system 10 depicted in FIG. 1 further includes a display device 20 coupled to the processing / control system 18. For instance, the processing / control system 18 may comprise reading or CAD software 12, which includes a worklist graphical user interface (GUI) and an image reading page GUI. The worklist GUI, or worklist may include a multitude of data, including subject name and ID, exam type (e.g., chest CT), ordering physician, priority, status (e.g., scheduled, in progress, completed), assigned radiologist, exam location (e.g., department, facility), exam notes, and results or report status, with each row of the worklist comprising a subject case / exam to be reviewed. Artificial intelligence implemented in the processing / control system 18 may process and analyze the image and meta data to localize, delineate, and / or classify diseased structures or organs. The display device 20 also enabled user-inputted information (e.g., via an input device, such as a mouse, keyboard, and / r touch or voice activated device), including annotations of region of interest features, such as inputted by a clinician. In one embodiment, the processing / control system 18 enables, via user interaction with the display device 20, tagging of features in the image or labeling, such as to label false responses generated on input data by the artificial intelligence system (e.g., an Al model or engine).
[0021] The medical imaging and processing system 10 is communicatively coupled to one or more endpoint device 22. For instance, endpoint device 22 may be associated with one or more recipients of information provided by the processing / control system 18. Endpoint devices 22 may include a laptop, phone, pager or any other device of medical staff to inform them of, say, a detected medical condition, and / or reporting information (e.g., statistics of the failure clusters generated by the processing / control system 18).
[0022] Note that in some embodiments, endpoint devices 22 may represent local and / or remotely-accessed (e.g., via a wide or local area network) resources of information that may be used for acquiring training data sets, image data, patient data, etc. For instance, endpoint devices 22 may provide a source of medical records of one or more hospital information systems HIS (hospital information system) or PACS (picture archive communication system), or medical data repositories (e.g., Oncology information system (“OIS”)), and / or imagery relevant to a region of interest (e.g., lungs, heart, etc.) and type of medical condition may be retrieved based on relevant information in DICOM header files of the image data or in annotations, possibly combined with information held in patient data records.
[0023] Note that functionality of the medical imaging and processing system 10 may be implemented in one or more devices that are local to each other or located in different locations and in communication over a network (e.g., wide or local area network).
[0024] The processing / control system 18 may be implemented using a dedicated or general purpose computing system, including one or more co-located or network - connected processors. For instance, the processing / control system 18 may include one or more multi-core processors such as GPU (graphics processor unit) or TPU (tensor processing unit). A single one or a plurality of computing units may be used in the processing / control system 18, and in one embodiment, may be implemented as one or more server(s) communicatively coupled in a communication network such as in a Cloud architecture. For instance, the computing for the machine learning or artificial intelligence functionality component may be a local computing resource, or implemented remotely via one or more servers. In some embodiments, one or more of the functionality of the processing / control system 18 may be implemented using hardware components such as FPGAs or ASICS.
[0025] Having generally described an example medical imaging and processing system 10, attention is now directed to FIG. 2, which illustrates an embodiment of an example processing / control system 18 that comprises Al-assisted medical imaging diagnostics and Al-model result monitoring. In the depicted example, shown is the medical imaging system 12 that provides input data24 to the processing / control system18. It should be appreciated that the input data24 may be provided to the processing / control system 18 via an intermediary device or system (e.g., image and meta data storage, etc.). Also shown is the display device 20 with a graphical user interface (GUI) 26, which may be rendered via software in the processing / control system 18 to present a reconstructed medical image that has been segmented and classified via implementation of an Al model on the input data 24. Tagging by a user (e.g., clinician) of the reconstructed medical image rendered in the GUI 26 may be enabled via the use of input device 28, which may be user-manipulated (e.g., mouse, keyboard, toggle, voice-activated, etc.).
[0026] In the depicted embodiment of FIG. 2, the processing / control system 18 comprises an Al engine 30 implementing a trained Al model. As explained above, in some embodiments, the Al engine 30 may be implemented locally (e.g., co-located in the facility in which the medical imaging system 12 is used), or in a device implemented as part of a remote server farm or cloud platform. The Al engine 30 may be implemented as a neural-network architecture, for instance a convolutional neural- network architecture (CNN), though in some embodiments, other machine learning models may also be used, such as support vector machines, regression models, decision trees and others.
[0027] Note that the terms Al engine 30 and Al model are used interchangeably herein, with the understanding that the Al engine provides a computational resource for implementing an Al algorithm or model. The Al model is trained on a training data set. The Al engine 30 (via the Al model) provides an Al result 32, which may be a disease classification and localization of, say, a chest x-ray image. The Al result 32 may be presented on the GUI 26 of the display device 20.
[0028] The processing / control system 18 further includes a monitoring system 34, which is explained further below. The monitoring system 34 provides an Al algorithm or model) (e.g., such as implemented by the Al engine 30) with a method to monitor its own performance within a clinical application and to self-detect classification issues, which allows post-market surveillance beyond the training data set on which the Al model is trained. With the monitoring result, the Al-output (i.e., the Al result 32) may be annotated or labelled as “to be checked” or similar message that provides an indicationof confidence of applicability of the Al model to the input data, since error issues are known for similar cases based on a library of failure clusters). In some embodiments, the monitoring system 34 map provide an overview of failure cases that can be communicated (e.g., to recipients associated with endpoint devices 22 and / or the clinician running the Al model on the present input data 24 for a given reconstructed image) to assist in understanding when the Al model is (e.g., highly) accurate and when it is not.
[0029] With continued reference to FIG. 2, reference is made to FIGS. 3A-3D, which illustrate an example library of one or more clusters, and a training data set, in a similarity space of failure cases for an example processing system (e.g., the processing / control system 18 of FIGS. 1-2). As shown in these example illustrations, an embodiment of the monitoring system 34 comprises a similarity space 36. The similarity space 36 comprises (in FIGS. 3A-3D) a training data cluster 38, a cluster failurel 40 (FIGS. 3B-3D), a cluster failure2 42 (FIGS. 3C-3D), and a cluster failures 44 (FIG. 3D) based on different iterations and content of input data 24 (e.g., respectively, data(i), data(ii), data(iii), data(iv)). Note that the input data may be in any one or combination of image data, meta data, text, audio, video material, and / or data from multiple sensors and sources. The clusters 40-44 are collected in a library, or, Al failures library 46. The transition in the cluster quantity from FIGS. 3B-3D are indicative of input data 24 for which similarity measures or scores cause the instantiation or generation by the monitoring system 34 of an additional cluster. Note that the illustrations in FIGS. 3A-3D are merely for illustrative example, with the understanding that there may be many more failure clusters. As more failure cases are collected, the strength of the out-of- distribution clusters increases. Assuming data security concerns are properly dealt with, the failure cases may be collected on more than one hospital (e.g., of a hospital chain like Asklepios). The monitoring system 34 is likely to lead to clearer and more stable results much faster with the increase in failure clusters.
[0030] As noted, the monitoring system 34 receives the input data 24 and receives and responds to the Al result 32.
[0031] Explaining further, as explained above, the processing / control system 18 (e.g., the Al system portion of the processing / control system 18) implements a taggingsystem in conjunction with the GUI 26 and input device 28. Through the use of the tagging system, a clinician can label false responses. A false response is saved, together with the input data (e.g., an X-ray image) and available meta-data (e.g., age, gender, size, weight, co-morbidities, true disease classification, etc.) in the library 46 (also referred to herein as an Al failures library).
[0032] Additionally, a similarity measure (also referred to herein as similarity score) is defined for the data in the Al failures library 46. It is expected that, at least in some cases for the training data, this similarity measure may be only partially applicable (e.g., for classification of disease on, say, chest X-rays). The training data set may contain images and, potentially, meta data (e.g., age or gender of the patients), but may not contain disease classifications that are not represented in the Al model or other comorbidities. Thus, the similarity measure for the training data set may be missing a few dimensions if similarity to the training data set is computed. In some embodiments, the images may be combined with meta data (e.g., age, gender, etc.) by extracting image feature vectors from the Al model. The combined vector of image features and meta data may then be used to derive a similarity measure. Examples for the similarity measure may include L1 norm, L2 norm, or cosine similarity.
[0033] The similarity measure is used to compare the Al failure cases to the original training data, and with each other. For the chest X-ray example, this means computing an image similarity measure, and a similarity measure for the available metadata vector in one embodiment. It is likely, that some failure cases are very close to the training data. However, it is also likely, that many of the failure cases, in some sense, are “out-of-distribution” of the training data on at least some properties. One aim of clustering of all the failure cases is to differentiate between in-distribution failure cases and “out-of-distribution” failure cases or clusters. For sufficiently large out-of-distribution clusters, properties may be deduced. In some embodiment, dimensionality reduction, and latent space representation and analysis may be used.
[0034] As to the training data and data used for the similarity measures, a few points are worth mention. In some instances, the original training data may be represented by averages, atlases, constructed example images, or distributions, and hence there is no need to disclose the training data, and the transfer of this distributionfor the similarity measure is feasible. The same is true for “out-of-distribution” clusters computed by the monitoring system 34: coarse representations as distributions or atlas- like-data may be shared. In some implementations, a team of clinicians may be able to classify the difference in human understandable terms (e.g., the population subgroup or disease classification where Al fails), and then the data itself may not be needed for actions. As an additional measure, reference data may be stored encrypted and / or in a latent space representation (e.g., from an autoencoder, where the back-transform is unique to a customer or patient and access rights to the back-transform(s) may be managed flexibly, possibly providing only a subset of the original training data dimensionality to certain ownership groups (e.g., physicians, researchers, FDA, etc.), depending on their purpose and permissions. The autoencoder / encryption algorithms may include de-identification measures.
[0035] Additionally, meta data on patients may, in some instances, be challenging to extract automatically for the Al engine 30, as the meta data may be available in a different clinical system. However, one option is to request metadata from the clinician (e.g., who reviews and rejects the Al result 32, such as via a form, where a true classification and additional data might be asked).
[0036] Referring to FIG. 2, a normal workflow follows the path of input data 24 to Al engine 30 to Al result 32 and acceptance or rejection (e.g., via a clinician viewing the results on the GUI 26) of the Al result 32. With the addition of the monitoring system 34, and referring also to FIGS. 3A-3D, failure cases are collected and clustered (e.g., clusters 40-44) in the similarity space 36 of failure cases. Any new case is projected onto that space 36 to determine, if it is close to the training data cluster 38, or close to a known “out-of-distribution” cluster (e.g., 40 42, 44). In the latter case, an alert or warning is issued (e.g., audio and / or visual, such as via GUI 26).
[0037] As an example illustration of operation of the monitoring system 34, assuming a first clustering (e.g., cluster failurel 40) and a new case (e.g., new input data 24) for the Al model, a similarity measure is applied to determine if the new case is similar to the training data set 38 (also, training data cluster), or similar to one of the “out-of-distribution” clusters (e.g., cluster failurel 40), or none. If the case is most similar to a known “out-of-distribution” cluster (e.g., cluster failurel 40), the Al result 32 may bepresented via GUI 26 with a warning (or alert) to alert a clinician to check this result carefully. Accordingly, if most of the failure cases are flagged with a warning, the clinicians can trust or have confidence in the Al results in other cases.
[0038] Referring to FIG. 4, shown is the monitoring system 34, including the library 46 of failure clusters 40-44, and the training data cluster 38. In some embodiments, the monitoring system 34 outputs a statistical report 48 (or simply, report) on the out-of-distribution clusters (40-44) for which the Al model is likely to produce failures. The report (e.g., in some embodiments, generated automatically) on the “out- of-distribution” clusters 40-44 may be made available at regular intervals, and / or in some embodiments, on demand. This report 48 may be reviewed by a clinician, or depending on data protection measures, even by scientific communities, a manufacturer (e.g., of the Al model), or regulatory bodies like the FDA. Potential consequences of such a review may be as follows:
[0039] (a) an identification of clear “clusters”, where the Al algorithm is not applicable and may be de-activated or ignored. This action may, for example, reveal bias in the model for a certain population (e.g., defined by certain demographic properties). Some examples may include a model trained on Caucasian western patients, which may not work well on Asians, or a model trained on elderly adults, which may not work well on young adults.
[0040] (b) an identification of a typical error of the Al-algorithm, which in itself is consistent (e.g., the Al algorithm always identifies disease A, when in fact property X and disease B is present. This information may be used to identify disease B (e.g., by informing the staff of this correlation).
[0041] (c) re-training of the Al-algorithm to include those “out-of-distribution” clusters, or training of a new Al-algorithm to deal with the “out-of-distribution” clusters.
[0042] Note that since the monitoring system 34 does not change the Al result, and does not provide any clinical recommendation, it may be easier to add the monitoring system 34 as a check on an already certified Al-solution. The updated Al- solution or a new Al-solution covering “out-of-distribution” clusters may require new certification, yet the monitoring system 34 may even support a subsequent certification.
[0043] FIG. 5 is a block diagram of an embodiment of example processing / control system 18. In the depicted embodiment, the processing / control system 18 is illustrated as a computing system, though as explained above, the processing / control system may be a sub-component(s) of a computing system, or functionality of the same may be distributed among multiple devices or computing systems. In some embodiments, the Al engine 30 may be a separate component that interfaces with the monitoring system (e.g., via a hardware or communications interface). Referring now to FIG. 5, shown is an embodiment of the example processing / control system 18 that may be used to implement the Al engine 30 and monitoring system 34. In the depicted embodiment, functionality of the Al engine 30 and monitoring system 34 is implemented as a system of co-located software and hardware components collectively embodied as a computing device (which may include a medical device) and plural sub-systems, wherein the computing device is operatively connected to the plural sub-systems. In some embodiments, the computing device functionality may be embedded in one of the subsystems, or one or more of the herein-described sub-system functionality may be integrated into fewer devices. It should be appreciated that, in some embodiments, functionality of the processing / control system 18 may be implemented via plural computing devices that are networked at a similar location or situated in different locations, potentially remote from each other, and connected via one or more networks. Note that reference to control software for image acquisition is omitted here for brevity. In some embodiments, the control software for image acquisition may be implemented in a separate device or system.
[0044] The processing / control system 18 includes one or more processors 50 (e.g., 50A...50N), input / output interface(s) 52, and memory 54, etc. coupled to one or more data busses, such as data bus 56. The processor(s) 50 may be embodied as a custom-made or commercially available processor, including a single or multi-core central processing unit (CPU), tensor processing unit (TPU), graphics processing unit (GPU), vector processing unit (VPU), or an auxiliary processor among several processors, a semiconductor-based microprocessor (in the form of a microchip), a macroprocessor, one or more application specific integrated circuits (ASICs), field programmable gate arrays (FPGUs), a plurality of suitably configured digital logic gates,and / or other existing electrical configurations comprising discrete elements both individually and in various combinations to coordinate the overall operation of the processing / control system 18.
[0045] The I / O interfaces 52 comprise hardware and / or software to provide one or more interfaces to various sub-systems, including to one or more user interface(s) 58 (which may include GUI 26 and display device 20) and the medical imaging system 12. The I / O interfaces 52 may also include additional functionality, including a communications interface for network-based communications. For instance, the I / O interfaces 52 may include a cable and / or cellular modem, and / or establish communications with other devices or systems via an Ethernet connection, hy brid / fi ber coaxial (HFC), copper cabling (e.g., digital subscriber line (DSL), asymmetric DSL, etc.), using one or more of various communication protocols (e.g., TCP / IP, UDP, etc.). In general, the I / O interfaces 52, in cooperation with a communications module (not shown), comprises suitable hardware to enable communication of information via PSTN (Public Switched Telephone Networks), POTS, Integrated Services Digital Network (ISDN), Ethernet, Fiber, DSL / ADSL, Wi-Fi, cellular (e.g., 3G, 4G, 5G, Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), etc.), Bluetooth, near field communications (NFC), Zigbee, among others, using TCP / IP, UDP, HTTP, DSL.
[0046] The user interface(s) 58 may include a keyboard, scroll-wheel, mouse, microphone, immersive head set, display device(s), GUI, etc., which enable input and / or output by or to a user, and / or visualization to a user. In some embodiments, the user interface(s) 58 may cooperate with associated software to enable augmented reality or virtual reality. The user interface(s) 58, when comprising a display device, enables the display of segmented anatomical structure(s), including organs, nodules, tumor(s), abnormal growth tissue (e.g., collectively referred to herein as a region or regions of interest) according to an Al result, and provision for tagging and alerts from the monitoring system 34. In some embodiments, the user interface(s) 52 may be coupled directly to the data bus 56.
[0047] The medical imaging system 12 includes one or more medical imaging systems and / or image storage sub-systems that are used to enable visualization ofanatomical structures and regions of interest associated with the Al result 32. The medical imaging system 12 may include ultrasound imaging, fluoroscopy, magnetic resonance, computed tomography, and / or positron emission tomography (PET). In some embodiments, the images may be retrieved from a picture archiving and communication systems (PACS) or any other suitable imaging component or delivery system. In some embodiments, the images provided by the medical imaging system 12 may be segmented according to existing segmentation algorithms.
[0048] The memory 54 may include any one or a combination of volatile memory elements (e.g., random-access memory RAM, such as DRAM, and SRAM, etc.) and nonvolatile memory elements (e.g., ROM, Flash, solid state, EPROM, EEPROM, hard drive, tape, CDROM, etc.). The memory 54 may store a native operating system, one or more native applications, emulation systems, or emulated applications for any of a variety of operating systems and / or emulated hardware platforms, emulated operating systems, etc. In some embodiments, a separate storage device (STOR DEV) may be coupled to the data bus 56 or as a network-connected device (or devices) via the I / O interfaces 52 and one or more networks. The storage device may be embodied as persistent memory (e.g., optical, magnetic, and / or semiconductor memory and associated drives).
[0049] In the embodiment depicted in FIG. 8, the memory 54 comprises an operating system 60 (OS) (e.g., LINUX, macOS, Windows, etc.), the library 46, the Al engine 30 comprising the Al model, the monitoring system 34, which includes similarity measure / score functionality 62, clustering functionality 64, and reporting functionality 66, and GUI software 68 (e.g., for tagging). In one embodiment, the Al engine 30 and the monitoring system functionality (e.g., including similarity measure / score functionality 62 and clustering functionality 64 and reporting functionality 66) is as described above in association with FIGS. 2-4, and may comprise a plurality of modules (e.g., executable code) co-hosted on the computing device (processing / control system 18), though in some embodiments, the modules may be distributed across various systems or subsystems over one or more networks, including implementation using a cloud computing platform. In some embodiments, there may be fewer or additional modules. For instance, functionality of some modules may be combined, or additional functionalitymay be implemented yet not shown, such as a communication module, security module, etc.
[0050] Functionality of one or more of the various modules is briefly explained here. The segmentation module 118 provides image segmentation of anatomical structures according to existing segmentation functionality. In some embodiments, segmentation may be performed elsewhere, and the segmented imaging transferred to the CAD software 116.
[0051] The Al engine 30 may utilize any of existing Al functionality that creates, trains, and deploys machine learning algorithms that emulate logical decision making, and includes linear or logical regression algorithms, decision trees, Bayes, K-nearest neighbors, support vector machines, and / or neural or deep neural networks. The Al engine 30 may be trained to perform disease classification and localization and delineation.
[0052] The monitoring system 34, which includes similarity measure / score functionality 62, clustering functionality 64, and reporting functionality 66, performs similarity measures / scores between input data and the training data set and between the clusters, and also clusters data based on the similarity measures / scores. The reporting functionality 66 provides for statistical reporting of the clusters, which may be communicated locally or to a remote location via a network communication delivered via I / O 52. Additional information on clustering may be found in K. Bindra and A. Mishra, "A detailed study of clustering algorithms," 2017 6th International Conference on Reliability, Infocom Technologies and Optimization (Trends and Future Directions) (ICRITO), Noida, India, 2017, pp. 371-376, doi: 10.1109 / ICRITO.2017.8342454. Additional information on similarity measures may be found in G. P. Penney, J. Weese, J. A. Little, P. Desmedt, D. L. G. Hill and D. J. hawkes, "A comparison of similarity measures for use in 2-D-3-D medical image registration," in IEEE Transactions on Medical Imaging, vol. 17, no. 4, pp. 586-595, Aug. 1998, doi: 10.1109 / 42.730403. Additional information on deep-learning for imaging may be found in Baltruschat, I. M., Nickisch, H., Grass, M., Knopp, T., & Saalbach, A. (2019). Comparison of deep learning approaches for multilabel chest X-ray classification. Scientific reports, 9(1), 1-10.
[0053] Note that the memory 54 and storage device may each be referred to herein as a non-transitory, computer readable storage medium or the like.
[0054] Execution of the one or more modules of the processing / control system 18 may be implemented by the one or more processors 50 under the management and / or control of the operating system 60.
[0055] When certain embodiments of the processing / control system 18 are implemented at least in part with software (including firmware), it should be noted that the software can be stored on a variety of non-transitory computer-readable (storage) medium for use by, or in connection with, a variety of computer-related systems or methods. In the context of this document, a computer-readable medium may comprise an electronic, magnetic, optical, or other physical device or apparatus that may contain or store a computer program (e.g., executable code or instructions) for use by or in connection with a computer-related system or method. The software may be embedded in a variety of computer-readable mediums for use by, or in connection with, an instruction execution system, apparatus, or device, such as a computer-based system, processor-containing system, or other system that can fetch the instructions from the instruction execution system, apparatus, or device and execute the instructions.
[0056] When certain embodiments of the processing / control system 18 are implemented at least in part with hardware, such functionality may be implemented with any or a combination of the following technologies, which are all already existing in the art: a discrete logic circuit(s) having logic gates for implementing logic functions upon data signals, an application specific integrated circuit (ASIC) having appropriate combinational logic gates, a programmable gate array(s) (PGA), a field programmable gate array (FPGA), TPUs, GPUs, and / or other accelerators / co-processors, etc.
[0057] One having ordinary skill in the art should appreciate in the context of the present disclosure that the example processing / control system 18 is merely illustrative of one embodiment, and that some embodiments of computing devices may comprise fewer or additional components, and / or some of the functionality associated with the various components depicted in FIG. 5 may be combined, or further distributed among additional modules or computing devices, in some embodiments. It should beappreciated that certain well-known components of computer systems are omitted here to avoid obfuscating more relevant features of the processing / control system 18.
[0058] FIG. 6 is a flow diagram of an example method of monitoring a medical diagnostic result of an Al model, in accordance with an embodiment of the invention. In one embodiment, the method, denoted as method 70, and which may be implemented by the processing / control system 18 or components thereof, includes receiving input data, the input data comprising image data and meta data (72); receiving the medical diagnostic result from the Al model based on the input data, the Al model trained on a training data set (74); determining a similarity score between the input data and the training data set and the input data and one or more clusters of a library (76); and providing an alert associated with use of the Al model based on the determined similarity scores (78).
[0059] While the invention has been illustrated and described in detail in the drawings and foregoing description, such illustration and description are to be considered illustrative or exemplary and not restrictive; the invention is not limited to the disclosed embodiments. Other variations to the disclosed embodiments can be understood and effected by those skilled in the art in practicing the claimed invention, from a study of the drawings, the disclosure, and the appended claims. Note that various combinations of the disclosed embodiments may be used, and hence reference to an embodiment or one embodiment is not meant to exclude features from that embodiment from use with features from other embodiments. In the claims, the word “comprising” does not exclude other elements or steps, and the indefinite article “a” or “an” does not exclude a plurality. A single processor or other unit may fulfill the functions of several items recited in the claims. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage. A computer program may be stored / distributed on a suitable medium, such as an optical medium or solid-state medium supplied together with or as part of other hardware, but may also be distributed in other forms. Any reference signs in the claims should be not construed as limiting the scope.
Claims
CLAIMS1 . A monitoring system (34) for a medical diagnostic Al model trained on a training data set (38), the monitoring system comprising: a memory (54) comprising executable instructions and information corresponding to the training data set and a library (46) comprising one or more clusters (40); and a processor (50), configured by the executable instructions to: receive input data (24); receive a medical diagnostic result (32) from an Al model based on the input data; determine a similarity score between the input data and the training data set and the input data and the one or more clusters; and provide an indication of confidence of applicability of the Al model to the input data based on the determined similarity scores.
2. The monitoring system of the preceding claim, wherein the processor is further configured by the executable instructions to either: output to one of the one or more clusters of the library the similarity score and the medical diagnostic result based on a similarity to the one of the one or more clusters; or instantiate a new cluster to the library and associate the medical diagnostic result, the similarity score, and the input data to the new cluster based on the similarity score.
3. The monitoring system of any one of the preceding claims, wherein the processor is further configured by the executable instructions to store user-inputted annotations associated with the medical diagnostic result and image data and meta data associated with the input data to the library.
4. The monitoring system of any one of the preceding claims, wherein the user- inputted annotations comprises a label conveying that the Al model is providing a false response, and the meta data comprises patient data.
5. The monitoring system of any one of the preceding claims, wherein the processor is further configured by the executable instructions to determine the similarity score for image data and meta data using a feature vector.
6. The monitoring system of any one of the preceding claims, wherein the one or more clusters comprise a plurality of clusters that are categorized as either an indistribution failure case or out-of-distribution failure case, the in-distribution failure case closer to the training data set than the out-of-distribution failure case.
7. The monitoring system of any one of the preceding claims, wherein the indication comprises alert to a user in using the Al result for the input data.
8. The monitoring system of any one of the preceding claims, wherein the processor is further configured by the executable instructions to provide a report (48) of out-of-distribution failure cases corresponding to the one or more clusters.
9. The monitoring system of any one of the preceding claims, wherein the processor is further configured by the executable instructions to increase an amount of clusters in the library over time.
10. The monitoring system of any one of the preceding claims, wherein the input data comprises patient imaging data and meta data associated with a patient and the patient imaging data.
11. A method (70) of monitoring a medical diagnostic result of an Al model, the method comprising: receiving input data, the input data comprising image data and meta data (72); receiving the medical diagnostic result from the Al model based on the input data, the Al model trained on a training data set (74);determining a similarity score between the input data and the training data set and the input data and one or more clusters of a library (76); and providing an alert associated with use of the Al model based on the determined similarity scores (78).
12. The method according to claim 11 , further comprising either: outputting to one of the one or more clusters of the library the similarity score and the medical diagnostic result based on a similarity to the one of the one or more clusters; or instantiate a new cluster to the library and associate the medical diagnostic result, the similarity score, and the input data to the new cluster based on the similarity score.
13. The method of any one of the claims 11 -12, further comprising: storing user-inputted annotations associated with the medical diagnostic result and the image data and the meta data to the library; and categorizing the one or more clusters as either an in-distribution failure case or out-of-distribution failure case, the in-distribution failure case closer to the training data set than the out-of-distribution failure case.
14. The method of any one of the claims 11 -13, further comprising providing a report of out-of-distribution failure cases corresponding to the one or more clusters.
15. A non-transitory, computer readable medium (54) comprising executable instructions that, when executed by a processor, causes the processor to implement any one of the claims 11 -14.
Citation Information
Patent Citations
Machine learning inference system
US20210125104A1
Computer-implemented natural language understanding of medical reports
WO2020214683A1