System and method for monitoring trained machine learning model in healthcare context
The system addresses the challenge of monitoring AI model performance in healthcare by extracting input data and user feedback characteristics to detect deviations, ensuring reliable and compliant operation without PHI, facilitating real-time adjustments.
Patent Information
- Application Number
- JP2024194069
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-22
- Filing Date
- 2024-11-06
- Publication Date
- 2025-07-08
AI Technical Summary
The challenge of monitoring AI model performance in healthcare environments is hindered by regulatory constraints and privacy concerns, necessitating a solution that can assess model performance without using protected health information (PHI) or clinically sensitive data.
A system and method that extracts characteristics of input data, model performance, and user feedback to generate surrogate data, comparing these with baseline data to detect deviations and send warnings, thus monitoring model performance in real-time without accessing PHI.
Enables continuous and compliant monitoring of AI model performance, ensuring reliability and regulatory compliance by using statistical surrogates for clinical data, allowing timely corrective actions.
Smart Images

Figure 2025102656000001_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention generally relate to the fields of artificial intelligence (AI) and machine learning, and more particularly, to systems and methods for monitoring the performance of AI models and applications in the field of medical imaging.
Background Art
[0002] As the introduction of AI in healthcare progresses, the need to monitor the performance of AI models in real-world scenarios has become essential. However, due to regulatory constraints and privacy issues, access to actual clinical data (such as medical images and personal health information (PHI)) is often restricted. Data scientists and AI developers need to understand how their models are functioning in the field and make the necessary adjustments to improve the performance and reliability of the models. Therefore, the above limitations pose a major challenge for data scientists and AI developers.
[0003] Conventionally, when monitoring the performance of an AI model in a healthcare environment, patient clinical data has been used and the AI model has made inferences based on the patient's clinical data. However, this approach requires obtaining patient consent, overcoming region-specific constraints, and managing the data lifecycle. This process is not only time-consuming and complex but also raises privacy and security concerns. Therefore, there is a need for a solution that can monitor the performance of an AI model without using PHI and other highly clinically sensitive data.
Summary of the Invention
[0004] In one embodiment, a method for monitoring the performance of a trained machine learning model is to extract characteristics of input data supplied to the trained machine learning model to generate input data characteristic information, where the input data includes medical images, generating, while mapping the medical images to an output, extracting characteristics of the performance of the trained machine learning model to generate model performance characteristic information, extracting characteristics of received user feedback based on the output of the trained machine learning model to generate user feedback characteristic information, determining an input data deviation by comparing the input data characteristic information with a plurality of previously determined input data characteristic information, determining a model performance deviation by comparing the model performance characteristic information with a plurality of previously determined model performance characteristic information, determining a user feedback deviation by comparing the user feedback characteristic information with a plurality of previously determined user feedback characteristic information, and sending a warning to a user device in response to one or more of the input data deviation exceeding an input data deviation threshold, the model performance deviation exceeding a model performance deviation threshold, and the user feedback deviation exceeding a user feedback deviation threshold.
[0005] In another embodiment, the system is a first device disposed at a deployment site, the first device including a user input device, a first non-transitory memory including a trained machine learning model and instructions, and a first processor, which, when executing the instructions, causes the first processor to extract characteristics of input data supplied to the trained machine learning model in the first device to generate input data characteristic information, the input data including medical images, generate mapping the medical images to an output, extract characteristics of the performance of the trained machine learning model during the mapping to generate model performance characteristic information, extract characteristics of user feedback received through the user input device based on the output of the trained machine learning model to generate user feedback characteristic information, and transmit the input data characteristic information, the model performance characteristic information, and the user feedback characteristic information to a second device. A first device including: and a second device disposed remotely from the first device, the first device and the second device being communicatively coupled, the second device including a second non-transitory memory including instructions and a second processor, which, when executing the instructions, causes the second processor to receive, in the second device, the input data characteristic information, the model performance characteristic information, and the user feedback characteristic information from the first device, determine an input data deviation by comparing the input data characteristic information with a plurality of previously determined input data characteristic information, determine a model performance deviation by comparing the model performance characteristic information with a plurality of previously determined model performance characteristic information, determine a user feedback deviation by comparing the user feedback characteristic information with a plurality of previously determined user feedback characteristic information, and transmit a warning to a user device in response to one or more of the input data deviation exceeding an input data deviation threshold, the model performance deviation exceeding a model performance deviation threshold, and the user feedback deviation exceeding a user feedback deviation threshold.
[0006] In yet another embodiment, a method of monitoring the performance of a trained machine learning model includes, in response to a medical image of an imaging subject being input into the trained machine learning model, extracting characteristics of the imaging subject's medical oxygen to generate input data characteristic information, where the input data characteristic information excludes the intensity values of individual pixels or voxels and the positions of individual pixels or voxels; generating the input data characteristic information; extracting performance characteristics of the trained machine learning model while mapping the medical image to an output to generate model performance characteristic information; determining an input data deviation by comparing the input data characteristic information with a plurality of previously determined input data characteristic information; determining a model performance deviation by comparing the model performance characteristic information with a plurality of previously determined model performance characteristic information; and sending a warning to a user device in response to one or more of the input data deviation exceeding an input data deviation threshold and the model performance deviation exceeding a model performance deviation threshold.
[0007] Thus, the present disclosure can monitor the performance of a deployed model without using PHI or other underlying clinically sensitive data by extracting characteristics of input data, model performance, and user feedback and using these characteristics as a substitute for the underlying clinical data. The extracted characteristics can be used to detect model performance deviations of the deployed model relative to baseline model performance (e.g., model performance regarding a test dataset, a validation dataset, and a training dataset), and thus, data scientists and developers can directly access the underlying clinical data and understand in real time regarding the performance of the deployed model.
[0008] It should be understood that the above summary is provided to introduce in a simplified form a part of the concepts further described in the mode for carrying out the invention. This is not intended to identify the important features or essential features of the claimed subject matter, and the scope of the claimed subject matter is defined solely by the claims. Furthermore, the claimed subject matter is not limited to implementations that solve the above disadvantages or any of the disadvantages mentioned elsewhere in this disclosure.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Figure 14
Figure 15
Mode for Carrying Out the Invention
[0010] The present disclosure describes a system and method for monitoring the performance of a trained machine learning model in the field of medical imaging. As the adoption of AI in the medical field progresses, it is necessary to monitor the real-world performance of AI models without compromising patient privacy or violating regulatory constraints. The present system and method can monitor performance by extracting the characteristics of input data, model output, and user feedback when used clinically. These characteristics serve as a substitute for the underlying clinical data and enable an understanding of the model's performance without the need to access protected medical information or clinically sensitive data. The extracted characteristics are compared to an expected baseline obtained from training data or past model inferences to detect deviations in the performance of the deployed model. When the deviation exceeds a predetermined threshold, a warning is generated and can be investigated and addressed in a timely manner. By using statistical surrogates instead of actual clinical data, the present disclosure enables continuous and easy monitoring of performance while maintaining regulatory compliance and protecting highly confidential personal information, ensuring patient safety and model reliability. The present system and method provide a solution for quality assurance of AI in a real-world medical environment.
[0011] In one example, the image processing device 102 (shown in FIG. 1) disposed at the deployment site can monitor the performance of one or more trained machine learning models by extracting characteristics from the input medical image data, model performance, and user feedback. The image processing device 102 transmits the extracted characteristics to the remote model monitoring device 122. The remote model monitoring device 122 encodes the characteristics as one or more feature vectors and compares the feature vectors with a previously determined characteristic information vector to identify deviations in the input data, model performance, or user feedback. The system shown in FIG. 1 performs one or more operations of FIG. 2 to extract characteristics, determine deviations by comparison with previously determined characteristic information, and send a warning when the deviation exceeds one or more predetermined thresholds. The medical image input can be characterized using method 300 (shown in FIG. 3) by extracting metadata, pixel statistics, and ontology tags. Further, 3D medical image input can be characterized using method 400 shown in FIG. 4, and the 3D image is characterized using metadata, intensity histogram, and clinical ontology tags. The input data characteristics extracted by method 300 and / or 400 can be encoded as a vector and compared with the previously determined input data characteristic information as shown by method 500 of FIG. 5.
[0012] Similarly, the model performance can be characterized according to one or more operations of method 600 shown in FIG. 6, which includes extracting model confidence scores, uncertainty indicators, and other model performance characteristics. In embodiments where the model output includes a segmentation mask, one or more operations of method 700 shown in FIG. 7 can be used to extract model performance characteristics (such as analyzing the characteristics of the segmentation mask and / or identified landmarks). FIG. 8 shows method 800, in which the model performance characteristics generated by method 600 and / or 700 are encoded into a feature vector and compared with a previously determined model performance characteristic information vector to identify model performance deviations.
[0013] User feedback can be characterized by performing one or more operations of method 900 shown in FIG. 9, which includes incorporating user feedback (such as model evaluation and corrections provided by the user). In embodiments where the model output includes a press prescription for the scan plane, user feedback can be characterized according to one or more operations of method 1000 shown in FIG. 10, which includes recording parameters of the scan plane adjusted by the user. The user feedback characteristics determined by method 900 and / or 1000 are encoded as a feature vector, compared with previously determined user feedback characteristic information, and a deviation can be detected by performing one or more operations of method 1100 shown in FIG. 11.
[0014] FIG. 12 shows a method 1200 for extracting input data characteristics from training data to establish an expected baseline model performance. FIG. 13 shows details of a method 1300 for extracting characteristics of model performance during training. FIG. 14 shows an overview of the technique of a method 1400 for extracting ground truth characteristics from training data. Finally, FIG. 15 shows a method 1500 for accumulating characteristic information of deployment data to update the performance baseline specific to the deployment site over time. In this way, using a statistical surrogate of the underlying clinical data, the performance and usage of a trained machine learning model at the deployment site can be monitored substantially in real time.
[0015] First, referring to FIG. 1, a model monitoring system 100 is shown. The model monitoring system 100 is configured to obtain a medical image of an object to be photographed using an imaging device 116, process the obtained medical image using an image processing device 102, and extract, for example, characteristics of the medical image, performance of a trained machine learning model 108, and user feedback received through a user input device 112. The processed information can be displayed to the user by a display device 114. User input can be received by the image processing device 102 through the user input device 112, and one or more of obtaining an image and processing the image can be adjusted based on the received user input. In some embodiments, the user input device 112 can be used to provide user feedback (such as corrections to the output of the trained machine learning model 108, evaluation of the output, and other types of feedback) to the image processing device 102.
[0016] The image processing device 102 includes a processor 104 configured to execute machine-readable instructions stored in a non-transitory memory 106. The processor 104 may be a single-core or a multi-core, and the program executed by the processor may be configured to enable parallel processing or distributed processing. In some embodiments, the processor 104 may optionally include individual components distributed across two or more devices, and these two or more devices may be remotely located and / or configured to enable cooperative processing. In some embodiments, one or more aspects of the processor 104 may be virtualized and executed by a remotely accessible network computing device configured in a cloud computing configuration.
[0017] The non-transitory memory 106 can store the trained machine learning model 108 and the characteristic information module 110. The trained machine learning model 108 can include instructions for mapping a medical image to an output. The characteristic information module 110 can include instructions for extracting the characteristics of the input data supplied to the trained machine learning model 108, the characteristics of the performance of the trained machine learning model 108 during mapping the medical image to the output, and the characteristics of the user feedback received through the user input device 112 based on the output of the trained machine learning model 108. The trained machine learning model 108 is stored in the non-transitory memory 106 and includes instructions that can map a medical image to an output. This output can be a diagnosis, a prediction, or other medical-related information obtainable from the medical image. The trained machine learning model 108 can use various machine learning algorithms (especially, deep learning, convolutional neural network, or support vector machine, etc.). The characteristic information module 110 is also stored in the non-transitory memory 106 and includes instructions for extracting various characteristics related to the input data, the performance of the trained machine learning model 108, and the user feedback. The characteristics of the input data can include metadata from the medical image, statistics of pixels or voxels from the medical image, and one or more tags from the appearance ontology and clinical ontology for the medical image. The metadata does not include personally identifiable information, and the statistics of pixels or voxels do not hold the intensity values or positions of individual pixels or voxels. The appearance ontology and the clinical ontology provide additional context (such as the type of disease state and the organs present in the image) regarding the medical image. The performance characteristics of the trained machine learning model 108 can include the output of the model 108, one or more intermediate outputs generated by one or more hidden layers of the model 108, the confidence score of the output of the model 108, and one or more uncertainty indicators of the output of the model 108. These performance characteristics provide information on how well the model 108 is functioning and can be used to identify potential problems or areas for improvement.The user feedback characteristics can include a model output evaluation received through the user input device 112 and a user correction received through the user input device 112, and the user correction modifies the output of the trained machine learning model 108. This feedback provides information on how the user interacts with the output of the model 108, and the feedback can be used to further refine and improve the model 108.
[0018] The image processing device 102 can transmit input data characteristic information, model performance characteristic information, and user feedback characteristic information to the model monitoring device 122. The model monitoring device 122 is remotely located from the image processing device 102, and the image processing device 102 and the model monitoring device 122 are communicatively coupled.
[0019] The model monitoring device 122 includes a processor 124 configured to execute machine-readable instructions stored in the non-transitory memory 126. The processor 124 can be a single-core or a multi-core, and the program executed by the processor can be configured to enable parallel processing or distributed processing. In some embodiments, the processor 124 can optionally include individual components distributed across two or more devices, and these two or more devices can be remotely located and / or configured to enable cooperative processing. In some embodiments, one or more aspects of the processor 124 are virtualized and executable by a remotely accessible network computing device configured in a cloud computing configuration.
[0020] The non-volatile memory 126 can store the deviation module 128, the characteristic information database 130, and the warning module 132. The deviation module 128 can include instructions for determining the input data deviation by comparing the input data characteristic information with a plurality of previously determined input data characteristic information, instructions for determining the model performance deviation by comparing the model performance characteristic information with a plurality of previously determined model performance characteristic information, and instructions for determining the user feedback deviation by comparing the user feedback characteristic information with a plurality of previously determined user feedback characteristic information.
[0021] The characteristic information database 130 can store a plurality of previously determined input data characteristic information, model performance characteristic information, and user feedback characteristic information. These characteristic information may be obtained from the training data set used to train the machine learning model, or may be obtained from the previous inferences of the trained machine learning model at the deployment site. The characteristic information database 130 can also store the distribution statistics related to the plurality of input data characteristic information, model performance characteristic information, and user feedback characteristic information. These distribution statistics can be used to dynamically adjust each deviation threshold based on the distribution of the previously determined characteristic information. The characteristic information database 130 can further store a plurality of predetermined feature vectors corresponding to the plurality of previously determined input data characteristic information, model performance characteristic information, and user feedback characteristic information. By calculating the distance between the current feature vector and each feature vector of the plurality of predetermined feature vectors in the multi-dimensional space using these predetermined feature vectors, it is possible to determine how similar the current characteristic information is to the previously determined characteristic information. The characteristic information database 130 can be continuously updated when new characteristic information is received from the deployment site, and the system can adapt the monitoring to changes in the input data, the performance of the model, or the user feedback.
[0022] The warning module 132 can include instructions for sending a warning to the user device 140 in response to one or more of an input data deviation exceeding an input data deviation threshold, a model performance deviation exceeding a model performance deviation threshold, and a user feedback deviation exceeding a user feedback deviation threshold.
[0023] The user input device 134 can include one or more of a touch screen, a keyboard, a mouse, a trackpad, a motion sensing camera, or other devices configured to enable a user to interact with and manipulate the data within the model monitoring device 122.
[0024] The display device 136 can include one or more display devices that utilize substantially any type of technology. The display device 136 can be configured to visually present information to a user. This information can include, for example, input data characteristic information, model performance characteristic information, and user feedback characteristic information. The display device 136 can also present a deviation of each characteristic information from a previously determined respective characteristic information. If one or more of the plurality of deviations exceed their respective thresholds, the display device 136 can present a warning generated by the warning module 132. This warning can indicate which deviation exceeded its threshold (whether the input data deviation exceeded the input data deviation threshold, the model performance deviation exceeded the model performance deviation threshold, or the user feedback deviation exceeded the user feedback deviation threshold, etc.). The display device 136 can present this warning in various forms (such as a pop-up notification, highlighted text, or an audible alarm, etc.) according to the user's preferences and the urgency of the situation. In addition to presenting the warning, the display device 136 can also display various other types of information related to the operation of the trained machine learning model. The display device 136 may be combined with the processor 124, the non-transitory memory 126, and / or the user input device 134 in a shared housing, or may be a peripheral display device, and can include a monitor, a touch screen, a projector, or other display devices known in the art, and the user can view the MRI generated by the MRI system and / or interact with various data stored in the non-transitory memory 126.
[0025] It should be understood that the image processing device 102 and the model monitoring device 122 shown in FIG. 1 are for illustrative purposes and not for limitation. Another suitable image processing system and model monitoring device may include more components, fewer components, or different components.
[0026] Referring to FIG. 2, a method 200 for monitoring the performance of a trained machine learning model is shown. This method 200 is executed by an image processing system, which can extract the characteristics of input data, model performance, and user feedback, and determine the deviation by comparing these characteristics with previously determined characteristics.
[0027] In step 202, the image processing system extracts the characteristics of the input data supplied to the trained machine learning model and generates input data characteristic information. By this extraction, the system can identify the characteristics important for monitoring in the input data. In one embodiment, when extracting input data characteristics, it includes using metadata from medical images such as subject-related characteristics (e.g., age range, gender, BMI), machine-related characteristics (e.g., machine model name), acquisition-related characteristics (e.g., magnetic field strength, coil, scan type, orientation, number of slices), and clinical data characteristics (e.g., average intensity, central intensity, intensity histogram). The metadata does not include personally identifiable information, and the pixel or voxel statistics do not hold the intensity values or positions of individual pixels or voxels. Appearance ontology and clinical ontology provide additional background information (such as the type of pathology and the organs present in the image) regarding medical images without exposing protected health information (PHI).
[0028] In step 204, while mapping the input data to the output, the image processing system extracts the characteristics of the performance of the trained machine learning model and generates model performance characteristic information. By this extraction, the system can monitor the performance of the machine learning model. In one embodiment, extracting model performance characteristics includes taking in the output of the trained machine learning model together with one or more intermediate outputs generated by one or more hidden layers of the trained machine learning model, determining the confidence score of the output of the trained machine learning model, and determining one or more uncertainty indicators of the output of the trained machine learning model. In another embodiment, in step 204, the image processing system can take in model performance characteristics (such as classification, segmentation, image regression, point regression, line regression, and plane regression metrics).
[0029] In operation 206, the image processing system extracts the characteristics of user feedback to generate user feedback characteristic information. By this extraction, the system can monitor the interaction between the output of the machine learning model and the user. In one embodiment, the extraction of user feedback characteristics includes recording the model output evaluation received through the user input device, recording the user correction received through the user input device (the user correction modifies the output of the trained machine learning model), and incorporating additional user comments or user actions indicating that the user approves or modifies the prediction of the model. Using this feedback, it can be evaluated whether the model meets the user's expectations and clinical workflow requirements.
[0030] In operation 208, the image processing system determines the input data deviation by comparing the input data characteristic information with a plurality of previously determined input data characteristic information. By this determination, the system can monitor the change of the input data over time (such as the change of demographic values of the subject, device settings, and collection parameters that may affect the performance of the model).
[0031] In operation 210, the image processing system determines the model performance deviation by comparing the model performance characteristic information with a plurality of previously determined model performance characteristic information. By this determination, the system can monitor the change of the performance of the machine learning model over time (including classification accuracy, segmentation quality, and deviation in regression prediction), thereby indicating the need for retraining or adjustment of the model.
[0032] In operation 212, the image processing system determines a user feedback deviation by comparing the user feedback characteristic information with a plurality of previously determined user feedback characteristic information. By this determination, the system can monitor the temporal change in the interaction between the output of the machine learning model and the user, and obtain information regarding whether the model is useful and acceptable in a clinical environment.
[0033] In step 214, the image processing system evaluates whether one or more of the input data deviation, the model performance deviation, and the user feedback deviation exceed their respective thresholds. If any of the deviations exceeds the threshold, the image processing system proceeds to step 218 and sends a warning indicating that the deviation has exceeded the threshold to the user device. The warning can include details of the deviation and suggestions for corrective measures. If none of the deviations exceed their respective thresholds, the image processing system continues to monitor the model inference in step 216. After step 218 or step 216, method 200 ends.
[0034] Accordingly, method 200 can at least partially address the technical problem of monitoring and maintaining the performance of a machine learning model in the field of medical images without directly accessing protected health information (PHI) while complying with strict privacy regulations. The technical improvements provided by method 200 include the ability to extract surrogate data characteristic information (which functions as a non-invasive surrogate for the underlying clinical data) from the input data, the model performance, and the user feedback. This approach can detect the model performance deviation in real time by comparing these characteristic information with a baseline established from the previously determined characteristic information. When the deviation exceeds a predetermined threshold, the system generates a warning to facilitate prompt corrective measures to ensure the reliability and accuracy of the model. Due to this technical effect, the machine learning model can surely continue to function optimally in a medical environment, provide reliable support to medical experts, and improve the treatment outcomes of patients.
[0035] Referring to FIG. 3, a method 300 for characterizing a medical image is shown. This method 300 is executed by an image processing system, can extract the characteristics of a medical image, and can use these characteristics to monitor the performance of a trained machine learning model.
[0036] In step 302, the image processing system extracts metadata from the medical image. The metadata can include information about the medical image that does not include personally identifiable information (PII) and protected health information (PHI). For example, the metadata can include information about the subject of the image (whether the subject is a child), the laterality protocol, the pixel spacing, the slice thickness, etc. In some embodiments, the metadata can include information about the subject of the image, such as an age range, gender, and body mass index (BMI) calculated from weight, height, and gender. Further, the metadata can include machine-related characteristics (such as the machine model name of the device that generates the medical image). This metadata provides a high-level overview of the medical image and can be used to characterize the input data supplied to the trained machine learning model.
[0037] In step 304, the image processing system aggregates the statistics of the pixels or voxels from the medical image. These statistics do not maintain the intensity values or positions of the individual pixels or voxels, but rather represent an overview of the overall intensity distribution within the image. For example, the image processing system can generate an intensity histogram, intensity profiles of anatomical planes (such as anteroposterior (AP), superior-inferior (SI), and left-right (LR)), and other aggregated statistics (such as average intensity, median intensity, minimum intensity, and maximum intensity). These statistics can provide information about the overall quality and characteristics of the medical image and are useful for monitoring the performance of the trained machine learning model.
[0038] In operation 306, the image processing system determines one or more tags from the appearance ontology of the medical image. The appearance ontology can include a set of predefined tags that describe various visual characteristics or visual features of the medical image. For example, the tags can describe the type of pathology present in the image, the organs visible in the image, and other visual features (such as the presence of artifacts like shadows or blurring). These tags provide a way to classify and describe the visual content of the medical image in a structured and standardized manner.
[0039] In operation 308, the image processing system determines one or more tags from the clinical ontology of the medical image. The clinical ontology can include a set of predefined tags that describe various clinical features or clinical characteristics of the medical image. For example, the tags can describe the type of medical imaging modality used to acquire the image (e.g., MR, CT, ultrasound), the type of clinical task associated with the image (e.g., classification, segmentation, image regression), and other clinical characteristics (such as the presence of a specific organ within the field of view (FoV)). These tags provide a way to classify and describe the clinical content of the medical image in a structured and standardized manner.
[0040] After operation 308, method 300 ends. The extracted metadata, pixel or voxel statistics, and ontology tags, taken together, form the input data characteristic information of the medical image. This input data characteristic information can be used to monitor the performance of a trained machine learning model by comparing it with previously determined input data characteristic information.
[0041] Referring to FIG. 4, a method 400 for characterizing a three-dimensional (3D) medical image is shown. This method 400 can be executed by an image processing system to extract the characteristics of the 3D medical image and use them to generate input data characteristic information. This characteristic information can be used, in particular, to monitor the performance of a trained machine learning model in the context of medical images.
[0042] In step 402, the image processing system extracts metadata from the 3D medical image. This metadata can include information about the image that does not directly represent the image content (such as pixel spacing and slice thickness). Pixel spacing represents the physical distance between the centers of each pixel of the patient and is measured in millimeters. The slice thickness represents the thickness of the 3D slice represented in the image, and the slice thickness is also measured in millimeters. This metadata provides important information for understanding the image, and the metadata can be used to normalize or standardize the image data so that further processing can be performed.
[0043] In step 404, the image processing system determines the intensity histogram of the 3D medical image. The intensity histogram graphically represents the distribution of the pixel intensities of the image. In a 3D medical image, the pixel intensity can represent different tissue types or structures, and the intensity histogram can provide an overview of the distribution of different tissues or structures in the image. The intensity histogram can be used to identify features or patterns of the image data related to the trained machine learning model.
[0044] In step 406, the image processing system determines one or more tags from the clinical ontology of the 3D medical image. Clinical ontology is a structured set of terms or tags that represent concepts in the clinical field. These tags can be used to annotate or describe the content of the 3D medical image in a standardized way. For example, the tags can indicate what organs or anatomical structures are visible in the image, what type of imaging modality was used to acquire the image, or what types of pathological structures are present in the image. These tags can provide a high-level overview of the image content that can be used to segment or classify the image for further processing.
[0045] After operation 406, method 400 ends. The extracted metadata, intensity histogram, and clinical ontology tags together form the input data characteristic information of the 3D medical image. This characteristic information can be used to monitor the performance of the trained machine learning model by comparing it with the previously determined input data characteristic information.
[0046] Referring to FIG. 5, method 500 is shown. Method 500 encodes metadata, pixel or voxel statistics, tags from an appearance ontology, and tags from a clinical ontology as a feature vector, and compares the feature vector with a plurality of predetermined feature vectors corresponding to a plurality of previously determined input data characteristic information. Method 500 is executed by an image processing system, determines a deviation in the input data characteristic information, and this deviation can be used to monitor the performance of the trained machine learning model.
[0047] Referring to FIG. 5, method 500 is shown. Method 500 encodes metadata, pixel or voxel statistics, tags from an appearance ontology, and tags from a clinical ontology as a feature vector, and compares the feature vector with a plurality of predetermined feature vectors corresponding to a plurality of previously determined input data characteristic information. Method 500 is executed by an image processing system, determines a deviation in the input data characteristic information, and this deviation can be used to monitor the performance of the trained machine learning model.
[0048] In Project 502, the image processing system encodes metadata, statistics of pixels or voxels, tags from an appearance ontology, and tags from a clinical ontology as a feature vector. This encoding enables the system to represent the input data characteristic information in a form that can be easily compared with other input data characteristic information. In one embodiment, the encoding includes converting metadata, statistics of pixels or voxels, and tags into a numerical format that can be represented as vectors in a multi-dimensional space. The dimensions of this space can correspond to different types of encoded characteristics such as metadata, statistics of pixels or voxels, and tags. Also, the encoding can include normalizing the values of the characteristics to ensure they are on a similar scale, thereby preventing a particular characteristic from influencing the comparison result due to scale. The encoding can be performed using various techniques (such as one-hot encoding of categorical data, scaling of numerical data, etc.). Further, the encoding can utilize specific medical-related metadata (such as DICOM metadata), and can include characteristics of pixels or voxels obtained from medical images (such as average intensity, histogram, or entropy, etc.) without including information that can identify a patient, thereby ensuring the anonymization of the data.
[0049] In operation 504, the image processing system compares the feature vector with a plurality of predetermined feature vectors corresponding to a plurality of previously determined input data characteristic information. By this comparison, the system can determine to what extent the current input data characteristic information is similar to the previously determined input data characteristic information. In one embodiment, comparing includes calculating the distance between the feature vector and each of the predetermined feature vectors in the multi-dimensional space. The distance can be calculated using various distance metrics (Euclidean distance, Manhattan distance, cosine similarity, etc.). Next, the system can determine the deviation by comparing the calculated distance with a threshold. If the distance exceeds the threshold, this indicates that the current input data characteristic information is significantly different from the previously determined input data characteristic information, indicating potential problems with the performance of the trained machine learning model. The system can send a warning to the user device in response to such a deviation, as explained in method 200 shown in FIG. 2. After operation 504, method 500 ends.
[0050] In operation 504, the image processing system compares the feature vector with a plurality of predetermined feature vectors corresponding to a plurality of previously determined input data characteristic information. Through this comparison, the system can determine to what extent the current input data characteristic information is similar to the previously determined input data characteristic information. In one embodiment, comparing includes calculating the distance between the feature vector and each of the predetermined feature vectors in the multi-dimensional space. The distance can be calculated using various distance metrics (such as Euclidean distance, Manhattan distance, cosine similarity, etc.). Next, the system can determine the deviation by comparing the calculated distance with a threshold value. If the distance exceeds the threshold value, this indicates that the current input data characteristic information is significantly different from the previously determined input data characteristic information, indicating potential problems in the performance of the trained machine learning model. The system can send a warning to the user device in response to such a deviation, as described in method 200 shown in FIG. 2. The warning can be generated by the warning module 132 and sent to the user device 140. The user device 140 can be part of a hospital or data scientist's monitoring system. After operation 504, method 500 ends.
[0051] Referring to FIG. 6, a method 600 for characterizing the output of a trained machine learning model in a medical imaging environment is shown. Method 600 is executed by an image processing system, captures the output of the trained machine learning model, determines a confidence score for the output, and determines one or more uncertainty metrics for the output. The confidence score and the uncertainty metrics can overall characterize the performance of the trained machine learning model when mapping the input data to the desired output.
[0052] In operation 602, the image processing system takes in the output of the trained machine learning model, along with one or more intermediate outputs generated by one or more hidden layers of the trained machine learning model. By this operation, the system can monitor the performance of the model at various stages of its operation. The output of the model can include, for example, the identification of anatomical landmarks in 3D medical images, the characteristics of the scan plane determined based on the anatomical landmarks, and a list of the anatomical landmarks identified in the 3D medical image by the trained machine learning model. The intermediate outputs can include, for example, the results of individual layers or groups of layers within the model, thereby providing information about the internal operation of the model and identifying parts that may need improvement or correction.
[0053] In operation 604, the image processing system determines a confidence score for the output of the trained machine learning model. The confidence score provides a quantitative measure of the certainty or reliability in the output of the model. The confidence score can be determined based on various factors, such as the internal parameters of the model, the quality of the input data, and the complexity of the task currently being worked on. For example, if the quality of the input data is high and the task is relatively simple, the confidence score will be high, and if the quality of the input data is low or the task is complex, the confidence score will be low. The confidence score can be used to evaluate the performance of the model or to determine whether to trust its output or seek additional information or confirmation.
[0054] In operation 606, the image processing system determines one or more uncertainty metrics of the output of the trained machine learning model. The uncertainty metric provides a quantitative measure of the uncertainty or variability in the output of the model. The uncertainty metric can be determined based on various factors (such as internal parameters of the model, quality of the input data, complexity of the task currently being worked on, etc.). For example, when the input data is of low quality and the task is complex, the uncertainty metric will be high, and when the input data is of high quality and the task is relatively simple, the uncertainty metric will be low. The uncertainty metric can be used to evaluate the performance of the model, or the uncertainty metric can be used to determine whether to trust its output or to request additional information or confirmation.
[0055] After operation 606, method 600 ends. By this method, the performance of the trained machine learning model in medical images can be continuously monitored and evaluated, which helps to ensure the accuracy and reliability of the model's output, and ultimately can improve the quality of patient care.
[0056] Referring to FIG. 7, method 700 is shown, and method 700 generates characteristic information of the model output by determining the characteristics of one or more segmentation masks generated by the trained machine learning model and identifying anatomical landmarks in the 3D medical image. Method 700 is executed by the image processing system and can extract the characteristics of the segmentation masks used to identify anatomical landmarks in the 3D medical image.
[0057] In operation 702, the image processing system extracts the characteristics of one or more segmentation masks generated by a trained machine learning model that identifies anatomical landmarks in 3D medical images. The segmentation masks are generated by a trained machine learning model and are used to identify and separate anatomical landmarks within the 3D medical images. The characteristics of the segmentation masks include, but are not limited to, the size of the segmentation mask, the center of gravity of the segmentation mask, and the intensity profile within the segmentation mask. These characteristics provide valuable information about the anatomical landmarks and can be used to further characterize the input data of the machine learning model.
[0058] In one embodiment, the size of the segmentation mask can be determined by counting the number of pixels or voxels within the mask, thereby obtaining an estimate of the volume or area of the anatomical landmark represented by the mask. The center of gravity of the segmentation mask is the geometric center of the mask and can be calculated by averaging the coordinates of all the pixels or voxels within the mask. The intensity profile within the segmentation mask can be determined by calculating statistical measures such as the average value, median value, standard deviation, and histogram of the intensities of the pixels or voxels within the mask. Additionally, the characteristics of the segmentation mask can include an uncertainty score that quantifies the reliability of the segmentation, and the type of the segmentation mask (a bounding box with dimensions, a polygon with vertices, a mask encoded in Base-64 format representing the segmented region).
[0059] In operation 704, the image processing system extracts characteristics of the scan plane determined based on anatomical landmarks. The scan plane is a two-dimensional slice determined based on the positions of anatomical landmarks in a 3D medical image. The characteristics of the scan plane include, but are not limited to, the orientation of the plane, the position of the plane within the 3D medical image, and the intensity of pixels or voxels within the plane. These characteristics can provide additional information regarding the anatomical landmarks and surrounding tissues or structures.
[0060] In one embodiment, the orientation of the scan plane can be represented by three direction cosines (cosines of the angles formed by the normal vector of the plane and the axes of the image coordinate system). The position of the scan plane within the three-dimensional medical image can be represented by the coordinates of the center point of the plane. The intensity of pixels or voxels within the plane can be characterized by calculating statistics (such as the mean, median, standard deviation, and histogram of the intensity). Further, in some embodiments, the characteristics of the scan plane can include the extent of the plane, which can be represented by dimensions in the x, y, and z directions, and the center, which can be represented by coordinates in the image coordinate system.
[0061] In operation 706, the image processing system records a list of identified anatomical landmarks of the 3D medical image by a trained machine learning model. This list can include the name or label of the landmark, the position within the 3D medical image, and a related confidence score indicating the certainty of each landmark identified by the model. This list can be used to further characterize the input data of the machine learning model and to monitor the performance of the model. After operation 706, method 700 ends.
[0062] Referring to FIG. 8, a method 800 for determining a deviation of model performance characteristics from previously determined multiple model performance characteristic information is shown. Method 800 is executed by a model monitoring system and can determine a deviation in the performance of a machine learning model based on the characteristics of the model output.
[0063] In step 802, the model monitoring system encodes the output of the trained machine learning model as a feature vector, together with one or more intermediate outputs, a confidence score, and one or more uncertainty indicators. This encoding enables the system to obtain a comprehensive representation of the model's performance. The output of the trained machine learning model includes the final prediction or decision obtained by the model (such as the identification of anatomical landmarks in a medical image). The intermediate outputs include the outputs of one or more hidden layers of the model, which can provide information about the internal operation of the model. The confidence score can represent the certainty of the model's output, and the uncertainty indicator can quantify the uncertainty or variability of the model's output.
[0064] In step 802, the model monitoring system encodes the output of the trained machine learning model as a feature vector, together with one or more intermediate outputs, a confidence score, and one or more uncertainty indicators. This encoding enables the system to obtain a comprehensive representation of the model's performance. The output of the trained machine learning model includes the final prediction or decision obtained by the model (such as the identification of anatomical landmarks in a medical image, the classification of a medical image into a predetermined category, the segmentation of a specific structure within a medical image, or a regression analysis for predicting a continuous variable from a medical image). The intermediate outputs include the outputs of one or more hidden layers of the model, which can provide information about the internal operation of the model. The confidence score can represent the certainty of the model's output, and the uncertainty indicator can quantify the uncertainty or variability of the model's output and can include uncertainty scores related to predicted masks, points, lines, or planes within a medical image.
[0065] In operation 804, the model monitoring system compares the feature vector with a plurality of predetermined feature vectors corresponding to a plurality of previously determined model performance characteristic information. Through this comparison, the system can determine how much the current performance of the model deviates from the past performance or the expected performance of the model. The predetermined feature vectors may be derived from the training dataset used to train the machine learning model, or may be derived from previous inferences of the model at the deployment site. The comparison can include calculating a measure of distance or similarity (such as Euclidean distance, cosine similarity, or other appropriate distance or similarity measure) between the feature vector and each of the predetermined feature vectors. As a result of the comparison, if it is found that the model performance deviation exceeds the threshold of the model performance deviation, the model monitoring system can respond and send a warning to the user device. The warning indicates that the performance of the model is significantly deviated from the expected performance of the model, and can suggest actions that can be taken to address the deviation (such as retraining the model, adjusting the parameters of the model, or investigating the cause of the deviation). After operation 804, method 800 ends.
[0066] Referring to FIG. 9, a method 900 for characterizing user feedback received based on the output from a trained machine learning model is shown. Method 900 is executed by an image processing system and can extract the characteristics of user feedback received based on the output of the trained machine learning model and generate user feedback characteristic information.
[0067] In operation 902, the image processing system records a model output evaluation received through a user input device. This model output evaluation can be a numerical or categorical value representing the user's evaluation of the quality or accuracy of the output of a trained machine learning model. The model output evaluation is based on various factors, such as the perceived accuracy of the output, the usefulness of the output in the user's workflow, the user's confidence in the output, and so on. The model output evaluation can be input by the user through a user input device (such as a keyboard, mouse, touch screen, voice command interface, or other suitable input device).
[0068] In operation 904, the image processing system records a user correction received through a user input device, and the user correction modifies the output of the trained machine learning model. The user correction is an adjustment or modification made by the user to the output of the trained machine learning model. Examples of user corrections can include adjusting the position or orientation of a segmentation mask, adding or removing landmarks, adjusting the parameters of a recommended scan plane, and so on. The user correction can be input by the user through a user input device and can be recorded by the image processing system as part of the user feedback characteristic information.
[0069] The user feedback characteristic information generated by method 900 provides valuable information about how a trained machine learning model is operating in the real-world environment of the user's workflow. By analyzing the user feedback characteristic information, the image processing system can identify parts of the trained machine learning model that need improvement and use this information to update or retrain the model, thereby improving the performance and usefulness of the model in a healthcare environment. Following operation 904, method 900 ends.
[0070] Referring to FIG. 10, a method 1000 for characterizing user feedback received based on the output of a trained machine learning model is shown. Method 1000 may be executed by an image processing system and adopted by the image processing system to determine user feedback characteristic information, which can be used when monitoring the performance of AI models and applications in the field of medical imaging.
[0071] In step 1002, the image processing system determines the central coordinates of the scan plane used to perform a diagnostic scan of the imaging object. By this determination, the system can identify the final scan plane approved / adjusted by the user used to perform the diagnostic scan.
[0072] In step 1004, the image processing system determines three direction cosines that uniquely identify the orientation of the scan plane. After step 1004, method 1000 ends.
[0073] The user feedback characteristic information includes the central coordinates of the scan plane and three direction cosines that uniquely identify the orientation of the scan plane, and provides information about the interaction between the AI model and the user. Using this information, the performance of the AI model can be monitored and deviations from the expected behavior can be identified. If such a deviation is detected, a warning can be sent to the user device and intervention and correction can be made in a timely manner.
[0074] Referring to FIG. 11, a method 1100 for determining a user feedback deviation using a plurality of pre-determined user feedback characteristic information is shown. Method 1100 may be executed by an image processing system and can compare the user feedback characteristic information with a plurality of previously determined user feedback characteristic information.
[0075] In operation 1102, the image processing system encodes model output evaluation and user correction as a feature vector. The model output evaluation and user correction are received through a user input device and are part of the user feedback characteristic information. The model output evaluation can be a numerical or categorical value indicating the user's evaluation of the quality or accuracy of the output of the trained machine learning model. User correction is the correction made by the user to the output of the trained machine learning model (such as adjustment of the position or orientation of the scanned surface, change of the identified landmark or segmentation mask, etc.). By encoding the model output evaluation and user correction as a feature vector, the user feedback characteristic information can be represented compactly and efficiently, and this user feedback characteristic information can be easily compared with other feature vectors representing previously determined user feedback characteristic information.
[0076] In operation 1104, the image processing system compares the feature vector with a plurality of predetermined feature vectors corresponding to a plurality of previously determined user feedback characteristic information. Through this comparison, the image processing system can determine a user feedback deviation, which is a measure of the difference between the current user feedback characteristic information and the previously determined user feedback characteristic information. The user feedback deviation may be calculated as the distance in the characteristic space between the feature vector and the closest feature vector among the predetermined feature vectors, or may be calculated as a statistical measure of the difference between the distribution of the feature vector and the predetermined feature vectors. The user feedback deviation provides an indication of how much the current user feedback deviates from the typical or expected user feedback, and by using the user feedback deviation, the performance of the trained machine learning model can be monitored and anomalies or problems in the operation of the model can be detected.
[0077] In some embodiments, the image processing system can send a warning to the user device in response to the user feedback deviation exceeding a user feedback deviation threshold. The user feedback deviation threshold may be a predetermined value or a value dynamically adjusted based on the distribution of the previously determined characteristic evaluation of the user feedback. The warning can include information about the user feedback deviation and suggestions for corrective actions (such as retraining the machine learning model, adjusting the parameters of the machine learning model, or providing additional guidance to the user). After operation 1104, method 1100 ends.
[0078] Referring to FIG. 12, a method 1200 for characterizing input data from a training dataset used to train a machine learning model is shown. Method 1200 can be executed by an image processing system and can extract input data characteristics from the input data of each pair of a plurality of training data pairs to generate a plurality of input data characteristic information, and the input data characteristic information can then be used as a comparison baseline for determining the deviation of the input data characteristic information for the input data received at the deployment site.
[0079] In step 1202, the image processing system obtains a training dataset used to train a machine learning model. The training dataset includes a plurality of training data pairs, and each training data pair includes input data and corresponding ground truth data. The input data can include medical images (such as MR images, CT images, ultrasound images, etc.), and the ground truth data can include annotations or labels indicating the presence or absence of specific features or medical conditions in the images.
[0080] In operation 1204, the image processing system extracts input data characteristics from the input data of each pair of a plurality of training data pairs to generate a plurality of input data characteristic information. This extraction process can include one or more of the plurality of operations described in detail above with respect to methods 300 and 400. The characteristics can include metadata, statistics of pixels or voxels, and tags from appearance ontology and clinical ontology. The metadata can include information regarding the imaging target (whether the subject is a child), laterality protocol, pixel spacing, slice thickness, and the like. The statistics of pixels or voxels can include intensity histograms, intensity profiles, noise, and other statistical measures derived from pixel values or voxel values within the image. The tags from appearance ontology and clinical ontology may include information regarding the type of medical condition, organs present in the image, and other high-level features that cannot be easily derived from the pixel data alone.
[0081] In operation 1206, the image processing system encodes the plurality of input data characteristic information to generate a plurality of input data characteristic information vectors. This encoding process includes converting the extracted characteristics into a numerical format that can be easily compared and analyzed. For example, the metadata, statistics of pixels or voxels, and ontology tags can be encoded as feature vectors, and each element of the vector corresponds to a specific characteristic.
[0082] In operation 1208, the image processing system determines the distribution statistics of the plurality of input data characteristic information vectors. These statistics can include measures of central tendency (such as the mean or median), measures of dispersion (such as variance or standard deviation), and other statistical measures that describe the distribution of the input data characteristic information vectors. These statistics provide an overview of the characteristics of the input data of the training data set, and using these statistics, the characteristics of the input data of new data can be compared with the characteristics of the training data.
[0083] In operation 1210, the image processing system stores a plurality of input data characteristic information vectors and distribution statistics in a non - transitory memory. This stored information is later used to monitor the performance of the trained machine learning model and detect deviations or drift amounts in the input data, enabling comparison of the characteristics of new input data with those of the training data. After operation 1210, method 1200 ends.
[0084] Referring to FIG. 13, a method 1300 for extracting model performance characteristics during training of a machine learning model is shown. Method 1300 can be executed by an image processing system to extract the model performance characteristics output for each of a plurality of training data pairs and generate a plurality of model performance characteristic information.
[0085] In operation 1302, the image processing system obtains a training data set used to train a machine learning model. The training data set includes a plurality of training data pairs, and each training data pair includes input data and corresponding ground - truth data. The input data can include medical images (such as MR scans or CT scans, ultrasound images, or other types of medical image data). The ground - truth data can include annotations or labels indicating the presence or absence of specific characteristics or medical conditions in the input data (such as the presence of a specific medical condition, the location of specific anatomical landmarks, or the correct classification of the input data).
[0086] During operation 1304, while training a machine learning model on a training dataset, the image processing system extracts model performance characteristics output for each of a plurality of training data pairs. This extraction process includes taking in the output of the trained machine learning model along with one or more intermediate outputs generated by one or more hidden layers of the trained machine learning model. The model performance characteristics can include, for example, a confidence score of the output of the trained machine learning model and one or more uncertainty metrics of the output of the trained machine learning model. The confidence score can indicate the degree of certainty of a prediction made by the model, and the uncertainty metric can provide a measure of the uncertainty or variability of the model in its predictions.
[0087] During operation 1306, the image processing system encodes a plurality of model performance characteristic information to generate a plurality of model performance characteristic information vectors. This encoding process includes converting the model performance characteristics into a form that can be easily compared and analyzed. For example, the model performance characteristics can be encoded as a feature vector, where each element of the vector corresponds to a different model performance characteristic.
[0088] During operation 1308, the image processing system determines distribution statistics of the plurality of model performance characteristic information vectors. These distribution statistics can include measures of central tendency (such as mean or median) and measures of dispersion (such as variance or standard deviation). The distribution statistics provide an overview of the overall performance of the machine learning model during the training phase.
[0089] During operation 1310, the image processing system stores the plurality of model performance characteristic information vectors and the distribution statistics in non-transitory memory. This stored information is used later to monitor the performance of the machine learning model by comparing the performance of the model on new data to the performance characteristics observed during training. After operation 1310, method 1300 ends.
[0090] Referring to FIG. 14, a method 1400 for extracting ground truth characteristics from a training data set used to train a machine learning model is shown. Method 1400 is executed by an image processing system and can extract ground truth characteristics from the ground truth data of each training data pair of a plurality of training data pairs to generate a plurality of ground truth data characteristic information.
[0091] In step 1402, the image processing system obtains a training data set used to train a machine learning model. The training data set includes a plurality of training data pairs, and each training data pair includes input data and corresponding ground truth data. The ground truth data can include information such as the correct classification or output for the corresponding input data. This ground truth data functions as reference data for the machine learning model during the training process, and the model can learn the correct mapping from the input data to the desired output.
[0092] In step 1404, the image processing system extracts ground truth characteristics from the ground truth data of each training data pair of the plurality of training data pairs to generate a plurality of ground truth data characteristic information. This extraction process can include various techniques depending on the nature of the ground truth data. For example, if the ground truth data includes labels for image classification, the ground truth characteristics can include the distribution of the labels in the data set. If the ground truth data includes bounding boxes for object detection, the ground truth characteristics can include the size, shape, and position of the bounding boxes. If the ground truth data includes segmentation masks for image segmentation, the ground truth characteristics can include the area, perimeter, and centroid of the segmentation masks.
[0093] In operation 1406, the image processing system encodes a plurality of ground truth feature information to generate a plurality of ground truth data feature information vectors. This encoding process includes converting the ground truth features into a numerical format that can be easily processed. For example, the ground truth features can be normalized, scaled, or otherwise transformed to generate the ground truth data feature information vectors.
[0094] In operation 1408, the image processing system determines distribution statistics of the plurality of ground truth data feature information vectors. These distribution statistics can include measures such as the mean, median, mode, variance, standard deviation, skewness, kurtosis, and other statistical measures of the ground truth data feature information vectors. These distribution statistics provide an overview of the ground truth data feature information vectors and enable the image processing system to understand the overall characteristics of the ground truth data in the training dataset.
[0095] In operation 1410, the image processing system stores the plurality of ground truth data feature information vectors and the distribution statistics in non - volatile memory. This stored information can be used later for various purposes, such as monitoring the performance of a trained machine learning model, detecting drift in the input data, or adjusting the training process of the machine learning model. After operation 1410, method 1400 ends.
[0096] Referring to FIG. 15, a method 1500 for accumulating input data feature information, model performance feature information, and user feedback feature information from a specific deployment site is shown. Method 1500 is executed by a model monitoring system and can establish baselines for various model performance metrics at a specific deployment site.
[0097] In Project 1502, the model monitoring system receives input data characteristic information, model performance characteristic information, and user feedback characteristic information from the deployment site. The input data characteristic information can include characteristics of the input data supplied to the trained machine learning model (such as metadata from medical images, statistics of pixels or voxels from medical images, and one or more tags from the appearance ontology and clinical ontology of medical images). The model performance characteristic information can include characteristics of the performance of the trained machine learning model while mapping medical images to the output (such as the output of the trained machine learning model, one or more intermediate outputs generated by one or more hidden layers of the trained machine learning model, the confidence score of the output of the trained machine learning model, and one or more uncertainty indicators of the output of the trained machine learning model). The user feedback characteristic information can include characteristics of the user feedback received based on the output of the trained machine learning model (such as model output evaluation and user corrections that modify the output of the trained machine learning model).
[0098] In Project 1504, the model monitoring system determines the deviation of each characteristic information based on a plurality of previously determined characteristic information from the deployment site. This determination includes identifying the deviation by comparing the received characteristic information with the previously determined characteristic information. The previously determined characteristic information may be derived from the training dataset used to train the trained machine learning model or from previous inferences of the trained machine learning model at the deployment site.
[0099] In Project 1506, the model monitoring system checks whether any of the determined deviations exceed their respective thresholds. If none of the deviations exceed their respective thresholds, the model monitoring system proceeds to Project 1510, where the received characteristic information is added to a plurality of previously determined characteristic information from the deployment site. Thereby, the model monitoring system continuously updates the "normal" operation of the trained machine learning model that the model monitoring system understands, and the model monitoring system can adapt its monitoring to changes in input data, model performance, or user feedback.
[0100] In Project 1506, if one or more deviations exceed their respective thresholds, the model monitoring system proceeds to Project 1508, where a warning indicating that the deviation threshold has been exceeded is sent to the user device. The warning can include information about the nature of the deviation, the magnitude of the deviation, and possible corrective actions. The warning can be sent via various communication channels (such as email, text message, push notification, or other appropriate communication methods). After Project 1510 or 1508, Method 1500 ends.
[0101] By the above-disclosed system and method for monitoring the performance of a trained machine learning model in a healthcare environment, the computational efficiency when monitoring the performance of a deployed model in real time can be improved. By extracting surrogate data characteristic evaluations from input data, model performance, and user feedback, the system eliminates the need to process and analyze the entire volume of highly confidential medical images and associated PHI. In this approach, the characteristic evaluations are lightweight compared to the original data and can be processed quickly so that performance deviations can be detected, reducing the computational load on the system. The ability of the system to encode these characteristic evaluations as feature vectors further enhances computational efficiency and allows for quick comparison with a baseline. This streamlined process not only saves computational resources but also enables real-time monitoring and rapid response to deviations, ensuring that the machine learning model operates reliably within a pre-determined range of performance parameters, e.g., within a pre-determined range of deviation from the performance achieved during training of the model. The computational efficiency of the above-disclosed system and method is particularly advantageous in a healthcare environment where timely and accurate model performance is of utmost importance and computational resources may be limited.
[0102] The present disclosure also provides support for a method. The method includes extracting characteristics of input data supplied to a machine learning model to generate input data characteristic information, where the input data includes a medical image, generating the input data characteristic information, extracting characteristics of the performance of the trained machine learning model while mapping the medical image to an output to generate model performance characteristic information, extracting characteristics of received user feedback based on the output of the trained machine learning model to generate user feedback characteristic information, determining an input data deviation by comparing the input data characteristic information with a plurality of previously determined input data characteristic information, determining a model performance deviation by comparing the model performance characteristic information with a plurality of previously determined model performance characteristic information, determining a user feedback deviation by comparing the user feedback characteristic information with a plurality of previously determined user feedback characteristic information, and transmitting a warning to a user device in response to one or more of the input data deviation exceeding an input data deviation threshold, the model performance deviation exceeding a model performance deviation threshold, and the user feedback deviation exceeding a user feedback deviation threshold. In a first embodiment of the method, extracting characteristics of the input data supplied to the trained machine learning model to generate the input data characteristic information includes extracting metadata from the medical image, where the metadata does not include personally identifiable information, extracting the metadata, aggregating statistics of pixels or voxels from the medical image, where the statistics of the pixels or voxels do not hold the intensity values or positions of individual pixels or voxels, aggregating the statistics, determining one or more tags from an appearance ontology of the medical image, and determining one or more tags from a clinical ontology of the medical image.In a second embodiment of the method, optionally including the first embodiment, determining an input data deviation by comparing the input data characteristic information with a plurality of previously determined input data characteristic information includes encoding the metadata, the statistics of the pixels or voxels, one or more tags from the appearance ontology, and one or more tags from the clinical ontology as a feature vector, and comparing the feature vector with a plurality of predetermined feature vectors corresponding to the plurality of previously determined input data characteristic information. In a third embodiment of the method, optionally including one or both of the first and second embodiments, extracting characteristics of the performance of the trained machine learning model to generate model performance characteristic information while mapping a medical image to an output includes capturing the output of the trained machine learning model together with one or more intermediate outputs generated by one or more hidden layers of the trained machine learning model, determining a confidence score for the output of the trained machine learning model, and determining one or more uncertainty indicators for the output of the trained machine learning model. In a fourth embodiment of the method, optionally including one or more of the first to third embodiments or each embodiment, determining a model performance deviation by comparing the model performance characteristic information with a plurality of previously determined model performance characteristic information includes encoding the output of the trained machine learning model as a feature vector together with the one or more intermediate outputs, the confidence score, and the one or more uncertainty indicators, and comparing the feature vector with a plurality of predetermined feature vectors corresponding to the plurality of previously determined model performance characteristic information. In a fifth embodiment of the method, optionally including one or more of the first to fourth embodiments or each embodiment, extracting characteristics of received user feedback to generate user feedback characteristic information based on the output of the trained machine learning model includes a model output evaluation received through a user input device and a user correction received through the user input device, where the user correction modifies the output of the trained machine learning model. including recording one or more of them. In a sixth embodiment of the method, optionally including one or more or each of the first to fifth embodiments, based on the output of the trained machine learning model, extracting the characteristics of the received user feedback to generate user feedback characteristic information includes receiving comments from the user through the user input device and determining the sentiment score of the output of the trained machine learning model based on the comments. In a seventh embodiment of the method, optionally including one or more or each of the first to sixth embodiments, determining a user feedback deviation by comparing the user feedback characteristic information with a plurality of previously determined user feedback characteristic information includes encoding the model output evaluation and the user correction as feature vectors and comparing the feature vectors with a plurality of predetermined feature vectors corresponding to a plurality of previously determined user feedback characteristic information.
[0103] The present disclosure also provides support for a system. The system includes a first device disposed at a deployment site, the first device including a user input device, a first non-transitory memory including a trained machine learning model and instructions, and a first processor. When the instructions are executed, the first processor causes the first device to extract characteristics of input data supplied to the trained machine learning model to generate input data characteristic information, where the input data includes medical images; extract characteristics of the performance of the trained machine learning model to generate model performance characteristic information while mapping the medical images to an output; extract characteristics of user feedback received through the user input device based on the output of the trained machine learning model to generate user feedback characteristic information; and transmit the input data characteristic information, the model performance characteristic information, and the user feedback characteristic information to a second device. The system further includes a second device disposed remotely from the first device, where the first device and the second device are communicatively coupled. The second device includes a second non-transitory memory including instructions and a second processor. When the instructions are executed, the second processor causes the second device to receive the input data characteristic information, the model performance characteristic information, and the user feedback characteristic information from the first device; determine an input data deviation by comparing the input data characteristic information with a plurality of previously determined input data characteristic information; determine a model performance deviation by comparing the model performance characteristic information with a plurality of previously determined model performance characteristic information; determine a user feedback deviation by comparing the user feedback characteristic information with a plurality of previously determined user feedback characteristic information; and transmit a warning to a user device in response to one or more of the input data deviation exceeding an input data deviation threshold, the model performance deviation exceeding a model performance deviation threshold, and the user feedback deviation exceeding a user feedback deviation threshold. and a second device including the same. In a first embodiment of the system, characteristic information of a plurality of previously determined input data, model performance characteristic information of a plurality of previously determined models, and user feedback characteristic information of a plurality of previously determined models are derived from a training data set used to train the trained machine learning model. In a second embodiment of the system, optionally including the first embodiment, characteristic information of a plurality of previously determined input data, model performance characteristic information of a plurality of previously determined models, and user feedback characteristic information of a plurality of previously determined models are derived from previous inferences of the trained machine learning model at the deployment site. In a third embodiment of the system, optionally including one or both of the first and second embodiments, each of the input data characteristic information, the model performance characteristic information, and the user feedback characteristic information does not include a medical image, and when the instruction is executed, the first processor does not transmit the medical image to the second device. In a fourth embodiment of the system, optionally including one or more or each of the first to third embodiments, the warning includes indicating one or more of that the input data deviation exceeds the input data deviation threshold, that the model performance deviation exceeds the model performance deviation threshold, and that the user feedback deviation exceeds the user feedback deviation threshold.
[0104] The present disclosure also provides support for a method. The method is a method for monitoring the performance of a trained machine learning model, which, in response to a medical image of an imaging object being input into the trained machine learning model, extracts the characteristics of the imaging object's oxygen and generates input data characteristic information, wherein the input data characteristic information excludes the intensity values of individual pixels or voxels and the positions of individual pixels or voxels, generating, while mapping the medical image to an output, extracting the performance characteristics of the trained machine learning model and generating model performance characteristic information, determining an input data deviation by comparing the input data characteristic information with a plurality of previously determined input data characteristic information, determining a model performance deviation by comparing the model performance characteristic information with a plurality of previously determined model performance characteristic information, and transmitting a warning to a user device in response to one or more of the input data deviation exceeding an input data deviation threshold and the model performance deviation exceeding a model performance deviation threshold. In a first embodiment of the method, the method further includes extracting the characteristics of received user feedback and generating user feedback characteristic information based on the output of the trained machine learning model, determining a user feedback deviation by comparing the user feedback characteristic information with a plurality of previously determined user feedback characteristic information, and transmitting a warning to the user device in response to the user feedback deviation exceeding a user feedback deviation threshold. In a second embodiment of the method, optionally including the first embodiment, the user feedback includes a scan plane used to obtain a diagnostic medical image of the imaging object. In a third embodiment of the method, optionally including one or both of the first and second embodiments, extracting the characteristics of received user feedback and generating user feedback characteristic information based on the output of the trained machine learning model includes determining the central coordinates of the scan plane and determining three direction cosines that uniquely identify the orientation of the scan plane.In a fourth embodiment of the present method, optionally one or more of the first to third embodiments or each embodiment is included, the medical image includes a three-dimensional (3D) medical image, and the trained machine learning model is configured to identify anatomical landmarks of the 3D medical image for positioning a scan plane for obtaining a diagnostic medical image of a region of interest. In a fifth embodiment of the present method, optionally one or more of the first to fourth embodiments or each embodiment is included, and extracting characteristics of the medical image of the imaging subject to generate the input data characteristic information includes capturing metadata of the 3D medical image, including the pixel spacing and slice thickness of the 3D medical image, determining aggregated intensity statistics of the 3D medical image, including the intensity histogram, and determining a clinical ontology of the 3D medical image, including a list of anatomical regions incorporated into the 3D medical image. In a sixth embodiment of the present method, optionally one or more of the first to fifth embodiments or each embodiment is included, and extracting characteristics of the performance of the trained machine learning model to generate model performance characteristic information while mapping the medical image to the output includes extracting characteristics of one or more segmentation masks generated by the trained machine learning model that identifies anatomical landmarks in the 3D medical image, extracting characteristics of the scan plane determined based on the anatomical landmarks, and recording a list of the anatomical landmarks identified in the 3D medical image by the trained machine learning model.
[0105] When introducing elements of various embodiments of the present disclosure, the articles "a," "an," and "the" are intended to mean that there is one or more of such elements. Terms such as "first," "second," etc. do not indicate order, quantity, or importance, but are used to distinguish one element from another. "Comprising," "including," and "having" are intended to be inclusive and mean that additional elements other than the recited elements may exist. In this specification, when terms such as "connected to," "coupled to," etc. are used, one object (e.g., a material, element, structure, member, etc.) can be connected or coupled to another object, whether one object is directly connected or coupled to the other object, or whether there is one or more intervening objects between one object and the other object. Additionally, it should be understood that references to "one embodiment" or "an embodiment" of the present disclosure are not intended to be construed as excluding the existence of additional embodiments that also incorporate the recited features.
[0106] In addition to the modified forms shown above, many other variant structures and alternative structures can be conceived by those skilled in the art without departing from the spirit and scope of this description, and the claims are intended to cover such modified forms and structures. Therefore, although the above information is described in particular detail with respect to what is currently considered to be the most practical and preferred embodiments, it will be apparent to those skilled in the art that many modifications are possible in terms of form, function, method of operation, and use (without limitation), etc., without departing from the principles and concepts described herein. Also, in this specification, examples and embodiments are meant to be illustrative in every respect and should not be construed in any way as limiting.
Description of Reference Numerals
[0107] 100 Model monitoring system 102 Image processing device 104 Processor 106 Temporary Memory 108 Trained Machine Learning Model 110 Feature Information Module 112 User Input Device 114 Display Device 116 Imaging Device 122 Model Monitoring Device 124 Processor 126 Temporary Memory 128 Deviation Module 130 Feature Information Database 132 Warning Module 134 User Input Device 136 Display Device 140 User Device 200 Method
Claims
1. Extracting characteristics of input data supplied to a trained machine learning model to generate input data characteristic information, wherein the input data includes medical images, and generating; Extracting characteristics of the performance of the trained machine learning model while mapping the medical images to the output to generate model performance characteristic information; Extracting characteristics of received user feedback based on the output of the trained machine learning model to generate user feedback characteristic information; Determining an input data deviation by comparing the input data characteristic information with a plurality of previously determined input data characteristic information; Determining a model performance deviation by comparing the model performance characteristic information with a plurality of previously determined model performance characteristic information; Determining a user feedback deviation by comparing the user feedback characteristic information with a plurality of previously determined user feedback characteristic information, and Sending a warning to a user device in response to one or more of the input data deviation exceeding an input data deviation threshold, the model performance deviation exceeding a model performance deviation threshold, and the user feedback deviation exceeding a user feedback deviation threshold A method comprising.
2. Extracting characteristics of input data supplied to the trained machine learning model to generate the input data characteristic information includes: Extracting metadata from the medical images, wherein the metadata does not include personally identifiable information, and extracting; Aggregating statistics of pixels or voxels from the medical images, wherein the statistics of the pixels or voxels do not hold the intensity values or positions of individual pixels or voxels, and aggregating the statistics; Determining one or more tags from the appearance ontology of the medical images, and Determining one or more tags from the clinical ontology of the medical images The method according to claim 1, comprising.
3. Determining an input data deviation by comparing the input data characteristic information with a plurality of previously determined input data characteristic information includes: Encoding the metadata, the statistics of the pixels or voxels, the one or more tags from the appearance ontology, and the one or more tags from the clinical ontology as a feature vector, and Comparing the feature vector with a plurality of predetermined feature vectors corresponding to a plurality of input data characteristic information determined previously The method according to claim 2, comprising:
4. During mapping a medical image to an output, extracting characteristics of the performance of the trained machine learning model to generate model performance characteristic information, comprises taking in the output of the trained machine learning model together with one or more intermediate outputs generated by one or more hidden layers of the trained machine learning model, determining a confidence score of the output of the trained machine learning model, and determining one or more uncertainty indicators of the output of the trained machine learning model The method according to claim 1, comprising:
5. Determining a model performance deviation by comparing the model performance characteristic information with a plurality of previously determined model performance characteristic information, encoding the output of the trained machine learning model as a feature vector together with the one or more intermediate outputs, the confidence score, and the one or more uncertainty indicators, and comparing the feature vector with a plurality of predetermined feature vectors corresponding to a plurality of previously determined model performance characteristic information The method according to claim 4, comprising:
6. Extracting characteristics of received user feedback based on the output of the trained machine learning model to generate user feedback characteristic information, recording one or more of a model output evaluation received through a user input device, and a user correction received through the user input device, wherein the user correction corrects the output of the trained machine learning model The method according to claim 1, comprising:
7. Extracting characteristics of received user feedback based on the output of the trained machine learning model to generate user feedback characteristic information, comprises receiving a comment from a user through the user input device, and determining a sentiment score of the output of the trained machine learning model based on the comment The method according to claim 6, comprising:
8. Determining a user feedback deviation by comparing the user feedback characteristic information with a plurality of previously determined user feedback characteristic information, encoding the model output evaluation and the user correction as a feature vector, and comparing the feature vector with a plurality of predetermined feature vectors corresponding to a plurality of previously determined user feedback characteristic information The method according to claim 6, comprising:
9. A first device disposed at a deployment site, the first device comprising: a user input device; a first non-transitory memory including a trained machine learning model and instructions; and a first processor, which, when executing the instructions, causes the first processor to cause the first device to: extract characteristics of input data supplied to the trained machine learning model to generate input data characteristic information, the input data including a medical image; extract characteristics of the performance of the trained machine learning model while mapping the medical image to an output to generate model performance characteristic information; extract characteristics of user feedback received through the user input device based on the output of the trained machine learning model to generate user feedback characteristic information; transmit the input data characteristic information, the model performance characteristic information, and the user feedback characteristic information to a second device A first processor for causing the execution A first device comprising: A second device disposed remotely from the first device, the first device and the second device being communicatively coupled, the second device comprising: a second non-transitory memory including instructions; and a second processor, which, when executing the instructions, causes the second processor to cause the second device to: receive, from the first device, the input data characteristic information, the model performance characteristic information, and the user feedback characteristic information; determine an input data deviation by comparing the input data characteristic information with a plurality of previously determined input data characteristic information; determine a model performance deviation by comparing the model performance characteristic information with a plurality of previously determined model performance characteristic information; determine a user feedback deviation by comparing the user feedback characteristic information with a plurality of previously determined user feedback characteristic information; and responding to one or more of the input data deviation exceeding an input data deviation threshold, the model performance deviation exceeding a model performance deviation threshold, and the user feedback deviation exceeding a user feedback deviation threshold, sending a warning to the user device a second processor that causes the execution a second device including a system including
10. The system according to claim 9, wherein the characteristic information of a plurality of input data determined previously, the characteristic information of a plurality of model performance characteristics determined previously, and the characteristic information of a plurality of user feedback characteristics determined previously are derived from a training data set used for training the trained machine learning model
11. The system according to claim 9, wherein the characteristic information of a plurality of input data determined previously, the characteristic information of a plurality of model performance characteristics determined previously, and the characteristic information of user feedback characteristics determined previously are derived from previous inferences of the trained machine learning model at the deployment site
12. Each of the input data characteristic information, the model performance characteristic information, and the user feedback characteristic information does not include a medical image, and when the instruction is executed, the first processor does not send the medical image to the second device. The system according to claim 9
13. The system according to claim 9, wherein the warning includes indicating one or more of the input data deviation exceeding the input data deviation threshold, the model performance deviation exceeding the model performance deviation threshold, and the user feedback deviation exceeding the user feedback deviation threshold
14. A method for monitoring the performance of a trained machine learning model, comprising responding to a medical image of an imaging subject being input into a trained machine learning model extracting characteristics of the medical of the imaging subject to generate input data characteristic information, wherein the input data characteristic information excludes intensity values of individual pixels or voxels and positions of individual pixels or voxels, generating extracting characteristics of the performance of the trained machine learning model while mapping the medical image to an output to generate model performance characteristic information determining an input data deviation by comparing the input data characteristic information with a plurality of input data characteristic information determined previously Determining a model performance deviation by comparing the model performance characteristic information with a plurality of previously determined model performance characteristic information, and Sending a warning to the user device in response to one or more of the input data deviation exceeding an input data deviation threshold and the model performance deviation exceeding a model performance deviation threshold A method comprising.
15. The method includes Extracting characteristics of received user feedback based on the output of the trained machine learning model to generate user feedback characteristic information, Determining a user feedback deviation by comparing the user feedback characteristic information with a plurality of previously determined user feedback characteristic information, and Sending a warning to the user device in response to the user feedback deviation exceeding a user feedback deviation threshold The method according to claim 14, comprising.
16. The method according to claim 15, wherein the user feedback includes a scan plane used to acquire a diagnostic medical image of the imaging target.
17. Extracting characteristics of received user feedback based on the output of the trained machine learning model to generate user feedback characteristic information includes Determining the center coordinates of the scan plane, and Determining three direction cosines that uniquely identify the orientation of the scan plane The method according to claim 16, comprising.
18. The method according to claim 14, wherein the medical image includes a three-dimensional (3D) medical image, and the trained machine learning model is configured to identify anatomical landmarks of the 3D medical image for positioning a scan plane for acquiring a diagnostic medical image of a region of interest.
19. Extracting characteristics of the medical image of the imaging target to generate the input data characteristic information includes Capturing metadata of the 3D medical image, the metadata including pixel spacing and slice thickness of the 3D medical image, Determining an aggregated intensity statistic of the 3D medical image, the aggregated intensity statistic including the intensity histogram, and Determining a clinical ontology of the 3D medical image, the clinical ontology including a list of anatomical regions incorporated into the 3D medical image The method according to claim 18, comprising.
20. Extracting performance characteristics of the trained machine learning model and generating model performance characteristic information while mapping the medical image to output Extracting characteristics of one or more segmentation masks generated by the trained machine learning model that identifies anatomical landmarks in the 3D medical image Extracting characteristics of the scan plane determined based on the anatomical landmarks, and Recording a list of the anatomical landmarks identified in the 3D medical image by the trained machine learning model The method according to claim 18, comprising
Citation Information
Patent Citations
Medical information processing apparatus and medical information processing method
JP2020204970A
Machine learning program, machine learning method and machine learning device
JP2022105916A
Ai quality monitoring system
JP2022189462A
Computer system and determination method for model switching timing
JP2023032843A