System and method for monitoring trained machine learning models in healthcare environment

By extracting the characteristics of input data, model performance and user feedback, monitoring the performance of AI models, solving the complex monitoring problems due to privacy and regulatory restrictions in traditional methods, real-time performance monitoring without damaging privacy is achieved.

CN120029865APending Publication Date: 2025-05-23GE PRECISION HEALTHCARE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411619041.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-22
Filing Date
2024-11-13
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

When using AI models in healthcare, the need to monitor model performance is complicated by privacy and regulatory constraints, and traditional approaches require access to sensitive clinical data, resulting in time-consuming, complex and potentially raising privacy and security issues.

Method used

By extracting the characteristics of input data, model performance, and user feedback, using these characteristics as an alternative to basic clinical data, monitoring the deployed model performance. This method does not require the use of personally identifiable information or other basic clinically sensitive data.

Benefits of technology

Real-time monitoring of the performance of AI models without compromising patient privacy or overcoming regulatory limitations, ensuring the reliability and accuracy of the models while maintaining regulatory compliance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120029865A_ABST
    Figure CN120029865A_ABST
Patent Text Reader

Abstract

The present disclosure provides a system and method (200) for monitoring performance of an artificial intelligence (AI) model and applications in the field of medical imaging without accessing underlying clinical data. The system extracts characteristics of the input data at the contact point (202), extracts characteristics of model performance (204), and extracts characteristics of user feedback (206), and uses these characteristics as an alternative to the base input data. The invention further comprises a method (500, 800, 1100) for determining a deviation by comparing the extracted characteristic with a previously determined characteristic, and responding to the deviation exceeding a threshold by sending an alert to a user device (140).
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention generally relate to the fields of artificial intelligence (AI) and machine learning, and more specifically to a system and method for monitoring the performance of AI models and applications in the field of medical imaging. Background Art

[0002] With the increasing adoption of AI in healthcare, the need to monitor the performance of AI models in real-world scenarios is becoming more and more urgent. However, access to actual clinical data such as medical images and personally identifiable information (PHI) is often limited due to regulatory restrictions and privacy issues. This poses a significant challenge to data scientists and AI developers, who need to understand how their models perform in the field and make necessary adjustments to improve the performance and reliability of the models.

[0003] Traditionally, AI model performance monitoring in healthcare settings has leveraged the patient clinical data upon which the AI ​​model reasoning is based. However, this approach may require obtaining patient consent, overcoming region-specific restrictions, and managing the data lifecycle. This process is not only time-consuming and complex, but may also raise privacy and security issues. Therefore, there is a need for solutions that can monitor AI model performance without using PHI or other underlying clinically sensitive data. Summary of the invention

[0004] In one embodiment, a method for monitoring the performance of a trained machine learning model includes: extracting characteristics of input data fed into the trained machine learning model to produce an input data representation, wherein the input data includes a medical image; extracting characteristics of the performance of the trained machine learning model during mapping the medical image to an output to produce a model performance representation; extracting characteristics of user feedback received based on the output of the trained machine learning model to produce a user feedback representation; determining an input data deviation by comparing the input data representation with a plurality of previously determined input data representations; determining a model performance deviation by comparing the model performance representation with a plurality of previously determined model performance representations; determining a user feedback deviation by comparing the user feedback representation with a plurality of previously determined user feedback representations; and responding to one or more of an input data deviation exceeding an input data deviation threshold, a model performance deviation exceeding a model performance deviation threshold, and a user feedback deviation exceeding a user feedback deviation threshold by sending an alert to a user device.

[0005] In another embodiment, a system includes: a first device located at a deployment site, wherein the first device includes a user input device, a first non-volatile memory including a trained machine learning model and instructions, and a first processor, wherein when executing the instructions, the first processor causes the first device to: extract characteristics of input data fed into the trained machine learning model to produce an input data representation, wherein the input data includes a medical image; extract characteristics of the performance of the trained machine learning model during mapping the medical image to an output to produce a model performance representation; extract characteristics of user feedback received via the user input device based on the output of the trained machine learning model to produce a user feedback representation; and send the input data representation, the model performance representation, and the user feedback representation to a second device, the second device being located remotely from the first device, wherein the first device Communicatively coupled to a second device, and wherein the second device includes a second non-volatile memory including instructions, and a second processor, wherein when executing the instructions, the second processor causes the second device to: receive an input data representation, a model performance representation, and a user feedback representation from the first device; determine an input data deviation by comparing the input data representation with a plurality of previously determined input data representations; determine a model performance deviation by comparing the model performance representation with a plurality of previously determined model performance representations; determine a user feedback deviation by comparing the user feedback representation with a plurality of previously determined user feedback representations; and respond to one or more of an input data deviation exceeding an input data deviation threshold, a model performance deviation exceeding a model performance deviation threshold, and a user feedback deviation exceeding a user feedback deviation threshold by sending an alert to a user device.

[0006] In yet another embodiment, a method for monitoring the performance of a trained machine learning model includes responding to a medical image of an imaged subject being input into the trained machine learning model by: extracting characteristics of the medical image of the imaged subject to produce an input data representation, wherein the input data representation does not include individual pixel or voxel intensity values ​​and individual pixel or voxel positions; extracting characteristics of the performance of the trained machine learning model during mapping the medical image to an output to produce a model performance representation; determining an input data deviation by comparing the input data representation with a plurality of previously determined input data representations; determining a model performance deviation by comparing the model performance representation with a plurality of previously determined model performance representations; and responding to one or more of an input data deviation exceeding an input data deviation threshold and a model performance deviation exceeding a model performance deviation threshold by sending an alert to a user device.

[0007] In this way, the present disclosure enables monitoring of deployed model performance without using PHI or other underlying clinical sensitive data by extracting features of input data, model performance, and user feedback and using these features as a substitute for underlying clinical data. The extracted features can be used to detect deviations in deployed model performance relative to baseline model performance (e.g., model performance on test datasets, validation datasets, and training datasets), thereby allowing data scientists and developers to understand the performance of deployed models in real time by directly accessing underlying clinical data.

[0008] It should be understood that the above brief description is provided to introduce in a simplified form selected concepts that are further described in the detailed description. It is not meant to identify key features or essential features of the claimed subject matter, the scope of which is uniquely defined by the claims that follow the detailed description. Furthermore, the claimed subject matter is not limited to implementations that solve any disadvantages noted above or in any part of this disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 is a block diagram illustrating an image processing apparatus and a model monitoring apparatus according to an embodiment of the present disclosure;

[0010] Figure 2 is a flow chart illustrating a method for monitoring the performance of a trained machine learning model according to an embodiment of the present disclosure;

[0011] Figure 3 is a flow chart illustrating a method for characterizing a medical image according to an embodiment of the present disclosure;

[0012] Figure 4 is a flow chart illustrating a method for characterizing a three-dimensional (3D) medical image according to an embodiment of the present disclosure;

[0013] Figure 5 is a flow chart illustrating a method for determining deviations of input data representations using a plurality of predetermined input data representations according to an embodiment of the present disclosure;

[0014] Figure 6 is a flow chart illustrating a method for characterizing the output of a trained machine learning model according to an embodiment of the present disclosure;

[0015] Figure 7 is a flow chart illustrating a method for characterizing a segmentation mask produced by a trained machine learning model according to an embodiment of the present disclosure;

[0016] Figure 8 is a flow chart illustrating a method for determining bias of an output of a trained machine learning model using a plurality of predetermined model output representations according to an embodiment of the present disclosure;

[0017] Fig. 9 is a flow chart illustrating a method for characterizing user feedback of a model output according to an embodiment of the present disclosure;

[0018] Fig.10 is a flow chart illustrating a method for characterizing user feedback by determining user adjustments to the output of a trained machine learning model according to an embodiment of the present disclosure;

[0019] Fig.11 is a flow chart illustrating a method for determining deviations of user feedback using a plurality of predetermined user feedback representations according to an embodiment of the present disclosure;

[0020] Fig.12 is a flow chart illustrating a method for characterizing input data from a training data set for training a machine learning model according to an embodiment of the present disclosure;

[0021] Fig.13 is a flow chart illustrating a method for extracting model performance characteristics during training of a machine learning model according to an embodiment of the present disclosure;

[0022] Fig.14 is a flow chart illustrating a method for extracting ground truth features from a training data set according to an embodiment of the present disclosure; and

[0023] Fig.15 is a flow chart illustrating a method for accumulating characterizations from deployment sites according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0024] The present disclosure describes systems and methods for monitoring the performance of trained machine learning models in the field of medical imaging. With the increasing adoption of AI in healthcare, there is a need to monitor the real-world performance of AI models without compromising patient privacy or overcoming regulatory restrictions. The systems and methods of the present invention achieve performance monitoring by extracting features of input data, model output, and user feedback at the point of clinical use. These features act as substitutes for underlying clinical data, providing insights into model performance without access to protected health information or clinically sensitive data. The extracted features are compared with expected baselines derived from training data or past model reasoning to detect deviations in the performance of deployed models. An alert is generated when the deviation exceeds a predetermined threshold for timely investigation and intervention. By using statistical substitutes rather than actual clinical data, the present disclosure facilitates continuous performance monitoring to ensure patient safety and model reliability while maintaining regulatory compliance and protecting sensitive personal information. The system and method provide a solution for AI quality assurance in real-world healthcare environments.

[0025] In one example, the image processing device 102 located at the deployment site (at Figure 1 ) can monitor the performance of one or more trained machine learning models by extracting features from input medical image data, model performance, and user feedback. The image processing device 102 sends the extracted features to the remote model monitoring device 122, which encodes the features into one or more feature vectors and compares the feature vectors with previously determined characterization vectors to identify deviations from the input data, model performance, or user feedback. Figure 1 The system shown in can perform Figure 2 One or more operations to extract features, determine deviations by comparing with previously determined characterizations, and send an alert when the deviation exceeds one or more predetermined thresholds. Figure 3 ) by extracting metadata, pixel statistics, and ontology labels to characterize medical image input. In addition, method 400 can be used to characterize 3D medical image input, such as Figure 4 As shown in Figure 1, metadata, intensity histograms, and clinical ontology labels are used to characterize 3D images. Figure 5 As shown in method 500 in , input data features extracted by methods 300 and / or 400 may be encoded into vectors and compared with previously determined input data representations.

[0026] Similarly, according to Figure 6 One or more operations of the method 600 shown in can be used to characterize the model performance, including extracting model confidence scores, uncertainty measures, and other model performance characteristics. In embodiments where the model output includes a segmentation mask, the Figure 7 One or more operations of the method 700 shown in FIG. 7 are used to extract model performance characteristics, such as analyzing properties of the segmentation mask and / or the identified landmarks. Figure 8 A method 800 is shown for encoding the model performance characteristics produced by methods 600 and or 700 into a feature vector and comparing to a previously determined model performance characterization vector to identify model performance deviations.

[0027] User feedback can be provided by executing Fig. 9 , which includes capturing user feedback such as model ratings and user-provided corrections. In embodiments where the model output includes scan plane specifications, the scan plane specifications may be based on Fig.10 The user feedback characteristics determined by methods 900 and / or 1000 may be encoded as a feature vector and compared to a previously determined user feedback characterization to determine the user feedback characteristics by performing Fig.11One or more operations of the method 1100 shown in are used to detect deviations.

[0028] Fig.12 A method 1200 is shown for extracting input data characteristics from training data to establish expected / baseline model performance. Fig.13 A method 1300 for extracting model performance characteristics during training is described in detail. Fig.14 The techniques in method 1400 for extracting ground truth features from training data are summarized. Finally, Fig.15 A method 1500 for accumulating deployment data representations to update deployment site-specific performance baselines over time is shown. In this way, the performance and use of trained machine learning models at deployment sites can be monitored in substantially real time using statistical alternatives to the underlying clinical data.

[0029] First reference Figure 1 , a model monitoring system 100 is shown. The model monitoring system 100 can be configured to acquire medical images of an imaging subject using an imaging device 116, and to process the acquired medical images using an image processing device 102, for example to extract characteristics of the medical images, the performance of the trained machine learning model 108, and user feedback received via a user input device 112. The processed information can be displayed to the user via a display device 114. The image processing device 102 can receive user input via the user input device 112, wherein one or more of the image acquisition and image processing can be adjusted based on the received user input. In some embodiments, the user input device 112 can be used to provide user feedback to the image processing device 102, such as corrections to the output of the trained machine learning model 108, ratings of the output, and other types of feedback.

[0030] The image processing device 102 includes a processor 104 configured to execute machine-readable instructions stored in a non-transitory memory 106. The processor 104 may be a single-core or multi-core processor, and the program executed thereon may be configured for parallel processing or distributed processing. In some embodiments, the processor 104 may optionally include separate components distributed on two or more devices, which may be located at a distance and / or configured for coordinated processing. In some embodiments, one or more aspects of the processor 104 may be virtualized and performed by a remotely accessible networked computing device configured in a cloud computing configuration.

[0031] The non-transitory memory 106 may store a trained machine learning model 108 and a characterization module 110. The trained machine learning model 108 may include instructions for mapping a medical image to an output. The characterization module 110 may include instructions for extracting characteristics of input data fed into the trained machine learning model 108, characteristics of the performance of the trained machine learning model 108 during mapping the medical image to the output, and characteristics of user feedback received via a user input device 112 based on the output of the trained machine learning model 108. The trained machine learning model 108 stored in the non-transitory memory 106 includes instructions that enable mapping a medical image to an output. The output may be a diagnosis, a prediction, or other medically relevant information that can be derived from a medical image. The trained machine learning model 108 may employ various machine learning algorithms, such as deep learning, a convolutional neural network, or a support vector machine, among others. The characterization module 110, also stored in the non-transitory memory 106, includes instructions for extracting various characteristics related to the input data, the performance of the trained machine learning model 108, and user feedback. The input data characteristics may include metadata from the medical image, pixel or voxel statistics from the medical image, and one or more labels from the appearance ontology and clinical ontology of the medical image. The metadata does not include personally identifiable information, and the pixel or voxel statistics do not retain individual pixel or voxel intensity values ​​or positions. The appearance ontology and clinical ontology provide additional context about the medical image, such as the pathology type and organs present in the image. The performance characteristics of the trained machine learning model 108 may include the output of the model 108, one or more intermediate outputs produced by one or more hidden layers of the model 108, confidence scores for the outputs of the model 108, and one or more uncertainty metrics for the outputs of the model 108. These performance characteristics provide information about the performance status of the model 108 and can be used to identify any potential problems or areas for improvement. User feedback characteristics may include model output ratings received via the user input device 112 and user corrections received via the user input device 112, wherein the user corrections modify the output of the trained machine learning model 108. This feedback provides insight into how the user interacts with the output of the model 108 and can be used to further refine and improve the model 108.

[0032] Image processing device 102 may send the input data representation, the model performance representation, and the user feedback representation to model monitoring device 122. Model monitoring device 122 is located remotely from image processing device 102, wherein image processing device 102 is communicatively coupled to model monitoring device 122.

[0033] The model monitoring device 122 includes a processor 124 configured to execute machine-readable instructions stored in a non-transitory memory 126. The processor 124 may be a single-core or multi-core processor, and the program executed thereon may be configured for parallel processing or distributed processing. In some embodiments, the processor 124 may optionally include separate components distributed on two or more devices, which may be located remotely and / or configured for coordinated processing. In some embodiments, one or more aspects of the processor 124 may be virtualized and executed by a remotely accessible networked computing device configured in a cloud computing configuration.

[0034] The non-transitory memory 126 may store a deviation module 128, a representation database 130, and an alert module 132. The deviation module 128 may include instructions for determining an input data deviation by comparing an input data representation with a plurality of previously determined input data representations, determining a model performance deviation by comparing a model performance representation with a plurality of previously determined model performance representations, and determining a user feedback deviation by comparing a user feedback representation with a plurality of previously determined user feedback representations.

[0035] The characterization database 130 may store a plurality of previously determined input data characterizations, model performance characterizations, and user feedback characterizations. These characterizations may be derived from a training data set used to train a machine learning model, or derived from previous reasoning of a trained machine learning model at a deployment site. The characterization database 130 may also store distribution statistics associated with a plurality of input data characterizations, model performance characterizations, and user feedback characterizations. These distribution statistics may be used to dynamically adjust corresponding deviation thresholds based on the distribution of previously determined characterizations. The characterization database 130 may also store a plurality of predetermined feature vectors corresponding to a plurality of previously determined input data characterizations, model performance characterizations, and user feedback characterizations. These predetermined feature vectors may be used to determine the degree of similarity of a current characterization to a previously determined characterization by calculating the distance between the current feature vector and each of the predetermined feature vectors in a multidimensional space. The characterization database 130 may be continuously updated as new characterizations are received from a deployment site, thereby allowing the system to adapt its monitoring to changes in input data, model performance, or user feedback.

[0036] Alert module 132 may include instructions for responding to one or more of an input data deviation exceeding an input data deviation threshold, a model performance deviation exceeding a model performance deviation threshold, and a user feedback deviation exceeding a user feedback deviation threshold by sending an alert to user device 140 .

[0037] User input device 134 may include one or more of a touch screen, keyboard, mouse, trackpad, motion sensing camera, or other device configured to enable a user to interact with and manipulate data within model monitoring device 122 .

[0038] The display device 136 may include one or more display devices utilizing almost any type of technology. The display device 136 may be configured to present information visually to the user. The information may include, for example, input data characterization, model performance characterization, and user feedback characterization. The display device 136 may also present the determined deviations of these characterizations from their respective previously determined characterizations. In the case where one or more deviations in the deviations exceed their respective thresholds, the display device 136 may present an alarm generated by the alarm module 132. The alarm may include an indication of which deviation exceeds its threshold, such as an input data deviation exceeding an input data deviation threshold, a model performance deviation exceeding a model performance deviation threshold, or a user feedback deviation exceeding a user feedback deviation threshold. Depending on the user's preferences and the urgency of the situation, the display device 136 may present the alarm in various formats, such as a pop-up notification, a highlighted text, or a sound alarm. In addition to presenting an alarm, the display device 136 may also display various other types of information related to the operation of the trained machine learning model. The display device 136 may be combined in a shared housing with the processor 124, the non-volatile memory 126, and / or the user input device 134, or may be a peripheral display device and may include a monitor, touch screen, projector, or other display device known in the art that enables a user to view the MRI generated by the MRI system and / or interact with various data stored in the non-volatile memory 126.

[0039] It should be understood that Figure 1 The image processing device 102 and model monitoring device 122 shown in FIG. 1 are for illustration and not limitation. Another suitable imaging system and model monitoring device may include more, fewer, or different components.

[0040] refer to Figure 2 , a method 200 for monitoring the performance of a trained machine learning model is shown. The method 200 can be used by an image processing system to extract characteristics of input data, model performance, and user feedback, and determine deviations by comparing these characteristics with previously determined characteristics.

[0041] At operation 202, the image processing system extracts characteristics of the input data fed into the trained machine learning model to generate input data representations. This extraction allows the system to identify specific features within the input data of interest for monitoring. In one embodiment, the extraction of input data characteristics involves utilizing metadata from medical images, such as subject-related characteristics (e.g., age range, gender, BMI), machine-related characteristics (e.g., machine model name), acquisition-related characteristics (e.g., field strength, coil, scan type, orientation, number of slices), and clinical data characteristics (e.g., mean intensity, median intensity, intensity histogram). The metadata does not include personally identifiable information, and the pixel or voxel statistics do not retain individual pixel or voxel intensity values ​​or locations. The appearance ontology and clinical ontology provide additional context about the medical image, such as the pathology types and organs present in the image, without leaking protected health information (PHI).

[0042] At operation 204, the image processing system extracts characteristics of the performance of the trained machine learning model during mapping input data to output to produce a model performance characterization. This extraction allows the system to monitor the performance of the machine learning model. In one embodiment, the extraction of model performance characteristics involves capturing the output of the trained machine learning model and one or more intermediate outputs produced by one or more hidden layers of the trained machine learning model; determining a confidence score for the output of the trained machine learning model; and determining one or more uncertainty metrics for the output of the trained machine learning model. In another embodiment, at operation 204, the image processing system may capture model performance characteristics such as classification, segmentation, image regression, point regression, line regression, and surface regression metrics.

[0043] At operation 206, the image processing system extracts characteristics of user feedback to generate a user feedback representation. This extraction allows the system to monitor the user's interaction with the output of the machine learning model. In one embodiment, the extraction of user feedback characteristics involves recording model output ratings received via a user input device; recording user corrections received via a user input device, wherein the user corrections modify the output of the trained machine learning model; and capturing additional user comments or actions indicating that the user accepts or modifies the model predictions. This feedback can be used to evaluate the consistency of the model with user expectations and clinical workflow requirements.

[0044] At operation 208, the image processing system determines input data deviations by comparing the input data representation to a plurality of previously determined input data representations. This determination allows the system to monitor changes in the input data over time, such as changes in subject demographics, equipment settings, and acquisition parameters, which may affect the performance of the model.

[0045] At operation 210, the image processing system determines a model performance deviation by comparing the model performance characterization to a plurality of previously determined model performance characterizations. This determination allows the system to monitor changes in the performance of the machine learning model over time, including classification accuracy, segmentation quality, and deviations in regression predictions, which can signal the need for model retraining or adjustment.

[0046] At operation 212, the image processing system determines a user feedback deviation by comparing the user feedback representation to a plurality of previously determined user feedback representations. This determination allows the system to monitor changes in user interaction with the output of the machine learning model over time, thereby providing insight into the usefulness and acceptance of the model in a clinical setting.

[0047] At operation 214, the image processing system evaluates whether one or more of the input data deviation, the model performance deviation, and the user feedback deviation exceeds their respective thresholds. If any deviation exceeds their respective thresholds, the image processing system proceeds to operation 218, where an alert is sent to the user device indicating that the deviation exceeds the threshold. The alert may include details of the deviation and suggestions for corrective actions. If no deviation exceeds their respective thresholds, the image processing system continues to monitor model reasoning at operation 216. After operation 218 or operation 216, method 200 may end.

[0048] Thus, method 200 can at least partially address the technical challenges of monitoring and maintaining the performance of machine learning models in the field of medical imaging while complying with strict privacy regulations and without direct access to protected health information (PHI). The technical improvements provided by method 200 include the ability to extract alternative data representations from input data, model performance, and user feedback, which serve as non-invasive alternatives to underlying clinical data. The method enables real-time detection of deviations in model performance by comparing these representations with baselines established from previously determined representations. When the deviation exceeds a predetermined threshold, the system generates an alert so that corrective action can be taken in a timely manner to ensure the reliability and accuracy of the model. This technical effect ensures that machine learning models continue to function optimally in healthcare environments, thereby providing reliable support to medical professionals and enhancing patient care outcomes.

[0049] refer to Figure 3 , a method 300 for characterizing a medical image is shown. The method 300 may be employed by an image processing system to extract characteristics of a medical image, which may be used to monitor the performance of a trained machine learning model.

[0050] At operation 302, the image processing system extracts metadata from the medical image. The metadata may include information about the medical image that does not include personally identifiable information (PII) or protected health information (PHI). For example, the metadata may include information about the subject of the image, such as whether the subject is a pediatric subject, laterality protocol, pixel spacing, slice thickness, etc. In some embodiments, the metadata may include information about the subject of the image, such as age range, gender, and body mass index (BMI) calculated from weight, height, and gender. In addition, the metadata may include machine-related characteristics, such as the machine model name of the equipment that produced the medical image. The metadata provides a high-level overview of the medical image and can be used to characterize the input data fed into the trained machine learning model.

[0051] At operation 304, the image processing system aggregates pixel or voxel statistics from the medical image. These statistics do not retain individual pixel or voxel intensity values ​​or locations, but rather provide a summary of the overall intensity distribution within the image. For example, the image processing system can generate intensity histograms, intensity profiles across anatomical planes such as anterior-posterior (AP), superior-inferior (SI), and lateral rotation (LR), and other summary statistics such as mean intensity, median intensity, minimum intensity, and maximum intensity. These statistics can provide insight into the overall quality and characteristics of the medical image, which can be useful for monitoring the performance of a trained machine learning model.

[0052] At operation 306, the image processing system determines one or more labels from an appearance ontology of the medical image. The appearance ontology may include a collection of predefined labels that describe various visual features or characteristics of the medical image. For example, the labels may describe the type of pathology present in the image, the organs visible in the image, and other visual features, such as the presence of artifacts such as shadows or blurs. These labels provide a way to categorize and describe the visual content of the medical image in a structured and standardized manner.

[0053] At operation 308, the image processing system determines one or more tags from a clinical ontology of the medical image. The clinical ontology may include a collection of predefined tags that describe various clinical features or characteristics of the medical image. For example, the tags may describe the type of medical imaging modality used to capture the image (e.g., MR, CT, ultrasound), the type of clinical task associated with the image (e.g., classification, segmentation, image regression), and other clinical features, such as the presence of a particular organ within the field of view (FoV). These tags provide a way to categorize and describe the clinical content of the medical image in a structured and standardized manner.

[0054] After operation 308, method 300 may end. The extracted metadata, pixel or voxel statistics, and ontology labels together form an input data representation for the medical image. The input data representation may be used to monitor the performance of the trained machine learning model by comparing the input data representation with a previously determined input data representation.

[0055] refer to Figure 4 , a method 400 for characterizing a three-dimensional (3D) medical image is shown. The method 400 can be employed by an image processing system to extract characteristics of a 3D medical image, which can be used to generate an input data representation. Such a representation can be used to monitor the performance of a trained machine learning model, particularly in the case of medical imaging.

[0056] At operation 402, an image processing system extracts metadata from a 3D medical image. The metadata may include information about the image that does not directly represent the image content, such as pixel pitch and slice thickness. Pixel pitch refers to the physical distance between the centers of each pixel in the patient, measured in millimeters. Slice thickness refers to the thickness of the 3D slice represented by the image, also measured in millimeters. The metadata may provide important context for understanding the image and may be used to normalize or standardize the image data for further processing.

[0057] At operation 404, the image processing system determines an intensity histogram of the 3D medical image. An intensity histogram is a graphical representation of the distribution of pixel intensities in an image. In a 3D medical image, pixel intensities may represent different tissue types or structures, and the intensity histogram may provide a summary of the distribution of these different tissues or structures in the image. The intensity histogram may be used to identify features or patterns in the image data that may be relevant to the trained machine learning model.

[0058] At operation 406, the image processing system determines one or more tags of a clinical ontology of the 3D medical image. A clinical ontology is a collection of structured terms or tags representing concepts in a clinical field. These tags can be used to annotate or describe the content of a 3D medical image in a standardized manner. For example, a tag can indicate an organ or anatomical structure visible in an image, what type of imaging modality is used to acquire an image, or what type of pathology is present in an image. These tags can provide a high-level summary of the image content, which can be used to classify or categorize the image for further processing.

[0059] After operation 406, method 400 may end. The extracted metadata, intensity histograms, and clinical ontology labels together form an input data representation for the 3D medical image. The representation may be used to monitor the performance of the trained machine learning model by comparing the representation with a previously determined input data representation.

[0060] refer to Figure 5 , shows a method 500 for encoding metadata, pixel or voxel statistics, labels from an appearance ontology, and labels from a clinical ontology into a feature vector, and comparing the feature vector to a plurality of predetermined feature vectors corresponding to a plurality of previously determined input data representations. The method 500 may be employed by an image processing system to determine deviations of input data representations, which may be used to monitor the performance of a trained machine learning model.

[0061] refer to Figure 5 , shows a method 500 for encoding metadata, pixel or voxel statistics, labels from an appearance ontology, and labels from a clinical ontology into a feature vector, and comparing the feature vector to a plurality of predetermined feature vectors corresponding to a plurality of previously determined input data representations. The method 500 may be employed by an image processing system to determine deviations of input data representations, which may be used to monitor the performance of a trained machine learning model.

[0062] At operation 502, the image processing system encodes metadata, pixel or voxel statistics, labels from the appearance ontology, and labels from the clinical ontology into feature vectors. This encoding allows the system to represent input data representations in a form that can be easily compared with other input data representations. In one embodiment, encoding involves transforming metadata, pixel or voxel statistics, and labels into a numerical format that can be represented as a vector in a multidimensional space. The dimensions of this space can correspond to different types of characteristics being encoded, such as metadata, pixel or voxel statistics, and labels. Encoding can also involve normalizing the values ​​of the characteristics to ensure that they are at a similar scale, which can help prevent any one characteristic from dominating the comparison due to its scale. Various techniques can be used to perform encoding, such as unique hot encoding for categorized data, scaling for numerical data, etc. In addition, encoding can utilize specific healthcare-related metadata, such as DICOM metadata, and can include pixel or voxel characteristics derived from medical images, such as average intensity, histograms, or entropy, without including any patient identity information, thereby ensuring the de-identification of data.

[0063] At operation 504, the image processing system compares the feature vector to a plurality of predetermined feature vectors corresponding to a plurality of previously determined input data representations. The comparison allows the system to determine the degree of similarity of the current input data representation to the previously determined input data representation. In one embodiment, the comparison involves calculating the distance between the feature vector and each of the predetermined feature vectors in the multidimensional space. The distance can be calculated using a variety of distance metrics, such as Euclidean distance, Manhattan distance, cosine similarity, etc. The system can then determine the deviation by comparing the calculated distance to a threshold. If the distance exceeds the threshold, this may indicate that the current input data representation is significantly different from the previously determined input data representation, which in turn may indicate a potential problem with the performance of the trained machine learning model. The system can respond to such a deviation by sending an alert to the user device, such as Figure 2 , which is described in method 200 shown in . After operation 504 , method 500 may end.

[0064] At operation 504, the image processing system compares the feature vector to a plurality of predetermined feature vectors corresponding to a plurality of previously determined input data representations. The comparison allows the system to determine the degree of similarity of the current input data representation to the previously determined input data representation. In one embodiment, the comparison involves calculating the distance between the feature vector and each of the predetermined feature vectors in the multidimensional space. The distance can be calculated using a variety of distance metrics, such as Euclidean distance, Manhattan distance, cosine similarity, etc. The system can then determine the deviation by comparing the calculated distance to a threshold. If the distance exceeds the threshold, this may indicate that the current input data representation is significantly different from the previously determined input data representation, which in turn may indicate a potential problem with the performance of the trained machine learning model. The system can respond to such a deviation by sending an alert to the user device, such as Figure 2 The method 200 shown in FIG. 1 is described in detail. The alarm may be generated by the alarm module 132 and sent to the user device 140, which may be part of a monitoring setting of a hospital or data scientist. After operation 504, the method 500 may end.

[0065] refer to Figure 6 , a method 600 for characterizing the output of a trained machine learning model in the context of medical imaging is shown. The method 600 can be employed by an image processing system to capture the output of a trained machine learning model, determine a confidence score for the output, and determine one or more uncertainty metrics for the output, which together can characterize the performance of the trained machine learning model in mapping input data to a desired output.

[0066] At operation 602, the image processing system captures the output of the trained machine learning model and one or more intermediate outputs produced by one or more hidden layers of the trained machine learning model. This operation allows the system to monitor the performance of the model at various stages of its operation. The output of the model may include, for example, identification of anatomical landmarks in the 3D medical image, properties of scan planes determined based on the anatomical landmarks, and a list of anatomical landmarks identified in the 3D medical image by the trained machine learning model. The intermediate outputs may include, for example, results for individual layers or groups of layers in the model, which may provide insight into the inner workings of the model and help identify potential areas for improvement or correction.

[0067] At operation 604, the image processing system determines a confidence score for the output of the trained machine learning model. The confidence score provides a quantitative measure of the certainty or reliability of the model output. The confidence score can be determined based on various factors, such as the internal parameters of the model, the quality of the input data, and the complexity of the task at hand. For example, if the input data is of high quality and the task is relatively simple, the confidence score may be high, and if the input data is of low quality or the task is complex, the confidence score may be low. The confidence score can be used to evaluate the performance of the model and decide whether to trust its output or seek additional information or confirmation.

[0068] At operation 606, the image processing system determines one or more uncertainty metrics for the output of the trained machine learning model. The uncertainty metric provides a quantitative measure of the uncertainty or variability of the model output. The uncertainty metric can be determined based on various factors, such as the internal parameters of the model, the quality of the input data, and the complexity of the task at hand. For example, if the input data is of low quality or the task is complex, the uncertainty metric may be high, and if the input data is of high quality and the task is relatively simple, the uncertainty metric may be low. The uncertainty metric can be used to evaluate the performance of the model and decide whether to trust its output or seek additional information or confirmation.

[0069] After operation 606, method 600 may end. The method allows for continuous monitoring and evaluation of the performance of trained machine learning models in the context of medical imaging, which can help ensure the accuracy and reliability of the model output and ultimately improve the quality of patient care.

[0070] refer to Figure 7 , a method 700 for generating a model output representation by determining properties of one or more segmentation masks generated by a trained machine learning model and identifying anatomical landmarks in a 3D medical image is shown. The method 700 can be employed by an image processing system to extract properties of a segmentation mask that are used to identify anatomical landmarks in a 3D medical image.

[0071] At operation 702, the image processing system extracts properties of one or more segmentation masks generated by the trained machine learning model that identify anatomical landmarks in the 3D medical image. The segmentation masks are generated by the trained machine learning model and are used to identify and isolate specific anatomical landmarks within the 3D medical image. The properties of these segmentation masks may include, but are not limited to, the size of the segmentation mask, the centroid of the segmentation mask, and the intensity profile within the segmentation mask. These properties provide valuable information about the anatomical landmarks and can be used to further characterize the input data of the machine learning model.

[0072] In one embodiment, the size of the segmentation mask can be determined by counting the number of pixels or voxels within the mask, which can provide an estimate of the volume or area of ​​the anatomical landmark represented by the mask. The centroid of the segmentation mask, which is the geometric center of the mask, can be calculated by averaging the coordinates of all pixels or voxels within the mask. The intensity profile within the segmentation mask can be determined by calculating statistics such as the mean, median, standard deviation, and histogram of the pixel or voxel intensities within the mask. In addition, the segmentation mask attributes can include an uncertainty score that quantifies the segmentation confidence, and the segmentation mask type, which can be a bounding box with dimensions, a polygon with vertices, or a base-64 encoded mask representing the segmented region.

[0073] At operation 704, the image processing system extracts properties of a scan plane determined based on the anatomical landmark. A scan plane is a two-dimensional slice of a 3D medical image determined based on the location of the anatomical landmark. The properties of the scan plane may include, but are not limited to, the orientation of the plane, the location of the plane within the 3D medical image, and the intensity of pixels or voxels within the plane. These properties may provide additional information about the anatomical landmark and surrounding tissue or structure.

[0074] In one embodiment, the orientation of the scan plane can be represented by three direction cosines, which are the cosines of the angles between the normal vector of the plane and the axes of the image coordinate system. The position of the scan plane in the 3D medical image can be represented by the coordinates of the center point of the plane. The pixel or voxel intensity in the plane can be characterized by calculating statistics such as the mean, median, standard deviation and histogram of the intensity. In addition, in some embodiments, the scan plane attributes can include the plane range that can be represented by the dimensions in the x, y and z directions, and the center that can be represented by the coordinates in the image coordinate system.

[0075] At operation 706, the image processing system records a list of anatomical landmarks identified in the 3D medical image by the trained machine learning model. The list may include the names or labels of the landmarks, their locations within the 3D medical image, and their associated confidence scores, which indicate the certainty of the model in the identification of each landmark. The list may be used to further characterize the input data of the machine learning model and monitor the performance of the model. After operation 706, method 700 may end.

[0076] refer to Figure 8 , a method 800 for determining a deviation of a model performance representation from a plurality of previously determined model performance representations is shown. The method 800 may be employed by a model monitoring system to determine a deviation of the performance of a machine learning model based on characteristics of a model output.

[0077] At operation 802, the model monitoring system encodes the output of the trained machine learning model as well as one or more intermediate outputs, confidence scores, and one or more uncertainty metrics into a feature vector. This encoding allows the system to capture a comprehensive representation of the model's performance. The output of the trained machine learning model may include the final prediction or decision made by the model, such as the identification of anatomical landmarks in a medical image. The intermediate outputs may include the output of one or more hidden layers of the model, which provide insight into the inner workings of the model. The confidence score may represent the certainty of the model output, and the uncertainty metric may quantify the uncertainty or variability of the model output.

[0078] At operation 802, the model monitoring system encodes the output of the trained machine learning model as well as one or more intermediate outputs, confidence scores, and one or more uncertainty metrics as a feature vector. Such encoding allows the system to capture a comprehensive representation of the model's performance. The output of the trained machine learning model may include a final prediction or decision made by the model, such as identification of anatomical landmarks in a medical image, classification of a medical image into a predetermined category, segmentation of a specific structure within a medical image, or regression analysis for predicting a continuous variable from a medical image. Intermediate outputs may include outputs of one or more hidden layers of the model that provide insight into the inner workings of the model. Confidence scores may represent the certainty of the model output, and uncertainty metrics may quantify the uncertainty or variability of the model output, including uncertainty scores associated with predicted masks, points, lines, or planes in a medical image.

[0079] At operation 804, the model monitoring system compares the feature vector with a plurality of predetermined feature vectors corresponding to a plurality of previously determined model performance characterizations. The comparison allows the system to determine the deviation of the current performance of the model from its past performance or expected performance. The predetermined feature vector may be derived from a training data set for training the machine learning model, or derived from a previous reasoning of the model at the deployment site. The comparison may involve calculating a distance or similarity measure between the feature vector and each of the predetermined feature vectors, such as Euclidean distance, cosine similarity, or other suitable distance or similarity measures. If the comparison shows that the deviation of the model performance exceeds a model performance deviation threshold, the model monitoring system may respond by sending an alert to the user device. The alert may indicate that the performance of the model deviates significantly from its expected performance, and may suggest potential actions to resolve the deviation, such as retraining the model, adjusting the parameters of the model, or investigating the cause of the deviation. After operation 804, method 800 may end.

[0080] refer to Fig. 9 , a method 900 for characterizing user feedback received based on output from a trained machine learning model is shown. The method 900 can be employed by an image processing system to extract characteristics of user feedback received based on output from a trained machine learning model to generate a user feedback characterization.

[0081] At operation 902, the image processing system records a model output rating received via a user input device. The model output rating can be a numerical or categorized value representing a user's assessment of the quality or accuracy of the output of the trained machine learning model. The model output rating can be based on various factors, such as the perceived accuracy of the output, the usefulness of the output in the user's workflow, the user's confidence in the output, etc. The model output rating can be input by the user via a user input device such as a keyboard, a mouse, a touch screen, a voice command interface, or any other suitable input device.

[0082] At operation 904, the image processing system records user corrections received via the user input device, wherein the user corrections modify the output of the trained machine learning model. The user corrections may be adjustments or modifications made by the user to the output of the trained machine learning model. The user corrections may include, for example, adjusting the position or orientation of the segmentation mask, adding or removing markers, adjusting parameters of the recommended scan planes, etc. The user corrections may be input by the user via the user input device and may be recorded by the image processing system as part of the user feedback representation.

[0083] The user feedback characterization generated by method 900 provides valuable information about the performance of a trained machine learning model in the real-world environment of a user workflow. By analyzing the user feedback characterization, an image processing system can identify areas where the trained machine learning model may need improvement and can use this information to update or retrain the model, thereby improving its performance and usefulness in a healthcare environment. After operation 904, method 900 may end.

[0084] Reference Fig.10 , shows method 1000 for characterizing user feedback received based on the output of a trained machine learning model. Method 1000 can be employed by an image processing system to determine a user feedback characterization that can be used to monitor the performance of an AI model and for applications in the field of medical imaging.

[0085] At operation 1002, the image processing system determines the central coordinates of the scan plane for performing a diagnostic scan of an imaging subject. This determination allows the system to identify the scan plane for final user approval / adjustment for performing the diagnostic scan.

[0086] At operation 1004, the image processing system determines three direction cosines that uniquely identify the orientation of the scan plane. After operation 1004, method 1000 may end.

[0087] The user feedback characterization provides information about the user's interaction with the AI model, and the user feedback characterization includes the central coordinates of the scan plane and the three direction cosines that uniquely identify the orientation of the scan plane. This information can be used to monitor the performance of the AI model and identify any deviations from the expected behavior. If such a deviation is detected, an alert can be sent to the user device for timely intervention and correction.

[0088] Reference Fig.11 , shows method 1100 for determining a deviation in user feedback using multiple predetermined user feedback characterizations. Method 1100 can be employed by an image processing system to compare the user feedback characterization with multiple previously determined user feedback characterizations.

[0089] At operation 1102, the image processing system encodes the model output rating and the user correction as a feature vector. The model output rating and the user correction are received via a user input device and are part of the user feedback representation. The model output rating can be a numerical or categorized value indicating a user's assessment of the quality or accuracy of the output of the trained machine learning model. The user correction can be a modification made by the user to the output of the trained machine learning model, such as an adjustment to the position or orientation of a scan plane, or a change to an identified landmark or segmentation mask. Encoding the model output rating and the user correction as a feature vector allows a compact and efficient representation of the user feedback representation, which can be easily compared with other feature vectors representing previously determined user feedback representations.

[0090] At operation 1104, the image processing system compares the feature vector to a plurality of predetermined feature vectors corresponding to a plurality of previously determined user feedback representations. The comparison allows the image processing system to determine a user feedback deviation, which is a measure of the difference between the current user feedback representation and the previously determined user feedback representation. The user feedback deviation may be calculated as the distance between the feature vector and the closest one of the predetermined feature vectors in the feature space, or as a statistical measure of the difference between the feature vector and a distribution of predetermined feature vectors. The user feedback deviation provides an indication of the deviation of the current user feedback from typical or expected user feedback, which may be used to monitor the performance of the trained machine learning model and detect any anomalies or problems in its operation.

[0091] In some embodiments, the image processing system can respond to the user feedback deviation exceeding the user feedback deviation threshold by sending an alert to the user device. The user feedback deviation threshold can be a predetermined value or a value dynamically adjusted based on a distribution of previously determined user feedback characterizations. The alert can include information about the user feedback deviation and suggestions for corrective actions, such as retraining the machine learning model, adjusting its parameters, or providing additional guidance to the user. After operation 1104, method 1100 can end.

[0092] refer to Fig.12 , a method 1200 for characterizing input data from a training data set for training a machine learning model is shown. The method 1200 can be employed by an image processing system to extract input data features from input data of each of a plurality of training data pairs to produce a plurality of input data characterizations, which can then be used as a baseline for comparison to determine deviations of the input data characterizations of input data received at a deployment site.

[0093] At operation 1202, the image processing system acquires a training data set for training a machine learning model. The training data set includes a plurality of training data pairs, where each training data pair includes input data and corresponding ground truth data. The input data may include medical images, such as MR, CT, or ultrasound images, and the ground truth data may include annotations or labels indicating the presence or absence of certain features or conditions in the image.

[0094] At operation 1204, the image processing system extracts input data features from the input data of each training data pair in the plurality of training data pairs to generate a plurality of input data representations. The extraction process may involve one or more operations in the operations described in more detail above with respect to methods 300 and 400. Features may include metadata, pixel or voxel statistics, and labels from appearance and clinical ontology. Metadata may include information about the imaging subject, such as whether the subject is a pediatric subject, laterality protocol, pixel spacing, slice thickness, etc. Pixel or voxel statistics may include intensity histograms, intensity profiles, noise, and other statistical measures derived from the pixel or voxel values ​​in the image. Labels from appearance and clinical ontology may include information about pathology types, organs present in the image, and other high-level features that cannot be easily derived from pixel data alone.

[0095] At operation 1206, the image processing system encodes the plurality of input data representations to generate a plurality of input data representation vectors. The encoding process may involve converting the extracted features into numerical form that can be easily compared and analyzed. For example, metadata, pixel or voxel statistics, and ontology labels may be encoded as feature vectors, where each element of the vector corresponds to a particular feature.

[0096] At operation 1208, the image processing system determines distribution statistics for the plurality of input data characterization vectors. These statistics may include central tendency measures (such as mean or median), dispersion measures (such as variance or standard deviation), and other statistical measures describing the distribution of the input data characterization vectors. These statistics provide a summary of the characteristics of the input data in the training data set and can be used to compare the input data characteristics of new data with the characteristics of the training data.

[0097] At operation 1210, the image processing system stores the plurality of input data characterization vectors and distribution statistics in a non-transitory memory. This stored information can later be used to compare the characteristics of the new input data with the characteristics of the training data in order to monitor the performance of the trained machine learning model and detect any deviations or drifts in the input data. After operation 1210, method 1200 can end.

[0098] refer to Fig.13, a method 1300 for extracting model performance characteristics during training of a machine learning model is shown. The method 1300 can be employed by an image processing system to extract model performance characteristic outputs for each training data pair in a plurality of training data pairs to generate a plurality of model performance representations.

[0099] At operation 1302, the image processing system acquires a training data set for training a machine learning model. The training data set includes a plurality of training data pairs, where each training data pair includes input data and corresponding ground truth data. The input data may include medical images, such as MR or CT scans, ultrasound images, or other types of medical imaging data. The ground truth data may include annotations or labels that indicate the presence or absence of certain features or conditions in the input data, such as the presence of a specific pathology, the location of certain anatomical landmarks, or the correct classification of the input data.

[0100] At operation 1304, during training of the machine learning model based on the training data set, the image processing system extracts model performance characteristic outputs for each of the plurality of training data pairs. The extraction process may involve capturing the outputs of the trained machine learning model and one or more intermediate outputs produced by one or more hidden layers of the trained machine learning model. The model performance characteristics may include, for example, confidence scores for the outputs of the trained machine learning model, and one or more uncertainty metrics for the outputs of the trained machine learning model. The confidence scores may indicate the degree of certainty with which the model makes its predictions, while the uncertainty metrics may provide a measure of the uncertainty or variability of the model's predictions.

[0101] At operation 1306, the image processing system encodes the plurality of model performance characterizations to generate a plurality of model performance characterization vectors. This encoding process may involve transforming the model performance characteristics into a format that can be easily compared and analyzed. For example, the model performance characteristics may be encoded as a feature vector, where each element of the vector corresponds to a different model performance characteristic.

[0102] At operation 1308, the image processing system determines distribution statistics for the plurality of model performance characterization vectors. These distribution statistics may include central tendency measures, such as mean or median, and dispersion measures, such as variance or standard deviation. The distribution statistics provide an overview of the overall performance of the machine learning model during its training phase.

[0103] At operation 1310, the image processing system stores the plurality of model performance characterization vectors and distribution statistics in a non-transitory memory. This stored information can later be used to monitor the performance of the machine learning model when the machine learning model is deployed in a real-world environment by comparing the performance of the model on new data with the performance characteristics observed during training. After operation 1310, method 1300 can end.

[0104] refer to Fig.14 , a method 1400 for extracting ground truth features from a training data set for training a machine learning model is shown. The method 1400 can be employed by an image processing system to extract ground truth features from ground truth data of each training data pair in a plurality of training data pairs to generate a plurality of ground truth data representations.

[0105] At operation 1402, the image processing system acquires a training data set for training a machine learning model. The training data set includes a plurality of training data pairs, where each training data pair includes input data and corresponding ground truth data. The ground truth data may include information such as the correct classification or output for the corresponding input data. The ground truth data is used as a reference for the machine learning model during the training process, so that the model can learn the correct mapping from the input data to the desired output.

[0106] At operation 1404, the image processing system extracts ground truth features from the ground truth data of each training data pair in the plurality of training data pairs to generate a plurality of ground truth data representations. Depending on the nature of the ground truth data, the extraction process may involve various techniques. For example, if the ground truth data includes labels for image classification, the ground truth features may include the distribution of the labels in the data set. If the ground truth data includes a bounding box for object detection, the ground truth features may include the size, shape, and position of the bounding box. If the ground truth data includes a segmentation mask for image segmentation, the ground truth features may include the area, perimeter, and centroid of the segmentation mask.

[0107] At operation 1406, the image processing system encodes the plurality of ground truth representations to generate a plurality of ground truth data representation vectors. The encoding process may involve transforming the ground truth characteristics into a numerical format that can be easily processed. For example, the ground truth characteristics may be normalized, scaled, or otherwise transformed to generate the ground truth data representation vectors.

[0108] At operation 1408, the image processing system determines distribution statistics of the plurality of ground truth data characterization vectors. These distribution statistics may include metrics such as mean, median, mode, variance, standard deviation, skewness, kurtosis, and other statistical metrics of the ground truth data characterization vectors. These distribution statistics provide an overview of the ground truth data characterization vectors, thereby enabling the image processing system to understand the overall characteristics of the ground truth data in the training data set.

[0109] At operation 1410, the image processing system stores the plurality of ground truth data characterization vectors and distribution statistics in a non-transitory memory. This stored information may be used later for various purposes, such as monitoring the performance of a trained machine learning model, detecting drift in input data, or adjusting the training process of a machine learning model. After operation 1410, method 1400 may end.

[0110] refer to Fig.15 , a method 1500 for accumulating input data representations, model performance representations, and user feedback representations from a particular deployment site is shown. The method 1500 can be employed by a model monitoring system to establish a baseline for various model performance metrics for a particular deployment site.

[0111] At operation 1502, the model monitoring system receives input data representations, model performance representations, and user feedback representations from a deployment site. The input data representations may include characteristics of the input data fed into the trained machine learning model, such as metadata from a medical image, pixel or voxel statistics from the medical image, and one or more labels from an appearance ontology and a clinical ontology of the medical image. The model performance representations may include characteristics of the performance of the trained machine learning model during mapping of the medical image to an output, such as an output of the trained machine learning model, one or more intermediate outputs produced by one or more hidden layers of the trained machine learning model, confidence scores for the outputs of the trained machine learning model, and one or more uncertainty metrics for the outputs of the trained machine learning model. The user feedback representations may include characteristics of user feedback received based on the outputs of the trained machine learning model, such as model output ratings and user corrections that modify the outputs of the trained machine learning model.

[0112] At operation 1504, the model monitoring system determines deviations for each representation based on a plurality of previously determined representations from the deployment site. The determination involves comparing the received representations to the previously determined representations to identify any deviations. The previously determined representations may be derived from a training data set used to train the trained machine learning model, or from previous inferences of the trained machine learning model at the deployment site.

[0113] At operation 1506, the model monitoring system checks whether any of the determined deviations exceed corresponding thresholds. If no deviations exceed their respective thresholds, the model monitoring system proceeds to operation 1510, where the received representation is added to a plurality of previously determined representations from the deployment site. This allows the model monitoring system to continuously update its understanding of the "normal" operation of the trained machine learning model and adapt its monitoring to changes in input data, model performance, or user feedback.

[0114] If at operation 1506, one or more of the deviations exceed their respective thresholds, the model monitoring system proceeds to operation 1508, where an alert is sent to the user device indicating that the deviation exceeds the threshold. The alert may include information about the nature of the deviation, the magnitude of the deviation, and potential corrective actions. The alert may be sent via various communication channels, such as email, text message, push notification, or other suitable communication methods. After operation 1510 or 1508, method 1500 may end.

[0115] The systems and methods disclosed above for monitoring the performance of trained machine learning models in a healthcare environment enable improved computational efficiency for real-time monitoring of deployed model performance. By extracting alternative data representations from input data, model performance, and user feedback, the system does not need to process and analyze all sensitive medical images and associated PHI. This approach reduces the computational load on the system because the representations are lightweight compared to the original data and can be processed quickly to detect performance deviations. The system's ability to encode these representations into feature vectors further improves computational efficiency for rapid comparison with a baseline. This streamlined process not only saves computational resources, but also enables real-time monitoring and rapid response to deviations, thereby ensuring that the machine learning model operates within predetermined performance parameters (e.g., within a predetermined deviation from the performance achieved during model training). The computational efficiency of the systems and methods disclosed above may be particularly advantageous in healthcare environments, where timely and accurate model performance is extremely important and computational resources may be limited.

[0116] The present disclosure also provides support for a method comprising: extracting characteristics of input data fed into a trained machine learning model to produce an input data representation, wherein the input data comprises a medical image; extracting characteristics of the performance of the trained machine learning model during mapping the medical image to an output to produce a model performance representation; extracting characteristics of user feedback received based on the output of the trained machine learning model to produce a user feedback representation; determining an input data deviation by comparing the input data representation with a plurality of previously determined input data representations; determining a model performance deviation by comparing the model performance representation with a plurality of previously determined model performance representations; determining a user feedback deviation by comparing the user feedback representation with a plurality of previously determined user feedback representations; and responding to one or more of an input data deviation exceeding an input data deviation threshold, a model performance deviation exceeding a model performance deviation threshold, and a user feedback deviation exceeding a user feedback deviation threshold by sending an alert to a user device. In a first example of the method, extracting characteristics of input data fed into a trained machine learning model to produce an input data representation includes: extracting metadata from a medical image, wherein the metadata does not include personally identifiable information; aggregating pixel or voxel statistics from the medical image, wherein the pixel or voxel statistics do not retain individual pixel or voxel intensity values ​​or positions; determining one or more labels from an appearance ontology of the medical image; and determining one or more labels from a clinical ontology of the medical image. In a second example of the method, optionally including the first example, determining input data deviation by comparing the input data representation with a plurality of previously determined input data representations includes: encoding the metadata, the pixel or voxel statistics, the one or more labels from the appearance ontology, and the one or more labels from the clinical ontology into a feature vector, and comparing the feature vector to a plurality of predetermined feature vectors corresponding to the plurality of previously determined input data representations. In a third example of the method, optionally including one or both of the first and second examples, extracting characteristics of the performance of a trained machine learning model during mapping a medical image to an output to produce a model performance representation includes: capturing the output of the trained machine learning model and one or more intermediate outputs produced by one or more hidden layers of the trained machine learning model; determining a confidence score for the output of the trained machine learning model; and determining one or more uncertainty metrics for the output of the trained machine learning model. In a fourth example of the method, optionally including one or more or each of the first to third examples, determining a model performance deviation by comparing the model performance representation with a plurality of previously determined model performance representations includes: encoding the output of the trained machine learning model and one or more intermediate outputs, confidence scores, and one or more uncertainty metrics into a feature vector, and comparing the feature vector with a plurality of predetermined feature vectors corresponding to a plurality of previously determined model performance representations.In a fifth example of the method, optionally including one or more or each of the first to fourth examples, extracting characteristics of user feedback received based on the output of the trained machine learning model to generate a user feedback representation includes: recording one or more of the following: a model output rating received via a user input device and a user correction received via the user input device, wherein the user correction modifies the output of the trained machine learning model. In a sixth example of the method, optionally including one or more or each of the first to fifth examples, extracting characteristics of user feedback received based on the output of the trained machine learning model to generate a user feedback representation also includes: receiving comments from the user via a user input device, and determining an emotional score of the output of the trained machine learning model based on the comments. In a seventh example of the method, optionally including one or more or each of the first to sixth examples, determining a user feedback deviation by comparing the user feedback representation with a plurality of previously determined user feedback representations includes: encoding the model output rating and the user correction into a feature vector, and comparing the feature vector with a plurality of predetermined feature vectors corresponding to a plurality of previously determined user feedback representations.

[0117] The present disclosure also provides support for a system comprising: a first device located at a deployment site, wherein the first device comprises a user input device, a first non-volatile memory comprising a trained machine learning model and instructions, and a first processor, wherein when executing the instructions, the first processor causes the first device to: extract characteristics of input data fed into the trained machine learning model to generate an input data representation, wherein the input data comprises a medical image; extract characteristics of the performance of the trained machine learning model during mapping the medical image to an output to generate a model performance representation; extract characteristics of user feedback received via the user input device based on the output of the trained machine learning model to generate a user feedback representation; and send the input data representation, the model performance representation, and the user feedback representation to a second device, the second device being located away from the first device, wherein the first processor causes the first device to: extract characteristics of input data fed into the trained machine learning model to generate an input data representation, wherein the input data comprises a medical image; extract characteristics of the performance of the trained machine learning model during mapping the medical image to an output to generate a model performance representation; extract characteristics of user feedback received via the user input device based on the output of the trained machine learning model to generate a user feedback representation; and send the input data representation, the model performance representation, and the user feedback representation to a second device, the second device being located away from the first device, A device is communicatively coupled to a second device, and wherein the second device includes a second non-transitory memory including instructions, and a second processor, wherein when the instructions are executed, the second processor causes the second device to: receive input data representations, model performance representations, and user feedback representations from the first device; determine input data deviations by comparing the input data representations with a plurality of previously determined input data representations; determine model performance deviations by comparing the model performance representations with a plurality of previously determined model performance representations; determine user feedback deviations by comparing the user feedback representations with a plurality of previously determined user feedback representations; and respond to one or more of the input data deviation exceeding an input data deviation threshold, the model performance deviation exceeding a model performance deviation threshold, and the user feedback deviation exceeding a user feedback deviation threshold by sending an alert to a user device. In a first example of the system, the plurality of previously determined input data representations, the plurality of previously determined model performance representations, and the plurality of previously determined user feedback representations are derived from a training data set used to train a trained machine learning model. In a second example of the system, optionally including the first example, the plurality of previously determined input data representations, the plurality of previously determined model performance representations, and the plurality of previously determined user feedback representations are derived from previous inferences of a trained machine learning model at a deployment site. In a third example of the system, optionally including one or both of the first and second examples, the input data representation, the model performance representation, and the user feedback representation each do not include a medical image, and wherein when the instruction is executed, the first processor does not send the medical image to the second device. In a fourth example of the system, optionally including one or more or each of the first to third examples, the alert includes an indication that the input data deviation exceeds an input data deviation threshold, the model performance deviation exceeds a model performance deviation threshold, and the user feedback deviation exceeds a user feedback deviation threshold.

[0118] The present disclosure also provides support for a method for monitoring the performance of a trained machine learning model, the method comprising: responding to a medical image of an imaged subject being input into the trained machine learning model by the following steps: extracting characteristics of the medical image of the imaged subject to produce an input data representation, wherein the input data representation does not include individual pixel or voxel intensity values ​​and individual pixel or voxel positions; extracting characteristics of the performance of the trained machine learning model during mapping the medical image to an output to produce a model performance representation; determining an input data deviation by comparing the input data representation with a plurality of previously determined input data representations; determining a model performance deviation by comparing the model performance representation with a plurality of previously determined model performance representations; and responding to one or more of an input data deviation exceeding an input data deviation threshold and a model performance deviation exceeding a model performance deviation threshold by sending an alert to a user device. In a first example of the method, the method further comprises: extracting characteristics of user feedback received based on the output of the trained machine learning model to produce a user feedback representation; and determining a user feedback deviation by comparing the user feedback representation with a plurality of previously determined user feedback representations; and responding to a user feedback deviation exceeding a user feedback deviation threshold by sending an alert to a user device. In a second example of the method, optionally including the first example, the user feedback includes a scan plane for acquiring a diagnostic medical image of an imaging subject. In a third example of the method, optionally including one or both of the first and second examples, extracting characteristics of user feedback received based on the output of the trained machine learning model to generate a user feedback representation includes: determining the center coordinates of the scan plane, and determining three direction cosines that uniquely identify the orientation of the scan plane. In a fourth example of the method, optionally including one or more or each of the first to third examples, the medical image includes a three-dimensional (3D) medical image, and wherein the trained machine learning model is configured to identify anatomical landmarks in the 3D medical image for positioning the scan plane for acquiring a diagnostic medical image of a region of interest. In a fifth example of the method, optionally including one or more or each of the first to fourth examples, extracting characteristics of a medical image of an imaged subject to produce an input data representation includes: capturing metadata of the 3D medical image, including pixel spacing and slice thickness of the 3D medical image; determining aggregate intensity statistics of the 3D medical image, including an intensity histogram; and determining a clinical ontology of the 3D medical image, including a list of anatomical regions captured in the 3D medical image.In a sixth example of the method, optionally including one or more or each of the first to fifth examples, extracting characteristics of the performance of the trained machine learning model during mapping a medical image to an output to produce a model performance characterization includes: extracting properties of one or more segmentation masks generated by the trained machine learning model that identify anatomical landmarks in the 3D medical image; extracting properties of scanning planes determined based on the anatomical landmarks; and recording a list of anatomical landmarks identified in the 3D medical image by the trained machine learning model.

[0119] When introducing the elements of various embodiments of the present disclosure, the articles "one", "a kind of" and "the" are intended to mean that there are one or more such elements. The terms "first", "second", etc. do not represent any order, amount or importance, but are used to distinguish one element from another element. The terms "include", "comprises", and "have" are intended to be inclusive, and mean that additional elements may also exist in addition to the listed elements. As used herein, the terms "connected to", "coupled to", etc., an object (e.g., material, element, structure, member, etc.) may be connected to or coupled to another object, regardless of whether the one object is directly connected or coupled to another object, or whether there are one or more intervening objects between the one object and another object. In addition, it should be understood that reference to "an embodiment" or "embodiment" of the present disclosure is not intended to be interpreted as excluding the existence of additional embodiments that are also combined with the cited features.

[0120] In addition to any previously indicated modifications, those skilled in the art may devise numerous other variations and alternative arrangements without departing from the spirit and scope of the present specification, and the appended claims are intended to cover such modifications and arrangements. Therefore, although the information has been described in detail and in detail as described above in conjunction with what are currently considered to be the most practical and preferred aspects, it will be apparent to those of ordinary skill in the art that many modifications, including but not limited to form, function, mode of operation, and use, may be made without departing from the principles and concepts set forth herein. Likewise, as used herein, the examples and embodiments are intended to be illustrative in all respects only and should not be construed as limiting in any way.

Claims

1. A method (200), comprising: extracting characteristics of input data fed into a trained machine learning model (108) to produce an input data representation (202), wherein the input data comprises a medical image; extracting characteristics of the performance of the trained machine learning model (108) in mapping the medical image to an output to produce a model performance representation (204); extracting characteristics of user feedback received based on the output of the trained machine learning model (108) to produce a user feedback representation (206); determining an input data deviation by comparing the input data representation to a plurality of previously determined input data representations (500); determining a model performance deviation by comparing the model performance characterization to a plurality of previously determined model performance characterizations (800); determining a user feedback deviation by comparing the user feedback representation to a plurality of previously determined user feedback representations (1100); and Responding to one or more of the input data deviation exceeding an input data deviation threshold, the model performance deviation exceeding a model performance deviation threshold, and the user feedback deviation exceeding a user feedback deviation threshold is performed by sending an alert to a user device (140).

2. The method (200) of claim 1, wherein extracting characteristics of input data fed into the trained machine learning model (108) to generate the input data representation (202) comprises: extracting metadata from the medical image (302), wherein the metadata does not include personally identifiable information; aggregating pixel or voxel statistics (304) from the medical image, wherein the pixel or voxel statistics do not retain individual pixel or voxel intensity values ​​or positions; determining one or more labels from an appearance ontology of the medical image (306); as well as One or more labels from a clinical ontology of the medical image are determined (308).

3. The method (200) of claim 2, wherein determining the input data deviation (500) by comparing the input data representation with the plurality of previously determined input data representations comprises: encoding the metadata, the pixel or voxel statistics, the one or more labels from the appearance ontology, and the one or more labels from the clinical ontology into a feature vector (502); and The feature vector is compared to a plurality of predetermined feature vectors corresponding to the plurality of previously determined input data representations (504).

4. The method (200) of claim 1, wherein extracting characteristics of the performance of the trained machine learning model (108) in mapping the medical image to the output to produce the model performance characterization (204) comprises: capturing the output of the trained machine learning model (108) and one or more intermediate outputs (602) produced by one or more hidden layers of the trained machine learning model (108); determining a confidence score (604) for the output of the trained machine learning model (108); and One or more uncertainty metrics (606) for the output of the trained machine learning model (108) are determined.

5. The method (200) of claim 4, wherein determining the model performance deviation (800) by comparing the model performance characterization with the plurality of previously determined model performance characterizations comprises: encoding the output of the trained machine learning model (108) and the one or more intermediate outputs, the confidence score, and the one or more uncertainty metrics into a feature vector (802); and The feature vector is compared to a plurality of predetermined feature vectors corresponding to the plurality of previously determined model performance characterizations (804).

6. The method (200) of claim 1, wherein extracting characteristics of user feedback received based on the output of the trained machine learning model (108) to generate the user feedback representation (206) comprises: Record one or more of the following: a model output rating (902) received via a user input device (112, 134); and User corrections (904) received via the user input device (112, 134), wherein the user corrections modify the output of the trained machine learning model (108).

7. The method (200) of claim 6, wherein extracting characteristics of user feedback received based on the output of the trained machine learning model (108) to generate the user feedback representation (206) further comprises: receiving comments from a user via the user input device (112, 134); as well as A sentiment score for the output of the trained machine learning model (108) is determined based on the reviews.

8. The method (200) of claim 6, wherein determining the user feedback deviation (1100) by comparing the user feedback representation with the plurality of previously determined user feedback representations comprises: encoding the model output rating and the user correction into a feature vector (1102); as well as The feature vector is compared to a plurality of predetermined feature vectors corresponding to the plurality of previously determined user feedback representations (1104).

9. A system (100), comprising: A first device (102), the first device being located at a deployment site, wherein the first device (102) comprises: User input device (112); a first non-transitory memory (106) comprising a trained machine learning model (108) and instructions; and A first processor (104), wherein when executing the instructions, the first processor (104) causes the first device (102): extracting characteristics of input data fed into the trained machine learning model (108) to produce an input data representation, wherein the input data comprises a medical image; extracting characteristics of the performance of the trained machine learning model (108) in mapping the medical image to an output to produce a model performance representation; extracting characteristics of user feedback received via the user input device (112) based on the output of the trained machine learning model (108) to produce a user feedback representation; and sending the input data representation, the model performance representation, and the user feedback representation to a second device (122); The second device (122) is located remotely from the first device (102), wherein the first device (102) and the second device (122) are communicatively coupled, and wherein the second device (122) includes: a second non-transitory memory (126), the second non-transitory memory comprising instructions; and a second processor (124), wherein when executing the instructions, the second processor (124) causes the second device (122): receiving the input data representation, the model performance representation, and the user feedback representation from the first device (102); determining an input data deviation by comparing the input data representation to a plurality of previously determined input data representations; determining a model performance deviation by comparing the model performance representation to a plurality of previously determined model performance representations; determining a user feedback deviation by comparing the user feedback representation to a plurality of previously determined user feedback representations; and Responding to one or more of the input data deviation exceeding an input data deviation threshold, the model performance deviation exceeding a model performance deviation threshold, and the user feedback deviation exceeding a user feedback deviation threshold is performed by sending an alert to a user device (140).

10. The system (100) of claim 9, wherein the plurality of previously determined input data representations, the plurality of previously determined model performance representations, and the plurality of previously determined user feedback representations are derived from a training dataset used to train the trained machine learning model (108).

11. The system (100) of claim 9, wherein the plurality of previously determined input data representations, the plurality of previously determined model performance representations, and the plurality of previously determined user feedback representations are derived from previous inferences of the trained machine learning model (108) at the deployment site.

12. The system (100) of claim 9, wherein the input data representation, the model performance representation, and the user feedback representation each do not include the medical image, and wherein when executing the instructions, the first processor (104) does not send the medical image to the second device (122).

13. The system (100) of claim 9, wherein the alert comprises an indication of one or more of the input data deviation exceeding the input data deviation threshold, the model performance deviation exceeding the model performance deviation threshold, and the user feedback deviation exceeding the user feedback deviation threshold.

14. The system (100) of claim 9, wherein when executing the instructions, the first processor (104) is configured to extract characteristics of the input data fed into the trained machine learning model (108) to generate the input data representation by: extracting metadata from the medical image, wherein the metadata does not include personally identifiable information; aggregating pixel or voxel statistics from the medical image, wherein the pixel or voxel statistics do not retain individual pixel or voxel intensity values ​​or positions; determining one or more labels from an appearance ontology of the medical image; as well as One or more labels from a clinical ontology of the medical image are determined.

15. The system (100) of claim 14, wherein when executing the instructions, the first processor (104) is configured to determine the input data deviation by comparing the input data representation with the plurality of previously determined input data representations by: encoding the metadata, the pixel or voxel statistics, the one or more labels from the appearance ontology, and the one or more labels from the clinical ontology into a feature vector; and The feature vector is compared to a plurality of predetermined feature vectors corresponding to the plurality of previously determined input data representations.