Method and apparatus for analyzing performance of machine learning model

A computing device monitors and optimizes machine learning models in medical imaging by detecting drift and adjusting thresholds, ensuring reliable and fair performance.

WO2026095394A1PCT designated stage Publication Date: 2026-05-07LUNIT
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
LUNIT
Filing Date
2025-10-01
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

The performance of machine learning models used in medical imaging varies due to factors like patient population distribution, imaging equipment diversity, and model updates, necessitating continuous monitoring and evaluation to ensure accurate and reliable diagnostic results.

Method used

A computing device analyzes performance changes of machine learning models over time, detects drift, and outputs information for optimization, including real-time threshold adjustments and retraining recommendations to maintain accuracy and fairness.

Benefits of technology

Ensures the reliability and fairness of medical information by continuously optimizing the performance of machine learning models, addressing model drift and regulatory compliance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025015661_07052026_PF_FP_ABST
    Figure KR2025015661_07052026_PF_FP_ABST
Patent Text Reader

Abstract

A computing device, according to one aspect, comprises: at least one memory in which at least one instruction is stored; and at least one processor operating according to the at least one instruction, wherein the at least one processor: receives performance evaluation data including prediction data generated as at least one medical image is analyzed by a machine learning model; calculates at least one metric indicating performance of the machine learning model on the basis of the performance evaluation data; analyzes a performance change of the machine learning model over time on the basis of the performance evaluation data and the calculation result; and outputs information related to the machine learning model on the basis of the analysis result.
Need to check novelty before this filing date? Find Prior Art

Description

Method and device for analyzing the performance of machine learning models

[0001] The present disclosure relates to a method and apparatus for analyzing the performance of a machine learning model. Specifically, the present disclosure relates to a method and apparatus for monitoring the performance of a model across multiple medical institutions while preserving privacy. Furthermore, the present disclosure relates to an adaptive drift detection and predictive performance management system utilizing medical images.

[0002] Recently, technologies are being developed to predict medical information about subjects by analyzing medical images through machine learning models. Representative examples include machine learning models that diagnose patient diseases (e.g., cancer) by analyzing medical images.

[0003] However, the performance of machine learning models can vary depending on factors such as the distribution of patient populations by medical institution, the diversity of imaging equipment, the quality of medical images, and updates to the machine learning model. Therefore, to derive accurate and reliable diagnostic results, it is necessary to develop technology capable of continuously monitoring and evaluating the performance of machine learning models by reflecting the unique characteristics of medical images and actual clinical requirements.

[0004] The present disclosure provides a method and apparatus for monitoring, evaluating, and optimizing the performance of a machine learning model that analyzes medical images. Additionally, the present disclosure provides a computer-readable recording medium storing a program for executing the above-described method on a computer.

[0005] Furthermore, the present disclosure is not limited to merely monitoring the performance of machine learning models. For example, the present disclosure relates to techniques for managing machine learning models, including detecting changes in data distribution, evaluating fairness, and autonomous optimization. In particular, medical imaging-based machine learning models can be combined with post-market surveillance requirements in regulatory environments, such as clinical safety and regulatory requirements.

[0006] The technical challenges to be solved are not limited to those mentioned above, and other technical challenges may exist.

[0007] A computing device according to one aspect comprises: at least one memory in which at least one instruction is stored; and at least one processor that operates according to the at least one instruction. The at least one processor receives performance evaluation data including prediction data generated as at least one medical image is analyzed by a machine learning model, calculates at least one indicator representing the performance of the machine learning model based on the performance evaluation data, analyzes the performance change of the machine learning model over time based on the performance evaluation data and the calculation result, and outputs information related to the machine learning model based on the analysis result.

[0008] A method for analyzing the performance of a machine learning model according to other aspects comprises: receiving performance evaluation data including prediction data generated as at least one medical image is analyzed by the machine learning model; calculating at least one indicator representing the performance of the machine learning model based on the performance evaluation data; analyzing the change in performance of the machine learning model over time based on the performance evaluation data and the calculation result; and outputting information related to the machine learning model based on the analysis result.

[0009] A computer-readable recording medium according to another aspect includes a recording medium that records a program for executing the above-described method on a computer.

[0010] FIG. 1 is a diagram illustrating an example in which a medical image is analyzed based on a machine learning model according to one embodiment.

[0011] FIG. 2a is a configuration diagram illustrating an example of a user terminal according to one embodiment.

[0012] FIG. 2b is a configuration diagram illustrating an example of a server according to one embodiment.

[0013] FIG. 3 is a flowchart illustrating an example of a method for analyzing the performance of a machine learning model according to one embodiment.

[0014] FIG. 4 is a diagram illustrating an example of performance evaluation data received by a computing device according to one embodiment.

[0015] FIG. 5 is a diagram illustrating an example of a processor calculating performance indicators according to one embodiment.

[0016] FIG. 6 is a drawing for explaining an example of output to a display device according to one embodiment.

[0017] FIG. 7 is a diagram illustrating an example of a reference distribution and a first distribution according to one embodiment.

[0018] FIG. 8 is a flowchart illustrating an example of operation in which a processor according to one embodiment operates when a change in the performance of a machine learning model is confirmed.

[0019] FIGS. 9 to 16 are drawings illustrating examples in which information related to a machine learning model according to one embodiment is output.

[0020] A computing device according to one aspect comprises: at least one memory in which at least one instruction is stored; and at least one processor that operates according to the at least one instruction. The at least one processor receives performance evaluation data including prediction data generated as at least one medical image is analyzed by a machine learning model, calculates at least one indicator representing the performance of the machine learning model based on the performance evaluation data, analyzes the performance change of the machine learning model over time based on the performance evaluation data and the calculation result, and outputs information related to the machine learning model based on the analysis result.

[0021] The terms used in the embodiments have been selected to be as close as possible to currently widely used general terms; however, these may vary depending on the intent of those skilled in the art, case law, the emergence of new technologies, etc. Additionally, in specific cases, terms have been selected at the applicant's discretion, and in such cases, their meanings will be described in detail in the relevant description section. Therefore, terms used in the specification must be defined not merely by their names, but based on their meanings and the content throughout the specification.

[0022] When a part of the specification is described as "comprising" a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components. Furthermore, terms such as "unit" and "module" as used in the specification refer to a unit that performs at least one function or operation, and this may be implemented in hardware or software, or as a combination of hardware and software.

[0023] Additionally, terms including ordinal numbers, such as "first" or "second," used in the specification may be used to describe various components, but said components shall not be limited by said terms. Such terms may be used for the purpose of distinguishing one component from another.

[0024] In the following, "medical information" may refer to any medically meaningful information or clinical information of a patient that can be extracted from medical images. Medical images may include not only pathology slide images but also radiographic images (X-ray, CT, MRI, PET, etc.). For example, medical information may include at least one of an immune phenotype, genotype, expression type, biomarker, tumor purity, information regarding RNA, tumor microenvironment, cancer regimen expressed in the pathology slide image, survival information, treatment response, treatment outcome, genetic characteristics, and medical records.

[0025] In addition, medical information may also include anatomical structural information extracted from medical images, types of lesions, locations and sizes of lesions, morphological features of lesions (e.g., boundaries, texture, density), functional indicators (e.g., blood flow, metabolic activity), abnormal findings of organs, indicators related to treatment prognosis obtained from medical images, information regarding findings obtained by analyzing medical images using artificial intelligence models, abnormality scores of said findings, reliability of said findings, and image biomarkers (radiomic features).

[0026] In addition, medical information may include findings such as the presence or absence of nodules in the medical image, signs of pneumonia, the presence of pneumothorax, the location and type of fractures, the location, size, shape, and boundary characteristics of masses, the distribution of microcalcifications, asymmetry, and breast tissue density and structural distortion. Such findings may be calculated along with, but are not limited to, an abnormality score or risk score for the relevant image.

[0027] Additionally, medical information may include, but is not limited to, the area, location, and size of specific tissues (e.g., cancer tissue, cancer stromal tissue, etc.) and / or specific cells (e.g., tumor cells, lymphocytes, macrophages, endothelial cells, fibroblasts, etc.) within the medical image, diagnostic information of cancer, information related to the patient's probability of developing cancer, and / or medical conclusions related to cancer treatment.

[0028] In addition, medical information may include not only quantified values ​​obtainable from medical images but also information visualizing the values, predictive information based on the values, image information, statistical information, etc. For example, medical information may be provided to a user terminal or output through a display device.

[0029] Embodiments are described in detail below with reference to the attached drawings. However, embodiments may be implemented in various different forms and are not limited to the examples described herein.

[0030] FIG. 1 is a diagram illustrating an example in which a medical image is analyzed based on a machine learning model according to one embodiment.

[0031] Referring to FIG. 1, a computing device (20) can analyze a medical image (10) and output medical information (30). The medical image (10) may be an image of various modalities. For example, the medical image (10) may include a pathology slide image, a CT image, an X-ray image, a mammography image, an MRI image, a PET image, etc., but is not limited thereto.

[0032] A computing device (20) can generate medical information (30) by analyzing a medical image (10) using a machine learning model. For example, the machine learning model may be a model that analyzes a medical image (or, region of interest) and outputs medical information (30) about an object included in the medical image (or, region of interest).

[0033] A machine learning model refers to a statistical learning algorithm implemented based on the structure of a biological neural network, or a structure that executes such an algorithm. For example, a machine learning model may represent a model capable of problem-solving, in which nodes—artificial neurons that form a network through synaptic connections as in biological neural networks—learn by repeatedly adjusting the weights of the synapses to reduce the error between the correct output corresponding to a specific input and the inferred output. For example, a machine learning model may include arbitrary probability models, neural network models, etc., used in artificial intelligence learning methods such as deep learning.

[0034] For example, a machine learning model can be implemented as a multilayer perceptron (MLP) composed of multiple layers of nodes and connections between them. The machine learning model according to the present embodiment can be implemented using one of various artificial neural network model structures including an MLP. For example, the machine learning model may be composed of an input layer that receives an input signal or data from the outside, an output layer that outputs an output signal or data corresponding to the input data, and at least one hidden layer located between the input layer and the output layer, which receives a signal from the input layer, extracts a feature, and transmits it to the output layer. The output layer receives a signal or data from the hidden layer and outputs it to the outside.

[0035] Accordingly, the machine learning model can be trained to extract medical information (30) about one or more objects (e.g., cells, tissues, structures, etc.) included in the medical image.

[0036] Meanwhile, the performance of machine learning models can vary due to various factors. The phenomenon where the predictive performance of a machine learning model gradually declines in a real-world operating environment is called model drift. Types of model drift include data drift, where the distribution of input data changes, and concept drift, where the relationship between input and output changes.

[0037] Concept drift refers to a change in the relationship between input variables and output variables (targets), signifying a situation where the previously learned patterns of a machine learning model are no longer valid.

[0038] Data drift refers to a phenomenon in which the predictive performance of a machine learning model deteriorates over time due to changes in the statistical attributes of input data (i.e., medical images and metadata analyzed by the machine learning model). This involves changes in the probability distribution of input variables and can be caused, for example, by (i) fluctuations in demographic distributions (age, gender, race, etc.), (ii) replacement or updates of imaging equipment (modality, vendor, software version) or changes in acquisition parameters, (iii) revisions of imaging protocols, clinical guidelines, and diagnostic criteria, and (iv) regional / temporal shifts and seasonal changes in disease prevalence and severity distributions. Due to the nature of medical imaging machine learning models, these changes directly induce deviations in the distribution of outlier scores and the prediction-ground truth relationship, thereby affecting clinical reliability and safety.

[0039] Accordingly, the present invention compares the distribution of the detection period relative to the reference period and determines the presence of drift through multiple statistical tests, and ensures fairness in the comparison through up / down sampling when the difference in sample sizes between the reference and detection periods is large. Conventionally, there were limitations in maintaining the reliability of the output of machine learning models due to the lack of technology to monitor and evaluate the performance of machine learning models.

[0040] A computing device (20) according to one embodiment analyzes performance changes of a machine learning model that analyzes medical images in a time-series manner and evaluates and outputs information related to the machine learning model based on this. That is, if the performance of the machine learning model deteriorates, the computing device (20) automatically provides a report and alert to the user. In addition, it outputs model-related information so that the computing device (20) can continuously optimize the performance of the machine learning model. When drift is detected, the computing device (20) identifies causal variables (age, gender, vendor, acquisition parameters, etc.) and performs an operational threshold reset simulation and verification pipeline to recommend a retraining and verification process for the machine learning model. Through this, the performance of the machine learning model can be improved, and the accuracy, reliability, and fairness of medical information (30) can be ensured in clinical operations.

[0041] Additionally, the computing device (20) performs a simulation based on a change in a threshold related to the performance of a machine learning model. For example, the computing device (20) provides (i) a target sensitivity fixed mode and (ii) a target specificity fixed mode depending on the purpose of operation.

[0042] The target sensitivity fixed mode can be used in scenarios where minimizing false negatives (FN) is a priority, such as screening. For example, the computing device (20) can explore a threshold to satisfy the user (or pre-defined) target sensitivity (e.g., 95%).

[0043] The target specificity fixed mode can be used in diagnostic scenarios where reducing the reading burden and minimizing unnecessary re-examinations are important, for example, the computing device (20) can calculate a threshold that satisfies the target specificity (e.g., 90%).

[0044] In the target sensitivity fixed mode and target specificity fixed mode, threshold-based sensitivity, specificity, PPV, NPV, F1 Score, and AUC are immediately updated and displayed, thereby providing reliability. For example, in the case of breast cancer screening triage, with specificity fixed at 90%, users can monitor changes in sensitivity, PPV, and NPV and select a threshold that aligns with the clinician's operational objectives. Consequently, users can monitor the performance of the machine learning model in real time and adjust it according to clinical requirements in the field of radiology. Specifically, the performance of the machine learning model can be tuned and optimized through threshold simulation, and changes in the input data distribution can be monitored in real time through the detection of data drift.

[0045] In addition, the computing device (20) can analyze the cause of data drift, and accordingly, the performance of the machine learning model for the analysis of medical images (10) can be continuously improved. In addition, regulatory compliance in the medical field becomes easier, and clinical reliability can be improved.

[0046] Hereinafter, with reference to FIGS. 2a to 16, examples of analyzing the performance of a machine learning model in which a computing device (20) analyzes input data (10) will be described.

[0047] For example, the computing device (20) may be a user terminal or a server. In other words, the operations performed by the computing device (20) may be performed by a user terminal or a server. Alternatively, some of the operations performed by the computing device (20) may be performed by a user terminal, and the remainder may be performed by a server.

[0048] A user terminal may be an electronic device comprising a display device and a device for receiving user input (e.g., a keyboard, a mouse, etc.), and including memory and a processor. Additionally, the display device may be implemented as a touch screen to perform the function of receiving user input. For example, the user terminal may include, but is not limited to, notebook PCs, desktop PCs, laptops, tablet computers, smartphones, etc.

[0049] A server may be a device that communicates with external devices (e.g., user terminals). For example, a server may be a device that stores various data, including medical information and information about machine learning models. Alternatively, a server may be an electronic device that includes memory and a processor and possesses its own computing capabilities. For example, a server may be a cloud server or an on-premise server.

[0050] Hereinafter, examples of a user terminal and a server will be described with reference to FIGS. 2a and 2b.

[0051] FIG. 2a is a configuration diagram illustrating an example of a user terminal according to one embodiment.

[0052] Referring to FIG. 2a, the user terminal (100) includes a processor (110), memory (120), an input / output interface (130), and a communication module (140). For convenience of explanation, FIG. 2a only illustrates components related to the present invention. Accordingly, other general-purpose components may be included in the user terminal (100) in addition to the components illustrated in FIG. 2a. Furthermore, it is obvious to those skilled in the art that the processor (110), memory (120), input / output interface (130), and communication module (140) illustrated in FIG. 2a may be implemented as independent devices.

[0053] The processor (110) can process instructions of a computer program by performing basic arithmetic, logic, and input / output operations. Here, instructions may be provided from memory (120) or an external device (e.g., a server (200), etc.). Additionally, the processor (110) can control the overall operation of other components included in the user terminal (100).

[0054] The processor (110) receives performance evaluation data including prediction data generated as at least one medical image is analyzed by a machine learning model.

[0055] For example, the data for performance evaluation may further include, in addition to the prediction data, at least one of at least one medical image, electronic medical record (EMR) data corresponding to the medical image, or metadata corresponding to the medical image.

[0056] Metadata may include personal information about the patient (e.g., age, gender, ethnicity, etc.), characteristics of the patient's disease (e.g., clinical condition, prognostic indicators, etc.), conditions for acquiring the medical image (e.g., parameters related to image acquisition, orientation of the image, date of the image, etc.), and information about the equipment used to acquire the medical image (modality, vendor, software version). For example, metadata may be stored based on DICOM tags.

[0057] Predictive data may include various outputs generated by a machine learning model analyzing medical images. For example, predictive data may include, but is not limited to, clinical indicators such as an abnormality score, the location of a lesion, the type of lesion, and whether a target lesion (e.g., cancer, nodule, etc.) is included in the medical image. The abnormality score may be represented as the probability of abnormal findings or lesions existing in a medical image, which is output by the machine learning model analyzing the medical image (e.g., the probability value may be output as a continuous value between 0 and 1 or a continuous value between 0 and 100, but is not limited thereto).

[0058] Additionally, the predictive data may include at least one of an immune phenotype, genotype, expression type, biomarker, tumor purity, RNA information, tumor microenvironment, cancer regimen, survival information, treatment response, treatment outcome, and genetic characteristics derived from medical imaging.

[0059] An example of the processor (110) receiving data for performance evaluation is described later with reference to step 310.

[0060] The processor (110) calculates at least one metric representing the performance of a machine learning model based on data for performance evaluation. For example, at least one metric representing the performance of a machine learning model may include sensitivity, specificity, accuracy, positive predictive value (PPV), negative predictive value (NPV), F1 score, ROC (Receiver Operating Characteristic) curve and AUC (Area Under Curve), or PR (Precision-Recall) curve and AUC, etc.

[0061] For example, the processor (110) can compute an indicator based on prediction data and ground truth data corresponding to at least one medical image. Additionally, when the processor (110) receives user input, it can recompute the indicator according to said input. For example, the user input may be an input that changes a threshold related to the performance of a machine learning model.

[0062] Correct answer data may be obtained based on a physician's diagnostic report, diagnostic information obtained from the patient's tissue (biopsy specimen) or blood sample, etc. More specifically, correct answer data may include at least one of pathological confirmation, cytology results, molecular diagnostic results, genetic testing results, radiology reports, clinical diagnosis reports, the patient's Electronic Medical Record (EMR), outcome data obtained through follow-up, or a combination thereof. Additionally, correct answer data may be provided by a single medical professional or verified by multiple experts (e.g., consensus reading).

[0063] For example, the processor (110) can obtain correct data from medical data included in the performance evaluation data and, by comparing the correct data with the predicted data, calculate at least one indicator representing the performance of the machine learning model that outputs the predicted data. For example, the processor (110) can obtain as correct data the result of confirming that the lesion is cancer through a pathological examination of the tissue obtained from the patient's lesion.

[0064] For example, the processor (110) may obtain correct data corresponding to at least one medical image from an external server, map it to prediction data of a machine learning model for the image, and store it in an internal memory (120) or an external server. Afterward, the processor (110) may retrieve the stored correct data and compare it with the prediction data to calculate at least one indicator representing the performance of the machine learning model that output the prediction data. For example, the processor (110) may obtain the diagnosis results of a medical professional regarding a patient from the patient's electronic medical records included in the PACS (Picture Archiving and Communication System) server as correct data.

[0065] For example, when the processor (110) is requested to perform an metric operation representing the performance of a machine learning model (e.g., when receiving a user request input or reaching a predetermined period), it may access internal memory (120) or external storage to obtain the patient's medical data and derive correct answer data therefrom. For example, the processor (110) may use a large-scale language model (LLM) to extract correct answer data from the patient's medical data. At this time, the source document, prompt, model / version, confidence score, and human verification status may be recorded as metadata to ensure traceability and regulatory compliance.

[0066] An example of the processor (110) calculating at least one indicator is described later with reference to step 320.

[0067] The processor (110) can analyze changes in the performance of the machine learning model based on the result of the metric calculation.

[0068] As an example, if the difference between the result of the metric calculation and a preset threshold metric value exceeds a predetermined allowable range, the processor (110) may determine that the performance of the machine learning model has deteriorated.

[0069] Additionally, the processor (110) can analyze changes in the performance of a machine learning model over time by considering the performance evaluation data and the results of the metric calculation together. In this process, the processor (110) is not limited to simply determining whether there is a change in performance, but can also determine whether the performance change is due to a specific data drift.

[0070] For example, the processor (110) may obtain, as prediction data, the result of a machine learning model recently analyzing multiple medical images to determine whether each medical image is positive or negative. The processor (110) may calculate an indicator by comparing the prediction data with the correct answer data. Based on the result of the indicator calculation, the processor (110) may determine whether there is a change in the performance of the machine learning model and analyze the change in the performance of the machine learning model by comparing the first distribution of anomaly scores with the reference distribution. Here, the reference distribution may be generated based on anomaly scores obtained by analyzing multiple medical images during a reference period. Additionally, the first distribution may be generated based on anomaly scores during a detection period after the reference period.

[0071] An example of the processor (110) analyzing the performance change of a machine learning model is described later with reference to step 330.

[0072] The processor (110) outputs information related to the machine learning model based on the analysis results.

[0073] For example, the processor (110) can output a warning signal based on the degree of change in the performance of the machine learning model. Here, the warning signal may be pre-set to a plurality of levels, and the level of the warning signal may be determined according to the degree of change in the performance of the machine learning model.

[0074] Additionally, the processor (110) can output a strategy for improving the performance of a machine learning model based on the degree of change in the performance of the machine learning model. For example, the processor (110) can analyze the change in the performance of the machine learning model based on various factors and provide a strategy for improving the performance.

[0075] Additionally, the processor (110) can update the machine learning model based on the analysis results. For example, the processor (110) can identify factors that degrade the performance of the machine learning model and perform updates in a direction that improves performance. For example, the processor (110) can change thresholds related to the performance of the machine learning model.

[0076] An example of the processor (110) outputting information related to the machine learning model and updating the machine learning model is described later with reference to step 340.

[0077] The processor (110) may be implemented as an array of multiple logic gates, or as a combination of a general-purpose microprocessor and memory storing a program that can be executed on the microprocessor. For example, the processor (110) may include a general-purpose processor, a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a controller, a microcontroller, a state machine, etc. In some environments, the processor (110) may include an application-specific integrated circuit (ASIC), a programmable logic device (PLD), a field programmable gate array (FPGA), etc. For example, the processor (110) may refer to a combination of processing devices such as a combination of a digital signal processor (DSP) and a microprocessor, a combination of multiple microprocessors, a combination of one or more microprocessors combined with a digital signal processor (DSP) core, or any other combination of such configurations.

[0078] The memory (120) may include any non-transient computer-readable recording medium. As an example, the memory (120) may include a permanent mass storage device such as a random access memory (RAM), read-only memory (ROM), disk drive, solid state drive (SSD), or flash memory. As another example, a permanent mass storage device such as a ROM, SSD, flash memory, or disk drive may be a separate permanent storage device distinct from the memory. Additionally, the memory (120) may store an operating system (OS) and at least one program code (e.g., code for the processor (110) to perform an operation described later with reference to FIGS. 3 through 16).

[0079] These software components may be loaded from a computer-readable recording medium separate from the memory (120). This separate computer-readable recording medium may be a recording medium that can be directly connected to a user terminal (100), and may include, for example, a computer-readable recording medium such as a floppy drive, disk, tape, DVD / CD-ROM drive, or memory card. Alternatively, the software components may be loaded into the memory (120) via a communication module (140) that is not a computer-readable recording medium. For example, at least one program may be loaded into the memory (120) based on a computer program (e.g., a computer program for the processor (110) to perform the operation described later with reference to FIGS. 3 to 16) which is installed by files provided through the communication module (140) by developers or a file distribution system that distributes installation files for the application.

[0080] The input / output interface (130) may be a means for interfacing with a device for input or output (e.g., keyboard, mouse, etc.) that may be connected to or included in the user terminal (100). In FIG. 2a, the input / output interface (130) is shown as an element configured separately from the processor (110), but is not limited thereto, and the input / output interface (130) may be configured to be included in the processor (110).

[0081] The communication module (140) may provide a configuration or function for the server (200) and the user terminal (100) to communicate with each other via a network. Additionally, the communication module (140) may provide a configuration or function for the user terminal (100) to communicate with other external devices. For example, control signals, commands, data, etc. provided under the control of the processor (110) may be transmitted to the server (200) and / or external devices via the communication module (140) and the network.

[0082] Meanwhile, although not illustrated in FIG. 2a, the user terminal (100) may further include a display device. Alternatively, the user terminal (100) may be connected to an independent display device via wired or wireless communication to transmit and receive data to and from each other. For example, medical images, medical information, information related to machine learning models, etc., may be provided to the user through the display device.

[0083] FIG. 2b is a configuration diagram illustrating an example of a server according to one embodiment.

[0084] Referring to FIG. 2b, the server (200) includes a processor (210), memory (220), and a communication module (230). For convenience of explanation, FIG. 2b shows only the components related to the present invention. Accordingly, other general-purpose components may be included in the server (200) in addition to the components shown in FIG. 2b. Furthermore, it is obvious to those skilled in the art that the processor (210), memory (220), and communication module (230) shown in FIG. 2b may be implemented as independent devices.

[0085] The processor (210) may receive performance evaluation data from at least one of a memory (220), a user terminal (100), and other external devices. The processor (210) may calculate an indicator representing the performance of a machine learning model based on the performance evaluation data, analyze a change in the performance of a machine learning model based on at least one of the performance evaluation data and the result of calculating the indicator, or output information related to the machine learning model based on the analysis result. Alternatively, the processor (210) may transmit data (or a report) containing information related to the machine learning model to the user terminal (100).

[0086] In other words, at least one of the operations of the processor (110) described above with reference to FIG. 2a can be performed by the processor (210). In this case, the user terminal (100) can output information transmitted from the server (200) through a display device.

[0087] Meanwhile, since the implementation example of the processor (210) is the same as the implementation example of the processor (110) described above with reference to FIG. 2a, a detailed description is omitted.

[0088] Various data, such as data generated according to the operation of the processor (210), can be stored in the memory (220). Additionally, an operating system (OS) and at least one program (e.g., a program required for the processor (210) to operate) can be stored in the memory (220).

[0089] Meanwhile, since the implementation example of the memory (220) is the same as the implementation example of the memory (120) described above with reference to FIG. 2a, a detailed description is omitted.

[0090] The communication module (230) may provide a configuration or function for the server (200) and the user terminal (100) to communicate with each other via a network. Additionally, the communication module (230) may provide a configuration or function for the server (200) to communicate with other external devices. For example, control signals, commands, data, etc. provided under the control of the processor (210) may be transmitted to the user terminal (100) and / or external devices via the communication module (230) and the network.

[0091] FIG. 3 is a flowchart illustrating an example of a method for analyzing the performance of a machine learning model according to one embodiment.

[0092] The method illustrated in FIG. 3 consists of steps processed chronologically in the computing device (20) or processor (110, 210) illustrated in FIG. 1 to 2b. Therefore, even if details are omitted below, the details described above regarding the computing device (20) or processor (110, 210) may also be applied to the method illustrated in FIG. 3. Additionally, as described above with reference to FIG. 2b, at least one of the steps performed by the processor (110) below may be processed by the processor (210).

[0093] In step 310, the processor (110) receives performance evaluation data including prediction data generated as at least one medical image is analyzed by a machine learning model.

[0094] For example, the data for performance evaluation may include not only prediction data but also at least one of medical images, electronic medical record data corresponding to medical images, or metadata corresponding to medical images. Hereinafter, an example of data for performance evaluation will be described with reference to FIG. 4.

[0095] FIG. 4 is a diagram illustrating an example of performance evaluation data received by a computing device according to one embodiment.

[0096] Referring to FIG. 4, a computing device (420) receives performance evaluation data (410). For example, the performance evaluation data (410) may include at least one of prediction data (411), medical images (412), electronic medical record data (413), or metadata (414).

[0097] For example, a computing device (420) can receive medical images (412) through a picture archiving communication system (PACS). Here, the medical images (412) may be images in which the patient's personal information, etc., has been de-identified. For example, the medical images (412) may include pathology slide images, CT images, X-ray images, mammography images, MRI images, PET images, etc., but are not limited thereto.

[0098] Additionally, the computing device (420) can receive electronic medical record data (413) corresponding to the medical image (412). For example, the electronic medical record data (413) may include a pathology report.

[0099] Additionally, the computing device (420) may receive metadata (414) corresponding to the medical image (412). For example, the metadata may include personal information about the patient (e.g., age, gender, ethnicity, etc.), characteristics of the patient's disease (e.g., clinical condition, prognostic indicators, etc.), conditions for acquiring the medical image (e.g., parameters, orientation, date, etc.), and information about the equipment that took the medical image (e.g., modality, vendor, software version for processing the captured medical image, etc.). For example, the metadata may be stored based on DICOM tags.

[0100] Additionally, the computing device (420) may receive prediction data (411), which is the result of an analysis of a machine learning model on a medical image (412). For example, the prediction data may include an abnormality score, the location of a lesion, the type of lesion, whether a target lesion (e.g., cancer, nodule, etc.) is included in the medical image, whether a specific disease has manifested, etc., but is not limited thereto, and may include any medically meaningful information or clinical information of a patient that can be extracted from the medical image.

[0101] Meanwhile, the computing device (420) can preprocess the performance evaluation data (410) to form integrated data that complies with security requirements. For example, the computing device (420) can store the integrated data, in which the performance evaluation data (410) has been preprocessed, in a security database. Here, the security database may be a database that complies with the requirements of HIPAA (Health Insurance Portability and Accountability Act), but is not limited thereto.

[0102] Referring again to FIG. 3, in step 320, the processor (110) calculates at least one performance indicator representing the performance of a machine learning model based on data for performance evaluation.

[0103] Performance metrics can be measures indicating how accurate the output of a machine learning model based on medical images is. Machine learning models analyze medical images to output medical information. This medical information can be used to diagnose a patient's disease (e.g., cancer), predict the patient's response to treatment, or determine the prognosis of treatment. Therefore, the accuracy of the medical information output by a machine learning model can be related to the model's performance.

[0104] For example, performance metrics can be calculated based on predictive data output by a machine learning model analyzing medical images and ground truth data corresponding to those medical images. Here, predictive data refers to the output data of a machine learning model based on medical images.

[0105] Hereinafter, with reference to FIG. 5, an example of a performance indicator calculated by the processor (110) will be described.

[0106] FIG. 5 is a diagram illustrating an example of a processor calculating performance indicators according to one embodiment.

[0107] Figure 5 illustrates a matrix based on predicted data and ground truth data. Predicted data can be generated by a machine learning model analyzing a medical image. Additionally, ground truth data may be the accurate diagnosis result for the patient (or subject within the patient) that is the subject of the medical image.

[0108] For example, the predicted data and the ground truth data can each be classified as positive and negative. Here, a positive sign indicates that a specific disease has manifested, and a negative sign indicates that the disease has not manifested.

[0109] As illustrated in Fig. 5, the prediction data produced by a machine learning model analyzing a medical image can indicate the patient's condition as positive or negative. Additionally, the correct data representing the actual patient's condition for the same medical image can also be indicated as positive or negative. At this time, when comparing the prediction data and the correct data, four cases can be distinguished.

[0110] 1) True Positive (TP): A case where the machine learning model's prediction is positive and the actual correct data is also positive. In other words, it is a case where a patient with the disease is accurately identified as positive.

[0111] 2) True Negative (TN): A case where the machine learning model's prediction is negative and the actual correct answer data is also negative. In other words, it is a case where a patient without disease is accurately classified as negative.

[0112] 3) False Positive (FP): A case where a machine learning model predicts a positive result, but the actual ground truth data is negative. In other words, it is a situation where a patient without the disease is mistakenly classified as positive, as if they have the disease.

[0113] 4) False Negative (FN): A case where the machine learning model predicted a negative result, but the actual correct data was positive. In other words, it is a case where a patient with the disease was missed and incorrectly judged as negative.

[0114] Therefore, depending on the combination of predicted data and correct data, the case can be classified into one of TP, TN, FP, or FN.

[0115] The processor (110) can calculate at least one indicator using data classified into TP, TN, FP, and FN. For example, the processor (110) can generate a confusion matrix by comparing the predicted data and the correct data, and calculate at least one indicator based on this.

[0116] For example, the metric may be sensitivity, specificity, accuracy, ROC Curve (receiver operating characteristic curve), AUC (area under the curve), NPV (negative predictive value), precision (PPV), negative predictive value (NPV), recall, or F1 score.

[0117] Sensitivity is the ratio of predicted data that is also positive among the correct data that is positive, and can be calculated according to the following mathematical formula 1.

[0118]

[0119] Specificity is the ratio of predicted data that is also negative among the correct data that is negative, and can be calculated according to the following mathematical formula 2.

[0120]

[0121] Accuracy is the ratio of accurately predicted (i.e., predicted exactly as the actual) data out of the total predicted data, and can be calculated according to the following mathematical formula 3.

[0122]

[0123] The ROC curve refers to a graph representing the relationship between sensitivity and specificity according to various threshold values. Meanwhile, AUC is the area under the curve and is an indicator representing the overall performance of a machine learning model.

[0124] Precision is the ratio of the correct answer data to the prediction data that is positive, and can be calculated according to the following mathematical formula 4.

[0125]

[0126] The voice prediction rate is the ratio of the correct answer data to the predicted data that is voiced, and can be calculated according to the following mathematical formula 5.

[0127]

[0128] Recall refers to the proportion of positive predicted data points among the positive ground data points. Additionally, the F1 score is the harmonic mean of precision and recall, serving as an indicator of the overall performance of a machine learning model.

[0129] For example, the processor (110) can calculate a performance matrix using some of the indicators described above. The processor (110) can provide confidence intervals for the matrix and can guarantee the accuracy of the performance evaluation. Additionally, the processor (110) may calculate indicators for the entire prediction data, or it may calculate indicators by subdividing them by subgroup (e.g., patient age, gender, image modality, etc.).

[0130] As described above, the processor (110) can track changes in the performance of a machine learning model over time in a time series by calculating model performance indicators at predetermined time intervals (e.g., daily, weekly, monthly, yearly, or predetermined time intervals). Accordingly, the processor (110) can detect long-term stability and performance degradation of the machine learning model. For example, the processor (110) can determine that model performance degradation has occurred if the model performance indicators continuously decline or if the model performance indicators exceed a threshold value. Such determination can be made not only with the entire data but also at detailed units such as by institution, by equipment, by modality, and by patient age group.

[0131] Additionally, the processor (110) may further calculate a correction matrix that evaluates the probability estimation accuracy of the machine learning model. Accordingly, the prediction performance of the machine learning model can be analyzed more precisely.

[0132] Meanwhile, the processor (110) can visualize the calculated indicators in various ways and output them through a display device. Hereinafter, with reference to FIG. 6, an example of an indicator and related information being output will be described.

[0133] FIG. 6 is a drawing for explaining an example of output to a display device according to one embodiment.

[0134] Figure 6 illustrates an example of an image output to a display device. Referring to Figure 6, images (610, 620, 630) in which calculated indicators are visualized in various ways may be output to the display device. Additionally, a menu (640) may be output to the display device, which allows the user to select an element (e.g., a disease) to check.

[0135] Meanwhile, the processor (110) may recalculate the indicator upon receiving user input that changes the threshold related to the performance of the machine learning model. For example, the threshold may refer to a reference value for classifying the patient's condition as positive or negative based on an abnormality score calculated by the machine learning model.

[0136] For example, the user can change the threshold value through the UI (650) displayed on the display device. As the threshold value changes, the processor (110) can output changes in the ROC curve (610), confusion matrix (620), and the positive / negative case distribution based on abnormality scores (630) of the actual data in real time.

[0137] The recalculation of metrics based on user threshold adjustments (i.e., threshold simulation) enables users to define and adjust performance thresholds in real time. For example, sensitivity is the rate at which a machine learning model correctly identifies actual positive cases (e.g., abnormal cases where lesions are found in medical images) as positive, and can be equivalent to the true positive rate (TPR). Meanwhile, specificity refers to the rate at which a machine learning model correctly identifies actual negative cases (i.e., normal cases where lesions are not found in medical images) as negative. Additionally, the false positive rate (FPR) is the complementary value of specificity and can be defined as FPR = 1 - Specificity.

[0138] In one embodiment, the processor (110) can perform multi-purpose optimization with clinical situation-specific weights. For example, it can dynamically search for and recommend an optimal threshold value according to a predefined clinical scenario, such as prioritizing sensitivity in an emergency room environment and emphasizing specificity in a general screening.

[0139] The processor (110) can check the impact of various settings on the performance of the machine learning model by adjusting the threshold value based on the aforementioned indicators. As shown in FIG. 6, the user can visually check the relationship between TPR and FPR according to changes in the threshold value as well as the ROC curve. Through this, the user can optimize the performance of the machine learning model and adjust the performance in real time to meet clinical needs. Alternatively, the process of receiving user input can be omitted, and the processor (110) can automatically adjust the threshold value to a value that optimizes the performance of the machine learning model.

[0140] Additionally, if a distribution different from the targeted specificity, sensitivity, etc. based on the threshold is detected, the processor (110) may immediately provide a notification to the user. Accordingly, the user can monitor important clinical data and changes in the performance of the machine learning model in real time and quickly take necessary measures.

[0141] Additionally, the processor (110) can calculate a confusion matrix and analyze the performance of the current machine learning model. Accordingly, the processor can set an NPV value. Additionally, if a change in the performance of the machine learning model that differs from the expected value and a potential bias are detected, the processor (110) may generate a warning and provide a response strategy.

[0142] Referring again to FIG. 3, in step 330, the processor (110) can analyze the performance change of the machine learning model based on at least one of the performance evaluation data and the computation result.

[0143] For example, the processor (110) can determine whether there is a change in the performance of the machine learning model or the extent thereof by referring to the result of the metric calculation. Specifically, if the difference between the calculated metric value and the preset threshold metric value exceeds a predetermined allowable range, the processor (110) can determine that the performance of the machine learning model has deteriorated.

[0144] Additionally, the processor (110) can analyze performance changes more precisely by considering the performance evaluation data and the results of the metric calculation together. In this process, it is not limited to simply determining whether there is a performance change, but can also determine whether the change is due to data drift. To this end, the processor (110) can compare the prediction data included in the performance evaluation data with the correct answer data and determine whether data drift has occurred based on the result.

[0145] For example, the processor (110) can obtain abnormality scores from the prediction data of a machine learning model included in the performance evaluation data. Additionally, the input data may include prediction data accumulated from the past (i.e., the analysis results of the machine learning model). Accordingly, the processor (110) can also obtain past abnormality scores from the performance evaluation data.

[0146] The processor (110) can analyze changes in the performance of a machine learning model by comparing a reference distribution and a first distribution of abnormal scores corresponding to medical images. Here, the reference distribution can be generated based on abnormal scores during a reference period. Meanwhile, the first distribution can be generated based on abnormal scores during a detection period after the reference period. The reference distribution can reflect a point in time that represents the normal and stable performance of the machine learning model, and can effectively evaluate whether there has been a change in model performance by comparing it with the detection period distribution during the subsequent detection period.

[0147] Hereinafter, the reference distribution and the first distribution will be explained with reference to FIG. 7.

[0148] FIG. 7 is a diagram illustrating an example of a reference distribution and a first distribution according to one embodiment.

[0149] As illustrated in FIG. 7, the reference distribution and the first distribution can be visualized as a histogram or a cumulative distribution function (CDF) overlay. For example, FIG. 7 shows a graph in which the reference distribution of anomaly scores is shown as a solid line and the first distribution of anomaly scores is shown as a dotted line. The processor (110) can derive the reference distribution using anomaly scores accumulated during a predetermined time interval in the past (i.e., reference period). Additionally, the processor (110) can derive the first distribution using anomaly scores accumulated during a detection period after the reference period (e.g., including the present).

[0150] The processor (110) can detect whether data drift has occurred based on the difference between the reference distribution and the first distribution.

[0151] The processor (110) can perform multiple statistical tests on the performance evaluation data actually input into the machine learning model through the reference distribution and the first distribution. Accordingly, the processor (110) continuously monitors changes in the statistical distribution of the performance evaluation data. To this end, the processor (110) can set a reference period and derive the reference distribution during the reference period. The processor (110) can track fluctuations in the distribution of the output data occurring over time through statistical analysis.

[0152] For example, the processor (110) can analyze changes in the performance of a machine learning model by utilizing various techniques, such as performing an analysis that statistically tests differences between data distributions (e.g., Chi-squared Test, Kolmogorov-Smirnov Test, Cramer-von Mises, Fisher's Exact Test, etc.), performing an analysis that detects drift between data distributions (e.g., Maximum Mean Discrepancy, Least-Squares Density Difference, etc.), or performing an analysis that detects changes in data distributions using a classifier (e.g., Learned Kernel, Classifier-based drift detector Lopez-Paz and Oquab, Spot-the-diff drift detector, Model Uncertainty drift detector, etc.). The processor (110) can select multiple techniques suitable for the characteristics of the data to be analyzed. The processor (110) can obtain test results (e.g., p-value, etc.) by applying the selected techniques in parallel. The processor (110) can perform meta-analysis by synthesizing the test results. Accordingly, the reliability of the final drift judgment is maximized, and changes in the performance of the machine learning model can be analyzed.

[0153] The Kolmogorov-Smirnov Test (hereinafter referred to as the KS Test) is a technique that compares the distribution of abnormal scores during a standard period with the distribution of abnormal scores during a detection period based on demographic variables (e.g., age). For example, it is assumed that the distribution of abnormal scores is divided by patient age groups. Additionally, it is assumed that the abnormal scores of the 50-60 age group were concentrated between 40 and 60 during the standard period, but shifted to the 70-90 range during the detection period. In this case, the processor (110) can detect that data drift has occurred through the KS Test. Through the KS Test, whether the distribution of abnormal scores is statistically significantly different can be evaluated through the D value of the maximum vertical distance between the two CDF values ​​of the two periods (i.e., the standard period and the detection period). If the P-value is significantly low, the processor (110) can determine that data drift has occurred due to demographic factors.

[0154] The Chi-Squared Test is a technique that allows comparing the distribution of abnormal scores using categorical variables (e.g., type of equipment, method of shooting) based on equipment-specific data. For example, the processor (110) can divide the abnormal scores into intervals (e.g., 0-10, 11-20, ...) based on data (i.e., medical images) taken with equipment A and equipment B, and then compare the distribution of abnormal scores by equipment during the standard period and the detection period. Then, the processor (110) can evaluate whether the distribution of abnormal scores for each equipment is the same during the two periods. If the P-value is low, the processor (110) can determine that data drift has occurred due to changes in equipment.

[0155] The Cramer-von Mises Test is a technique for evaluating changes in the distribution of abnormal scores based on input data of a machine learning model. For example, the processor (110) can compare the cumulative distribution of abnormal scores during a standard period and a detection period by considering the patient's lesion information (e.g., nodule, cardiomegaly, etc.). By evaluating changes across the cumulative distribution, the processor (110) can detect the actual influence that lesion information has on the distribution of abnormal scores.

[0156] Maximum Mean Discrepancy (MMD) is a technique that evaluates the difference in anomaly scores between a standard period and a detection period based on input data including various demographic variables and equipment-specific data. For example, the processor (110) can compare the average embedding vectors during two periods (standard period and detection period) by considering the patient's age, gender, imaging equipment, etc. The processor (110) can determine whether data drift has occurred by detecting a change in the data distribution based on the average embedding vector.

[0157] The Least-Square Density Difference (LSDD) is a technique for measuring the difference in distribution density of anomaly scores between a standard period and a detection period. Specifically, it can be determined how the distribution density of anomaly scores has changed based on demographic variables such as specific age groups or gender groups. For example, if the p-value is significantly lower than the difference in distribution density between the standard period and the detection period, the processor (110) can determine that data drift has occurred in a specific demographic group.

[0158] Fisher's Exact Test is a technique that can evaluate changes in the distribution of anomaly scores during the standard period and detection period when the sample size of the data per equipment is small. For example, when the number of samples taken with equipment A is small, the processor (110) can use Fisher's Exact Test to determine whether the distribution of anomaly scores corresponding to equipment A has changed significantly.

[0159] A Classifier-Based Drift Detector is a technique that evaluates whether data from a standard period and a detection period can be distinguished by inputting data from the input data into a normal / abnormal binary classification model. For example, if a machine learning model can distinguish data from the two periods based on specific demographic variables or data by equipment, this may indicate that there is a difference in distribution. The processor (110) can determine that data drift has occurred as the accuracy of the classification model increases.

[0160] The Spot-the-Diff Drift Detector is a technique for evaluating the relationship between input data and anomaly scores. Through the Spot-the-Diff Drift Detector, the difference in output between a machine learning model during a standard period and a detection period can be evaluated. For example, if the machine learning model mainly predicted anomaly scores in the range of 80-100 during the standard period, but the prediction range changed to 60-80 during the detection period, the processor (110) can determine that the performance of the machine learning model has deteriorated.

[0161] A Model Uncertainty Drift Detector is a technique for detecting data drift by evaluating the uncertainty when a machine learning model predicts based on input data. For example, if the prediction of an anomaly score in a specific piece of equipment is more uncertain during the detection period, the processor (110) may determine that data drift has occurred due to a change in equipment.

[0162] The Feature-Wise Two-Sample Kolmogorov-Smirnov Test and Chi-Squared Test are techniques that apply the KS Test to continuous variables (i.e., outlier scores) and the Chi-Squared Test to categorical variables (e.g., equipment type, patient gender, etc.) to evaluate the difference in distribution between the two periods. The processor (110) can use these techniques to evaluate whether data drift occurs according to various variables related to outlier scores.

[0163] The Context-Aware Maximum Mean Discrepancy Drift Detector is a technique that detects changes in the distribution of abnormal scores by considering the patient's age, gender, equipment information, etc. For example, the processor (110) can analyze how the distribution of abnormal scores has changed during the standard period and the detection period according to the patient's age and gender. In addition, the processor (110) can detect data drift more precisely with abnormal scores, not only the difference in the distribution of abnormal scores.

[0164] FIG. 8 is a flowchart illustrating an example of operation in which a processor according to one embodiment operates when a change in the performance of a machine learning model is confirmed.

[0165] In steps 810 and 820, the processor (110) analyzes the performance change of the machine learning model and checks whether the performance change has been confirmed.

[0166] For example, the processor (110) can determine whether data drift has occurred through the various methods described above with reference to FIG. 7. If data drift has occurred, the processor (110) can determine that the performance of the machine learning model has changed.

[0167] For example, the processor (110) can determine whether the performance of the machine learning model has deteriorated based on time series analysis of key indicators (e.g., AUC, sensitivity, specificity, etc.), verification of performance changes through statistical significance testing, and preset thresholds.

[0168] If a change in the performance of the machine learning model is confirmed, the processor (110) performs steps 830 through 850. If no change in the performance of the machine learning model is confirmed, the procedure is terminated. That is, the processor (110) continuously monitors whether there is a change in the performance of the machine learning model without performing any separate action.

[0169] In step 830, the processor (110) may output a warning signal based on the degree of change in the performance of the machine learning model. For example, the warning signal may be determined and output as one of a plurality of levels.

[0170] As described above with reference to FIG. 7, the processor (110) analyzes the trend of the performance of the machine learning model through regression analysis and predicts changes in the relationship between the input data and the machine learning model. In particular, the processor (110) can analyze the performance difference of the machine learning model according to various factors such as the patient's age, gender, and ethnicity. Accordingly, the processor (110) can evaluate how reliably the machine learning model analyzes the input data corresponding to various patient groups.

[0171] If a predefined threshold is reached or if performance degradation of the machine learning model occurs, the processor (110) may automatically generate a warning signal. For example, the warning signal may be classified into High, Intermediate, and Low depending on the degree of performance degradation of the machine learning model. Specifically, a High warning signal may be generated if performance indicators drop by more than 20% or if clinically significant false positives or false negatives increase; an Intermediate warning signal may be generated if performance indicators drop by 10-20% or if performance degradation of the machine learning model is detected when analyzing medical images of patients included in a specific subgroup; and a Low warning signal may be generated if performance indicators drop by 5-10% or if minor data drift is detected.

[0172] Additionally, the processor (110) may automatically generate periodic performance reports. For example, the performance reports may be provided to users (e.g., medical professionals, managers of machine learning models, etc.) via email or an application.

[0173] That is, the processor (110) can provide real-time notifications as well as convey the severity of the problem (i.e., performance degradation of the machine learning model) to the user through a report. Accordingly, the user can immediately take action regarding the problem.

[0174] In step 840, the processor (110) can provide a performance improvement strategy based on changes in the performance of the machine learning model.

[0175] For example, if performance degradation of a machine learning model is detected, the processor (110) can perform root cause analysis and identify the root cause of the problem. The processor (110) can analyze the cause of performance degradation based on various factors, such as changes in the demographics of the patient whose medical image was taken (e.g., race, nationality, age, etc.), changes in the modality of which the medical image was acquired, and replacement of imaging equipment.

[0176] For example, the analysis of the cause of performance degradation can be performed through the analysis of changes in the distribution of input data (i.e., medical images), the analysis of demographic changes in medical image objects (e.g., age distribution, gender distribution, ethnic distribution, BMI, Breast Density, etc.), the analysis of changes in the quality and characteristics of medical images (e.g., PGMI, contrast, sharpness, noise, etc.), the analysis of the modality and equipment used to acquire the medical images, the analysis of changes in the clinical environment (e.g., changes in diagnostic criteria or protocols, emergence of new types of lesions or diseases), and the analysis of machine learning models and related issues (e.g., review of AI model versions and update history, analysis of model weights and biases, and analysis of whether there is overfitting or underfitting).

[0177] And, the processor (110) can suggest improvement strategies based on the analysis results, such as retraining the machine learning model, improving the input data collection strategy, and mitigating analysis bias.

[0178] The improvement strategy may be an improvement in the data collection strategy for the machine learning model (e.g., collecting additional data for insufficient data types or subgroups, or suggesting oversampling). Alternatively, the improvement strategy may be a retraining of the model (e.g., retraining and fine-tuning the machine learning model with new training data that reflects the latest distribution, or suggesting domain adaptation by data time point / institution / equipment and vendor-specific correction).

[0179] Alternatively, the improvement strategy may be the application of ensemble techniques (e.g., combining multiple different machine learning models to propose a combined prediction result). As an example, the processor (110) may propose a strategy of applying a majority voting technique to determine the result predicted by the most machine learning models among the results predicted by multiple machine learning models as the final result value. As another example, the processor (110) may propose a strategy of applying high weights to machine learning models with high performance and low weights to machine learning models with low performance. As yet another example, the processor (110) may propose a strategy of randomly selecting multiple training data samples and training independent machine learning models for each sample to provide predictions that reflect the diversity of the data. Alternatively, the processor may propose a strategy of utilizing bagging / boosting to prevent overfitting of machine learning models, or selecting training data by assigning weights to data where the previous model had a high error, and training subsequent models to predict those data cases more accurately to gradually improve model performance. As another example, the processor (110) may propose a strategy of inputting the prediction results of each machine learning model into a meta model to derive a final prediction.

[0180] For example, the aforementioned warning signals and improvement strategies can be provided in various ways, such as graphics, sound, or text, so that users can intuitively recognize them.

[0181] In step 850, the processor (110) can update the machine learning model automatically or semi-automatically. For example, the processor (110) can perform automated updates through a connection with the machine learning model provider, depending on the cause of the performance degradation. Specifically, if the performance degradation of the machine learning model is detected, the processor (110) can report to the provider using a predefined improvement strategy. And, it can receive automated updates through the provider's system.

[0182] Referring again to FIG. 3, at step 340, the processor (110) outputs information related to the machine learning model based on the analysis results.

[0183] Here, the statement that the processor (110) outputs information includes controlling the display device so that information is output through the display device. Specifically, the processor (110) visualizes information related to a machine learning model, and the display device can output the visualized result to the screen.

[0184] Information related to a machine learning model may include information about the machine learning model itself (e.g., version, etc.), information about input data and metadata, information about indicators representing the performance of the machine learning model, information about changes in the performance of the machine learning model, etc. In other words, the processor (110) may output input data, the machine learning model, the analysis results of the machine learning model, information about the performance of the machine learning model, etc.

[0185] Hereinafter, examples of information related to a machine learning model being output will be described with reference to FIGS. 9 to 16.

[0186] FIGS. 9 to 16 are drawings illustrating examples in which information related to a machine learning model according to one embodiment is output.

[0187] Referring to FIG. 9, information (910, 920) related to changes in the performance of a machine learning model may be output to the display device. For example, a reference distribution and a first distribution of anomaly scores may be visualized and output. The reference distribution and the first distribution may be output as a graph (910), or the content of the analysis of changes in the performance of the machine learning model may be output as text (920). At this time, the text (920) may include a reference period, the number of image samples in the reference period, a detection period, the number of image samples in the detection period, a statistical technique used to detect data drift, and various data derived according to the application of the technique.

[0188] Additionally, a UI (930) that allows the user to adjust analysis conditions may be displayed on the display device. Here, analysis conditions refer to factors that affect the machine learning model performance analysis results.

[0189] For example, the UI (930) may include a menu for setting a reference period and / or a detection period. Additionally, the UI (930) may include a menu for selecting metadata (e.g., patient's age, gender, ethnicity, information about the imaging equipment (manufacturer of the modality, etc.). Accordingly, the user can easily understand the analysis results of the machine learning model based on various conditions.

[0190] To analyze changes in the performance of a machine learning model, the distribution of prediction data generated by the machine learning model during the detection period can be compared with the distribution during the reference period. FIG. 10 shows an example of a screen that allows filtering data or setting analysis conditions in more detail for such comparison. As illustrated in FIG. 10, the user can directly specify demographic characteristics of the analysis target, characteristics of the equipment that acquired the medical image, the data collection period, etc. The processor (110) can use the sampled data based on such user input as data for performance evaluation.

[0191] Accordingly, as shown in the UI (930) of FIG. 9, analysis conditions may be selected through a simple menu within the screen displaying the analysis results, or as shown in FIG. 10, a screen may be provided where various analysis conditions can be selected in more detail.

[0192] Referring to FIGS. 11 and 12, various information such as time series trends regarding the core performance of machine learning models and performance comparisons between machine learning models can be output to the display device.

[0193] The processor (110) can provide a performance matrix of machine learning models, performance trends of machine learning models, and a dashboard for comparison between machine learning models. Accordingly, the user can intuitively check the distribution of analysis results over time and real-time trends through the dashboard.

[0194] Additionally, the processor (110) can output the analysis results of the detection period and the analysis results of the reference period together. Accordingly, the user can easily determine whether the distribution of anomaly scores output by the machine learning model is increasing or decreasing.

[0195] Additionally, the processor (110) can provide analysis results by metadata in a graph. The metadata may include the patient's age, gender, ethnicity, information about the equipment (e.g., manufacturer of the modality, etc.), output values ​​of the machine learning model (e.g., abnormal score, lesion location, lesion type, etc.), and types of medical images (e.g., PA / AP view - Chest X-ray, MLO (Left, Right) / CC (Left, Right) view, etc.). Accordingly, the user can easily identify whether the machine learning model's analysis indicators are increasing or decreasing and the trends of change.

[0196] Additionally, referring to FIG. 13, the processor (110) may output information about changes in the performance of a machine learning model as a report. The report may include statistical techniques used to analyze changes in performance and numerical values ​​derived from the application of said techniques.

[0197] FIG. 14 shows an example of a settings screen for providing the report illustrated in FIG. 13. In the 'Reference Period' area at the top, the user can set a reference period, specify a sample size (e.g., 1000), or directly select a period.

[0198] In the 'Finding types' area, the user can select items to include in the report. For example, the user can choose to include a summary of all findings in the report, or to include only specific items (e.g., fracture, atelectasis, calcification, cardiac hypertrophy, pleural effusion, tuberculosis screening, etc.) in the report. The processor (110) may also output a report containing information on changes in the performance of a machine learning model that detects selected findings (Finding types) from medical images.

[0199] In the 'Delivery Date & Time' area, users can set the report delivery cycle. For example, users can select annual, semi-annual, or quarterly, and once the reservation is complete, the report is automatically generated at the specified time.

[0200] Through 'Email Recipients', users can enter an email address to receive reports. Alternatively, reports can be viewed in the system's report list even if the user does not enter an email address. For example, if the user presses the 'Save' button, the settings are saved and the schedule is activated.

[0201] In this way, users can regularly check the monitoring results regarding changes in the distribution of input data by simply setting the reference period, analysis items, reporting cycle, and recipients.

[0202] FIG. 15 illustrates an example of a screen that summarizes and outputs the results of a data drift analysis. The processor (110) can compare the data distribution of the reference period with the data distribution of the detection period to summarize and display the occurrence and degree of drift for each column. For example, the screen may display the ratio and number of columns in which drift is detected among all columns. Through this, the user can intuitively grasp the degree of distribution change occurring throughout the data.

[0203] For example, the distribution of anomaly scores for the reference period and the detection period can be superimposed on the screen as a histogram of the same interval. For instance, the average value for each period can be represented by a vertical dashed line. Users can check the mean, standard deviation, and sample size through the tooltip on the dashed line. Additionally, drift analysis results can be displayed on the screen as a table. For instance, the table may display the name of the feature, data type (such as categorical or numerical), drift score, and statistical analysis value.

[0204] For example, drift judgments can be recorded as Detected / Not Detected. The 'Drift Score' can be calculated by applying weights to the statistical figures of multiple statistical analyses. When a user clicks the header of the table, criteria sorting and threshold filtering of the drift scores can be performed. Additionally, when a user clicks a row of the table, a detail panel is expanded to provide the distribution by sub-level and the change in the ratio by level of the corresponding column. Figure 15 illustrates that drift was detected by changes in the statistical distribution of two columns (Patient Age, Breast Density, 40%) out of the five columns (Patient Age, Modality, Breast Density, Acquisition Site, Sex) of the table.

[0205] As described above with reference to FIG. 15, whether drift occurs can be reliably determined based on statistical figures for each characteristic of the reference period and the detection period.

[0206] For example, for a column in the table shown in FIG. 15 where drift is detected, the processor (110) can calculate a warning level and send a notification. Additionally, the processor (110) can suggest improvement strategies (e.g., checking equipment calibration, simulating thresholds, etc.). Additionally, the processor (110) can automatically record logs (e.g., model version, period, sample size, threshold used, applied test, judgment result, etc.). Accordingly, the user can easily track the cause of the change and take follow-up actions to meet clinical needs. Additionally, if necessary, the user can perform model retraining through a provider. Accordingly, performance degradation of the machine learning model can be prevented.

[0207] Meanwhile, as illustrated in FIG. 16, the processor (110) can not only monitor the distribution of input data to analyze changes in the performance of the machine learning model, but also analyze and display the time taken for the machine learning model to analyze medical images.

[0208] For example, a time-series graph showing the trend of task processing time over a period can be presented at the top of the screen. The graph displays the average processing time and the maximum processing time separately, and when a user hovers the mouse over a specific point in time, the average processing time and maximum processing time values ​​for that time can be displayed as a popup. In the example illustrated in Fig. 16, the average processing time is displayed as approximately 10ms and the maximum processing time as of 23:00 on June 9, 2024. Through this visualization, the user can grasp the fluctuations in processing performance over time at a glance.

[0209] For example, processing times per analysis stage can be displayed as dots at the bottom of the screen. Each dot represents the processing time of an individual task, and normal ranges and outliers can be distinguished by color. This allows users to easily identify cases where processing times are abnormally long or fluctuate significantly at a specific analysis stage.

[0210] A processing time analysis screen as illustrated in FIG. 16 can be provided by a processor (110) collecting the start and end times of each task from server logs or work records, and calculating and visualizing indicators such as average, maximum, and variance based on this. Through this screen, the operator can monitor not only changes in the data input distribution but also the stability of the system processing speed and whether bottlenecks occur, and based on this, can perform subsequent actions such as adjusting resource allocation, optimizing performance, and sending warning notifications.

[0211] Meanwhile, although not illustrated in FIGS. 3 to 16, the computing device (20) may include an interactive interface that allows the user to select a specific period, a version of the machine learning model, etc.

[0212] Additionally, the processor (110) can record versions of the machine learning model and analyze and record the performance of each version in real time. For example, the processor (110) can assign a unique identifier to each version and record training data sets, hyperparameters, architecture, etc. for each version.

[0213] Additionally, whenever a machine learning model is updated or changes occur, the processor (110) can automatically monitor changes in the model's performance. Additionally, the processor (110) can record when a new version of the machine learning model is deployed. Accordingly, when data drift occurs, the processor (110) can easily track the cause of the data drift.

[0214] Additionally, the processor (110) can evaluate the performance of a machine learning model based on user input in real time and provide results comparing the performance between different versions. Accordingly, the user can check how much the performance of the updated model has improved compared to the existing model.

[0215] Additionally, the computing device (20) may perform federated learning among machine learning models applied to various sites. Accordingly, performance improvement can be achieved not only for a machine learning model applied to a single site but also for multiple machine learning models.

[0216] As described above, the performance of a medical imaging diagnostic system based on a machine learning model can be continuously monitored and optimized. Consequently, the accuracy, reliability, and fairness of medical imaging-based disease screening (e.g., cancer) can be improved. Furthermore, since threshold simulation and data drift detection are possible, the performance of the machine learning model can be adjusted according to clinical requirements in the field of radiology. In other words, the performance of the machine learning model can be adjusted and optimized through threshold simulation, and changes in the data distribution for performance evaluation can be monitored in real time through the detection of data drift. Therefore, performance degradation of the machine learning model can be prevented.

[0217] Furthermore, as the causes of data drift are analyzed, changes in the data input to machine learning models, such as demographic shifts and changes in medical imaging equipment and methods, can be analyzed. Through this, the performance of machine learning models can be continuously improved, compliance with clinical regulations can be facilitated, and reliability in clinical practice can be enhanced.

[0218] Meanwhile, the above-described method can be written as a program executable on a computer and can be implemented on a general-purpose digital computer that operates the program using a computer-readable recording medium. In addition, the structure of the data used in the above-described method can be recorded on a computer-readable recording medium through various means. The computer-readable recording medium includes storage media such as magnetic storage media (e.g., ROM, RAM, USB, floppy disk, hard disk, etc.) and optical reading media (e.g., CD-ROM, DVD, etc.).

[0219] A person skilled in the art related to the present embodiment will understand that it may be implemented in modified forms without departing from the essential characteristics of the description above. Therefore, the disclosed methods should be considered in an illustrative rather than a restrictive sense, and the scope of rights is defined in the claims rather than the description above, and should be interpreted to include all differences within the scope of equivalence.

Claims

1. At least one memory in which at least one instruction is stored; and It includes at least one processor that operates according to the above at least one instruction; and The above-mentioned at least one processor is, A computing device that receives performance evaluation data including prediction data generated as at least one medical image is analyzed by a machine learning model, calculates at least one performance indicator representing the performance of the machine learning model based on the performance evaluation data, analyzes the change in performance of the machine learning model over time based on at least one of the performance evaluation data and the calculation result, and outputs information related to the machine learning model based on the analysis result.

2. In Paragraph 1, The above performance evaluation data is, A computing device further comprising at least one of the above-mentioned at least one medical image, electronic medical record (EMR) data corresponding to the above-mentioned at least one medical image, or metadata corresponding to the above-mentioned at least one medical image.

3. In Paragraph 1, The above-mentioned at least one processor is, Calculate the at least one performance indicator based on the above prediction data and the correct answer data corresponding to the at least one medical image, and If performance degradation of the machine learning model is detected based on the above calculation result, the cause of the performance degradation is analyzed based on the performance evaluation data, A computing device comprising at least one of the above cause analysis, which includes an analysis of demographic changes in medical image objects, an analysis of changes in the quality and characteristics of medical images, an analysis of changes in the modality of acquired medical images, an analysis of changes in the clinical environment, and an analysis of machine learning models and related issues.

4. In Paragraph 1, The above-mentioned at least one processor is, A computing device that recalculates at least one indicator upon receiving user input that changes a threshold value related to the performance of the machine learning model.

5. In Paragraph 1, The above-mentioned at least one processor is, A computing device that analyzes changes in the performance of a machine learning model by comparing a first distribution of abnormality scores corresponding to at least one medical image analyzed by the machine learning model during a detection period with a reference distribution of abnormality scores obtained by analyzing a plurality of medical images during a reference period.

6. In Paragraph 5, The above-mentioned at least one processor is, Calculate a statistical value representing the difference between the first distribution and the reference distribution using at least one statistical technique for each characteristic of the medical image object, and A computing device that determines that performance degradation has occurred due to data drift in the corresponding characteristic when the above statistical value exceeds a threshold value.

7. In Paragraph 1, The above-mentioned at least one processor is, A computing device that outputs a warning signal determined as one of a plurality of levels based on the degree of change in performance of the machine learning model.

8. In Paragraph 1, The above-mentioned at least one processor is, A computing device that generates a report including the analysis results based on the degree of change in the performance of the machine learning model.

9. In Paragraph 1, The above-mentioned at least one processor is, A computing device that updates the machine learning model by performing at least one of retraining or parameter adjustment based on the above analysis results.

10. A step of receiving performance evaluation data including predictive data generated as at least one medical image is analyzed by a machine learning model; A step of calculating at least one performance indicator representing the performance of the machine learning model based on the above performance evaluation data; A step of analyzing the performance change of the machine learning model based on at least one of the performance evaluation data and the computation result; and A method for analyzing the performance of a machine learning model, comprising the step of outputting information related to the machine learning model based on the above analysis results.

11. In Paragraph 10, The above-mentioned calculation step is, The method includes the step of calculating the at least one performance indicator based on the above prediction data and the correct answer data corresponding to the at least one medical image, The step of analyzing the performance change of the above machine learning model is, If performance degradation of the machine learning model is detected based on the above calculation result, the method includes a step of analyzing the cause of performance degradation based on the performance evaluation data. The step of analyzing the cause of the above performance degradation is, A method comprising the step of performing at least one of an analysis of demographic changes in medical image objects, an analysis of changes in the quality and characteristics of medical images, an analysis of changes in the modality of acquired medical images, an analysis of changes in the clinical environment, and an analysis of machine learning models and related issues.

12. In Paragraph 10, The above-mentioned analysis step is, A method for analyzing a change in the performance of a machine learning model by comparing a first distribution of abnormality scores corresponding to at least one medical image analyzed by the machine learning model during a detection period with a reference distribution of abnormality scores obtained by analyzing a plurality of medical images during a reference period.

13. In Paragraph 12, The above-mentioned analysis step is, A step of calculating a statistical value representing the difference between the first distribution and the reference distribution using at least one statistical technique according to the characteristics of the medical image object; and A method comprising the step of determining that performance degradation has occurred due to data drift in the corresponding characteristic when the above statistical value exceeds a threshold value.

14. In Paragraph 10, A method further comprising the step of updating the machine learning model by performing at least one of retraining or parameter adjustment based on the above analysis results.

15. A computer-readable recording medium storing a program for executing the method of paragraph 10 on a computer.

Citation Information

Patent Citations

  • Fire suppression system and firefighting aircraft having the same

    KR102599693B1

  • Ai data management system and method based on data drift detection network

    KR102613177B1

  • Control method of clutch actuator for Automated manual transmission

    KR102777554B1

  • KR20200092447A

  • KR20240065711A