System and method for adding interpretability to and evaluating bias of ECG analysis model

By extracting interpretable criteria and evaluating biases from the ECG dataset, the interpretability and bias issues of AI-based ECG analysis models are addressed, improving clinicians' understanding and trust in the model outputs and increasing their application value in clinical settings.

CN121601203APending Publication Date: 2026-03-03GE PRECISION HEALTHCARE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511091056.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-08-27
Filing Date
2025-08-05
Publication Date
2026-03-03

AI Technical Summary

Technical Problem

AI-based ECG analysis models lack interpretability and are biased, limiting their acceptability and use in clinical settings.

Method used

By extracting interpretability criteria and evaluating model biases on the ECG dataset, interpretability and bias assessments are output to users to increase the model's usability in clinical settings.

Benefits of technology

This improved clinicians' understanding and reliance on the output of the ECG analysis model, and increased the model's usability in the clinical setting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121601203A_ABST
    Figure CN121601203A_ABST
Patent Text Reader

Abstract

Systems and methods for adding interpretability to and evaluating the bias of an ECG analysis model are provided herein. In one example, a method includes obtaining a diagnostic output on an ECG data set from an AI-based ECG analysis model (302); extracting an interpretable criterion from the ECG dataset for a target of a diagnostic output of the AI-based ECG analysis model to predict an output of the ECG dataset based on the extracted criterion; determining one or more characteristics of the extracted criteria (304); evaluating a bias from the output of the ECG analysis model based on a comparison between the output of the ECG analysis model and the predicted output (404); and outputting the one or more characteristics and bias evaluations to the user equipment (362).
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-references to related applications

[0002] This application claims priority to Greek Patent Application No. 20240100584, filed on August 21, 2024. The entire contents of the above application are hereby incorporated by reference for all purposes. Technical Field

[0003] The implementation schemes of the subject matter disclosed herein relate to ECG analysis, and more specifically to systems and methods for adding interpretability to ECG analysis models and evaluating the biases of ECG analysis models. Background Technology

[0004] ECG is a graphical representation of the heart's electrical activity and is typically represented as waveforms. Traditional rule-based ECG analysis models receive an ECG, determine a set of ECG features (e.g., findings), compare this feature set with corresponding criteria of a set of rules, determine a diagnosis based on the criteria satisfied by the ECG features, and generate a diagnostic interpretation of the ECG. Physicians can review the ECG and / or diagnostic interpretation to assess a patient's cardiac activity.

[0005] However, AI-based ECG analysis models (such as neural networks) do not utilize the same feature and standard set in a predictable manner. Therefore, such AI-based ECG analysis models often lack interpretability and are susceptible to confounding biases, limiting their acceptability and use in clinical settings. Summary of the Invention

[0006] In one example, one approach includes: obtaining diagnostic output from an AI-based ECG analytics model on an ECG dataset; extracting interpretable criteria from the ECG dataset to predict the diagnostic output from the AI-based ECG analytics model, targeting the diagnostic output; evaluating biases in the AI-based ECG analytics model to assess its suitability for deployment; and outputting the interpretable criteria and bias assessment to a user device. By outputting interpretable criteria that explain the diagnostic output and bias assessment of the AI-based ECG analytics model, users can more easily understand the reasoning behind the diagnostic output.

[0007] It should be understood that the above brief description is provided to introduce selected concepts further described in the detailed embodiments in a simplified form. This is not intended to identify key or essential features of the claimed subject matter, the scope of which is uniquely defined by the claims following the detailed embodiments. Furthermore, the claimed subject matter is not limited to specific implementations that address any shortcomings mentioned above or in any part of this disclosure. Attached Figure Description

[0008] The invention will be better understood by referring to the following description of non-limiting embodiments, in which:

[0009] Figure 1 This is a diagram of an example system used for interpreting electrocardiograms (ECGs);

[0010] Figure 2 yes Figure 1 A diagram showing example components of the device in the example system;

[0011] Figure 3A This is a flowchart illustrating a method for adding interpretability to AI-based ECG analysis models for ECG datasets that lack ground truth.

[0012] Figure 3B This is a flowchart illustrating a method for evaluating the bias of an AI-based ECG analysis model;

[0013] Figure 4 This is a flowchart illustrating a method for evaluating the bias of an AI-based ECG analysis model;

[0014] Figure 5 This is a flowchart illustrating a method for using an AI model to extract standards for adding interpretability to an AI-based ECG analysis model;

[0015] Figure 6 This is a diagram illustrating the standard extraction example;

[0016] Figure 7 This is a graph of example predicted scores; and

[0017] Figure 8 This is an example chart used to determine a diagnosis. Detailed Implementation

[0018] The following describes various implementation schemes related to the interpretation of electrocardiograms (ECGs). Specifically, systems and methods are provided for adding interpretability and evaluating the biases of artificial intelligence (AI)-based ECG analysis models, such as machine learning models, neural networks, etc. Traditional rule-based ECG analysis models can use ECG features and fully interpretable criteria developed by subject matter experts. Therefore, rule-based models can identify specific features and criteria to output a diagnosis or diagnostic interpretation.

[0019] In contrast, AI-based ECG analysis models lack this interpretability. Specifically, AI-based ECG analysis models may output diagnoses or diagnostic interpretations that cannot be fully explained to clinicians (e.g., physicians or technicians).

[0020] Furthermore, AI-based ECG analysis models are often biased. Specifically, model training can cause the model to bias towards certain outputs rather than others within the context of the same input. This lack of interpretability and instances of bias in ECG analysis models limit their usability in clinical settings, as clinicians are less likely to trust the outputs when they cannot explain why they were produced.

[0021] Therefore, this paper proposes a system and method for adding interpretability and evaluating the bias of AI-based ECG analysis models. By extracting fully interpretable criteria, particularly targeting the diagnostic output of the AI-based ECG analysis model, on an ECG dataset where the model is deployed, the model's output can be predicted. Thus, interpretable criteria can explain the diagnostic output from the AI-based ECG analysis model, thereby adding interpretability to the model. Furthermore, based on the extracted criteria, bias in the output can be evaluated. The interpretability and bias assessment can be presented to users (e.g., clinicians, physicians, care providers, ECG technicians, etc.), enabling them to more reliably evaluate the output of the ECG analysis model, thereby increasing the usability of such models in clinical settings.

[0022] The systems and methods disclosed herein will now be described by way of example with reference to the accompanying drawings, wherein... Figure 1 A diagram of an example ECG analysis system is shown. Figure 2 It shows Figure 2 The example system shown in Figures 3 to 4 illustrates example components of the device. Figure 5 The diagram illustrates a method for adding interpretability to an ECG analysis model, evaluating the biases of the ECG analysis model, and extracting criteria for determining a diagnosis from the ECG analysis model. Figure 6 A diagram illustrating an example of extracting criteria is shown. Figure 7 A graph showing the predicted scores is displayed, and Figure 8 A graph depicting the bias of the ECG analysis model is shown.

[0023] from Figure 1 The diagram of an example ECG analysis system 100 is shown at the beginning. The ECG analysis system 100 can be configured to extract interpretability criteria from the diagnostic output of an AI-based ECG analysis model and add interpretability to the ECG analysis model. For example... Figure 1 As shown, system 100 may include an ECG device 110, multiple electrodes 128, an ECG analysis device 112, an AI-based ECG analysis model 114, an interpretability and bias module 116, a platform 118, an AI model 120, a user device 122, a database 124, one or more medical data repositories 130, and a network 126.

[0024] ECG device 110 can be configured to generate ECG lines for a patient. For example, ECG device 110 can be a stand-alone ECG device, a portable ECG device, a multi-vital sign monitoring device, etc. The ECG device can receive cardiac electrical signals via multiple electrodes 128 and can generate ECG lines based on these signals. The ECG can be a single-lead ECG, a 3-lead ECG, a 5-lead ECG, a 6-lead ECG, a 12-lead ECT, etc. ECG device 110 can include any number of electrodes 128. For example, ECG device 110 can include electrodes for generating a 12-lead ECG. In this example, the leads can include leads I, II, III, aVF, aVR, aVL, V1, V2, V3, V4, V5, and V6. ECG device 110 can be configured to use a subset of standard 12 leads, or alternatively use ECGs with non-standard lead placements (e.g., Holter monitor lead placements), or use synthetic ECG leads from the actual acquired lead set (e.g., GE HealthCare's 12RL algorithm for assembling 12-lead ECGs from a reduced set of leads).

[0025] ECG analysis device 112 can be configured to receive ECGs from ECG device 110 and use an AI-based ECG analysis model 114 to output diagnoses. In some examples, the AI-based ECG analysis model 114 is a machine learning model. For example, the AI-based ECG analysis model 114 can be a deep neural network (DNN), convolutional neural network (CNN), recurrent neural network (RNN), or other types of machine learning model. The AI-based ECG analysis model 114 can be trained on a training dataset to generate and output diagnoses based on features identified by the ingested ECGs. In some examples, the training data used to train the AI-based ECG analysis model 114 may be unknown. In other examples, such as when the AI-based ECG analysis model 114 is generated internally, the training data used to train the AI-based ECG analysis model 114 may be known.

[0026] The interpretability and bias module 116 may include instructions for adding interpretability to the AI-based ECG analysis model 114 based on interpretability criteria determined using platform 118 and AI model 120. Furthermore, the interpretability and bias module 116 may include instructions for evaluating biases in the output of the AI-based ECG analysis model 114. In some examples, bias evaluation may be performed in part by subject matter experts.

[0027] Platform 118 can be configured to use AI model 120 to extract interpretable criteria for diagnoses determined by AI-based ECG analysis model 114. For example, platform 118 can be a server, cloud computing system, etc. The extracted interpretable criteria can include one or both of single ECG criteria and longitudinal criteria.

[0028] AI model 120 can be configured to determine interpretability criteria for diagnoses determined by AI-based ECG analysis model 114. For example, AI model 120 can be a decision tree (e.g., a classification tree or regression tree), a linear regression model, a neural network (e.g., a deep neural network (DNN), a convolutional neural network (CNN), or a recurrent neural network (RNN)), a logistic regression model, a support vector machine, etc.

[0029] User equipment 122 may be configured to receive criteria for a diagnosis determined by AI-based ECG analysis model 114, and to provide the criteria for display. For example, user equipment 122 may be a smartphone, laptop computer, desktop computer, wearable device, medical device, etc. User equipment 122 may be configured to display the diagnosis determined by AI-based ECG analysis model 114, as well as the model's characteristics (e.g., prediction score) and bias assessment.

[0030] Database 124 can be configured to store ECGs, ECG feature sets, ECG diagnoses, ECG diagnostic interpretations, modification information, patient information associated with ECGs, ECG classification targets, etc. Furthermore, database 124 can be configured to store criteria used to determine diagnoses by an AI-based ECG analysis model 114, including single ECG criteria and / or longitudinal criteria. For example, database 124 can be a hierarchical database, a network database, a relational database, etc. In some examples, the criteria stored in database 124 may initially be generated by subject matter experts. For example, database 124 could be a Marquette database. TM 12SL ECG analysis program or other similar program database.

[0031] One or more medical data repositories 130 may include one or more of the following: an electronic medical record (EMR) database, an ECG database, a radiology information system (RIS), a picture archiving and communication system (PACS), or other types of databases configured to store medical records for multiple patients. For example, an ECG performed on a patient may be stored in an ECG database along with associated data, including individual ECG features, corresponding diagnostic interpretations, acquisition dates, ordering providers, etc.

[0032] Network 126 can be configured to allow communication between devices of system 100. For example, network 126 can be a cellular network (e.g., fifth-generation (5G) network, long-term evolution (LTE) network, third-generation (3G) network, code division multiple access (CDMA) network, etc.), public land mobile network (PLMN), local area network (LAN), wide area network (WAN), metropolitan area network (MAN), telephone network (e.g., public switched telephone network (PSTN)), private network, ad hoc network, intranet, Internet, fiber-based network, etc., and / or combinations of these or other types of networks.

[0033] As an example, the interpretability and bias module 116 can communicate with the ECG analysis device 112 to obtain its output via network 126. Furthermore, the ECG analysis device 112 can access data from database 124 and / or one or more medical data repositories 130 via network 126. In this way, system 100 can be an interconnected system through which modules communicate with each other to perform various processes and / or methods.

[0034] Figure 1 The number and arrangement of devices in the system 100 shown are provided as an example. In practice, the system 100 may include additional devices, fewer devices, different devices, or devices connected to... Figure 1 The devices shown are arranged differently. Additionally or alternatively, a group of devices in system 100 (e.g., one or more devices) may perform one or more functions described as being performed by another group of devices in system 100.

[0035] Figure 2 It shows Figure 1 The diagram illustrates example components of device 200 in the example ECG analysis system 100. Device 200 may correspond to ECG device 110, ECG analysis device 112, platform 118, user device 122, interpretability and bias module 116, and / or database 124. Figure 2 As shown, device 200 may include bus 210, processor 220, memory 230, storage component 240, input component 250, output component 260 and communication interface 270.

[0036] Bus 210 may include components that allow communication between components of device 200. Processor 220 may be implemented using hardware, firmware, or a combination of hardware and software. Processor 220 may be a central processing unit (CPU), graphics processing unit (GPU), accelerated processing unit (APU), microprocessor, microcontroller, digital signal processor (DSP), field-programmable gate array (FPGA), application-specific integrated circuit (ASIC), or another type of processing component. Processor 220 may include one or more processors capable of being programmed to perform functions. Specifically, processor 220 may include one or more processors 220 configured to perform the operations described herein. Alternatively, multiple processors 220 may be collectively configured to perform the operations described herein, and each of the multiple processors 220 may be configured to perform a subset of the operations described herein. For example, a first processor 220 may perform a first subset of the operations described herein, a second processor 220 may be configured to perform a second subset of the operations described herein, and so on.

[0037] Memory 230 may include non-transitory memory, random access memory (RAM), read-only memory (ROM), and / or another type of dynamic or static storage device (e.g., flash memory, magnetic memory, and / or optical memory) that stores information and / or instructions for use by processor 220. For example, memory 230 may store executable instructions thereon that can be executed by processor 220 to perform the operations described herein.

[0038] Storage component 240 may store information and / or software related to the operation and use of device 200. For example, storage component 240 may include hard disk (e.g., magnetic disk, optical disk, magneto-optical disk and / or solid-state disk), compact disc (CD), digital versatile disc (DVD), floppy disk, cassette, magnetic tape and / or another type of non-transitory computer-readable medium and corresponding drives.

[0039] Input component 250 may include components that allow device 200 to receive information such as via user input (e.g., a touchscreen display, keyboard, keypad, mouse, buttons, switches, camera, and / or microphone for receiving reference audio input and / or visual input). Additionally or alternatively, input component 250 may include sensors for sensing information (e.g., a Global Positioning System (GPS) component, accelerometer, gyroscope, and / or actuator). Output component 260 may include components that provide output information from device 200 (e.g., a display, a speaker for outputting sound at an output sound level, and / or one or more light-emitting diodes (LEDs)).

[0040] Communication interface 270 may include transceiver-like components (e.g., a transceiver and / or separate receiver and transmitter) that enable device 200 to communicate with other devices, such as via a wired connection, a wireless connection, or a combination of wired and wireless connections. Communication interface 270 may allow device 200 to receive information from and / or transmit information to another device. For example, communication interface 270 may include an Ethernet interface, an optical interface, a coaxial interface, an infrared interface, a radio frequency (RF) interface, a universal serial bus (USB) interface, a Wi-Fi interface, a cellular network interface, etc.

[0041] Device 200 can execute one or more processors and methods described herein. Device 200 can perform these processes and / or methods based on processor 220 executing software instructions stored in a non-transitory computer-readable medium such as memory 230 and / or storage component 240. Computer-readable medium may be defined herein as a non-transitory memory device. A memory device may include memory space within a single physical storage device or memory space distributed across multiple physical storage devices.

[0042] Software instructions may be read into memory 230 and / or storage component 240 via communication interface 270 from another computer-readable medium or from another device. When executed, the software instructions stored in memory 230 and / or storage component 240 may cause processor 220 to perform one or more processes and / or methods described herein. Additionally or alternatively, hardwired circuitry may be used in place of or in combination with software instructions to perform one or more methods and / or processes described herein. Therefore, the specific implementations described herein are not limited to any particular combination of hardware circuitry and software.

[0043] Figure 2 The number and arrangement of components shown are provided as an example. In practice, device 200 may include additional components, fewer components, different components, or components related to... Figure 2 The components shown are arranged differently. Additionally or alternatively, a group of components of device 200 (e.g., one or more components) may perform one or more functions described as being performed by another group of components of device 200.

[0044] See now Figure 3A A flowchart illustrating method 300 for adding interpretability to an AI-based ECG analysis model and evaluating its biases is shown. For example, the AI-based ECG analysis model could be relative to... Figure 1The AI-based ECG analysis model 114 is described. Therefore, the AI-based ECG analysis model can be a DNN, CNN, or other type of machine learning model. Method 300 can be executed by one or more processors (e.g., processor 220) according to instructions stored in non-transitory memory (e.g., memory 230).

[0045] At 302, method 300 includes obtaining diagnostic output from an AI-based ECG analysis model on an ECG dataset from a general population with unknown ground truth values. The diagnostic output can be obtained during the model development phase, thus serving as an exemplary output given for model evaluation before actual deployment. The ECG dataset may include one or more ECGs obtained from a general population (e.g., a patient population with multiple diagnoses, medical histories, and ECG characteristics). The AI-based ECG analysis model can output a diagnosis and / or diagnostic interpretation for each ECG in the ECG dataset. Therefore, the obtained diagnostic output may include output for each ECG in the ECG dataset.

[0046] Each ECG in the ECG dataset may include a feature set (e.g., ECG features). For example, the feature set may include the amplitude and duration of the P wave, the amplitude and duration of the Q wave, the amplitude and duration of the R wave, the PR interval, the amplitude and duration of the S wave, the duration of the QRS complex, the amplitude and duration of the T wave, the QT interval, etc. Additionally or alternatively, the feature set may be defined by data from a single lead (e.g., a single-lead specific P wave amplitude or P wave duration), or by data from multiple leads (e.g., the global QT duration from the earliest Q onset to the latest T offset across all leads, or QT dispersion, which is the difference between the shortest and longest single-lead QT interval measurements across all leads), or by data from spatial relationships across multiple leads (e.g., spatial QRS-T angles or ventricular gradients).

[0047] At 304, method 300 includes extracting interpretable criteria on an ECG dataset to predict and interpret the diagnostic output of the AI-based ECG analysis model, targeting the diagnostic output from the model. An exemplary method for extracting the criteria is relative to... Figure 5 Further description. The extracted criteria may include combinations and / or sub-combinations of ECG features (e.g., parameters), which together lead to an interpretation endpoint. As noted, the interpretation endpoint may correspond to the diagnostic output of the ECG analysis model. The interpretation endpoint may be a diagnostic interpretation, diagnosis, etc.

[0048] In some examples, one or more characteristics of the extracted criteria can also be identified, which provide further context to the diagnostic output of the AI-based ECG analysis model. For example, the predicted value of each criterion in the criteria can be determined. In some examples, criteria with high positive or negative predicted values ​​can be provided as interpretive context for the output of the ECG analysis model for any ECG acquired in the future. Therefore, extracting criteria and identifying those with high predicted values ​​can add interpretability to neural networks such as AI-based ECG analysis models.

[0049] At 306, method 300 includes deploying an AI-based ECG analysis model for evaluation. As noted, interpretable criteria are extracted during the development phase for evaluating the AI-based ECG analysis model. Once the evaluation is performed, the model can be deployed on newly acquired ECGs. The newly acquired ECGs can be directly from ECG devices (e.g., Figure 1 ECG data obtained from an ECG device 110, or from a medical data repository, such as an ECG database (e.g., ECG data from an ECG device 110). Figure 1 ECG obtained from one or more medical data repositories 130.

[0050] At 308, method 300 includes outputting a diagnosis and corresponding interpretability criteria for a newly acquired ECG from an evaluated AI-based ECG analysis model to a user device. As an example, a subset of interpretability criteria applicable to the extraction of the newly acquired ECG can be determined based on features of the newly acquired ECG or based on the output of the AI-based ECG analysis model. The diagnosis and corresponding interpretability criteria can be output to a user device relative to... Figure 1 The user equipment 122 described herein is communicatively coupled to an AI-based ECG analysis model, an ECG device, and a module configured with instructions for adding interpretability as described in method 300 herein. As an example, interpretability criteria can be displayed on the user equipment along with the diagnostic output of the AI-based ECG analysis model in the form of lists, charts, panels, etc. In this way, the user can view the displayed interpretability criteria as well as the displayed diagnostic output from the AI-based ECG analysis model, and in some examples, view the ECG tracing lines themselves, thereby allowing for easy viewing and comparison between them to enhance understanding of the diagnostic output.

[0051] In this way, when the ground truth of the ECG dataset is unknown, interpretability can be added to AI-based ECG analysis models by extracting criteria from the ECG dataset during model development and evaluation. Therefore, when deployed for ECG, interpretable criteria can also be output to explain the output of the AI-based model. For example, criteria can be provided as an interpretable context at deployment time, allowing users to better understand and act on the diagnoses of the AI-based ECG analysis model. Furthermore, biases can be evaluated based on the output and the predicted classification output generated based on the extracted criteria, as described below.

[0052] Figure 3B A flowchart illustrating method 350 for evaluating biases in an AI-based ECG analysis model is shown. For example, the AI-based ECG analysis model can be relative to... Figure 1 The AI-based ECG analysis model 114 is described. Therefore, the AI-based ECG analysis model can be a DNN, CNN, or other type of machine learning model. Method 300 can be executed by one or more processors (e.g., processor 220) according to instructions stored in non-transitory memory (e.g., memory 230). In some examples, method 350 can be executed during the development phase of the AI-based ECG analysis model.

[0053] At 352, method 350 includes obtaining the extracted interpretability criteria for a general population ECG dataset with unknown ground truth. As described above relative to method 300, the diagnostic output of an AI-based ECG analysis model on a general population ECG dataset with unknown ground truth can be obtained. Using the diagnostic output as the target, the interpretability criteria for the general population ECG dataset can be determined. Interpretability criteria can be extracted from an ECG dataset with known ground truth based on the diagnostic output of the AI-based ECG analysis model (e.g., where the diagnostic output of the AI-based ECG analysis model is the target or interpretability endpoint).

[0054] At 354, method 350 includes extracting criteria on an ECG dataset with known ground truth for the diagnostic output of the AI-based ECG analysis model. In some examples, the ECG dataset may be a dataset specific to a particular selected population, and the specific selected population may provide interpretive endpoints, thus giving the ground truth of the dataset. In some examples, the interpretive endpoints of such a selected population may be confirmed via other procedures, such as echocardiography. In other examples, the ECG dataset may be a general population dataset with simple or common endpoints (e.g., interpretive endpoints common in the general population), for which ground truth can be obtained from the interpretation of an expert reviewer. In this case, predicting the ground truth may refer to the process of extracting interpretable criteria that predict and / or characterize the ground truth. This set of criteria may be compared with criteria characterizing the predictions of the AI-based ECG analysis model for the general population ECG dataset (e.g., interpretable criteria extracted in method 300), as described below. Any bias due to overfitting can be determined based on this comparison.

[0055] At 356, method 350 includes comparing the extraction criteria from a general population ECG dataset with unknown ground truth, as obtained at 352, with the extraction criteria from an ECG dataset with known ground truth. At 358, method 350 may then include determining whether there is a difference between the extraction criteria from the general population ECG dataset with unknown ground truth and the extraction criteria from the ECG dataset with known ground truth. The difference between them may indicate the presence of bias in the neural network. For example, an ML-based process can be trained on a target to extract interpretable criteria (e.g., ECG patterns) representing that target. The target may be the ground truth of an ECG dataset, or it may be a prediction from an AI model. The comparison between the two sets of extraction criteria can be a very deep probe of potential biases in an AI-based ECG analysis model. If a difference exists (yes at 358), method 350 proceeds to 360 to output a notification indicating the presence of neural network bias to a user, for example, via a user device (e.g., user device 122). The bias may be due to overfitting, distribution shift, etc. If there is no difference (no at 358), then method 350 proceeds to 362.

[0056] At 362, method 350 includes outputting extracted interpretable criteria from a general population ECG dataset with unknown ground truth to add interpretability to the AI-based ECG analysis model. For example, similar to those described with respect to method 300, the extracted criteria can be used during the model development phase and then output along with diagnostics from the deployment of the AI-based ECG analysis model on newly acquired ECGs. As described with respect to method 300, interpretable criteria determined for the general population ECG dataset to interpret the model's diagnostic output can be output along with the model's diagnostic output to a user device (e.g., user device 122). Thus, the user can view the interpretable criteria to be able to interpret and understand the diagnostic output of the AI-based ECG analysis model.

[0057] In this way, interpretability can be added to AI-based ECG analysis models even when the ground truth of the ECG dataset is unknown. For example, interpretability can be added by extracting criteria from the ECG dataset for which the ECG analysis model provides an interpretability endpoint for a given diagnostic output during model development. Then, once the interpretability criteria are extracted, the model can be deployed, and the corresponding criteria can be provided to interpret the model's output.

[0058] Now go to Figure 4 The diagram illustrates a flowchart of method 400 for evaluating biases in an AI-based ECG analysis model. In some examples, method 400 may be performed in conjunction with method 300 or method 350, depending on whether the ground truth of the ECG dataset is known. For example, the AI-based ECG analysis model could be relative to... Figure 1 The AI-based ECG analysis model 114 is described. Therefore, the AI-based ECG analysis model can be a DNN, CNN, or other type of machine learning model. Method 300 can be executed by one or more processors (e.g., processor 220) according to instructions stored in non-transitory memory (e.g., memory 230).

[0059] At 402, method 400 includes extracting interpretable criteria from the ECG dataset from the general population. For example, as relative to... Figure 3A The described method obtains diagnostic output from an AI-based ECG analysis model on an ECG dataset from a general population. In some examples, the ground truth may not be available for datasets from a general population and is therefore not necessary for Method 400. Thus, the general population ECG dataset may be referred to in this paper as having unknown ground truth. Utilizing the diagnostic output of the AI-based ECG analysis model as an interpretation endpoint, interpretable criteria can be extracted from the ECG dataset, such as relative to... Figure 5Further description: The extracted interpretable criteria can therefore interpret the diagnostic output from the AI-based ECG analysis model.

[0060] At 404, method 400 includes evaluating the bias in the AI-based ECG analysis model based on extracted interpretability criteria. The likelihood of bias in the AI-based ECG analysis model can be assessed based on interpretability criteria that explain its output, such as relative to... Figure 3B Other techniques for evaluating bias can include comparisons with expectations from the literature, as determined by human experts, and / or comparisons with standards that interpret the output of an AI-based ECG analysis model extracted from its training dataset, such as in examples where the ground truth is known. Therefore, the ability to evaluate on general population ECG datasets accesses previously undiscovered forms of bias. For example, this approach can be used to detect biases that may be caused by distribution shifts. In general, even external testing of AI-based ECG analysis models is performed on selected populations where the ground truth is known. Therefore, previously undiscovered forms of bias can be discovered by evaluating the output of AI-based ECG analysis models on general populations.

[0061] At 406, method 400 determines whether a bias is detected. Determining whether a bias is detected can indicate the suitability of the AI-based ECG analysis for deployment and / or indicate to the user whether the provided output can be trusted. If a bias is detected based on the techniques described above, method 400 proceeds to 408. If no bias is detected, method 400 terminates. When no bias is detected, it can be highly confident that the AI-based ECG model can be deployed without bias, and its diagnostic output can be trusted upon deployment. For example, interpretable criteria extracted for the AI-based ECG analysis model can be output as described above without making any changes to the AI-based ECG analysis model or its output.

[0062] At 408, when a bias is detected, method 400 includes modifying the use of the AI-based ECG analysis model. In some examples, modifying the use of the AI-based ECG analysis model may include updating the AI-based ECG analysis model based on the detected bias, as noted at 410. For example, the AI-based ECG analysis model may revert to the model development phase to reduce the likelihood that the model acquires the detected bias, which can be achieved by training on a dataset that may be more representative of the population on which the model is intended to be deployed. In such examples, the methods of this paper can be repeated to re-evaluate the bias to determine whether the updated model has removed the detected bias. In other examples, modifying the use of the AI-based ECG analysis model may include deploying the AI-based ECG analysis model with limited applicability, as noted at 412. For example, notifications may be output along with extracted interpretability criteria, as these are related to diagnostic outputs indicating that a bias has been detected, and therefore the applicability of the diagnostic outputs may be limited. For example, if the possibility of bias is detected in the presence of certain criteria, the model output can be suppressed in the presence of those criteria, or a warning conveying reduced confidence can be displayed as a notification to the user on the user's device.

[0063] Regardless of whether a bias is detected, the bias assessment can be output to the user device. For example, when no bias is detected, a notification indicating that no bias was detected can be displayed on the user device along with extracted interpretable criteria. When a bias is detected, a bias assessment indicating the presence of the bias and / or how the model should be updated (e.g., reverting to model development or deployment with limited applicability) can be displayed on the user device (e.g., user device 122). In some examples, the amount or level of the bias can be displayed as a comparison with a predefined threshold bias amount. For example, if a bias amount below a predefined threshold is detected, the bias assessment can be displayed to include the bias amount and the comparison with the predefined threshold. Therefore, both the bias assessment and the interpretable criteria can be displayed on the user device along with the diagnostic output of the AI-based ECG analysis model, allowing the user to analyze and understand the model's output more comprehensively.

[0064] Go to Figure 5 The diagram illustrates a flowchart of method 500 for extracting interpretable terms from an ECG dataset to add interpretability to an AI-based ECG analysis model (e.g., DNN, CNN, RNN, etc.). Method 500 can be executed by one or more processors (e.g., processor 220) according to instructions stored in non-transitory memory (e.g., memory 230). It should be understood that method 500 described herein is exemplary in nature, and other methods for extracting interpretability criteria to add interpretability to an AI-based ECG analysis model may be possible without departing from the scope of this disclosure.

[0065] At 502, method 500 includes receiving training data for training an AI model to extract interpretability criteria from an ECG dataset. The AI ​​model can be trained to extract interpretability criteria to add interpretability to AI-based ECG analysis models such as those disclosed herein. As an example, platforms (such as...) Figure 1 The ECG analysis system 100 platform 118 can receive data for training AI models (such as...). Figure 1 The training data of AI model 120.

[0066] As an example, training data may include ECG as input and criteria as targets. For example, ECG may be a waveform acquired via ECG device 110. Additionally or alternatively, the input to the training data may include a feature set of the ECG. For example, the feature set may include the amplitude and duration of the P wave, the amplitude and duration of the Q wave, the amplitude and duration of the R wave, the PR interval, the amplitude and duration of the S wave, the duration of the QRS complex, the amplitude and duration of the T wave, the QT interval, etc. Additionally or alternatively, the feature set may be defined by data from a single lead (e.g., a single lead-specific P wave amplitude or P wave duration), or by data from multiple leads (e.g., the global QT duration from the earliest Q onset to the latest T offset across all leads, or QT dispersion, which is the difference between the shortest and longest single-lead QT interval measurements across all leads), or by data from spatial relationships across multiple leads (e.g., spatial QRS-T angles or ventricular gradients).

[0067] The target of the training data can include a set of ECG criteria. ECG criteria can include combinations of one or more ECG features corresponding to a diagnosis. For example, the input ECG can include a first feature and a second feature, and the target criteria can be a combination of the first and second features that together correspond to a diagnosis. Therefore, for example, an AI model can be trained to ingest ECGs, identify the first and second features, and extract criteria from the combination of the first and second features.

[0068] Therefore, training data can include ECG diagnoses. A diagnosis can be the endpoint of the extraction. For example, diagnoses can include atrial pacing rhythm, ventricular pacing rhythm, atrial flutter, ectopic atrial tachycardia, sinus bradycardia, junctional bradycardia, atrial fibrillation, LBBB, LVH, septal infarction, non-ST-segment elevation myocardial infarction, ST-segment elevation myocardial infarction, etc. Furthermore, the diagnostic interpretation can be the interpretation of the ECG results determined by a physician. Additionally or alternatively, the diagnostic interpretation can be a diagnosis made by a physician using an ECG alone, or alternatively using another source of clinical information (e.g., high-sensitivity troponin levels, cardiac echo measurements, or angiographic findings), or a combination of ECG and non-ECG clinical information.

[0069] In addition, training data may include patient information associated with ECG. For example, patient information may identify whether a specific diagnosis exists in the patient's previous diagnostic interpretations, whether the physician has previously modified the diagnostic interpretations, the patient's demographic information, the patient's health status (e.g., previous diagnoses, comorbidities, etc.), medications prescribed to the patient, previous procedures performed on the patient, and the patient's previous diagnoses (e.g., resolved diagnoses), etc.

[0070] Furthermore, the training data can include ECG classification objectives. For example, classification objectives could be ECG diagnosis, ECG diagnostic interpretation, ECG modification information, etc., as noted in this paper.

[0071] At 504, method 500 includes training an AI model using training data of ECG features (input) and diagnosis (target). Specifically, the target is diagnosis, and the diagnostic criteria are extracted through optimization as the diagnostic target is achieved, using the procedures described herein. For example, Figure 1 The platform 118 of the ECG analysis system 100 can train the AI ​​model 120 based on training data. Alternatively, systems or devices other than the platform 118 can be used to generate and / or train the AI ​​model 120. For example, the system or device may include instructions for generating the AI ​​model 120 and / or instructions for training the AI ​​model 120. The system or device can provide the resulting trained AI model 120 to the platform 118 for use.

[0072] In some examples, AI model 120 may include a training phase, a deployment phase, and a monitoring phase. During the training phase, platform 118 may receive and process training data to generate a trained AI model 120 for extracting criteria used to interpret diagnostic standards through the ECG analysis model. The training data may include multiple training datasets, each comprising one or more of the following: ECG, ECG feature sets, ECG diagnoses, ECG diagnostic interpretations, modification information, patient information associated with ECGs, ECG classification targets, etc. Each of the multiple training datasets may be associated with a specific ECG and a specific diagnostic interpretation result.

[0073] Training data may be generated, received, or otherwise obtained from internal and / or external resources. For example, training data may be generated, received, or otherwise obtained from ECG device 110, ECG analysis device 112, user device 122, and / or database 124.

[0074] Typically, AI model 120 may include a set of variables (e.g., nodes, neurons, filters, etc.) tuned (e.g., weighted or biased) to different values ​​through the application of training data. According to one embodiment, the training process may employ supervised, unsupervised, semi-supervised, and / or reinforcement learning processes to train AI model 120. According to one embodiment, a portion of the training data may be retained during training and / or used to validate the trained AI model 120.

[0075] For a supervised learning process, training data may include labels or scores that can facilitate the training process by providing ground truth values. For example, labels or scores may indicate a classification target. Training can continue by feeding the training dataset into AI model 120. AI model 120 may have variables set to initial values ​​(e.g., randomly, based on Gaussian noise, based on pre-trained values, etc.). AI model 120 can generate outputs. The outputs can be compared with corresponding labels or scores (e.g., ground truth values) and then backpropagated through AI model 120 to adjust the values ​​of the variables. This process can be repeated for multiple samples, at least until the determined loss or error is below a predefined threshold. According to one implementation, some of the training data may be retained and used for further validation or testing of the trained AI model 120.

[0076] For unsupervised learning processes, training data may not include pre-assigned labels or scores to aid the learning process. Instead, unsupervised learning processes can include clustering, classification, etc., to identify patterns naturally present in the training data. As an example, training data can be clustered into groups based on identified similarities and / or patterns. K-means clustering or K-nearest neighbors can also be used, and these can be supervised or unsupervised. A combination of K-nearest neighbors and unsupervised clustering techniques can also be used. For semi-supervised learning, a combination of training data with pre-assigned labels or scores and training data without pre-assigned labels or scores can be used to train an AI model120.

[0077] When reinforcement learning is employed, an agent (e.g., an algorithm) can be trained to make decisions from training data through trial and error regarding whether a diagnostic interpretation should be modified. For example, based on the decisions made, the agent can then receive feedback (e.g., a positive reward for predicting a value above a predetermined threshold), adjust its next decision to maximize the reward, and repeat until the loss function is optimized.

[0078] After being trained, the trained AI model 120 can be stored and subsequently applied by platform 118 during the deployment phase. For example, during the deployment phase, the trained AI model 120, executed by platform 118, can receive input data and generate output data. During the monitoring phase, monitoring data can be analyzed along with the output and input data to determine the accuracy of the trained AI model 120. According to one implementation, based on the analysis, platform 118 can return to the training phase, where the values ​​of one or more variables of the AI ​​model 120 can be adjusted to improve the accuracy of the AI ​​model 120.

[0079] According to one implementation, AI model 120 can be a decision tree. In this case, platform 118 can use training techniques to generate the decision tree. For example, training techniques may include random forest, boosting tree, bootstrapping, rotating forest, etc. AI model 120 may include a set of nodes. For example, the set of nodes may include a root node, one or more intermediate nodes, and leaf nodes. Platform 118 can use attribute selection metrics to generate AI model 120. For example, attribute selection metrics may be information gain, gain ratio, Gini index, etc. Platform 118 can generate the decision tree and use pruning techniques to prune the decision tree. For example, pruning techniques may be cost complexity pruning, error reduction pruning, etc.

[0080] At point 506, method 500 includes extracting interpretable criteria from the general population ECG dataset. For example, Figure 1Platform 118 can use AI model 120 to extract criteria for diagnoses determined by AI-based ECG analysis model 114. In some examples, a general population ECG dataset can be a dataset on which an AI-based ECG analysis model is deployed. Interpretable criteria can be extracted relative to the diagnostic output of the AI-based ECG analysis model. For example, the diagnostic output of the AI-based ECG analysis model on a general population ECG dataset can be obtained and then used as an interpretation endpoint when extracting interpretable criteria.

[0081] As described herein, platform 118 can determine the decision branch of AI model 120 corresponding to a specific target classification. In this example, the specific target classification can be the diagnostic output of an AI-based ECG analysis model. For example, the target classification could be an ECG diagnosis, an ECG diagnostic interpretation result, ECG modification information, etc. The decision branch can include one or more nodes corresponding to a relevant criterion. For example, a decision branch can include a root node, one or more intermediate nodes, and leaf nodes.

[0082] Furthermore, in some examples, platform 118 can determine a decision branch of AI model 120, which includes one or more nodes associated with an attribute selection metric that satisfies a threshold. For example, the attribute selection metric could be information gain, gain ratio, Gini index, etc. As a specific example, platform 118 can determine a decision branch that includes leaf nodes corresponding to a criterion with a Gini index less than a threshold. According to one implementation, platform 118 can determine the decision branch based on a metric of the decision path. For example, the metric could be accuracy, positive prediction value, sensitivity, etc. Additionally or alternatively, platform 118 can determine the decision branch based on the performance of generalization to an external dataset.

[0083] Based on one or more examples, platform 118 can identify specific nodes among one or more nodes for criterion extraction. Platform 118 can identify specific nodes based on feature selection metrics. Feature selection metrics can correspond to the importance of features associated with the criteria corresponding to the node. According to one implementation, platform 118 can fit another AI model 120 (e.g., a decision tree) to a dataset separate from the training dataset. In this way, platform 118 can reduce the risk of overfitting and increase the statistical power of the technique.

[0084] In some examples, platform 118 can extract criteria from decision branches. For example, platform 118 can extract criteria belonging to a decision branch. Each criterion may include a corresponding feature and a corresponding threshold. In some examples, platform 118 can determine the threshold for the criterion. For example, platform 118 may determine the threshold based on optimization techniques, based on input from user device 122, etc.

[0085] Furthermore, platform 118 can generate rules that include the extracted set of criteria. For example, rules may include criteria from a specific decision branch. Alternatively, rules may include criteria from different decision branches. Platform 118 can generate rules using permutations and / or combinations of criteria. According to one implementation, platform 118 can determine the metrics of the rules. For example, metrics may be accuracy, positive predictive value, sensitivity, etc. Furthermore, platform 118 can determine that rules include metrics that satisfy thresholds.

[0086] Furthermore, in some examples, the predicted value of each criterion extracted can be determined, providing further context to the diagnostic output of the AI-based ECG analysis model. In some examples, as previously described, criteria with high positive or negative predicted values ​​can be provided as interpretive context for the ECG analysis model output for any ECGs acquired in the future. Therefore, extracting criteria and identifying those with high predicted values ​​adds interpretability to neural networks such as AI-based ECG analysis models.

[0087] At 508, method 500 includes an output-extracted interpretability criterion for adding interpretability to the AI-based ECG analysis model. As described above, the extracted interpretability criterion can be used to interpret the output of the AI-based ECG analysis model when the output is generated on the same dataset for which the extraction criterion was extracted. As an example, the criterion extraction presented herein can be incorporated into methods for adding interpretability and evaluating biases during the development phase of an AI-based ECG analysis model (e.g., into methods 300, 350, and 400 presented above).

[0088] It should be understood that method 500 is provided herein as an example of how to extract the standard, and other methods may be performed without departing from the scope of this disclosure.

[0089] Now go to Figure 6 Figure 600 illustrates example criteria used to determine outputs that can add interpretability to AI-based ECG analysis models. (Example: ...) Figure 6 As shown, platform 118 can generate a first group of standards 610, a second group of standards 620, a third group of standards 630, a fourth group of standards 640, and a fifth group of standards 650. Platform 118 can generate standard groups based on standards extracted from the corresponding decision branches of AI model 120. Platform 118 can group standards based on some commonalities or similarities, and can generate groups based on the grouping of standards.

[0090] As an example, the first group 610 may include criteria corresponding to non-monomorphic R waves in lateral leads, including peak amplitude of the S wave, peak amplitude in leads 5 and 6, etc. The second group 620 may include criteria corresponding to Q waves in lateral leads, including peak amplitude of the Q wave, Q wave duration, etc. The third group 630 may include criteria corresponding to short R wave peak times in lateral leads, including peak time of the R wave in leads 5 and 6. The fourth group 640 may include criteria corresponding to the lack of appropriate QRS-ST inconsistencies, including ST voltage greater than a threshold in various leads, etc. The fifth group 650 may include criteria corresponding to finite QRS duration, including QRS duration, R wave duration, S wave duration, etc. Therefore, the grouping of criteria can be performed based on an overall theme. Criteria can then be assigned to groups based on a satisfactory theme.

[0091] Figure 7 A graph 700 shows the prediction scores of AI-based ECG analysis models (such as AI-based ECG analysis model 114). A first histogram 704 shows the model prediction scores for a large, general population. Most samples are below the threshold prediction score 702, resulting in negative findings at the search endpoints. A second histogram 706 shows the model prediction scores in the presence of some combination of criteria for identification. The area of ​​the second histogram 706, as shown, is normalized for direct comparison with the first histogram 704.

[0092] When the identification criteria are met, the AI-based ECG analysis model can predict scores higher than the threshold prediction score of 702. Therefore, for those samples, and for other samples that meet those criteria, the criteria are more likely to predict and explain the model's classification output. Thus, in the future, criteria with high predictive values ​​could be provided as interpretive context for the model.

[0093] Now go to Figure 8 Figure 800 shows the diagnostic output rate of an AI-based ECG analysis model. Bias evaluation can be performed on the output of the ECG analysis model by parsing the criteria of the identifiers and their combinations to obtain an indication of spurious activity. Figure 8 In the example shown, the ECG analysis model can output positive predictions more frequently in the presence of a first diagnosis (e.g., atrial fibrillation) and a second diagnosis (e.g., pacing rhythm). For example, a first count 802 for the first diagnosis may correspond to a negative prediction, and a second count 804 may correspond to a positive prediction for the first diagnosis. Similarly, a first count 806 may correspond to a negative prediction, and a second count 808 may correspond to a positive prediction. The counts corresponding to positive predictions can be significantly higher than the counts corresponding to negative predictions.

[0094] The finding that the positive predictive value increases in the presence of these diagnoses could indicate a bias in AI-based ECG analysis models. For example, the higher positive predictive value in the presence of diagnoses suggests that the model associates those diagnoses with the target endpoint, but the correlation may be spurious or expected to differ between the training dataset and the general population for a specific endpoint targeted by the AI-based ECG analysis model.

[0095] The technical advantage of the system and method disclosed in this paper is that interpretability can be added to the AI-based ECG analysis model even when the ground truth of the ECG dataset on which the AI-based ECG analysis model is deployed is unknown. Adding interpretability increases the usability of the AI-based ECG analysis model in a clinical setting because users can more easily understand the reasons behind the model's diagnoses and their interpretations. Furthermore, biases can be assessed, and if present, the user can be notified of the presence of biases in the AI-based ECG analysis model, allowing for a more accurate evaluation of the output.

[0096] As used herein, elements or steps listed in the singular and beginning with the word "a" or "an" should be understood to not exclude multiple said elements or steps unless such exclusion is explicitly stated. Furthermore, references to "an embodiment" of the invention are not intended to be construed as excluding the existence of additional embodiments that also include the referenced features. Moreover, unless explicitly stated to the contrary, embodiments that "comprise," "include," or "have" elements or multiple elements having a particular characteristic may include additional such elements that do not have that characteristic. The terms "comprise" and "in" are used as concise linguistic equivalents to the corresponding terms "comprising" and "wherein." Furthermore, the terms "first," "second," and "third," etc., are used merely as notations and are not intended to impose numerical requirements or a particular order of position on their objects.

[0097] This written description uses examples to disclose the invention, including the best mode, and also enables those skilled in the art to practice the invention, including making and using any device or system and performing any included methods. The scope of patentability of the invention is defined by the claims, but may include other examples that would occur to those skilled in the art. Such other examples are intended to fall within the scope of the claims if they have structural elements that are not indistinguishable from the literal language of the claims, or if they include equivalent structural elements that have minor differences from the literal language of the claims.

Claims

1. A method for adding interpretability to an artificial intelligence (AI)-based electrocardiogram (ECG) analysis model, the method comprising: Diagnostic output is obtained from the AI-based ECG analysis model on an ECG dataset with unknown ground truth (302); Interpretable criteria are extracted from the ECG dataset using the interpretable endpoints of the diagnostic output from the AI-based ECG analysis model (304); Deploy the AI-based ECG analysis model (306) on the ECG; as well as The diagnostic output of the AI-based ECG analysis for the ECG is output to the user equipment, and the extracted interpretable criteria corresponding to the ECG are provided, wherein the interpretable criteria interpret the diagnostic output (308).

2. The method according to claim 1, wherein the AI-based ECG analysis model is a machine learning model (114).

3. The method of claim 1, wherein evaluating the bias of the AI-based ECG analysis model comprises: The interpretation endpoint for the diagnostic output of the AI-based ECG analysis model from the ECG dataset with known ground truth extracts criteria from the ECG dataset (354); The extracted criteria from the ECG dataset with known ground truth are compared with the extracted interpretable criteria from the ECG dataset with unknown ground truth (356); as well as Determine the difference between the extracted criteria from the ECG dataset with known ground truth and the extracted interpretable criteria from the ECG dataset with unknown ground truth (358).

4. The method of claim 3, wherein the ECG dataset having known ground truth is an ECG dataset of a selected population (354).

5. The method of claim 1, wherein the ECG dataset having unknown ground truth is a general population ECG dataset (302).

6. The method of claim 1, further comprising determining the presence of a bias in the AI-based ECG analysis model, and in response to detecting the presence of a bias, modifying the use of the AI-based ECG analysis model (408), wherein modifying the use includes one or more of the following: deploying the AI-based ECG analysis model with limited applicability (412), and returning the AI-based ECG analysis model to the development phase to address the bias (410).

7. The method of claim 1, wherein extracting the interpretable criteria further comprises determining a predicted value for each of the interpretable criteria, and outputting the extracted interpretable criteria comprises outputting criteria with high predicted values ​​(304).

8. The method of claim 7, wherein the criteria with high predictive values ​​are output as interpretive context (308) for the diagnostic output of the AI-based ECG analysis model.

9. An apparatus, the apparatus comprising: A memory (230) configured to store instructions; as well as One or more processors (220), said one or more processors being configured to operate based on instructions stored in memory: During the development phase of the AI-based electrocardiogram (ECG) analysis model, diagnostic outputs for general population ECG datasets are obtained from the AI-based ECG analysis model (302). Interpretable criteria are extracted from the general population ECG dataset to target the diagnostic output, and the predicted output of the AI-based ECG analysis model is generated based on the extracted interpretable criteria (304). The existence of bias in the AI-based ECG analysis model is determined based on the extracted interpretability criteria (404); During the deployment phase of the AI-based ECG analysis model, the diagnosis of the newly acquired ECG by the AI-based ECG analysis model is determined (306); as well as The diagnostics, the extracted interpretability criteria, and the determined bias are output to the user equipment that is communicatively coupled to the device (308).

10. The device of claim 9, wherein the newly acquired ECG is obtained from an ECG device (110) communicatively coupled to the device.

11. The device of claim 9, wherein the presence of bias is determined by comparison with an expectation from the literature (404).

12. The device according to claim 9, wherein, To determine the existence of the bias, the one or more processors are further configured to: The interpretation endpoint for the diagnostic output of the AI-based ECG analysis model from the ECG dataset with known ground truth extracts criteria from the ECG dataset with known ground truth (354); The extracted criteria from the ECG dataset with known ground truth are compared with the extracted interpretable criteria from the general population ECG dataset (356); as well as Determine the difference between the extracted criteria from the ECG dataset with known ground truth and the extracted interpretable criteria from the ECG dataset with unknown ground truth (358).

13. The device according to claim 12, wherein, The one or more processors are further configured to output a bias notification to the user equipment when, based on the comparison, a difference is determined between an extracted standard from the ECG dataset having known ground truth and an extracted interpretable standard from the general population ECG dataset, wherein the bias is due to one of overfitting and distribution shift (360).

14. The device according to claim 9, wherein, One or more processors are configured to, in response to the detection of the presence of a bias, modify the use of the AI-based ECG analysis model based on instructions stored in memory (408), wherein the modification of use includes one or more of the following: deploying the AI-based ECG analysis model with limited applicability (412), and returning the AI-based ECG analysis model to the development phase to address the bias (410).

15. The device according to claim 9, wherein the AI-based ECG analysis model (114) is one of a deep neural network (DNN), a convolutional neural network (CNN), and a recurrent neural network (RNN).