Systems and methods for selection of priority-wise artificially intelligent mechanisms per one or more characteristics
Patent Information
- Application Number
- GB2025014250
- Authority / Receiving Office
- GB · GB
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-02-16
- Filing Date
- 2024-02-15
- Publication Date
- 2025-12-24
AI Technical Summary
Radiologists face challenges in trusting and integrating Artificial Intelligence (AI) systems for medical image diagnosis due to subjectivity and variability in diagnosis, leading to slow adoption in clinical practices.
A system and method for selecting priority-wise artificially intelligent mechanisms based on characteristics, involving data input, parsing of scan and patient characteristics, cohort-based vectorization analysis, and feedback mechanisms to determine trust scores and prioritize AI systems.
Enhances the reliability and trust in AI systems by providing a confidence measure and informed decision-making for radiologists, improving the adoption and efficiency of AI in radiology practices.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[0001] SYSTEMS AND METHODS FOR SELECTION OF PRIORITY- WISE ARTIFICIALLY INTELLIGENT MECHANISMS PER ONE OR MORE CHARACTERISTICS
[0002] FIELD OF THE INVENTION:
[0003] This invention relates to the field of networking systems, computations systems, communication systems, and information systems.
[0004] Particularly, this invention relates to the field of healthcare technology, healthcare management, electronic medical records, electronic health records, decision support systems, healthcare information, healthcare reporting, and doctor-patient-interaction systems.
[0005] Specifically, this invention relates to systems and methods for selection of prioritywise artificially intelligent mechanisms per one or more characteristics.
[0006] BACKGROUND OF THE INVENTION:
[0007] Radiology is a medical discipline that uses medical imaging technologies to diagnose diseases. However, the diagnosis of one radiologist may differ from that of another radiologist.
[0008] Furthermore, even trained radiologists may overlook some critical findings.
[0009] Artificial intelligence (Al) can help overcome subjectivity and improve disease detection accuracy. In recent years, Al algorithms have made substantial advances in image recognition tasks. Al-enabled medical imaging solutions, provided by various Al vendors, enable radiologists to perform accurate and accessible disease screening.
[0010] Al systems can be deployed by hospitals, and they aid radiologists by highlighting suspicious regions of interest within the scans. Trained physicians and radiologists can interpret the scan by assessing the Al-generated outputs and reporting findings associated with the scan.
[0011] Although there is tremendous interest in employing Al solutions in medical image diagnosis, there is often apprehension when it comes to actually integrating them into clinical practices. Lack of trust, and a dissatisfying user experience, are some of the reasons for slow adoption of Al in clinical practices. Radiologists are, often, uncertain how much trust they should put in Al predictions. The successful implementation of Al in radiology practices ultimately depends on the trust of radiologists in Al outputs.
[0012] Therefore, there is a need for systems and methods which solve these aforementioned problems.
[0013] OBJECTS OF THE INVENTION:
[0014] An object of the invention is to determine efficiency / efficacy / trust quotient of Artificially Intelligent mechanisms.
[0015] Another object of the invention is to improve reliability of, and reliance on, Artificially Intelligent mechanisms. Yet another object of the invention is to improve adoption of Artificially Intelligent mechanisms
[0016] SUMMARY OF THE INVENTION:
[0017] According to this invention, there are provided systems and methods for selection of priority-wise artificially intelligent mechanisms per one or more characteristics, said method comprising the steps of:
[0018] - receiving images as data items;
[0019] - identifying an artificially intelligent system being used for determination of efficacy, each of the data items, being processed by one or more identified artificially intelligent systems;
[0020] - parsing, and outputting, scan characteristics and patient characteristics, from the data items, output of said parsing being first output (scan characteristics) and second output (patient characteristics);
[0021] - analysing to receive at least a first output and / or at least a second output and configured to receive feedback signal from a feedback model; and
[0022] - serving, as an output, upon analysing, a selection of a priority- wise-ranked artificially intelligent system, from the plurality of artificially intelligent systems, said selected system being per parsed scan characteristic and / or per parsed patient characteristic.
[0023] In at least an embodiment, said paring being selected from a group of parsing steps consisting of:
[0024] - parsing, and outputting, scan characteristics, from the data items, output of said scan parser being a first output (scan characteristics); and
[0025] - parsing, and outputting, patient characteristics, from the data items, output of said patient parser being a second output (patient characteristics); In at least an embodiment, said feedback signal being received from a first feedback module, per artificially intelligent system, in that, feedback relating to accuracy metrics of analysis, correlative to performance metrics of the selected artificially intelligent system, is recorded and fed back to the step of analysing.
[0026] In at least an embodiment, said feedback signal being received from a second feedback module, per artificially intelligent system, in that, feedback correlative to first output of the selected artificially intelligent system, is recorded and fed back to the step of analysing.
[0027] In at least an embodiment, said feedback signal being received from a third feedback module, per artificially intelligent system, in that, feedback correlative to second output of the selected artificially intelligent system, is recorded and fed back to the step of analysing.
[0028] In at least an embodiment, said step of analysing being a cohort-based vectorization analysis configured to determine, and form, cohort-based data sets in order to determine, and form, cohort-based characteristics, characterized, in that,
[0029] - sorting data items, basis its data and / or its metadata, into cohorts according to sorting rules, correlative to said first output and / or said second output, defined in a sorting rule engine;
[0030] - using, by a performance processor, one or more data items, with specific metadata, as sorted by said sorter, in order to compute a performance function for all data items belonging to individual cohorts and intersections of two or more cohorts; and - configured to extracting features from the cohorts, upon which a performance function is used by said performance processor, said performance function being correlative to one or more metrics relating to said first output, said performance function being applied to each cohort in order to obtain its feature, said feature processor being configured to collate a set of obtained features, per cohort, to form a feature vector X, for each data item, each vector being shaped by a first set (“k”) of data items and a second set (“n”) of features;
[0031] In at least an embodiment, said analysis engine being a cohort-based vectorization module configured to perform the steps of:
[0032] - receiving the first output and / or the second output and a feedback signal from a feedback model;
[0033] - computing cohorts based on sorting rules defined in a sorting rule engine, followed by grouping data items with common characteristics;
[0034] - processing cohorts, using a performance processor, to compute performance metrics for individual cohorts and their intersections;
[0035] - developing feature vectors based on performance metrics, for data items, sharing common characteristics;
[0036] - setting a ground truth for training artificially intelligent system using manual mechanisms or machine-fed mechanisms;
[0037] - calculating trust scores for data items using regression or classification models with weighted features, minimizing a loss function; and
[0038] - determining prioritization of artificially intelligent system based on threshold values and a rule engine considering the relative importance of characteristics identified by machine learning models. In at least an embodiment, said step of analysing cooperating with:
[0039] - generating data, from said first feedback module, correlative to said data items;
[0040] - outputting said first output from said data items configured to be analysed by said performance parser;
[0041] - outputting said second output from said data items configured to be analysed by said performance parser; and
[0042] - establishing Ground Truth data for each cohort, in order to obtain metrics per cohort per data item to be fed to a training module; and
[0043] - comparing said established ground truth per data item with its own analysed output per data item in order to determine updateable weights based on agreement between said ground truth and said analysed output.
[0044] In at least an embodiment, said step of analysing cooperating with:
[0045] - organizing output of said analysis engine in correlation with said first output and said second output; and
[0046] - generating a mapping between each considered artificially intelligent system and its corresponding trust score.
[0047] In at least an embodiment, said method comprising the steps of:
[0048] - gathering data items, including ground truth predictions, per data item, made by one or more of the existing artificially intelligent systems, first output, and said second output; and
[0049] - using said gathered data items, their ground truth predictions, in order to feed to a Training Module which trains internal weights (W) for said analysis engine. In at least an embodiment, a training module is active only when a ranking module is inactive.
[0050] According to this invention, there are provided systems and methods for selection of priority-wise artificially intelligent mechanisms per one or more characteristics, said system comprising:
[0051] - a data input module configured to receive images, as data items;
[0052] - an identifier configured to identify an artificially intelligent system being used for determination of efficacy, each of the data items, from the data input module, being processed by one or more identified artificially intelligent systems;
[0053] - a set of parsers, configured to parse, and output, scan characteristics and patient characteristics, from the data items, output of said parsers being first output (scan characteristics) and second output (patient characteristics);
[0054] - an analysis engine configured to receive at least a first output and / or at least a second output and configured to receive feedback signal from a feedback model; and
[0055] - an output module, cooperating with the analysis engine, serving a selection of a priority-wise-ranked artificially intelligent system, from the plurality of artificially intelligent systems, said selected system being per parsed scan characteristic and / or per parsed patient characteristic.
[0056] In at least an embodiment, said set of parsers being selected from a group of parsers consisting of: - a scan parser configured to parse, and output, scan characteristics, from the data items, output of said scan parser being a first output (scan characteristics); and
[0057] - a patient parser configured to parse, and output, patient characteristics, from the data items, output of said patient parser being a second output (patient characteristics);
[0058] In at least an embodiment, said feedback model comprising a first feedback module, per artificially intelligent system, in that, feedback relating to accuracy metrics of analysis, correlative to performance metrics of the selected artificially intelligent system, is recorded and fed back to the analysis engine.
[0059] In at least an embodiment, said feedback model comprising a second feedback module, per artificially intelligent system, in that, feedback correlative to first output of the selected artificially intelligent system, is recorded and fed back to the analysis engine.
[0060] In at least an embodiment, said feedback model comprising a third feedback module, per artificially intelligent system, in that, feedback correlative to second output of the selected artificially intelligent system, is recorded and fed back to the analysis engine.
[0061] In at least an embodiment, said analysis engine being a cohort-based vectorization module configured to determine, and form, cohort-based data sets in order to determine, and form, cohort-based characteristics, characterized, in that, - a sorter sorts data items, basis its data and / or its metadata, into cohorts according to sorting rules, correlative to said first output and / or said second output, defined in a sorting rule engine;
[0062] - a performance processor configured to use one or more data items, with specific metadata, as sorted by said sorter, in order to computes a performance function for all data items belonging to individual cohorts and intersections of two or more cohorts; and
[0063] - a feature processor configured to extract features from the cohorts, upon which a performance function is used by said performance processor, said performance function being correlative to one or more metrics relating to said first output, said performance function being applied to each cohort in order to obtain its feature, said feature processor being configured to collate a set of obtained features, per cohort, to form a feature vector X, for each data item, each vector being shaped by a first set (“k”) of data items and a second set (“n”) of features;
[0064] In at least an embodiment, said analysis engine being a cohort-based vectorization module configured to perform the steps of:
[0065] - receiving the first output and / or the second output and a feedback signal from a feedback model;
[0066] - computing cohorts based on sorting rules defined in a sorting rule engine, followed by grouping data items with common characteristics;
[0067] - processing cohorts, using a performance processor, to compute performance metrics for individual cohorts and their intersections;
[0068] - developing feature vectors based on performance metrics, for data items, sharing common characteristics; - setting a ground truth for training artificially intelligent system using manual mechanisms or machine-fed mechanisms;
[0069] - calculating trust scores for data items using regression or classification models with weighted features, minimizing a loss function; and
[0070] - determining prioritization of artificially intelligent system based on threshold values and a rule engine considering the relative importance of characteristics identified by machine learning models.
[0071] In at least an embodiment, said analysis engine cooperating with:
[0072] - a training module for generating data, from said first feedback module, correlative to said data items;
[0073] - a scan parser, from said set of parsers, in order to output said first output from said data items configured to be analysed by said performance parser;
[0074] - a patient parser, from said set of parsers, in order to output said second output from said data items configured to be analysed by said performance parser);
[0075] - a Ground Truth Module, in cooperation with said performance parser, establishing Ground Truth data for each cohort, in order to obtain metrics per cohort per data item to be fed to the training module; and
[0076] - said analysis engine comparing said established ground truth per data item with its own analysed output per data item in order to determine updateable weights based on agreement between said ground truth and said analysed output.
[0077] In at least an embodiment, said analysis engine cooperating with:
[0078] - a ranking module to organize output of said analysis engine in correlation with said first output and said second output; and said output module configured to generates a mapping between each considered artificially intelligent system and its corresponding trust score.
[0079] In at least an embodiment, said system comprising:
[0080] - said data input module configured to gather data items, including ground truth predictions, per data item, made by one or more of the existing artificially intelligent systems, first output, and said second output; and
[0081] - said performance parser configured to use said gathered data items, their ground truth predictions, in order to feed to a Training Module which trains internal weights for said analysis engine.
[0082] In at least an embodiment, a training module is active only when a ranking module is inactive.
[0083] BRIEF DESCRIPTION OF THE ACCOMPANYING DRAWINGS:
[0084] This invention will now be described in relation to the accompanying drawing, in which:
[0085] FIGURE 1 illustrates a schematic block diagram of the system of this invention;
[0086] FIGURE 2 illustrates a representation of the system of this invention;
[0087] FIGURE 3 illustrates a diagram illustrating the determination of efficacy;
[0088] FIGURE 4 illustrates a schematic block diagram for the training module used by this system and method;
[0089] FIGURE 5 illustrates a schematic block diagram for the ranking module used by this system and method;
[0090] FIGURE 6 illustrates a schematic block diagram for a first embodiment relating to interaction of the training module, of FIGURE 4, and the ranking module, of FIGURE 5, used by this system and method; and FIGURE 7 illustrates a schematic block diagram for a second embodiment relating to interaction of the training module, of FIGURE 4, and the ranking module, of FIGURE 5, used by this system and method.
[0091] DETAILED DESCRIPTION OF THE ACCOMPANYING DRAWINGS:
[0092] According to this invention, there are provided systems and methods for selection of priority-wise artificially intelligent mechanisms per one or more characteristics.
[0093] FIGURE 1 illustrates a schematic block diagram of the system of this invention.
[0094] FIGURE 2 illustrates a representation of the system of this invention.
[0095] In at least an embodiment of the present invention, a data input module (DIM) is configured to receive images, as data input / data items (DI), through this module (DIM). Data items (DI), in at least the form of medical images (with associated metadata), is obtained and passed to further modules of this system and method.
[0096] In at least an embodiment of the present invention, an identifier (IDR) is configured to identify an artificially intelligent system (AH, AI2, AI3, , Ain) being used for determination of efficacy. Each of the data items (DI), from the data input module (DIM), is processed by one or more identified artificially intelligent systems (AH, AI2, AI3, . , Ain).
[0097] In at least an embodiment, of the present invention, a scan parser (SP) is configured to parse, and output, scan characteristics, from the data input / data items (DI). This is a first output (scan characteristics) (Ol). In at least an embodiment, of the present invention, a patient parser (PRP) is configured to parse, and output, patient characteristics, from the data input / data items (DI). This is a second output (patient characteristics) (02).
[0098] In at least an embodiment, of the present invention, an analysis engine (AE) is configured to receive at least a first output (i.e. scan characteristics) and / or at least a second output (i.e. patient characteristics) and is further configured to receive feedback signal from a feedback model (FM1, FM2, FM3). An output module (OM), associated with the analysis engine (AE), provides a selection (S), or a choice therefor, of a priority-wise-ranked artificially intelligent system, from the plurality of artificially intelligent systems (All, AI2, AI3, , Ain), used by this system and method. The priority-wise ranked output ensures that the selected artificially intelligent system (All, AI2, AI3, , Ain) is one of the more efficient / accurate systems for a given set of data items. In other words, the priority-wise ranked output ensures that the selected artificially intelligent system (All, AI2, AI3, , Ain) is the most suitable system per parsed scan characteristic (01) and / or per parsed patient characteristic (02).
[0099] In at least an embodiment, of the analysis engine (AE), there is provided a first feedback module (FM1), per artificially intelligent system (All, AI2, AI3, , Ain) such that a feedback relating to accuracy of analysis, correlative to performance metrics of the selected artificially intelligent system, is recorded and fed back to the engine (AE). This records historical data per output. This first feedback module (FM1) may be configured to parse, and record, across data items, one or more of the following parameters:
[0100] - sensitivity; - specificity;
[0101] - AUROC;
[0102] - custom metric.
[0103] In at least an embodiment, of the analysis engine (AE), there is provided a second feedback module (FM2), per artificially intelligent system (All, AI2, AI3, , Ain) such that a feedback correlative to first output (scan characteristics) of the selected artificially intelligent system, is recorded and fed back to the engine (AE). This records scan data / scan metadata per output. This second feedback module (FM2) may be configured to parse, and record, across data items, one or more of the following parameters:
[0104] - modality;
[0105] - scanner manufacturer;
[0106] - scanner model;
[0107] In at least an embodiment, of the analysis engine (AE), there is provided a third feedback module (FM3), per artificially intelligent system (All, AI2, AI3, , Ain) such that a feedback correlative to second output (patient characteristics) of the selected artificially intelligent system, is recorded and fed back to the engine (AE). This records patient data / patient metadata per output. This third feedback module (FM3) may be configured to parse, and record, across data items, one or more of the following parameters:
[0108] - age;
[0109] - sex; others
[0110] In at least an embodiment, of the analysis engine (AE), a cohort-based vectorization module is configured to determine, and form, cohort-based data sets and to determine, and form, cohort-based characteristics. In at least an embodiment, of the analysis engine (AE), a sorter (SR) sorts data items (DI) into cohorts according to sorting rules defined in a sorting rule engine. These rules may relate to data / metadata. Accordingly, the scans are divided into cohorts, which are scans that have at least one characteristic in common with a current scan. For example, for n characteristics, cohort 1 could be view position (e.g., chest AP, chest PA, lateral), cohort 2 could be sex (e.g. M / F), cohort 3 could be the age group (i.e. 0-18, 18-35, 35-60, 60+), and so on.
[0111] In at least an embodiment, of the analysis engine (AE), a performance processor (PP) uses one or more data items (DI), with specific metadata, as sorted by the sorter (SR), and computes a performance function for all scans belonging to individual cohorts and intersections of two or more cohorts:
[0112] • Xi cohort 1
[0113] • X2. cohort 2
[0114] • X3'. cohort 3...
[0115] • xn: cohort n
[0116] • xn+i‘. intersection of cohort 1 and 2
[0117] • Xif. intersection of cohort i and j ...
[0118] • Xk'. intersection of cohort 1, 2, 3, ..., n In at least an embodiment, of the analysis engine (AE), a dataset comprising various data items (DI) is divided, by the sorter (SR), into cohorts, as taught above, based on patient and scan characteristics.
[0119] Typically, ‘cohorts’ are defined as sub-datasets where the data has at least one characteristic in common. As an implication of the definition, data with multiple common characteristics can be grouped into a single, unified cohort. For example'. If the data characteristics are gender (Male, Female), age group (0-18, 18+), view position (AP and PA), and scanner type (Scanner A, Scanner B), the scans will be grouped into cohorts in the following manner:
[0120] Cohort 1: Data where the patient is Male
[0121] Cohort 2: Data where the patient is Female
[0122] Cohort 3: Data where patient age is 0-18
[0123] Cohort 4: Data where patient age is above 18 years
[0124] Cohort 5: Data where scan view position is AP
[0125] Cohort 6: Data where scan view position is PA
[0126] Cohort 7: Data where scan is acquired with Scanner A
[0127] Cohort 8: Data where scan is acquired with Scanner B
[0128] In addition to these cohorts, the intersection of multiple characteristics is also considered. This is represented as follows:
[0129] Cohort 9: Data where the patient is Male AND below 18 years of age AND scanner view position is AP AND scan is acquired with Scanner A.
[0130] Although the combinations of cohorts are exhaustive, they are not being listed, as an addition of a single characteristic would change the number significantly. It is crucial to note the importance of using cohorts for the data processing stage. These cohorts accurately capture the effect of individual patient and scan characteristics on the data items. Moreover, they also capture the interaction of these characteristics amongst themselves, e.g., if an Al system is performing sub- optimally on males in the 18+ age group, this information will not be missed by the cohorts’ system.
[0131] In at least an embodiment, of the analysis engine (AE), a feature processor is configured in order to extract features from the cohorts, upon which a performance function is used vide the performance processor (PP). This performance function could be one or a combination of metrics, such as Sensitivity or Specificity or AUROC or Fl Score or a custom metric. A performance function is applied to each cohort to obtain its feature. A set of such features over all cohorts in the data point will form the feature vector X.
[0132] For example, consider a data point with the following characteristics - Male, above 18 years of age, AP view position, and Scanner A. A performance function will be applied to each cohort in the data point, i.e., Male, Male AND above 18 years, Male AND AP view, etc., as follows: xl = performanceFunctionl(Male patients in the dataset) x2 = performanceFunction2(Patients with age above 18 years in the dataset) x3= performanceFunction3 (Patients with scans acquired in AP view in the dataset) xn = performanceFunctionN(Some combination of characteristics in the dataset)
[0133] Typically, vector X, is developed, for each scan where, shape of vector is k * N
[0134] The feature vector X will be an array of each feature as follows -
[0135] X = [xl, x2, x3, , xn]
[0136] Based on this implementation, data points with the same patient and scan characteristics will have the same feature vectors. This is particularly important as the performance of the Al system over all Al predictions for all patient and scan characteristics is captured. Hence, given there are “k” scans with “n” features, the feature set will have the shape of (k x n).
[0137] In at least an embodiment, of the analysis engine (AE), the performance processor (PP) is configured to initialise with its baseline being ground truth. In order to train the model, of this invention, in order to be able to compute priorities, for determination of priority-wise selection, a ground truth has to be set. The ground truth is set based on feedback (whether manual, or machine-fed) where the ground truth is provided. If the prediction, of a particular artificially intelligent system (All, AI2, AI3, , Ain), matches the feedback, the ground truth Trust Score is set as 1. In case the prediction, of a particular artificially intelligent system (All, AI2, AI3, , Ain), does not match the feedback, the ground truth Trust Score is set as 0. This method of setting the ground truth captures trust in a particular artificially intelligent system (All, AI2, AI3, , Ain) over the entire data collection period. Therefore, the output feature set will have the shape of (k x 1).
[0138] As an intialiser, the ground truth for a weighted output, of the output module (OM) (denoted by y) prediction can be a value between 0-100 and is defined as a measure of agreement or disagreement between the artificially intelligent system (All, AI2, AI3, , Ain) and a person using the artificially intelligent system (All, AI2, AI3, , Ain) for a given one or more data items (DI).
[0139] Typically, this performance function, vide the performance processor (PP), is applied to each cohort.
[0140] Typically, this performance function, vide the performance processor (PP), could be one or a combination of common machine learning classification metrics, including but not limited to sensitivity, specificity, AUROC, Fl score or a custom metric; as recorded / determined by the first feedback module (FM1). The trust score will then be calculated as: where, f Rk-> / ? is a regression or classification model,
[0141] X = [xi, X2, ..., x / 7 is the input vector, and y is the predicted trust score of the scan.
[0142] The trust score will be optimized using a loss function g: / ? x / ? -> / ? such that g(y, y) is minimized.
[0143] In at least an embodiment, of the analysis engine (AE), the performance processor (PP) is configured to determine how weight affect loss function. Based on this dataset, the Trust Score (represented as y) will then be calculated as the following: y= f [x x2, ..., x ), where f Rk-> / ? is a regression or classification model, such as linear regression or logistic regression or bayesian regression or Support Vector Machines (SVM) or K-nearest neighbors (KNN) or Gaussian Naive Bayes or Multinomial Naive Bayes or Complement Naive Bayes or Bernoulli Naive Bayes or Decision Trees or Random Forests or Gradient Boosting or Neural Networks, which will be fitted using the input vector X =[xj, X2, ..., xn], and output y, which is the predicted trust score of the scan. This calculation is done by assigning importance, in the form of weights, to each feature. These weights can be randomly or manually initialized. During the course of training, these weights will be optimized using a loss function, g: / ? x / ? -> R, such as Binary Cross-Entropy or Hinge or Squared Hinge or Sigmoid CrossEntropy or Kullback-Leibler (KL) Divergence, Focal Loss or Weighted CrossEntropy Loss or Squared Error Loss or Huber Loss or Area Under Receiver Operating Characteristic (AUROC) curve Loss or Contrastive Divergence Loss or Center Loss or Gaussian Mixture Loss or Huber Loss or some custom loss function such that g(y, y) is minimized.
[0144] Once the parameters of the machine learning model are determined, it elucidates the relative importance of the different scan and / or patient characteristics. Using that information, the system and method of this invention is able to provide an explanation for low, medium, or high trust scores, for example, if the artificially intelligent system has previously not performed well for a particular age group.
[0145] In preferred embodiments, priority-wise-ranked selection of artificially intelligent systems is determined by applying threshold values, based on a rule engine. Once the parameters of the regression or classification model are determined, it elucidates relative importance of different scan and / or patient characteristics. Using that information, the system and method, of this invention, is able to provide an explanation for the priority-wise-ranked selection of artificially intelligent systems per one or more scan characteristic / s, per one or more patient characteristic / s, or the like one or more characteristic.
[0146] FIGURE 3 illustrates a diagram illustrating the determination of efficacy.
[0147] FIGURE 4 illustrates a schematic block diagram for the training module used by this system and method.
[0148] In the Training Module, one or more of the existing artificially intelligent systems (All, AI2, AI3, , Ain) generate Al data (DAI) for scans previously annotated by users of these systems. These predictions, along with the data of the input files to these artificially intelligent systems, are organized in the Training Data Module (DMT).
[0149] The Scan Parser (SP) and the Patient Parser (PRP) extract respective scan characteristics and patient characteristics from input files and generate input data (DI), which is analyzed by the performance parser (PP). The performance parser (PP) generates metrics such as Sensitivity, Specificity, AUROC, and custom metrics for each cohort using the Ground Truth data (DGT) generated by the Ground Truth Module (GTM). The metrics per cohort per scan generated by the Performance Parser (PP) serve as the Training Data (DT) and are passed to the Database Module (DBM) and the Training Module (TM). The algorithm undergoes training by comparing Ground Truth data (DGT) with the predictions from the existing artificially intelligent systems (All, AI2, AI3,....,AIn) in the training module (TM). During this process, the algorithm’s internal weights (W) are updated based on the agreement between the ground truth and the Al predictions. Following training, the algorithm with updated internal weights (W), is utilized within the Analysis Engine (AE).
[0150] FIGURE 5 illustrates a schematic block diagram for the ranking module used by this system and method.
[0151] Data generated (DAI) by one or more of the existing artificially intelligent systems (All, AI2, AI3, , Ain) is sent to the Ranking Data Module (DMR). This module organizes the predictions and associated scan and patient characteristics, forming the input data (DI). DI is then transmitted to the Analysis Engine (AE), where the trust scores for one or more of the existing artificially intelligent systems (All, AI2, AI3, , Ain) are evaluated. The Output Module (OM) then generates a mapping between each considered artificially intelligent system and its corresponding trust score. The selection module (S) orchestrates the ranking of these artificially intelligent systems based on their trust scores, providing a streamlined representation of their reliability. FIGURE 6 illustrates a schematic block diagram for a first embodiment relating to interaction of the training module, of FIGURE 4, and the ranking module, of FIGURE 5, used by this system and method.
[0152] During the data collection phase, the system gathers all data including ground truth, predictions made by one or more of the existing artificially intelligent systems (All, AI2, AI3, , Ain), scan characteristics, and patient characteristics. This data is then utilized by the Performance Parser (PP), followed by the Training Module in which the internal weights (W) of the Trust Score algorithm are updated. In this phase, the Training Module (TM) takes precedence and is active while the Ranking Data Module (DMR) is inactive. This is because the ranking of artificially intelligent systems cannot be determined during the data collection phase. The updated weights are stored by the Analysis Engine (AE).
[0153] FIGURE 7 illustrates a schematic block diagram for a second embodiment relating to interaction of the training module, of FIGURE 4, and the ranking module, of FIGURE 5, used by this system and method.
[0154] After the data collection phase, the Database Module (DBM) continues to store the data generated by the Performance Parser (PP) but the Training Module (TM) becomes inactive as the algorithm will have optimized its internal weights (W) using the data collected during the data collection phase. The Ranking Data Module (DMR) and Analysis Engine (AE) become active within the system process flow and the algorithm becomes eligible to evaluate and rank one or more of the existing artificially intelligent systems (All, AI2, AI3, . , Ain). With the use of this system and method, the output provides radiologists, using one of many artificially intelligent mechanisms, confidence while using the Al mechanism. The output also comprises an explanation in correlation with efficacy. This enables users to:
[0155] 1) Take an informed decision to agree or disagree with Al predictions based on its past performance.
[0156] 2) Determine the weight, significance, and importance to be assigned to the Al output.
[0157] 3) Understand the model better and predict its future behavior.
[0158] According to a first non-limiting exemplary embodiment, the system and method, of this invention was used to test performance of Tuberculosis- Al model by Vendor 1 (VI) on different cohorts.
[0159] TABLE 1
[0160] This table 1 illustrates past performances, quantified by an AUROC score, for an Al system designed to detect Tuberculosis in Chest X-rays. The performance of the Al system is consistent over various cohorts such as gender, age group, view position, and scanner type. Therefore, for a scan parsed through the TB detection Al system, a user can accept the Al prediction without much scrutiny.
[0161] According to a second non-limiting exemplary embodiment, the system and method, of this invention was used to test performance of Tuberculosis-AI model by Vendor 1 (V2) on different cohorts. TABLE 2
[0162] This table 2 illustrates past performances, quantified by the AUROC score, for an Al system designed to detect Tuberculosis in Chest X-rays. Notably, the Al system exhibits subpar performance for scans of patients in the 0-18 age group. Conversely, it demonstrates moderate performance for patients aged 60 and above and excellent performance for those in the 18-35 and 35-60 age groups. Here, the user would not have to scrutinize the Al prediction much if it belongs to a patient in the 18-35 or 35-60 age groups. However, they would have to focus more on the Al predictions for the scans for the patients in the 0-18 and 60+ age groups.
[0163] In the case of a chest X-ray from a patient in the 0-18 age group, the user may confidently rely on the prediction of the Al system from V 1 , given its robust performance. Conversely, when evaluating a PA scan acquired on scanner SI from a male (M) patient of age group 35-60, the user may opt for V2, given its superior AUROC score.
[0164] According to a third non-limiting exemplary embodiment, the system and method, of this invention was used to test performance of Al model that detect abnormalities in lateral Chest X-rays on different cohorts.
[0165]
[0166] TABLE 3
[0167] This table 3 illustrates past performances of an Al system, quantified by the AUROC score, for an Al system designed to detect abnormalities in Lateral chest X-rays. The performance of VI is consistently inferior across all age groups compared to V2. Notably, within the VI subset, the performance of the model is subpar for the patients in the 0-18 age group. This could be due to insufficient data during model training specifically for age group 0-18, or due to subop timal training parameters. These findings underscore the lack of reliability on VI for detection of abnormalities from lateral Chest X-rays. Based on the past performance of Al models by VI and V2, the user may rely on the predictions of V2 as compared to predictions of V 1.
[0168] The TECHNICAL ADVANCEMENT of this invention lies in systems and methods for selecting a priority-wise artificially-intelligent system based on one or more characteristics. This selection endorses trust values in various artificially- intelligent systems. The following discussion is intended to provide a brief, general description of Suitable computing environments in which the system and method may be implemented. Although not required, the disclosed embodiments will be described in the general context of computer-executable instructions, such as program modules, being executed by a single computer. In most instances, a “module' constitutes a software application.
[0169] Generally, program modules include, but are not limited to routines, Subroutines, software applications, programs, objects, components, data structures, etc., that perform particular tasks or implement particular abstract data types and instructions. Moreover, those skilled in the art will appreciate that the disclosed method and system may be practiced with other computer system configurations, such as, for example, hand-held devices, multi-processor systems, data networks, microprocessor-based or programmable consumer electronics, networked PCs, minicomputers, mainframe computers, servers, and the like.
[0170] For the purposes of this specification, term ‘module’, as utilized herein, may refer to a collection of routines and data structures that perform a particular task or implements a particular abstract data type. Modules may be composed of two parts: an interface, which lists the constants, data types, variable, and routines that can be accessed by other modules or routines, and an implementation, which is typically private (accessible only to that module) and which includes source code that actually implements the routines in the module. The term module may also simply refer to an application, such as a computer program designed to assist in the performance of a specific task, such as word processing, accounting, inventory management, etc. The interface, which is preferably a graphical user interface (GUI), can serve to receive inputs for engagement with other modules as also to display results, whereupon a user may supply additional inputs or terminate a particular session.
[0171] While this detailed description has disclosed certain specific embodiments for illustrative purposes, various modifications will be apparent to those skilled in the art which do not constitute departures from the spirit and scope of the invention as defined in the following claims, and it is to be distinctly understood that the foregoing descriptive matter is to be interpreted merely as illustrative of the invention and not as a limitation.
Claims
CLAIMS,1. A method for selection of priority-wise artificially intelligent mechanisms per one or more characteristics, said method comprising the steps of:- receiving (DIM) images as data items (DI);- identifying (IDR) an artificially intelligent system (All, AI2, AI3, , Ain) being used for determination of efficacy, each of the data items (DI), being processed by one or more identified artificially intelligent systems (All, AI2, AI3, , Ain);- parsing (SP, PRP), and outputting, scan characteristics and patient characteristics, from the data items (DI), output of said parsing (SP, PRP) being first output (scan characteristics) (Ol) and second output (patient characteristics) (02);- analysing (AE) to receive at least a first output and / or at least a second output and configured to receive feedback signal from a feedback model (FM1, FM2, FM3); and- serving, as an output (OM), upon analysing (AE), a selection (S) of a priority-wise-ranked artificially intelligent system, from the plurality of artificially intelligent systems (All, AI2, AI3, , Ain), said selected system being per parsed scan characteristic (01) and / or per parsed patient characteristic (02).
2. The method as claimed in claim 1 wherein, said paring being selected from a group of parsing steps consisting of:- parsing (SP), and outputting, scan characteristics, from the data items (DI), output of said scan parser (SP) being a first output (scan characteristics) (Ol);parsing (PRP), and outputting, patient characteristics, from the data items (DI), output of said patient parser (PRP) being a second output (patient characteristics) (02);3. The method as claimed in claim 1 wherein, said feedback signal being received from a first feedback module (FM1), per artificially intelligent system (All, AI2, AI3, , Ain), in that, feedback relating to accuracy metrics of analysis, correlative to performance metrics of the selected artificially intelligent system, is recorded and fed back to the step of analysing (AE).
4. The method as claimed in claim 1 wherein, said feedback signal being received from a second feedback module (FM2), per artificially intelligent system (All, AI2, AI3, , Ain), in that, feedback correlative to first output (01) of the selected artificially intelligent system, is recorded and fed back to the step of analysing (AE).
5. The method as claimed in claim 1 wherein, said feedback signal being received from a third feedback module (FM3), per artificially intelligent system (All, AI2, AI3, , Ain), in that, feedback correlative to second output (02) of the selected artificially intelligent system, is recorded and fed back to the step of analysing (AE).
6. The method as claimed in claim 1 wherein, said step of analysing (AE) being a cohort-based vectorization analysis configured to determine, and form, cohort-based data sets in order to determine, and form, cohort-based characteristics, characterized, in that,- sorting (SR) data items (DI), basis its data and / or its metadata, into cohorts according to sorting rules, correlative to said first output (01) and / or said second output (02), defined in a sorting rule engine;- using, by a performance processor (PP), one or more data items (DI), with specific metadata, as sorted by said sorter (SR), in order to compute a performance function for all data items (DI) belonging to individual cohorts and intersections of two or more cohorts; and- configured to extracting features from the cohorts, upon which a performance function is used by said performance processor (PP), said performance function being correlative to one or more metrics relating to said first output (01), said performance function being applied to each cohort in order to obtain its feature, said feature processor being configured to collate a set of obtained features, per cohort, to form a feature vector X, for each data item, each vector being shaped by a first set (“k”) of data items and a second set (“n”) of features;7. The method as claimed in claim 1 wherein, said analysis engine (AE) being a cohort-based vectorization module configured to perform the steps of:- receiving the first output (01) and / or the second output (02) and a feedback signal from a feedback model (FM1, FM2, FM3);- computing cohorts based on sorting rules defined in a sorting rule engine, followed by grouping data items (DI) with common characteristics;- processing cohorts, using a performance processor (PP), to compute performance metrics for individual cohorts and their intersections;- developing feature vectors (X) based on performance metrics, for data items (DI), sharing common characteristics;setting a ground truth for training artificially intelligent system (All, AI2,AI3, , Ain) using manual mechanisms or machine-fed mechanisms;- calculating trust scores for data items (DI) using regression or classification models with weighted features, minimizing a loss function; and- determining prioritization of artificially intelligent system (All, AI2, AI3, , Ain) based on threshold values and a rule engine considering the relative importance of characteristics identified by machine learning models.
8. The method as claimed in claim 1 wherein, said step of analysing (AE) cooperating with:- generating data, from said first feedback module (FM1), correlative to said data items (DI);- outputting (SP) said first output (01) from said data items (DI) configured to be analysed by said performance parser (PP);- outputting (PRP) said second output (02) from said data items (DI) configured to be analysed by said performance parser (PP); and- establishing Ground Truth data (DGT) for each cohort, in order to obtain metrics per cohort per data item (DI) to be fed to a training module; and- comparing said established ground truth per data item (DI) with its own analysed output per data item (DI) in order to determine updateable weights (W) based on agreement between said ground truth and said analysed output.
9. The method as claimed in claim 1 wherein, said step of analysing (AE) cooperating with:- organizing (DMR) output of said analysis engine in correlation with said first output (01) and said second output (02); and- generating a mapping between each considered artificially intelligent system (All, AI2, AI3, , Ain) and its corresponding trust score.
10. The method as claimed in claim 1 wherein, said method comprising the steps of:- gathering (DIM) data items (DI), including ground truth predictions, per data item (DI), made by one or more of the existing artificially intelligent systems (All, AI2, AI3, , Ain), first output (01), and said second output (02); and- using said gathered data items (DI), their ground truth predictions, in order to feed to a Training Module which trains internal weights (W) for said analysis engine (AE).
11. The method as claimed in claim 1 wherein, a training module is active only when a ranking module is inactive.
12. A system for selection of priority- wise artificially intelligent mechanisms per one or more characteristics, said system comprising:- a data input module (DIM) configured to receive images, as data items (DI);- an identifier (IDR) configured to identify an artificially intelligent system (All, AI2, AI3, , Ain) being used for determination of efficacy, each of the data items (DI), from the data input module (DIM), being processed by one or more identified artificially intelligent systems (All, AI2, AI3, , Ain);- a set of parsers (SP, PRP), configured to parse, and output, scan characteristics and patient characteristics, from the data items (DI), output of said parsers (SP, PRP) being first output (scan characteristics) (Ol) and second output (patient characteristics) (02);- an analysis engine (AE) configured to receive at least a first output and / or at least a second output and configured to receive feedback signal from a feedback model (FM1, FM2, FM3); and- an output module (OM), cooperating with the analysis engine (AE), serving a selection (S) of a priority-wise-ranked artificially intelligent system, from the plurality of artificially intelligent systems (All, AI2, AI3, , Ain), said selected system being per parsed scan characteristic (01) and / or per parsed patient characteristic (02).
13. The system as claimed in claim 1 wherein, said set of parsers being selected from a group of parsers consisting of:- a scan parser (SP) configured to parse, and output, scan characteristics, from the data items (DI), output of said scan parser (SP) being a first output (scan characteristics) (Ol); and- a patient parser (PRP) configured to parse, and output, patient characteristics, from the data items (DI), output of said patient parser (PRP) being a second output (patient characteristics) (02);14. The system as claimed in claim 1 wherein, said feedback model comprising a first feedback module (FM1), per artificially intelligent system (All, AI2, AI3, , Ain), in that, feedback relating to accuracy metrics of analysis, correlative to performance metrics of the selected artificially intelligent system, is recorded and fed back to the analysis engine (AE).
15. The system as claimed in claim 1 wherein, said feedback model comprising a second feedback module (FM2), per artificially intelligent system (All, AI2, AI3, , Ain), in that, feedback correlative to first output (01) of the selected artificially intelligent system, is recorded and fed back to the analysis engine (AE).
16. The system as claimed in claim 1 wherein, said feedback model comprising a third feedback module (FM3), per artificially intelligent system (All, AI2, AI3, , Ain), in that, feedback correlative to second output (02) of the selected artificially intelligent system, is recorded and fed back to the analysis engine (AE).
17. The system as claimed in claim 1 wherein, said analysis engine (AE) being a cohort-based vectorization module configured to determine, and form, cohort-based data sets in order to determine, and form, cohort-based characteristics, characterized, in that,- a sorter (SR) sorts data items (DI), basis its data and / or its metadata, into cohorts according to sorting rules, correlative to said first output (01) and / or said second output (02), defined in a sorting rule engine;- a performance processor (PP) configured to use one or more data items (DI), with specific metadata, as sorted by said sorter (SR), in order to computes a performance function for all data items (DI) belonging to individual cohorts and intersections of two or more cohorts; and- a feature processor configured to extract features from the cohorts, upon which a performance function is used by said performance processor (PP), said performance function being correlative to one or more metricsrelating to said first output (01), said performance function being applied to each cohort in order to obtain its feature, said feature processor being configured to collate a set of obtained features, per cohort, to form a feature vector X, for each data item, each vector being shaped by a first set (“k”) of data items and a second set (“n”) of features;18. The system as claimed in claim 1 wherein, said analysis engine (AE) being a cohort-based vectorization module configured to perform the steps of:- receiving the first output (01) and / or the second output (02) and a feedback signal from a feedback model (FM1, FM2, FM3);- computing cohorts based on sorting rules defined in a sorting rule engine, followed by grouping data items (DI) with common characteristics;- processing cohorts, using a performance processor (PP), to compute performance metrics for individual cohorts and their intersections;- developing feature vectors (X) based on performance metrics, for data items (DI), sharing common characteristics;- setting a ground truth for training artificially intelligent system (All, AI2, AI3, , Ain) using manual mechanisms or machine-fed mechanisms;- calculating trust scores for data items (DI) using regression or classification models with weighted features, minimizing a loss function; and- determining prioritization of artificially intelligent system (All, AI2, AI3, , Ain) based on threshold values and a rule engine considering the relative importance of characteristics identified by machine learning models.
19. The system as claimed in claim 1 wherein, said analysis engine (AE) cooperating with:- a training module for generating data, from said first feedback module (FM1), correlative to said data items (DI);- a scan parser (SP), from said set of parsers (SP, PRP), in order to output said first output (01) from said data items (DI) configured to be analysed by said performance parser (PP);- a patient parser (PRP), from said set of parsers (SP, PRP), in order to output said second output (02) from said data items (DI) configured to be analysed by said performance parser (PP);- a Ground Truth Module (GTM), in cooperation with said performance parser (PP), establishing Ground Truth data (DGT) for each cohort, in order to obtain metrics per cohort per data item (DI) to be fed to the training module; and- said analysis engine (AE) comparing said established ground truth per data item (DI) with its own analysed output per data item (DI) in order to determine updateable weights (W) based on agreement between said ground truth and said analysed output.
20. The system as claimed in claim 1 wherein, said analysis engine (AE) cooperating with:- a ranking module (DMR) to organize output of said analysis engine in correlation with said first output (01) and said second output (02); and- said output module (OM) configured to generates a mapping between each considered artificially intelligent system (All, AI2, AI3, , Ain) and its corresponding trust score.
21. The system as claimed in claim 1 wherein, said system comprising:- said data input module (DIM) configured to gather data items (DI), including ground truth predictions, per data item (DI), made by one or more of the existing artificially intelligent systems (All, AI2, AI3, , Ain), first output (01), and said second output (02); and- said performance parser (PP) configured to use said gathered data items (DI), their ground truth predictions, in order to feed to a Training Module which trains internal weights (W) for said analysis engine (AE).
22. The system as claimed in claim 1 wherein, a training module is active only when a ranking module is inactive.
Citation Information
Patent Citations
Training method and device and prediction method and device of medical image classification model
CN115457349A
Automatic detection of disease from analysis of echocardiographer findings in echocardiogram videos
US20180108125A1
Determining Appropriate Medical Image Processing Pipeline Based on Machine Learning
US20200311861A1
Artificial intelligence dispatch in healthcare
US20200388386A1
Category discovery and image auto-annotation via looped pseudo-task optimization
WO2017151759A1