Method and system for additive artificial intelligence model
Through automated image preprocessing and analysis, combined with the results of multiple AI models, an additive AI method is used to generate a comprehensive diagnostic report, which solves the problems of low image processing efficiency and low report generation efficiency in the prior art, and achieves efficient and accurate image diagnosis and report generation.
Patent Information
- Application Number
- CN202380071551.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-08-05
- Filing Date
- 2023-07-27
- Publication Date
- 2025-05-30
AI Technical Summary
Prior art requires manual identification and cropping of images when processing animal radiological images, or sending images to multiple AI processors for evaluation, resulting in inefficiency and high error rates. At the same time, the report generation efficiency of AI model diagnostic results is inefficient and cannot be effectively extended to explain the results of a large number of AI models.
A novel system is adopted to automate image preprocessing and analysis, including identifying and cropping images, creating sub-images, and evaluating cropped images based on target AI models. The system combines the results of multiple AI models to generate a comprehensive diagnostic report through additive AI methods.
The image processing is automated, efficiency and accuracy are improved, and can be effectively extended to interpret the results of a large number of AI models and generate high-quality diagnostic reports.
Smart Images

Figure CN120077385A_ABST
Abstract
Description
[0001] This application is a partial continuation of U.S. Invention Application Serial No. 17 / 134,990, filed on December 28, 2020. U.S. Invention Application Serial No. 17 / 134,990 is a continuation of International Application PCT / US20 / 66580, titled "Efficient artificial intelligence analysis of images with combined predictive modeling", filed on December 22, 2020 by inventors Seth Wallack, Ariel Ayaviri Omonte, Ruben Venegas, Yuan-Ching Spencer Teng, and Pratheev Sabaratnam Sreetharan. International Application PCT / US20 / 66580 claims the priority of U.S. Provisional Application Serial No. 62 / 954,046, filed on December 27, 2019, and the priority of No. 62 / 980,669, filed on February 24, 2020, both titled "Efficient Artificial Intelligence Analysis of Images", both filed by inventors Seth Wallack, Ariel Ayaviri Omonte, and Ruben Venegas; and claims the priority of U.S. Provisional Application Serial No. 63 / 083,422, titled "Efficient artificial intelligence analysis of images with combined predictive modeling", filed on September 25, 2020 by inventors Seth Wallack, Ariel Ayaviri Omonte, Ruben Venegas, Yuan-Ching Spencer Teng, and Pratheev Sabaratnam Sreetharan. And this application claims the benefit of U.S. Provisional Application Serial No. 63 / 395,525, titled "Additive AI classifiers", filed on August 5, 2022 by inventors Seth Wallack and Eric Goldman, each of which is incorporated herein by reference in its entirety. Background Art
[0002] Artificial Intelligence (AI) processors, such as trained neural networks, can be used to process radiological images of animals to determine the probability that the imaged animal has certain health conditions. Typically, individual AI processors are used to evaluate individual body regions (e.g., chest, abdomen, shoulder, forelimb, hindlimb, etc.) and / or specific orientations of each such body region (e.g., ventral dorsal (VD) view, lateral view, etc.). A specific AI processor determines the probability that a specific health condition exists in the body region under discussion for each body region and / or orientation. Each such AI processor includes a large number of trained models to evaluate the corresponding health conditions or organs within the imaged region. For example, for a lateral view of an animal's chest, the AI processor uses different models to determine the probability that the animal has certain health conditions related to the lungs (such as perihilar infiltration, pneumonia, bronchitis, lung nodules, etc.).
[0003] The amount of processing performed by each such AI processor and the amount of time required to complete such processing are huge. This task requires either (1) manually identifying and cropping each image before the specific AI processor evaluates the image to define a specific body region and orientation, or (2) feeding the image into each AI processor for evaluation. Different from human radiology where radiological studies are limited to specific regions, veterinary radiology typically includes multiple unlabeled images in a single study, and these images involve multiple body regions with unknown orientations.
[0004] In the traditional workflow for processing animal radiological images, the system assumes that the body region identified by the user is included in the image. Then, the images identified by the user are sent to specific AI processors, which, for example, use machine learning (ML) models to evaluate the probability of a medical condition existing in the specific body region. However, if the identified body region is incorrect or the image contains multiple regions, asking the user to identify the body region creates friction and leads to errors in the traditional workflow. In addition, when an image without a user identification of the body region is sent to the system, the traditional workflow becomes inefficient (or crashes). When this happens, the traditional workflow is inefficient because the unrecognized image is sent to a large number of AI processors that are not specific to the imaged body region. In addition, the traditional workflow is prone to producing incorrect results because incorrect region identification causes the image to be sent to an AI processor configured to evaluate a different body region.
[0005] The traditional workflow of using AI to analyze the diagnostic features of radiographs and prepare a diagnostic report based on an AI model generates an exponential number of possible output reports. The AI model diagnosis results provide a determination of normal or abnormal for a specific health condition. In some AI models, a determination of the severity of a specific health condition is also provided, such as normal, mild, moderate, or severe. The set of AI model diagnosis results determines which report is selected from a prefabricated report template. The process of creating and selecting a single report template from the set of AI model diagnosis results scales exponentially with the number of AI models. Six different AI model normal / abnormal diagnosis results require 64 different report templates (2 to the sixth power). Ten models require 1024 templates, and 16 models require 65536 templates. The scale change of AI models that detect severity is even worse. For example, 16 severity detection models, each with 5 possible severities, require more than 150 billion templates. Therefore, reports manually created for each combination of AI model diagnosis results do not scale well for simultaneously interpreting a large number of AI models.
[0006] Accordingly, there is a need for a novel system that has several fully automated image preprocessing and image analysis phases, including determining whether the received image includes a specific body region in a specific orientation (such as a side view, etc.); appropriately cropping the image; creating one or more sub-images from an original image that contains more than one body region or region of interest; labeling the original image and any created sub-images; and evaluating the cropped image and sub-images according to a target AI model. Additionally, there is a need for a novel system that analyzes and provides a radiologist report based on a large number of test results (including but not limited to AI model results).
[0007] AI models or classifiers in machine learning are continuously trained to improve model performance. Best practices for AI model deployment involve continuous training, which includes retraining the currently deployed AI model. Retraining is based on one or more of these common parameters, such as AI performance-based, data change-based triggers, or on-demand training. The goal of the current industry standard for AI model retraining is to replace the currently trained model with a new, "improved" AI model. Current AI technology is based on a single model result in a production environment, and thus the concepts of retraining and replacement have evolved.
[0008] The problem with the current method is that replacing the current AI model with a newly trained AI model works based on the assumption that the improvement in model performance lies in reducing unwanted "noise" or false positive data. Data detected as unwanted "noise" in the current AI model is deleted and replaced with the retrained AI model. However, it is incorrect to assume that "noise" is not valuable information in the overall system performance. The "noise" detected by the current model classifier helps to distinguish data points with similar but not exactly the same features. Therefore, data is lost by replacing the current AI model.
[0009] Therefore, a new method that combines both the new and old AI models is needed to continuously train the AI model. Summary of the Invention
[0010] One aspect of the invention described herein provides a method for obtaining additive AI results from a digital file, the method comprising: processing the digital file through a first artificial intelligence (AI) classifier and at least one second AI classifier to obtain a first evaluation result and at least one second evaluation result respectively; directing the first evaluation result and the at least one second evaluation result to at least one synthesis processor; and comparing the first evaluation result and the at least one second evaluation result with at least one data cluster to obtain additive AI results. The terms AI model and AI classifier may be used interchangeably and are defined as a machine learning algorithm for assigning class labels to data inputs.
[0011] Embodiments of the method further include measuring the distance from the additive AI results to example results from the data cluster to obtain an additive AI cluster identification. In an embodiment of the method, the data cluster further includes a matching written template. Embodiments of the method further include compiling the additive AI cluster identification and the matching written template to obtain a report. Embodiments of the method further include displaying the report to a user.
[0012] In an embodiment of the method, the second AI classifier is a derivative of the first AI classifier. In an embodiment of the method, at least a portion of the data used to train the first AI classifier is used to train the second AI classifier. For example, at least 99%, 95%, 90%, 85%, 80%, 75%, 70%, 66%, 60%, 55%, 50%, 45%, 40%, 35%, 30%, 25%, 20%, 15%, 10% or 5% of the data used to train the first AI classifier is used to train the second classifier. In an embodiment of the method, the second AI classifier is related to the first AI classifier. In some embodiments, the first AI classifier is a comprehensive classifier. In some embodiments, the second AI classifier is a specific classifier. In alternative embodiments, the first AI classifier is a specific classifier and the second AI classifier is a comprehensive classifier.
[0013] Embodiments of the method further include the steps of repeatedly bootstrapping and comparing a series of daisy-chained AI classifiers. Embodiments of the method further include comparing a first evaluation result with a second evaluation result for training the first and second AI classifiers, or for comparing AI results and testing expected performance.
[0014] Embodiments of the method further include adding the first evaluation result and the second evaluation result to a result database. Embodiments of the method further include obtaining a digital file prior to processing. Embodiments of the method further include converting an analog file to a digital file prior to processing. Embodiments of the method further include classifying the digital file prior to processing by performing at least one of tagging, cropping, editing, and orienting on the digital file.
[0015] Embodiments of the method further include heuristically adjusting the AI model result by applying a mathematical formula to the AI model result, the mathematical formula including at least one of addition, subtraction, multiplication, division, or any other standard mathematical formula.
[0016] One aspect of the invention described herein provides a system programmed to obtain additive AI results by any of the methods described herein, the system including: at least one first AI processor; at least one derived AI processor derived from the first AI processor; and an output device.
[0017] Embodiments of the system further include at least one database. Embodiments of the system further include a user interface. In some embodiments, the AI result is from only one derived classifier of one first AI classifier. In some embodiments, the result includes more than one derived classifier of only one first AI classifier. In some embodiments, the result includes one derived classifier of more than one first AI classifier. In some embodiments, the result includes more than one derived classifier of more than one first AI classifier. In some embodiments, the result includes one derived classifier for only some of the first AI classifiers and more than one derived classifier for some of the first AI classifiers. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1It is a schematic diagram of a conventional workflow for processing radiological images 102 of an animal. As is commonly represented in the veterinary field, the image 102 does not indicate parts of the animal. The image 102 is processed by each of a large number of AI processors 104a - 104l to determine whether that body region is present in the image and the probability that the animal represented in the image 102 has certain health conditions. Each of the AI processors 104a - 104l evaluates the image 102 by comparing the image with one or more machine learning models, each of which is trained to determine the probability that the animal has a specific health condition.
[0019] Figure 2 It is a schematic diagram of an embodiment of the system or method described herein. A radiological image preprocessor 106 is deployed to preprocess the image 102 to generate one or more sub - images 108, each sub - image corresponding to a specific view of a specific body region. Three sub - images 108a - c are generated, where one sub - image 108a is identified and cropped as a lateral view of the animal's chest, the second sub - image 108b is identified and cropped as a lateral view of the animal's abdomen; the third sub - image 108c is identified and cropped as a lateral view of the animal's pelvis.
[0020] As indicated, the sub - image 108a is processed only by the lateral chest AI processor 104a, the sub - image 108b is processed only by the lateral abdomen AI processor 104c, and the sub - image 108c is processed only by the lateral pelvis AI processor 104k. In some embodiments, the sub - images 108 are tagged to identify the body region and / or view represented by the sub - images 108.
[0021] Figure 3 It is a description of a set of computer operations performed by an embodiment of the system or method of the present invention for the novel workflow described herein. The image 302 is processed by using the radiological image preprocessor 106 and then processed using a subset of the AI processors 104 corresponding to the identified body region / view. Cropped images 304a, 304b of the respective body regions / views identified by the system are shown. The total time taken by the radiological image preprocessor to determine that the image 302 represents both a "lateral chest" image and a "lateral abdomen" image is 24 seconds, as reflected by the timestamp of the log entry corresponding to the bracket 306.
[0022] Figure 4 It is a set of conventional single - health - condition - based organ findings 401 to 407 in the lungs in a radiograph, followed by combinations of at least two single - health - condition - based organ findings. The permutations and combinations of the seven single - health - condition - based organ findings result in an exponential number of report templates.
[0023] Figures 5A - 5FIs a set of basic lung organ findings, which are classified into normal, mild, moderate, and severe according to severity and presented as separate AI model result templates. Boxes 501 to 557 represent individual line items in a specific AI report template. The line items are selected according to each AI model result template that matches the findings listed under the title "Code".
[0024] Figure 6 Is a collection (or AI model library) of individual binary models that are deployed to analyze radiological images to obtain probability results of the radiological images for a health condition or classification as negative or positive.
[0025] Figure 7 Is a lateral chest radiograph of a dog, and the image has been preprocessed, cropped, labeled, and recognized. The radiograph image is analyzed by the Figure 6 binary AI model library shown in.
[0026] Figure 8 Is a screenshot of the results of a single binary AI model obtained by analyzing a series of lateral radiological images similar to the images of Figure 7 through a specific binary AI model (such as a bronchitis AI model).
[0027] Figures 9A - 9E Is a set of screenshots showing the results of each AI model for the radiological images. For this specific case, the visual collection of the results of each individual AI model for each image and the mean of the AI model results for all images were evaluated. The average evaluation result for each model was created by compiling the individual image evaluation results and is displayed at the top of the screen in Figure 9A The results of each individual image and the AI model for that image are shown in Figures 9B - 9E . Figure 9A The timestamp 901 in shows that the AI analysis was completed in less than three minutes. Figure 9B Shows the results of individual AI models, such as perihilar infiltration, pneumonia, bronchitis, interstitial disease, lung lesions, tracheal hypoplasia, cardiomegaly, lung nodules, and pleural effusion. For each AI model, the label identifies the image as "normal" 902 or "abnormal" 903. In addition, the probability 904 that the image is "normal" or "abnormal" for a single health condition of the AI model is provided. Figure 9C Shows four images obtained by classifying and cropping a single radiological image. The timestamps 905 - 908 show that the AI analysis was completed in less than two minutes. Figure 9D and Figure 9E Show the results of each AI model for the radiological images, which include the label, probability, and view (909) of the radiological image, such as lateral, dorsal, anteroposterior, posteroanterior, ventral, dorsoventral, etc.
[0028] Figure 10 It is a screenshot of the AI case results displayed in the JavaScript Object Notation (JSON) format. The JSON format facilitates the copying of the average evaluation results of all models in the case, and this result can be transmitted to the AI evaluation tester for testing. The average evaluation result is evaluated by comparing it with the clustering result.
[0029] Figure 11 It is a screenshot of the graphical user interface that allows users to create K-means clusters. The user assigns the name 1101 to the new cluster under "Code". The user selects various parameters to create the cluster. The user selects the start case date 1102 and the end case date 1103 to select the case. The user selects the start case identifier (ID) 1104 and the end case ID 1105 to select the case. The user selects the maximum number of cases 1106 to be included in the cluster. The user selects the species 1107 of the cases to be included in the cluster, such as dogs, cats, dogs or cats, humans, etc. The user selects a specific diagnostic form 1108, such as X-ray, Computed Tomography (CT), Magnetic Resonance Imaging (MRI), blood analysis, urine analysis, etc., to be included in the creation of the cluster. The user specifies that the evaluation results be divided into a specific number of clusters. The number of clusters ranges from the minimum of one cluster to the maximum number of clusters, and the maximum number of clusters is only limited by the total number of cases entering the cluster.
[0030] Figure 12 It is a screenshot of the AI clustering results listed as a digital table. The leftmost column 1201 is the case ID. The next nine columns are the average evaluation results 1202 of each binary model for a specific case ID. The next column is the cluster label or cluster location 1203 that contains the specific case based on the set of evaluation results. The next four columns are the cluster coordinates and centroid coordinates. The last number is the case ID of the centroid or center of that specific cluster 1204. The radiologist's report for the most matching case ID is obtained. Then, this radiologist's report is used to generate the report for the new AI case. Compared with the traditional semi-manual report creation process, this process allows for unlimited scalability in terms of the number of combined AI models.
[0031] Figure 13 It is an example of a clustering graph. The clustering graph is created by dividing the average evaluation results into multiple different clusters according to the user-defined parameters 1102 - 1108. This example clustering graph is divided into 180 different clusters, and each cluster is represented by a set near the monochromatic points drawn on the graph.
[0032] Figure 14It is a screenshot of a user interface, showing an AI clustering model generated based on user-defined parameters 1102 - 1108. The first column from the left shows the cluster ID 1401, the second column shows the specified name of the cluster model 1402, the third column shows the number of different clusters 1403 into which the AI data is divided, and the fourth column shows the body area 1404 evaluated based on the cluster data results.
[0033] Figure 15 It is a screenshot of a user interface showing a screening evaluation configuration. The user interface allows a specific "cluster model" 1502 to be assigned to a specific "screening evaluation configuration" name 1501. The status 1503 of the screening evaluation configuration provides additional data about the configuration, such as whether the configuration is in real-time, test, or draft mode. Real-time mode is for production, and test mode is for development.
[0034] Figures 16A - 16C It is a set of screenshots of a user interface showing details of a specific cluster model. Figure 16A It shows a user interface displaying data for the cluster model 1601 Thorax 97. The types of AI evaluation classifiers 1602 included in the cluster are listed. The species or set of species specific to the cluster model 1603 is shown. The maximum number of cases 1604 with evaluation results used to generate the cluster is shown. The user interface shows the start and end dates 1605 of the cases used to create the cluster. It shows Figure 12 a link 1606 to a comma separated value (CSV) file that shows the cluster in tabular form of numbers. A part of the sub-clusters 1608 created according to the parameters 1602 - 1605 is listed. The total number of sub-clusters 1609 created for this cluster group is shown. For each sub-cluster, the centroid case ID 1610 is shown. A link to the log 1607 used to build the cluster is shown. Figure 16B It is a screenshot of the log created for the cluster model 1601 Thorax 97. Figure 16C It is a screenshot of a part of an AI evaluation model, including vertebral heart score, perihilar infiltration, pneumonia, bronchitis, interstitial disease, and diseased lungs.
[0035] Figures 17A - 17D It is a set of screenshots of a user interface of an AI evaluation tester. Figure 17A It shows a user interface (AI evaluation tester) where the Figure 10 values of the average evaluation results of all models in JSON format are imported 1701 to analyze the closest matching case / example results in the cluster from the case clusters made from the AI dataset using K-means clustering. Figure 17B It shows what is imported into the AI evaluation testerFigure 10 The value of the average evaluation result of all models in JSON format. Figure 17C and Figure 17D shows the evaluation results imported for a specific case. Figure 17D shows the screening evaluation type 1702 and the clustering model 1703 associated with the screening evaluation type selected by the user. By clicking on the test 1704, the analysis Figure 10 of the evaluation results shown therein and assigns them to the closest matching case / example result match in the cluster. The closest radiologist reports, top-ranked radiologist statements, and centroid radiologist reports of the example result match cluster are collected and displayed.
[0036] Figures 18A - 18E is a set of screenshots of the user interface. Figure 18A and Figure 18B is a set of screenshots of the user interface, showing the results displayed after clicking on the test 1704 on the AI evaluation tester. The diagnosis and conclusive findings 1801 of the evaluation results closest to the previously created cluster results from the radiologist reports are shown. The evaluation findings 1802 are selected from the radiologist reports in the evaluation result cluster and filtered based on the prevalence of specific statements in the findings section of a specific cluster. The recommendations 1803 in the radiologist reports of this cluster are selected based on the prevalence of each statement or similar statements in the recommendations section of this cluster. Interface 1804 shows the radiologist reports of the cluster, and interface 1805 shows the radiologist reports of the cluster centroid. Figure 18C is a screenshot of the user interface that lists the rankings of statements in the radiology report based on a specific cluster result. These statements include conclusion statements 1806, finding statements 1807, and recommendation statements 1808. Figure 18D and Figure 18E are screenshots of the user interface that allows the user to edit the radiology report by adding or deleting specific statements in the findings section 1809, conclusion section 1810, or recommendation section 1811.
[0037] Figure 19 is the radiologist report of the closest matching dataset case for generating a radiology report for a new case. The AI evaluation tester displays the radiologist report closest to the current AI evaluation result based on the similarity of the evaluation results between the new image AI evaluation result and the AI evaluation results within the cluster, as well as the radiologist report of the selected cluster centroid.
[0038] Figure 20A and Figure 20B is a set of radiographs. Figure 20A is the newly received radiograph being analyzed, Figure 20BIs the radiograph selected as the closest match based on the results evaluated by the AI using a clustering model. The clustering match is based on the AI evaluation results, rather than the image matching results.
[0039] Figure 21A and Figure 21B Are a set of schematic diagrams of the components in the AI radiograph processing unit. Figure 21A Is a schematic diagram showing that the radiography camera 2101 sends a radiological image to the desktop application 2102, and the desktop application 2102 guides the image to the web application 2103. The computer vision application 2104 and the web application guide the image to the image web application, which guides the image to the AI evaluation 2105. Figure 21B Is a schematic diagram of the components in the image matching AI processing. The image uploaded in the Local Interface to Online Network (LION) 2106 of the veterinary clinic is guided to the VetConsole 2107, and the VetConsole 2017 automatically rotates and crops the image to obtain a sub-image. The sub-image is guided to three locations. The first location is to the VetAI console 2108 to classify the image. The second location is to the image matching console 2109 to add the sub-image and the report to the image matching database. The third location is to the image database 2110 that stores the new image and the corresponding case ID number. The image matching console 2109 guides the image to the fine image matching console 2111 or the VetImageEditor console 2112 for further processing.
[0040] Figure 22A and Figure 22B Are a set of schematic diagrams of the server architecture for image matching. Figure 22A Is a schematic diagram of the server architecture currently used in AI radiograph analysis. Figure 22BIt is a schematic diagram of a server architecture for AI radiograph analysis, including preprocessing radiology images, analyzing images using an AI diagnostic processor, and preparing a report based on the clustering results. Images from a personal computer (PC) 2201 are directed to an NGINX load balancing server 2202, which directs the images to a V2 cloud platform 2203. Then, the images are directed to an image matching server 2204, a VetImages server 2205, and a database Microsoft SQL server 2207. The VetImages server directs the images to a VetAI server 2206, a database Microsoft Structured Query Language (MicrosoftSQL) server 2207, and a data storage server 2208.
[0041] Figures 23A to 23F It is a series of schematic diagrams of an artificial intelligence automatic cropping and evaluation workflow for images obtained for a subject. The workflow is divided into six columns according to platforms for completing tasks such as a clinic, a V2 end-user application, a VetImages web application, a VetConsole python script application, a VetAI machine learning application, an ImageMatch-directed python application, and an ImageMatch verification python application. In addition, based on processors for completing tasks such as a sub-image processor, an evaluation processor, and a synthesis processor, the tasks are shaded in different gray levels. The V2 application is an end-user application where the user interacts with the application and uploads images to be analyzed. The VetImages application processes the images to generate AI results or an AI report or evaluation results. VetConsole is a python script application that improves image quality and processes images in batches. VetAI is a machine learning application that creates an AI model and evaluates the images entering the system. ImageMatch-directed is a python application that searches its database for correctly oriented images similar to the input image. ImageMatch verification is a python application that searches its database for correctly classified images similar to the input image. The sub-image processor completes Figures 23A - 23C the tasks 2301 - 2332 listed in Figure 23D as well as Figure 23E and Figure 23F a part of Figure 23F and Figure 23E the tasks 2333 - 2346, 2356, and 2357 listed in a part of
[0042] Figure 24It is a schematic diagram of the current model, showing that the image classifiers are sequentially replaced, with the oldest model at the top and the latest model at the bottom.
[0043] Figure 25 It is the new model shown by the claimed method, where the derived image classifiers are used together instead of being replaced, where I indicates the same image classifier, a number or N indicates a derivative classifier, and N represents any positive integer.
[0044] Figure 26 It shows a specific embodiment where a text version can be associated with a daisy-chain of derivative classifiers to create an n:1 relationship.
[0045] Figure 27 It shows a specific embodiment where multiple text versions can be associated with a daisy-chain of derivative classifiers to create an n:n relationship.
[0046] Figure 28 It shows visual examples of both derived (same letter and subscript number) data and non-derived (same letter but different subscript letter) data. The database stores all the training data with relationships. The system allows data input (in this case I, T, and OI); single or clustered, AI or non-AI data, derived or non-derived data, to obtain example results from the database.
[0047] Figure 29A It shows a radiograph of a cat's chest with mild lung pathology. The results of the general lung classifier (GLC) are 0.68; the results of GLC2 are 0.69. These results are used together as a check and balance system to confirm the results of each classifier.
[0048] Figure 29B It shows a radiograph of a cat's chest with moderate lung pathology, especially a moderate bronchial pattern. Both GLC and its derived classifier GLC2 are trained to identify the bronchial pattern in the radiograph. The results of GLC are 0.85; the results of GLC2 are 0.78. These results are used together as a check and balance system to confirm that the classifier results are true positives. Detailed Description
[0049] One aspect of the invention described herein provides a method for analyzing diagnostic radiology images or subject images, the method comprising: automatically processing, using a processor, radiology images of a subject for classifying the images into one or more body regions and orienting and cropping the classified images to obtain at least one oriented, cropped, and labeled sub-image of each automatically classified body region; directing the sub-images to at least one artificial intelligence processor; and evaluating the sub-images by the artificial intelligence processor to thereby analyze the radiology images of the subject.
[0050] Embodiments of the method further comprise using the artificial intelligence processor to assess body regions and the presence or absence of a medical condition in the sub-images. Body regions are, for example: chest, abdomen, forelimb, hindlimb, etc. Embodiments of the method further comprise diagnosing a medical condition from the sub-images using the artificial intelligence processor. Embodiments of the method further comprise using the artificial intelligence processor to assess the sub-images to localize the subject. Embodiments of the method further comprise correcting the localization of the subject to an appropriate localization.
[0051] In embodiments of the method, the processor automatically and rapidly processes the radiology images to obtain the sub-images. In embodiments of the method, the processor processes the radiology images to obtain the sub-images in a time less than about 1 minute, less than about 30 seconds, less than about 20 seconds, less than about 15 seconds, less than about 10 seconds, or less than about 5 seconds. In embodiments of the method, the evaluation further comprises comparing the sub-images with multiple reference radiology images in at least one of a plurality of libraries. In embodiments of the method, each of the plurality of libraries comprises a corresponding plurality of reference radiology images.
[0052] In embodiments of the method, each of the plurality of libraries comprises a corresponding plurality of reference radiology images specific to or non-specific to an animal species. Embodiments of the method further comprise matching the sub-images with the reference radiology images to thereby assess the orientation and at least one body region. In embodiments of the method, the reference radiology images are oriented according to the Digital Imaging and Communication in Medicine (DICOM) standard hanging protocol.
[0053] In an embodiment of the method, cropping also includes isolating a specific body region in the sub-image. An embodiment of the method also includes classifying a reference radiological image according to veterinary radiological standard body region markings. In an embodiment of the method, orienting also includes adjusting the radiological image to a veterinary radiological standard Hanging Protocol. In an embodiment of the method, cropping also includes trimming the radiological sub-image to a standard aspect ratio. In an alternative embodiment of the method, cropping does not include trimming the radiological sub-image to a standard aspect ratio. In an embodiment of the method, classifying also includes identifying and marking body regions according to veterinary standard body region markings. In an embodiment of the method, classifying also includes comparing the radiological image with a library of sample standard radiological images.
[0054] An embodiment of the method also includes matching the radiological image with a sample standard image in the library so as to classify the radiological image into one or more body regions. In an embodiment of the method, cropping also includes identifying the boundaries in the radiological image that demarcate each classified body region. An embodiment of the method also includes extracting features of the radiological image before classification. In an embodiment of the method, the radiological image is from a radiological examination selected from the following: radiographs, i.e., X-rays, magnetic resonance imaging (MRI), magnetic resonance angiography (MRA), computed tomography (CT), fluoroscopy, mammography, nuclear medicine, positron emission tomography (PET), and ultrasound. In an embodiment of the method, the radiological image is a photograph.
[0055] In an embodiment of the method, the subject is selected from the following: mammals, reptiles, fish, amphibians, chordates, and birds. In an embodiment of the method, mammals are selected from the following: dogs, cats, rodents, horses, sheep, cows, goats, camels, alpacas, water buffalo, elephants, and humans. In an embodiment of the method, the subject is selected from the following: pets, farm animals, high-value zoo animals, wild animals, and research animals. An embodiment of the method also includes automatically generating at least one report with sub-image evaluation by an artificial intelligence processor.
[0056] One aspect of the invention described herein provides a system for analyzing a radiological image of a subject, the system including: a receiver for receiving a radiological image of the subject; at least one processor for automatically running image recognition and processing algorithms to identify, crop, orient, and mark at least one body region in the image to obtain a sub-image; at least one artificial intelligence processor for evaluating the sub-image; and a device for displaying the sub-image and the artificial intelligence result of the evaluation.
[0057] In an embodiment of the system, the processor automatically and rapidly processes radiological images to obtain sub-images. In an embodiment of the system, the processor processes radiological images to obtain labeled images within a time of less than 1 minute, less than 30 seconds, less than 20 seconds, less than 15 seconds, less than 10 seconds, or less than 5 seconds. Embodiments of the system further include a library of standard radiological images. In an embodiment of the system, the standard radiological images conform to hanging film protocols and body region markings that comply with veterinary specifications.
[0058] One aspect of the invention described herein provides a method for rapidly and automatically preparing radiological images of a subject for display, the method comprising: using a processor to process an unprocessed radiological image of the subject to algorithmically classify the image into one or more separate body region categories, obtaining an optimally matched orientation and body region marking by automatically cropping, extracting features, and comparing the cropped, oriented image features with a database of image features of known orientations and body regions; and presenting each prepared body region marked image on a display device for analysis.
[0059] One aspect of the invention described herein provides an improvement to a veterinary radiograph diagnostic image analyzer, the improvement comprising: using a processor to run a fast algorithm that preprocesses a radiograph image of a subject to automatically identify one or more body regions in the image; the processor further being configured to perform at least one of the following operations: automatically creating a separate sub-image for each identified body region, cropping and optionally normalizing the aspect ratio of each created sub-image, automatically labeling each sub-image as a body region, automatically orienting the body region in the sub-image, and the processor further automatically directing the diagnostic sub-image to at least one artificial intelligence processor specific to evaluating the cropped, oriented, and labeled diagnostic sub-image.
[0060] One aspect of the invention described herein provides a method for identifying and diagnosing the presence of a disease or health condition in at least one image of a subject, the method comprising: classifying the image into one or more body regions, labeling and orienting the image to obtain a classified, labeled, and oriented sub-image; directing the sub-image to at least one artificial intelligence (AI) processor to obtain an evaluation result, and comparing the evaluation result with a database or at least one data cluster having evaluation results and matching written templates to obtain at least one clustering result; measuring the distance between the clustering result and the evaluation result to obtain at least one clustering diagnosis; and compiling the clustering diagnosis to obtain a report for identifying and diagnosing the disease or health condition present in the subject. The evaluation result is synonymous with the AI result, and the AI processor result and the classification result may be used interchangeably.
[0061] Embodiments of the method further include: obtaining at least one radiological image or one data point of a subject before classification. Embodiments of the method further include: before comparison, compiling data clusters using a clustering tool selected from the following: K-means clustering, mean shift clustering, density-based spatial clustering, Expectation-Maximization (EM) clustering, and agglomerative hierarchical clustering. In an embodiment of the method, compiling further includes obtaining, processing, evaluating, and constructing a library of multiple identified and diagnosed data sets and corresponding medical reports, the corresponding medical reports being selected from the following: radiology reports, laboratory reports, histology reports, physical examination reports, and microbiology reports, the medical reports relating to a variety of known diseases or health conditions. The term "medical report" includes any type of medical data. "Radiologic" and "radiographic" shall have the same meaning.
[0062] In an embodiment of the method, processing further includes classifying the multiple identified and diagnosed data set images into body regions to obtain multiple classified data set images, and further includes orienting and cropping the multiple classified data set images to obtain multiple oriented, cropped, and labeled data set sub-images. In an embodiment of the method, evaluating further includes guiding the multiple oriented, cropped, and labeled data set sub-images and the corresponding medical reports to at least one AI processor to obtain at least one diagnosed AI processor result. In an embodiment of the method, guiding further includes classifying the multiple oriented, cropped, and labeled data set sub-images and the corresponding medical reports with at least one variable selected from species, breed, weight, gender, and location.
[0063] In an embodiment of the method, constructing a library of multiple identified and diagnosed data set images further includes creating at least one cluster of the diagnosed AI processor results to obtain at least one AI processor example result, thereby compiling data clusters. In some embodiments, the AI processor example result is an example case, example result, example point, or example. These terms are synonyms and can be used interchangeably. Embodiments of the method further include assigning at least one cluster diagnosis to the cluster of the diagnosed AI processor results. In an embodiment of the method, assigning a cluster diagnosis further includes adding reports and / or additional information written by an evaluator within the cluster. In an embodiment of the method, measuring further includes determining the distance between the cluster result and at least one selected from the evaluation result, data clusters, and the centroid of the cluster result.
[0064] Embodiments of the method also include selecting a result from among: the case with the closest match within the cluster, another case in the cluster, and the centroid case. In an embodiment of the method, the selection also includes an assessor adding result information of the clustering result to a report generated from the cluster. Embodiments of the method also include editing the report by deleting a portion of the cluster diagnostic report that is below a popularity threshold among multiple reports within the cluster. In an embodiment of the method, the report is generated from words, partial statements, statements, and paragraphs that are considered available for report generation. The words in the report are obtained from the example result case with the closest match. If the words available for report generation contain at least one identifier selected from among: subject name, date, reference to a previous study, or any other word that may render the generated report not generally applicable to all new cases that most closely match the example result, those words are excluded. The selection process is performed by Natural Language Processing (NLP) and language AI.
[0065] In an embodiment of the method, a prevalence threshold is specified by an assessor. The threshold can be set between 0.000001% and 99.999999%. In an embodiment of the method, the assessment results are processed quickly by a diagnostic AI processor to obtain a report. In an embodiment of the method, the diagnostic AI processor processes an image to obtain a report within a very short time interval: less than about 10 minutes, less than about 9 minutes, less than about 8 minutes, less than about 7 minutes, less than about 6 minutes, less than about 5 minutes, less than about 4 minutes, less than about 3 minutes, less than about 2 minutes, or less than about 1 minute. In an embodiment of the method, an image library of a dataset of identified and diagnosed cases with known diseases and health conditions is classified into at least one of multiple animal species.
[0066] Embodiments of the method also include identifying the results of the diagnosed AI processor with an identification label. Embodiments of the method also include selecting AI results from the subject's image and / or medical results and adding them to the database cluster while retaining the original image and the processed image.
[0067] One aspect of the invention described herein provides a system for diagnosing the presence of a disease or health condition in an image and / or medical outcome of a subject, the system comprising: a receiver for receiving an image and / or medical outcome of the subject; at least one processor for automatically running an image recognition and processing algorithm to identify, crop, orient, and label at least one body region in the image to obtain a sub-image; at least one artificial intelligence processor for evaluating the sub-image and / or medical outcome and obtaining an evaluation result; and at least one diagnostic artificial intelligence processor for automatically running a clustering algorithm to compare the evaluation results to obtain a clustering result, measuring the distance between the clustering result and a previously created clustering result in a specific dataset defined by one or more variables, evaluating the result to obtain a clustering diagnosis, and compiling a report.
[0068] In an embodiment of the method, the diagnostic AI processor automatically and rapidly processes the image and / or medical outcome to generate a report. In an embodiment of the method, the diagnostic AI processor processes the image and / or medical outcome to obtain a report within a time of: less than about 10 minutes, less than about 9 minutes, less than about 8 minutes, less than about 7 minutes, less than about 6 minutes, less than about 5 minutes, less than about 4 minutes, less than about 3 minutes, less than about 2 minutes, or less than about 1 minute. An embodiment of the method further includes a device for displaying the generated report.
[0069] One aspect of the invention described herein provides a method for diagnosing the presence of a disease or health condition in at least one image of a subject, the method comprising: classifying the image into at least one body region, labeling, cropping, and orienting the image to obtain at least one classified, labeled, cropped, and oriented sub-image; directing the sub-image to at least one artificial intelligence (AI) processor for processing and obtaining an evaluation result, and comparing the evaluation result with a database having a plurality of evaluation results and matching written templates or at least one data set clustering to obtain at least one clustering result; measuring the distance between the clustering result and the evaluation result to obtain at least one clustering diagnosis; and compiling the clustering diagnosis and the matching written template to obtain a report and displaying the report to identify and diagnose the presence of a disease or health condition in the subject.
[0070] An embodiment of the method further includes, after displaying, analyzing the report and confirming the presence of a disease or health condition. An alternative embodiment of the method further includes editing the written template. In an embodiment of the method, the processing time for obtaining the report is: less than about 5 minutes, less than about 2 minutes, or less than about 1 minute. In an embodiment of the method, the processing time for obtaining the report is: less than about 10 minutes, less than about 7 minutes, or less than about 6 minutes.
[0071] In an embodiment of the method, processing the sub-image further includes training an AI processor to diagnose the presence of a disease or a health condition in a subject image. In an embodiment of the method, training the AI processor further includes the steps of: passing a training image library to the AI processor, creating an AI model, storing the AI model in a database, testing the AI model using expected positive and negative data not used for training, and comparing the actual test set results with the expected test set results.
[0072] In an embodiment of the method, the training image library includes positive control training images and negative control training images. In an embodiment of the method, the positive control training images have the disease or health condition of the training images. In an embodiment of the method, the negative control training images do not have the disease or health condition of the training images. In various embodiments of the method, the negative control training images may have a disease or health condition other than the disease or health condition of the training images. In an embodiment of the method, the training image library further includes at least one of medical data, metadata, and auxiliary data.
[0073] One aspect of the present invention describes a novel system having several analysis stages, including determining whether a received image includes a specific body region in a specific orientation (such as a side view, etc.), appropriately cropping the image, and evaluating the cropped image by comparing it with a target AI model. In various embodiments, the newly received image is preprocessed to automatically identify and label one or more body regions and / or views represented in the image without user input or intervention. In some embodiments, the image is automatically cropped to generate one or more sub-images corresponding to the identified individual body regions / views. In some embodiments, the image and / or sub-images are selectively processed to a target AI processor configured to evaluate the identified body region / view without involving the remaining AI processors in the system.
[0074] In some embodiments, the radiology image pre-processor 106 additionally or alternatively tags the entire image 102 to identify the identified body regions and / or views within the image 102, and then passes the entire image 102 only to those AI processors 104 corresponding to the applied tags. Thus, in such embodiments, the AI processors 104 are responsible for cropping the image 102 to focus on the relevant regions for further analysis using one or more trained machine learning models or other methods. In some embodiments, in addition to tagging the image 102 with tags corresponding to specific body regions / views, the radiology image pre-processor 106 also additionally crops the image 102 to primarily focus on the regions within the image that actually represent animal parts and to remove as much of the black border around these regions as possible. In some embodiments, performing such a cropping step facilitates further cropping and / or other processing by the AI processors 104, and subsequently deploys the AI processors 104 to evaluate the specific body regions / views corresponding to the applied tags.
[0075] The radiology image pre-processor 106 is implemented in any of several ways. In some embodiments, for example, the radiology image pre-processor 106 employs one or more algorithms to identify one or more features indicative of one or more specific body regions and automatically crops the image 102 to focus on the regions including these features and / or the regions that actually represent the animal. In certain implementations, these algorithms are implemented using elements of the OpenCV-Python library. A description of the Open Source Computer Vision (“OpenCV”) library, as well as associated documentation and tutorials, can be found using the uniform resource locator (URL) for OpenCV. The entire content of the material accessible through the URL is incorporated herein by reference. In some embodiments, the radiology image pre-processor 106 additionally or alternatively employs image matching techniques to compare the image 102 and / or one or more of its cropped sub-images 108 with a library of stored images known to represent specific views of specific body regions and to determine that the image 102 and / or sub-image 108 represents the body region / view most strongly correlated with one or more of the stored images. In some embodiments, an AI processor trained to additionally or alternatively perform body region / view identification is employed within the radiology image pre-processor 106.
[0076] In some embodiments, one or more of the AI processors described herein are implemented using the TensorFlow platform. Descriptions, documentation, and tutorials for the TensorFlow platform can be found on the TensorFlow website. The entire content of the materials accessible through the website is incorporated herein by reference. Hope, Tom, et al. describe the TensorFlow platform and methods for building AI processors in Learning TensorFlow: A Guide to Building Deep Learning Systems. O'Reilly., 2017, the entire content of which is incorporated herein by reference.
[0077] In Figure 3 In the example shown, the preprocessing performed by the radiology image preprocessor 106 includes (1) an optional "general" automatic cropping step (reflected in the first five log entries delimited by bracket 306), according to which the image 302 is initially cropped to primarily focus on the image region representing the animal part and to remove as much of the black border around these regions as possible, (2) a "classification" automatic cropping step (reflected in log entries 6 to 9 within bracket 306), according to which a preliminary effort is made, for example using elements of the OpenCV-Python library, to identify a specific body region / view and crop the image 302 to focus on that body region / view, and (3) an AI region tagging step or "image matching" step (reflected in the last three log entries delimited by bracket 306), according to which the image 302 and / or one or more of its cropped sub-images 304a - 304b are compared with a library of stored images known to represent specific views of specific body regions. As indicated by the corresponding timestamps, it was observed that the general automatic cropping step was completed within two seconds, the classification automatic cropping step was completed within three seconds, and the image matching step was completed within nineteen seconds.
[0078] As Figure 3 As shown by the log entry delimited by bracket 308a, the lateral chest AI processor 104a took 4 seconds to determine whether the image 302 includes a lateral view of the animal's chest. Similarly, as shown by the log entry delimited by bracket 308b, the lateral abdomen AI processor 104c determined whether the image 302 includes a lateral view of the animal's abdomen within four seconds.
[0079] If the system needs to process the newly received image 302 using all possible AI processors 104a - 104l, rather than just the two AI processors corresponding to the body part / view identified by the radiology image preprocessor 106, then the time of the AI processors will be greatly extended and / or more processing resources will be consumed to complete the analysis. For example, in a system including thirty different AI processors 104, simply identifying the relevant AI model for determining the health status of the imaged animal will require at least one hundred and twenty seconds of processing time by the AI processors 104 (i.e., thirty AI processors, four seconds each), and the processing time can be much longer when each AI processor 104 considers multiple possible orientations of the image. On the other hand, by employing the radiology image preprocessor 106, it is observed that the identification of the relevant AI model only requires 8 seconds of processing time by the AI processors 104, plus 24 seconds of preprocessing time by the radiology image preprocessor 106.
[0080] This is useful for processing radiology images using artificial intelligence (AI) processors (such as trained neural networks) to identify and label image data and determine the probability that the imaged data has certain medical conditions. Typically, separate AI processors are used to evaluate individual body regions (e.g., chest, abdomen, shoulder, forelimb, hindlimb, etc.) and / or specific orientations of each such body region (e.g., ventrodorsal (VD) view, lateral view, etc.), and each such AI processor determines the probability that a specific health condition exists in the body region in question for the individual body region and / or orientation. Each such AI processor can include a large number of trained models to evaluate the corresponding health conditions or organs within the imaged region. For example, for a lateral view of an animal's chest, the AI processor can employ different models to determine the probability that the animal has certain health conditions related to the lungs, such as perihilar infiltration, pneumonia, bronchitis, lung nodules, etc.
[0081] Currently, radiology AI is practiced for the detection of a single disease condition, such as the presence or absence of pneumonia or pneumothorax. Compared with the single-disease detection of current radiology AI, human radiologists analyze radiographs in a holistic approach by simultaneously assessing the presence or absence of multiple health conditions. A limitation of the current AI process is that a separate AI detector must be used for each specific health condition. However, combinations of health conditions can lead to the diagnosis of a broader range of diseases. For example, in some cases, one or more diagnostic results obtained from radiological images are caused by several broader diseases. Determining the broader diseases present in a subject's radiograph requires the use of complementary diagnostic results in a process called differential diagnosis. In addition to radiological images, these complementary diagnostic results are also extracted from blood tests, patient history, biopsies, or other tests and procedures. The current AI process focuses on a single diagnostic result and cannot identify the broader diseases that require differential diagnosis. This article describes a novel AI process that can combine multiple diagnostic results to diagnose broader diseases.
[0082] The AI process currently uses specific-region, limited radiological images, which are common in radiological images of human subjects. In contrast, veterinary radiology typically includes multiple body regions in a single radiograph. This article describes a novel AI evaluation process for evaluating all body regions included in a study and providing a broader evaluation expected in veterinary radiology.
[0083] Figure 4 The traditional workflow of the current artificial intelligence reporting a single disease process is shown. Figure 4 The traditional single health condition report shown is insufficient for differential diagnosis of radiographs. In addition, using personalized rules for each combination of evaluation results is inefficient in creating reports and cannot meet the reporting standards of veterinary radiologists. Even for a single-disease process, determining the severity of a specific health condition, such as normal, mild, moderate, and severe, results in an exponential number of AI model result templates. The process of creating and selecting a single report template from a set of AI model diagnostic results changes in scale exponentially with the number of AI models. As Figures 5A - 5F shown, the number of AI models for a single-disease process results in 57 different templates for five different severities. Therefore, reports created manually for each combination of artificial intelligence model diagnostic results do not scale well for a large number of artificial intelligence models to be interpreted simultaneously.
[0084] AI Analysis Automation System
[0085] This document describes a novel system for analyzing images of a test subject animal, the system comprising: a receiver for receiving an image of the test subject animal; at least one sub-image processor for automatically identifying, cropping, orienting, and labeling at least one body region in the image to obtain sub-images; at least one artificial intelligence evaluation processor for evaluating whether the sub-images have at least one health condition; at least one synthesis processor for generating an overall result report based on at least one sub-image evaluation and optionally based on non-image data; and a device for displaying the sub-images and the overall synthesized diagnostic result report.
[0086] The system has made significant progress in veterinary diagnostic image analysis in the following ways: (1) using the sub-image processor to automatically extract sub-images, which is a task typically performed manually or with user assistance, and (2) using the synthesis processor to synthesize a large number of evaluation results and other non-image data points into a concise and coherent overall report.
[0087] A case includes a set of one or more images of the test subject animal and may include non-image data points such as, but not limited to, age, gender, location, medical history, and other medical test results. In an embodiment of the system, each image is sent to multiple sub-image processors, generating many sub-images of various views of multiple body regions. Each sub-image is processed by multiple evaluation processors, generating a large number of evaluation results for many different health conditions, findings, or other characteristics across multiple body regions. The synthesis processor processes all or a subset of the evaluation results and non-image data points to produce an overall synthesized diagnostic result report. In an embodiment of the system, multiple synthesis processors generate multiple synthesized diagnostic result reports based on different subsets of the evaluation results and non-image data points. These diagnostic reports are compiled with auxiliary data to create a final comprehensive diagnostic result report.
[0088] In an embodiment of the system, each synthesis processor operates on a subset of sub-images corresponding to a body region (such as the chest or abdomen) and non-image data points. Each comprehensive diagnostic report includes the body region, which is typical in veterinary radiology. The overall synthesized diagnostic result report includes descriptive data of the subject such as name, age, address, breed, and multiple sections corresponding to the output of each synthesis processor, such as a chest diagnostic result section and an abdominal diagnostic result section.
[0089] In an embodiment of the system, the test subject is selected from: mammals, reptiles, fish, amphibians, chordates, and birds. Mammals include dogs, cats, rodents, horses, sheep, cows, goats, camels, alpacas, buffalo, elephants, and humans. The test subject is a pet, a farm animal, a high-value zoo animal, a wild animal, and a research animal.
[0090] The images received by the system are images from radiological examinations such as X-rays (radiographs), magnetic resonance imaging (MRI), magnetic resonance angiography (MRA), computed tomography (CT), fluoroscopy, mammography, nuclear medicine, positron emission tomography (PET), and ultrasound. In some embodiments, the images are photographs.
[0091] In some embodiments of the system, the analysis of the subject's images generates and displays an overall composite result report within a very short time interval, which is: less than about 20 minutes, less than about 10 minutes, less than about 5 minutes, less than about one minute, less than about 30 seconds, less than about 20 seconds, less than about 15 seconds, less than about 10 seconds, or less than about 5 seconds.
[0092] Sub-image processor
[0093] The sub-image processor orients, crops, and labels at least one body region in the image to automatically and quickly obtain sub-images. The sub-image processor orients the image by rotating the image to a standard orientation according to a specific view. The orientation is determined by the veterinary radiograph standard hanging protocol. The sub-image processor crops the image by identifying the boundaries depicting one or more body regions in the image and creating a sub-image containing the image data within the identified boundaries.
[0094] In some embodiments, the boundaries have a consistent aspect ratio. In alternative embodiments, the boundaries do not have a consistent aspect ratio. The sub-image processor labels the sub-images by reporting the boundaries and / or positions of each body region contained within the sub-images. Body regions include, for example: chest, abdomen, spine, forelimbs, left shoulder, head, neck, etc. In some embodiments, the sub-image processor labels the sub-images according to veterinary radiology standard body region markings.
[0095] The sub-image processor matches the image with multiple reference images in at least one of multiple libraries to locate, crop, and label one or more sub-images. Each of the multiple libraries includes a corresponding multiple of reference images specific to or not specific to an animal species.
[0096] The sub-image processor extracts features of the image before orienting, cropping, and / or labeling the image, thereby allowing the image or sub-image to be quickly matched with similar reference images. The sub-image processor processes the image to obtain sub-images within the following short time intervals: for example, less than about 20 minutes, less than about 10 minutes, less than about 5 minutes, less than about 1 minute, less than about 30 seconds, less than about 20 seconds, less than about 15 seconds, less than about 10 seconds, less than about 5 seconds, less than about 4 seconds, less than about 3 seconds, less than about 2 seconds, less than about 1 second, less than about 0.5 seconds, and less than about 0.1 seconds.
[0097] Evaluation processor
[0098] The artificial intelligence evaluation processor determines whether a sub - image has a health condition, finding, or other feature. The evaluation processor reports the probability of the presence of a health condition, finding, or feature.
[0099] The evaluation processor diagnoses the presence of a medical condition from the sub - image. The evaluation processor determines non - medical features of the sub - image, such as the correct positioning of the subject. The evaluation processor generates instructions for correcting the subject's positioning.
[0100] Typically, the evaluation processor trains on negative control / normal and positive control / abnormal training sets related to health conditions, findings, or other features. The positive control / abnormal training set typically includes cases where the presence of a health condition, finding, or other feature is determined. The negative control / normal training set includes cases where there is no health condition, no finding, or no other feature that has been determined and / or cases that are considered completely normal. In some embodiments, the negative control / normal training set includes cases where the presence of other health conditions, findings, or features different from the presence of interest has been determined. Thus, the evaluation processor is robust.
[0101] The evaluation processor processes the sub - image to report the presence of a health condition within the following times: less than about 20 minutes, less than about 10 minutes, less than about 5 minutes, less than about 1 minute, less than about 30 seconds, less than about 20 seconds, less than about 15 seconds, less than about 10 seconds, less than about 5 seconds, less than about 4 seconds, less than about 3 seconds, less than about 2 seconds, less than about 1 second, less than about 0.5 seconds, and less than about 0.1 seconds.
[0102] Synthesis processor
[0103] The synthesis processor receives at least one evaluation from the evaluation processor and generates a comprehensive result report. The synthesis processor may include non - image data points, such as species, breed, age, weight, location, gender, medical test history including blood, urine, and fecal tests, radiology reports, laboratory reports, histology reports, physical examination reports, microbiology reports, or other medical and non - medical tests or results. An example result for a subject's case includes a set of at least one image, the associated evaluation processor results, and zero or more up - to - date non - image data points.
[0104] In an embodiment of the method, the synthesis processor uses the example case results to select results stored in the database.
[0105] In an embodiment of the method, it can be pre - written words, keywords, partial statements, complete statements, partial paragraphs, paragraphs, and / or more paragraphs to be output as an overall result report. The template is automatically customized based on the example case result elements to provide a customized overall result report.
[0106] The synthesis processor assigns the case example results of a subject to a cluster group. The cluster group contains other similar case example results from a reference library of case example results of other subjects. In some cases, the cluster group contains partial case example results, such as result reports. The reference library includes case example results from at least one of a variety of animal species, with or without known diseases and health conditions. New case example results are added to the reference library to improve the performance of the synthesis processor over time. The synthesis processor assigns coordinates representing the location of each case example result within the cluster group.
[0107] The synthesis processor assigns a single overall result report to the entire cluster group and designates the overall result report to the subject. In some embodiments, several overall result reports are assigned to various case example results within the cluster and / or to individual custom coordinates within the cluster, such as cluster centroids, without associated case example results. The coordinates of the subject's case example results are used to calculate the distance to the nearest or non-nearest case example result or custom coordinate with an associated overall result report, which is then assigned to the subject.
[0108] The overall result report is written by an experienced human assessor. In alternative embodiments, one or more overall result reports are generated from existing radiology reports. The content of the existing radiology reports is associated with example AI classifier results and modified by natural language processing (NLP) or AI to remove content that is not generally applicable, such as names, dates, references to previous studies, etc., to create a suitable overall result report. If a statement in the overall result report does not meet the prevalence threshold within the cluster, it is deleted or edited.
[0109] In some embodiments, a personalized output report is created in which the AI system is trained to best match the image evaluation and interpretation style of a single radiologist. In the personalized system, the radiologist's image ratings and nuances in word, statement, paragraph, grammar, and report format choices are stored in a database and used to create a personalized and stylized report using a standard AI classifier, personalized AI result thresholds, personalized AI result weighting, and a radiologist-specific AI language profile.
[0110] The synthesis processor outputs the designated overall result report of the subject, thereby identifying and diagnosing the presence of one or more findings, diseases, and / or health conditions within the subject. The following clustering tools are used to establish the cluster group from the reference library of case example results: K-means clustering, mean shift clustering, density-based spatial clustering, expectation maximization (EM) clustering, and agglomerative hierarchical clustering.
[0111] The synthesis processor processes case example results to generate an overall result report at the following times: less than about 20 minutes, less than about 10 minutes, less than about 9 minutes, less than about 8 minutes, less than about 7 minutes, less than about 6 minutes, less than about 5 minutes, less than about 4 minutes, less than about 3 minutes, less than about 2 minutes, less than about 1 minute, less than about 30 seconds, less than about 20 seconds, less than about 15 seconds, less than about 10 seconds, less than about 5 seconds, less than about 4 seconds, less than about 3 seconds, less than about 2 seconds, less than about 1 second, less than about 0.5 seconds, and less than about 0.1 seconds.
[0112] Clustering is an AI technique for grouping unlabeled examples based on the feature similarity of each example. This document describes a process for clustering patient studies based on AI processor diagnostic results and non-radiological and / or non-AI diagnostic results. The clustering process groups reports that share similar diagnoses or output reports, thereby facilitating the overall detection of health conditions or broader diseases in a scalable manner.
[0113] This document describes a novel system and method with multi-stage analysis that combines multiple AI predictive image analysis methods for radiograph images and a report library database, as well as the evaluation of newly received images, to accurately diagnose and report radiology cases. In various embodiments, the novel system described herein automatically detects the views and the regions covered by each radiology image.
[0114] In some embodiments, the system preprocesses the newly received radiology image 102 using a radiology image preprocessor 106 before AI evaluation to crop, rotate, flip, create sub-images, and / or normalize the image exposure. If more than one body region or view is identified, the system further crops the image 102 to generate one or more sub-images 108a, 108b, and 108c corresponding to the identified individual regions and views. In some embodiments, the system selectively processes the image and / or sub-images and directs them to a target AI processor configured to evaluate the identified region / view. Image 108a is only directed to AI processor 104a, which is a lateral chest AI processor. Image 108b is only directed to AI processor 104c, which is a lateral abdomen AI processor. Image 108c is only directed to AI processor 104k, which is a lateral pelvis AI processor. The image is not directed to non-target AI processors or the remaining AI processors in the system. For example, Figure 7 a chest image points to Figure 6 one or more AI processors for the diseases listed in , such as heart failure, pneumonia, bronchitis, interstitial disease, lung lesions, tracheal hypoplasia, cardiac enlargement, lung nodules, pleural effusion, gastritis, esophagitis, bronchiolectasis, pulmonary hyperinflation, pulmonary vascular enlargement, thoracic lymphadenopathy, etc.
[0115] In some embodiments, the AI model processor is a binary processor that provides a binary result of normal or abnormal. In various embodiments, the AI model processor provides a normal or abnormal diagnosis by determining the severity of a particular health condition. For example, the severity of the health condition can be classified as normal, mild, moderate, or severe.
[0116] In some embodiments, the newly received AI model processor results are displayed in the user interface. See Figures 9A - 9E .. The average AI model processor results for each model are collected and displayed from the evaluation results of a single image or sub-image. See Figure 9A .. The user interface displays a single image or sub-image and the AI model processor results for that image. See Figures 9B - 9E .. The AI analysis is completed in less than one minute, two minutes, or three minutes.
[0117] In some embodiments, the system uses the AI processor diagnostic results from a known radiology image library and the corresponding radiology report database to construct one or more clusters to develop the closest matching cases or AI processor "example results" for one or more AI processor results. The example results include at least one image, a set of associated evaluation processor results, and a set of zero or more non-image data points, such as age, gender, location, breed, medical test results, etc. The synthetic processor assigns coordinates representing the location of each case example result within the cluster group. Thus, if two cases have similar example results, then the diagnoses are similar or substantially the same, and a single overall result report applies to both cases. In some embodiments, a single example result is assigned to the entire cluster, and the subject cases within the cluster are assigned the example result. In some embodiments, multiple exemplary results are assigned to the cluster, and these results are either associated with a specific coordinate (e.g., centroid) within the cluster or with a specific dataset case. In some embodiments, the example results are written by a person or automatically generated based on the existing radiology reports associated with the cases.
[0118] In some embodiments, the user uses Figure 11The user interface specifies various parameters for creating clusters from a known radiology image library and a corresponding radiology report database. The user assigns a name 1101 to the new cluster under "Code". The user selects various parameters to create the cluster. The user selects a start case date 1102 and an end case date 1103 to select cases. The user selects a start case ID (1104) and an end case ID 1105 to select cases. The user selects the maximum number of cases 1106 to be included in the cluster. The user selects the species 1107 of the cases to be included in the cluster, such as dogs, cats, dogs or cats, humans, bird pets, farm animals, etc. The user selects a specific form of diagnosis 1108, such as X-ray, CT, MRI, blood analysis, urine analysis, etc., to be included in creating the cluster.
[0119] In various embodiments, the user specifies to divide the evaluation results into a specific number of clusters. The number of clusters ranges from a minimum of one cluster to a maximum number of clusters, and the maximum number of clusters is only limited by the total number of cases entering the cluster. In addition to the diagnostic results of the AI processor, the system also uses non-radiology and / or non-AI diagnostic results (such as blood tests, patient history, or other tests or procedures) to construct one or more clusters. As Figure 12 shown, these clusters are listed numerically in comma-separated value (CSV) file format. The CSV file lists the case IDs 1201 of the cases in the cluster. The average evaluation results 1202 of each binary model for a specific case ID are listed in the CSV file. The cluster label or cluster location 1203 of a specific case containing the set of evaluation results is listed in the CSV file. The CSV file lists the cluster coordinates. The case ID 1204 of the centroid or center of a specific cluster is listed in the CSV file.
[0120] In various embodiments, the clusters are represented by a cluster graph. See Figure 13 . The cluster graph is created by dividing the average evaluation results into multiple different clusters according to the user-defined parameters 1102 - 1108. The various different clusters are represented by a set of points plotted on the graph. Figure 13 The cluster graph of shows 180 clusters of different sizes.
[0121] In some embodiments, the user interface shows an AI cluster model generated based on the user-defined parameters 1102 - 1108. See Figure 14 . The user interface shows a screening evaluation configuration, where the user assigns a specific "cluster model" 1502 to a specific "screening evaluation configuration" name 1501. The status 1503 of the screening evaluation configuration provides additional information about the configuration, such as whether the configuration is in real-time, test, or draft mode. Real-time mode is for production, and test or draft mode is for development.
[0122] In some embodiments, the user interface describes details of a specific clustering model 1601 Thorax 97. See Figure 16A . In some embodiments, the user interface lists the AI evaluation classifier types 1602 included in the cluster. The user interface displays additional parameters used to construct the cluster, such as the cluster model species specific 1603, the maximum number of cases with evaluation results 1604, or the start and end dates of the cases used to create the cluster 1605. The user interface provides a link 1606 to a comma-separated values (CSV) file that displays the cluster in tabular form. The user interface lists the sub-clusters 1608 created according to the parameters 1602 - 1605. The user interface displays the total number of sub-clusters 1609 created for the cluster group. The user interface provides a centroid case ID 1610 for each sub-cluster. A log for constructing the cluster is provided in the user interface. See Figure 16B .
[0123] In various embodiments, the system utilizes one or more AI processors to evaluate newly received undiagnosed images and obtain newly received evaluation results. The system compares the newly received evaluation results with one or more clusters obtained from a known radiology image library and a corresponding radiology report database.
[0124] The user imports the newly received AI processor results into the AI evaluation tester. See Figure 17A . The user specifies the screening evaluation type 1702 and the corresponding cluster model 1703.
[0125] In addition to the newly received evaluation results, the system also compares non-radiology and / or non-AI diagnostic results with one or more clusters obtained from a known radiology image library, a corresponding radiology report database, and other available databases. The system measures the distance between the location of the newly received AI processor results and the cluster results and creates a radiologist report using one or more cluster results. In some embodiments, the system selects to use the entire radiologist report or a portion of the radiologist report from the known cluster results based on the location of the newly received AI processor results relative to the known cluster results. In various embodiments, the system selects to use the entire radiologist report or a portion of the radiologist report from other results in the same cluster. In some embodiments, the system selects to use the entire radiologist report or a portion of the radiologist report from the centroid of the cluster results.
[0126] The user interface displays the results of the AI evaluation tester. See Figure 18A。In various embodiments, the diagnosis and conclusive findings 1801 from a radiologist report are displayed, which is the radiologist report closest to the evaluation result based on the previously created clustering results. In some embodiments, the evaluation result 1802 is selected from the radiologist reports in the evaluation result clustering and filtered according to the prevalence of specific statements in the findings section of a specific cluster. In some embodiments, the recommendations 1803 from the radiologist reports in the cluster are selected based on the prevalence rate of each statement or similar statements in the recommendations section of the cluster. The user interface displays the clustered radiologist reports 1804 and the radiologist report of the centroid of the cluster 1805. The user interface allows the user to edit the report by adding or deleting specific statements in the findings section 1809, the conclusion section 1810, or the recommendations section 1811. See Figure 18D and Figure 18E 。The radiologist report closest to the database case is used to generate a radiology report for the newly received radiological image. The statements in the radiology report based on specific clustering results are ranked and listed according to the ranking and prevalence. See Figure 18A and Figure 18B 。
[0127] In various embodiments, the system utilizes one or more AI processors to evaluate the newly received undiagnosed images and obtain the newly received evaluation results. The system compares the newly received evaluation results with one or more clusters obtained from a known radiological image library and the corresponding radiological report database.
[0128] In addition to the newly received evaluation results, the system also compares non-radiological and / or non-AI diagnosis results with one or more clusters obtained from a known radiological image library, the corresponding radiological report database, and other available databases. The system measures the distance between the position of the newly received AI processor results and the clustering results and creates a radiologist report using one or more clustering results. In some embodiments, the system selects to use the entire radiologist report or a part of the radiologist report in the known clustering results according to the position of the newly received AI processor results relative to the known clustering results. In various embodiments, the system selects to use the entire radiologist report or a part of the radiologist report from other results in the same cluster. In some embodiments, the system selects to use the entire radiologist report or a part of the radiologist report from the centroid of the clustering results.
[0129] In some embodiments, one or more of the AI processors described herein are implemented using the TensorFlow platform. A description of the TensorFlow platform, as well as related documentation and tutorials, can be found on the TensorFlow website. The entire content of the materials accessible on the TensorFlow website is hereby incorporated by reference in its entirety.
[0130] In some embodiments, one or more of the clustering models described herein are implemented using the Plotly platform. A description of the Plotly platform, as well as related documentation and tutorials, can be found on the scikit-learn website. The entire content of the materials accessible on the scikit website is hereby incorporated by reference in its entirety. Methods for developing AI processors and clustering models using the TensorFlow platform and Scikit-learn are described in detail in the following references: Géron Aurélien. Hands-on Machine Learning with Scikit-Learn and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems. O'Reilly, 2019 (by Aurélien Géron, Hands-on Machine Learning with Scikit-Learn and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems, O'Reilly Media, 2019); Hope, Tom, et al. Learning TensorFlow: A Guide to Building Deep Learning Systems. O'Reilly., 2017 (by Tom Hope et al., Learning TensorFlow: A Guide to Building Deep Learning Systems, O'Reilly Media, 2017); and Sievert, Carson. Interactive Web-Based Data Visualization with R, Plotly, and Shiny. CRC Press, 2020 (by Carson Sievert, Interactive Web-Based Data Visualization: Based on R, Plotly, and Shiny, CRC Press, 2020). Each of these references is hereby incorporated by reference in its entirety.
[0131] If a radiologist attempts to evaluate each model and create rules based on the individual AI processor results found together, the creation of rules and reports is prohibitively time-consuming. Additionally, as the number of incorporated AI processor models increases, adding a single additional AI processor model in this scenario becomes exponentially more difficult. The problem of manually constructing reports when multiple AI processor results are found is addressed by adopting a novel workflow of AI processor result clustering or "exemplary result" comparison between new images and known data sets. Manually constructing reports and creating rules from individual AI processor results previously took months, while using the novel workflow takes only a short time.
[0132] In some embodiments of the system, components for AI evaluation are as Figure 21A shown. In various embodiments of the system, components for AI processing of image matching are as Figure 21B shown.
[0133] In various embodiments of the system, the server architecture for AI radiograph analysis includes preprocessing radiology images, analyzing the images using an AI diagnostic processor, and preparing a report based on the clustering results. See Figure 22B . In some embodiments, various servers including an NGINX load balancing server 2202, a V2 cloud platform 2203, a database Microsoft SQL server 2207, and a data storage server 2208 are used.
[0134] In various embodiments of the system, a user labels cases for training the AI system. In some embodiments of the system, the user labels cases if the radiology report needs to be edited due to inaccuracies in the radiology report, or if the report is insufficient, or if the case has a novel diagnosis and thus the radiology report requires new diagnostic language.
[0135] Figures 23A - 23FShows a series of schematic diagrams of the AI automatic drawing and evaluation workflow. The user accesses the V2 end-user application 2301 to upload an image to be analyzed by the system (in image formats such as Digital Imaging and Communications in Medicine (DICOM), Joint Photographic Experts Group (JPEG), Joint Image Group (JPG), Portable Network Graphics (PNG), etc.). In some embodiments, the image is directly uploaded 2305 in the VetImages application. The V2 processes 2302 the image, saves it to the data storage, and requests 2303 VetImages to further process the image. VetImages receives the request from V2 and starts 2304 asynchronous processing. VetImages accesses 2307 the image from the data storage and requests 2308 VetConsole to preprocess the image. VetConsole uses OpenCV 2309 to improve the quality of the image and automatically crops 2310 the image. The tasks after accessing the image from the data storage are performed by the sub-image processor.
[0136] VetConsole sends the automatically cropped image with improved quality to VetImages. VetImages analyzes the image and asks VetConsole to classify 2311 the chest, abdomen, and pelvis in the image. VetConsole classifies 2312 the chest, abdomen, and pelvis in the image and sends the coordinates to VetImages. VetImages sends the image and coordinates to ImageMatch for verification 2313. ImageMatch verification matches the image and coordinates with the correctly classified images in its database and sends 2314 the matching image distance and path to VetImages. The VetImages application receives the data of the matched image and confirms 2315 the body area using the database information. The next task is to determine the image orientation. The image is rotated and flipped 2317. After each rotation and flip, the image is sent 2318 to the ImageMatch orientation application, compared with the matched image, and the distance and image path between the matched image and the newly received image are measured. The ImageMatch orientation application sends the result 2319 with the distance and image path between the newly received image and the matched image. The orientation of the newly received image with the minimum distance from the matched image is selected 2320 by the VetImages application. The process of checking each orientation and each flip is repeated until the image is rotated 360 degrees and flipped at an appropriate angle. In some embodiments, the image with the selected orientation is sent to VetAI to detect 2321 the chest and 2323 the abdomen and obtain the coordinates for cropping the image to obtain sub-images with the chest 2322 and abdomen 2324. The process of obtaining the coordinates is coordinated with TensorFlow.
[0137] The VetImages application obtains coordinates from the ImageMatch verification and crops the 2325 image according to the coordinates to obtain a sub-image. The sub-image is sent 2326 to the ImageMatch verification application for matching. The database image is matched with the sub-image 2327, and the distance between the matched database image and the sub-image, as well as the image path, are sent to the VetImages application. The VetImages application receives 2328 the distance and image path data and uses the data received from the matched image to confirm the body area. The VetImages application analyzes each sub-image to check if each sub-image is valid. If the sub-image is invalid 2331, the regular cropped image from the VetConsole application is saved in the database or data storage 2332. If the sub-image is valid 2330, the sub-image is saved in the database or data storage 2332. The images saved in the database or data storage are the cropped images for further processing or analysis. The VetImages application saves the data of the sub-image or regular cropped image obtained from the VetConsole in the database 2332.
[0138] The following tasks are performed by the evaluation processor. The VetImages application sends 2333 the cropped image for localization evaluation to the VetAI application. The data 2334 received by the VetImages application from the VetAI application is saved in the database 2335, and a signal is sent to the V2 application to send an email to the clinic 2336. The VetImages application accesses 2337 the real-time AI model from the database. Based on the body area of the cropped image, the cropped image is sent 2339 to the appropriate AI model in the VetAI application. The appropriate AI model is predetermined for each body area. The VetAI application sends 2340 the AI evaluation label and the machine learning (ML) AI evaluation result to the VetImages application, and the VetImages saves this data 2341 in the database of the cropped image. The VetImages application calculates 2342 the label and probability of the image based on the AI evaluation result of the cropped image. The process of sending 2339 the cropped image to obtain the AI evaluation result is repeated 2342 until the predetermined AI model 2338 is processed.
[0139] The VetImages application analyzes whether 2344 each image in the case has been processed to obtain the AI evaluation result. If all the images in the case have not been processed, VetImages will return to the next image in the case for processing. If all the images in the case have been processed, VetImages calculates 2345 the label and probability of the entire case based on the label and the probability of each cropped image. Then, the VetImages application changes 2346 its status to the live state and filters the evaluation type from the database. After changing the status of the VetImages application to the live state, the task is executed by the synthesis processor. The VetImages application determines 2347 whether all the screening evaluations have been completed. If all the screening evaluations have not been completed, the VetImages application determines 2348 whether the screening evaluation must be completed through clustering. If the screening evaluation must be completed through clustering, the AI evaluation result of the processed image is sent 2349 to the VetAI application, and the best-matching clustering result is sent 2350 to the VetImage application, which generates and saves 2351 the screening result according to the result of the best-matching clustering in the database. If the VetImages application determines not to use clustering for the screening evaluation, it accesses 2352 the lookup rule and processes 2353 the AI evaluation result according to the lookup rule to obtain the screening result and save it in the database. The process of obtaining the screening result and saving it in the database is repeated 2354 until all the images in the case have been screened and evaluated and a complete result report is obtained.
[0140] The VetImages application determines 2355 whether the species of the subject has been identified and saved in the database. If the species has not been determined, the VetAI application evaluates 2357 the species of the subject and sends the species evaluation result to the VetImages application. The task of species evaluation 2356 - 2357 is executed by the evaluation processor. In some embodiments, the VetImages application determines 2358 whether the species is a canine. For example, if the species is confirmed to be a canine, the case is labeled 2359 and the evaluation is attached to the result report. The VetImages application notifies 2360 that the V2 case evaluation has been completed. The V2 application determines 2361 whether the case has been labeled. If the report has been labeled, the result report is saved 2362 in the case document and the result report is sent 2363 by email to the client. If the report has not been labeled, the result report is sent 2363 by email to the client without saving the report to the case document.
[0141] Clustering
[0142] Clustering is a semi-supervised learning method. A semi-supervised learning method is a method for extracting references from a dataset consisting of input data without labeled responses. Generally, clustering is used to find the meaningful structures inherent in a set of examples, the underlying processes that explain them, generate features, and groupings.
[0143] Clustering is the task of partitioning a population or data points into multiple groups such that data points in the same group are similar to other data points in the same group and different from data points in other groups. Thus, clustering is a method of grouping objects based on the similarity and dissimilarity between objects.
[0144] Clustering is an important process because it determines the inherent grouping among the data. There is no standard for good clustering as it depends on the user to select the criteria that are useful for meeting the user's purpose. For example, clustering is based on finding representatives of homogeneous groups (data reduction), finding "natural clusters" and describing their unknown properties ("natural" data types), finding useful and appropriate groupings ("useful" data classes), or finding outlier data objects (outlier detection). The algorithm makes assumptions about what constitutes the similarity of points, and each assumption gives rise to different and equally valid clusterings.
[0145] Clustering methods:
[0146] There are many clustering methods, as follows:
[0147] Density-based methods: These methods consider clusters as dense regions with certain similarities, separated from regions of lower density in space. These methods have good accuracy and the ability to merge two clusters. Examples of density-based methods are Density-Based Spatial Clustering of Applications with Noise (DBSCAN), Ordering Points to Identify Clustering Structure (OPTICS), etc.
[0148] Hierarchical methods: The clusters formed by this method are based on a hierarchical formation of a tree structure. New clusters are formed by using previously formed clusters. Hierarchical methods are divided into two categories: agglomerative (bottom-up method) and divisive (top-down method). Examples of hierarchical methods include Clustering Using Representatives (CURE), Balanced Iterative Reducing Clustering Hierarchies (BIRCH), etc.
[0149] Partitioning method: The partitioning method divides objects into k clusters, and each partition forms a cluster. This method is used to optimize the similarity function of the target criterion. Examples of the partitioning method include: K-Means, Clustering Large Applications based upon Randomized Search (CLARANS), etc.
[0150] Grid-based method: In the grid-based method, the data space is represented as a finite number of cells that form a grid-like structure. All clustering operations performed on these grids are fast and independent of the number of data objects. Examples of the grid-based method are: Statistical Information Grid (STING), Wave Clustering, Clustering In Quest (CLIQUE), etc.
[0151] K - Means Clustering
[0152] K-Means clustering is one of the unsupervised machine learning algorithms. Generally, unsupervised algorithms only make inferences from the dataset using the input vectors without referring to known or labeled results. The goal of K-Means is simply to group similar data points together and discover potential patterns. To achieve this goal, K-Means looks for a fixed number (k) of clusters in the dataset.
[0153] Clustering refers to a set of data points that are grouped together due to some similarities. The target number k refers to the number of centroids that the user needs in the dataset. A centroid is a virtual or real position that represents the center of the cluster. By reducing the sum of squares within the cluster, each data point is assigned to each of the clusters. The K-Means algorithm identifies K centroids and then assigns each data point to the nearest cluster while keeping the centroids as small as possible.
[0154] The "means" in K-Means refers to finding the average of the data used to find the centroids. To process the learning data, the K-Means algorithm in data mining starts with a first set of randomly selected centroids, which are used as the starting point for each cluster, and then performs iterative (repeated) calculations to optimize the positions of the centroids. When the centroids have stabilized and their values have not changed due to successful clustering or reaching the defined number of iterations, the algorithm stops creating and optimizing the clusters.
[0155] The K - means clustering method follows a simple approach to classify a given data set into a predefined number of clusters (assume k clusters). Each cluster defines k centers or centroids. The next step is to take each point belonging to the given data set and associate it with the nearest center. When there are no pending points, the first step is completed, and an initial grouping is done. The next step is to recalculate the k new centroids as the barycenters of the clusters produced in the previous step. After calculating the k new centroids, a new binding is performed between the same data set points and the nearest new centers, thus generating a loop. Due to this loop, the k centers change their positions step by step until the k centers stop changing positions. The K - means clustering algorithm aims to minimize an objective function called the squared error function, which is calculated by the following formula:
[0156]
[0157] where,
[0158] "||x i -v j ||" is the Euclidean distance between x i and v j .
[0159] "c i " is the number of data points in the i - th cluster.
[0160] "c" is the number of cluster centers.
[0161] Algorithm Steps of K - Means Clustering
[0162] The K - means clustering algorithm is as follows:
[0163] In K - means clustering, "c" cluster centers are randomly selected, the distance between each data point and the cluster centers is calculated, the data points are assigned to the nearest cluster centers, and the new cluster centers are recalculated using the following formula:
[0164]
[0165] where, "c i " is the number of data points in the i - th cluster,
[0166] X is the data point set {x 1 , x 2 , x 3 , ……, x n}, and
[0167] V is the center set {v 1 , v 2 , ……, v c}.
[0168] Measure the distance between each data point and the newly obtained cluster center. If a data point is reallocated, the process continues until no data points are reallocated.
[0169] AI Application Processing Flow
[0170] In some embodiments, Digital Imaging and Communications in Medicine (DICOM) images are submitted via LION and transferred to the V2 platform via Hypertext Transfer Protocol Secure (HTTPS) by the DICOM toolkit library (Offis.de DCMTK library). The DICOM images are temporarily stored in the V2 platform. In some embodiments, a DICOM record is created with limited information and the status is set to zero. In various embodiments, once the DICOM images are available in the temporary memory, the V2 PHP / Laravel application starts processing the DICOM images via a scheduled task job (Cron job).
[0171] In some embodiments, the scheduled task job (1) monitors V2 for new DICOM images, retrieves the DICOM images from temporary storage, extracts the tags, extracts the frames (single sub-image or multiple sub-images), saves the images and tags in a data store, and sets the processing status in the database. In some embodiments, the scheduled task job (1) converts and compresses the DICOM images into lossless JPG format using the Offis.de DCMTK library and sets the processing status to 1. In some embodiments, the scheduled task job (1) runs automatically every few minutes, such as every five minutes, every four minutes, every three minutes, every two minutes, or every minute. In some embodiments, the scheduled task job (1) saves the DICOM image metadata into a table named DICOM in a Microsoft SQL server, extracts the image / frame, and stores the image / frame in the directory of an image manager. In various embodiments, records are created during the processing, which contain additional information about the image and the case identifier (ID) associated with the image. These records contain other data, such as the physical examination results of the subject, the studies of the subject's multiple visits, the series of images obtained during each examination, and the hierarchy of the images within the case.
[0172] In various embodiments, DICOM images and metadata are processed by the VetImages application written in the PHP Laravel framework. V2 sends Representational State Transfer (REST) service requests to VetImages to process each image asynchronously. In some embodiments, VetImages immediately responds to V2 to confirm that the request has been received, and the cropping and evaluation process of the image will continue in the background. Since the images are processed in parallel, the entire process is executed at high speed.
[0173] In various embodiments, VetImages passes or transfers the images to a module called VetConsole, which is written in Python and uses the computer vision technology OpenCV to preprocess the images. VetConsole identifies body regions in the images, such as the chest, abdomen, pelvis, as a reserve in case the AI cropping server fails to classify the body regions in the images. VetImages rotates and flips the images until the correct orientation is achieved. In some embodiments, VetImages uses an image matching server to verify different angles and projections of the images. In various embodiments, the image matching server is written in Python and Elastic Search for identifying image matches. In some embodiments, the image database of the image matching server is carefully selected, and results are only returned when the image is accepted as being in the correct orientation and projection.
[0174] In various embodiments, after determining the orientation of the image, VetImages sends a REST Application Programming Interface (API) request to the Keras / TensorFlow server to classify and determine the regions of interest in the body regions of the image. The VetImages REST API request is verified using the image matching server to confirm that the returned regions of the image are classified as body regions, such as the chest, abdomen, pelvis, hind knee joint, etc. In some embodiments, if the evaluation result of the cropping is invalid, the VetConsole cropping result is verified and used.
[0175] In various embodiments, if VetImages determines that the image contains classified and verified body regions, an AI evaluation process for generating an AI report is initiated. In alternative embodiments, if VetImages determines that the image does not contain classified and verified body regions, the image cropping process ends without results and no report is generated.
[0176] In various embodiments, VetImages makes a REST service call to Keras / TensorFlow and sends the classified cropped images to a disease AI evaluation model hosted on a TensorFlow application server written in Python / Django. VetImages saves the results of the AI evaluation model for final evaluation and report generation.
[0177] VetImages also directs the chest cropped images to the TensorFlow server to determine if the images are well-positioned with respect to user-set parameters. VetImages sends the results of the AI evaluation model to the V2 platform to notify the clinic of the positioning evaluation results for each image.
[0178] In some embodiments, VetImages waits while processing the images in parallel until the TensorFlow server has cropped and evaluated all the images of the case. In some embodiments, after evaluating all the images of the case, VetImages processes all the results of the case using expert-defined rules to determine the content of the report in a more readable manner. In alternative embodiments, VetImages uses a user-created clustering model to determine the content of the AI report. In some embodiments, the AI report is compiled using radiologist reports previously used to build the clustering model in VetImages. In some embodiments, clustering is used to classify cases / images using the prediction results from other diagnostic models using scikit-learn.
[0179] In some embodiments, after determining the content of the AI report using expert rules or a clustering model, VetImages checks the species of the case. In some embodiments, a report is generated and sent to the V2 clinic only if VetImages determines that the species of the case is canine.
[0180] In some embodiments, VetImages sends a request to the V2 platform to notify the clinic that a new AI report has been sent to the clinic. In some embodiments, the V2 platform verifies the license of the clinic administrator user. In various embodiments, V2 attaches a copy of the report within the case file, or accompanies a copy of the report with the case file, such that the report can be accessed from the V1 platform if the clinic has a valid license. In some embodiments, V2 sends an email notification to the clinic email address that includes one or more links so that the email recipient can conveniently open the generated report immediately.
[0181] Additive AI Model
[0182] The practical standard for AI model retraining is to improve the performance of the current model by adding additional data, retraining, testing, and then replacing the current model with the updated AI model( Figure 24 ). According to the current standard, it is good practice to retain the old model for a specific window or period, or to retain the old model until the model meets a specific number of requirements. Provide the same data to the new model, analyze the results, and compare the results with the old model. Thus, the model with better performance is determined. If the performance of the new model is satisfactory, the new model is deployed. The current standard is a typical example of A / B testing, which ensures that the model is validated on upstream data. Another best practice is to automate the deployment of the retrained model. Kubernetes (K8s) is used to deploy machine learning models to the production environment. Kubernetes is an open-source system for automating the deployment, scaling, and management of containerized applications.
[0183] In the current method, it is assumed that the performance of the newly trained classifier is improved by reducing unwanted "noise" or false positive data. The assumption that the data detected by the current classifier is marked as unwanted "noise" and removed by the retrained classifier is caused by the wrong assumption that "noise" is not valuable information in the overall system performance.
[0184] Contrary to the above wrong assumption, the "noise" detected by the current classifier helps to distinguish data points with similar but not identical features. Compared with using any one of the AI classifiers alone, the AI classifier results after combined analysis of the current classifier and the new classifier can better identify and classify data.
[0185] The "noise" detected by the current model helps to distinguish data points with similar but not exactly the same features. Therefore, valuable data may be lost by replacing the current AI model. The learning concept in the human brain is based on knowledge and iterative learning steps. On the path of iterative learning, prior knowledge is not lost or forgotten, but is built upon and adjusted according to need, adding to the overall library of a person's knowledge and capabilities. Therefore, the concept of building based on prior knowledge is applied to the iteration of daisy chaining or related model deployment.
[0186] Embodiments of the method described herein allow for continuous training by iterative or related model deployment or by daisy chaining iterative or related models, rather than the current model retraining and replacement standard( Figure 25 ). These methods improve AI reporting performance and internal AI auditing of the system, thus moving towards artificial general intelligence (AGI) even if a fully operational general AI system has not been created.
[0187] For example, an AI classifier built to detect people with blonde hair can be very broad and detect anyone with any percentage of blonde hair. A retrained AI classifier can be more specific and only detect people with 80% or more blonde hair. According to current best practices in the industry, the AI classifier is removed from service and replaced with the retrained classifier.
[0188] Scenario 3 in Table 1 of this article shows that if the results of two AI classifiers are compared with the results of a single newly trained AI classifier, it is easy to detect incorrect results. In Scenario 3, the old classifier is more sensitive but less specific compared to the new classifier. Therefore, cases where the old classifier evaluates blonde hair as positive are also evaluated as positive by the new classifier. However, a negative result from the old or sensitive classifier and a positive result from the new or specific classifier is an incorrect result and will be immediately detected and become apparent in a multi-classifier system. Therefore, the method provided in this article includes rules that can be implemented within the system such that the system will continue to operate different versions of the same classifier to more quickly detect abnormal results.
[0189] Table 1: Potential Scenarios of Single Classifier vs. Double Classifier
[0190]
[0191] The method described in this article shows the additive implementation of a derived model into a production system. In the additive model approach, the user learns that in production, the first model has some limitations and the model is being retrained to improve the deficiencies. As subsequent AI models are built, the original or first iteration model is retained, and the results produced by the first and subsequent AI models are included in an AI cluster analysis as described in U.S. Patent Application Publication No. US2021 / 0202092A1, the entire content of which is incorporated herein by reference. The AI data results of the old and new models are added to a separate line in the results, and an AI fingerprint is created for the item being evaluated. Then, multiple AI fingerprints are clustered together to group similar results. This method of adding related or derived AI classifiers is defined as daisy-chaining related or iterative, or unrelated or non-iterative AI classifiers with accompanying words, phrases, statements, and paragraphs, and non-AI derived data ( Figure 26 and Figure 27 ).
[0192] Figure 28Shows visual examples of derived data (same letters and subscript numbers) and non-derived (same letters but different subscript letters) data. The database stores all the training data with relationships. The system allows data input (in this case I, T, and OI); individual or clustered, AI or non-AI data, derived or non-derived data, to obtain example results from the database. In Figure 28 it, the method is broken down into components A, B, and C, as shown by the bounding boxes.
[0193] The components in box A show visual examples that show how to group clustered data using derived or non-derived results (training data) and associate it with the training system (i.e., add information to the database), and as an output step, input the clustered derived groups of images (I) or text (T) into the system and bring back the corresponding results as the best example results from the database. In this case, the derived data is shown as the input in the images (I) and text (T).
[0194] The components in box B show visual examples that show the clustered input, derived data (input data), and the corresponding results (output extracted from the database) in a stacked visual arrangement.
[0195] The components in box C show visual examples that show the clustered input (input) of derived data and the corresponding results (output) in a horizontal visual arrangement. Only the derived classifier is not a prerequisite for the clustered input or output. In this example, OI in the output has the subscript letter OI A and OI B , indicating non-derived data. Boxes of different sizes represent different weights of the results. Additionally, the data represented by I, T, and OI should be considered building blocks that, when combined, form instructions for building the output. The output can be pixels, sets of pixels, images or sets of images grouped together, colors or specific color arrangements, words or several words, partial statements or complete statements, several statements together creating partial paragraphs, complete paragraphs or several paragraphs or complete templates, articles, papers, or other text-based outputs, measurement types, numbers, waveforms, equations, recipes, atomic structures.
[0196] Like DNA primers, AI and non-AI data, i.e., data that links derived and non-derived data together through clustering to form a unique fingerprint, are groupings of results that, when applied to complementary databases and specific user-defined presets, can serve as instructions for specific outputs. The novel concept here is to use structured, linear-format AI and non-AI derived input data as one or more "primer" instructions to generate inferential results from a large database of previously stored "primer" results and to build a computerized system around the data fingerprint as a transcription mechanism to create the output. This fingerprint or clustering "primer" method allows for an infinite number of possibilities for output results, thus creating an extremely robust system( Figure 28 ).
[0197] Additional sanity checks are performed on the results by daisy-chaining related or iterative classifiers into the AI evaluation( Figure 29A and Figure 29B ). The method described herein shows the additive implementation of a derived model into a production system. In the additive model approach, the user is aware that the first model has some limitations in production and is retraining the model to improve the deficiencies. As subsequent AI models are built, the original or first iteration model is retained, and the results generated by the first and subsequent AI models are included in the AI cluster analysis as described in U.S. Patent Application Publication No. US2021 / 0202092A1, the entire content of which is incorporated herein by reference. The AI data results of the old and new models are added to a separate line in the results, and an AI fingerprint is created for the item being evaluated. Multiple AI fingerprints are then clustered together to group similar results. This method of adding related or derived AI classifiers is defined as daisy-chaining related or iterative AI classifiers.
[0198] Additional sanity checks are performed on the results by daisy-chaining related or iterative classifiers into the AI evaluation. The sanity check allows the original classifier to cross-reference the results of subsequent classifiers. In addition, if the original or broader classifier has a negative result for a case while the specific or newly trained classifier has a positive result, the sanity check can also identify the AI evaluation newly introduced by the more specific classifier.
[0199] In Figure 29A and Figure 29BIn the example, the General Lung Classifier (GLC) is used together with a derived classifier (GLCn, or GLC2 in this specific example) as a confirmation of each other. GLC2 is a more specific bronchial pattern classifier. The GLC results are on the left and the GLC2 results are on the right. The general and specific classifier results are used as a mutual internal check on the findings. In these examples, it is expected that the classifier results will track each other because Example 8a has no abnormal findings, while Example 8b has an abnormal finding that is detected by the GLC and its derived classifier GLC2. Depending on the specific image search and classifier used, acceptable combinations for GLC and GLC2 are both negative for the finding (<0.5), both positive for the finding (>0.5), or GLC > 0.5 and GLC2 < 0.5. Any result with GLC < 0.5 and GLC2 > 0.5 is considered abnormal and a system error.
[0200] The methods of machine learning, clustering, and programming are comprehensively described in the following references: Shaw, Zed. Learn Python the Hard Way: A Very Simple Introduction to the Terrifyingly Beautiful World of Computers and Code. Addison-Wesley, 2017 (by Zed Shaw, "Learn Python the Hard Way: A Simple Introduction to the Beautiful World of Computers and Code", Addison-Wesley Publishing, 2017); Ramalho, Luciano. Fluent Python. O'Reilly, 2016 (by Luciano Ramalho, "Fluent Python", O'Reilly Media, 2016); Atienza, Rowel. Advanced Deep Learning with TensorFlow 2 and Keras: Apply DL, GANs, VAEs, Deep RL, Unsupervised Learning, Object Detection and Segmentation, and More. Packt, 2020 (by Rowel Atienza, "Advanced Deep Learning with TensorFlow 2 and Keras: Implement Deep Learning, Generative Adversarial Networks, Variational Autoencoders, Deep Reinforcement Learning, Unsupervised Learning, Object Detection and Segmentation, and More", Packt Publishing, 2020); Vincent, William S. Django for Professionals: Production Websites with Python & Django. StillRiver Press, 2020 (by William S. Vincent, "Django for Professionals: Building Websites with Python and Django", StillRiver Press, 2020); Bradski, Gary R., and Adrian Kaehler. Learning OpenCV. O'Reilly, 2011 (by Gary Bradski and Adrian Kaehler, "Learning OpenCV", O'Reilly Media, 2011); Battiti, Roberto, and Mauro Brunato. The LION Way: Machine Learning plus Intelligent Optimization, Version 2.0, April 2015. LIONlab Trento University, 2015 (by Roberto Battiti and Mauro Brunato, "The LION Way: Machine Learning and Intelligent Optimization" (2.Version 0, April 2015), LION Laboratory, University of Trento, 2015); Pianykh, Oleg S. Digital Imaging and Communications in Medicine (DICOM) a Practical Introduction and Survival Guide. Springer Berlin Heidelberg, 2012 (by Oleg S. Pianykh, "DICOM: Digital Imaging and Communications in Medicine - A Practical Introduction and Survival Guide", Springer Berlin Heidelberg, 2012); Busuioc, Alexandru. The PHP Workshop: a New, Interactive Approach to Learning PHP. Packt Publishing, Limited, 2019 (by Alexandru Busuioc, "The PHP Workshop: A New, Interactive Approach to Learning PHP", Packt Publishing, Limited, 2019); Stauffer, Matt. Laravel - Up and Running: a Framework for Building Modern PHP Apps. O'Reilly Media, Incorporated, 2019 (by Matt Stauffer, "Laravel Up and Running: A Framework for Building Modern PHP Apps", O'Reilly Media, Incorporated, 2019); Kassambara, Alboukadel. Practical Guide to Cluster Analysis in R: Unsupervised Machine Learning. STHDA, 2017 (by Alboukadel Kassambara, "Practical Guide to Cluster Analysis in R: Unsupervised Machine Learning", STHDA, 2017); and Wu, Junjie. Advances in K - Means Clustering: A Data Mining Thinking. Springer, 2012 (by Junjie Wu, "Advances in K - Means Clustering: A Data Mining Perspective", Springer, 2012). Each of these references is hereby incorporated by reference in its entirety.
[0201] It should be understood that any feature associated with any embodiment provided herein can be used alone, or in combination with other features described, or in combination with one or more features of any other embodiment or any combination of any other embodiment. Additionally, equivalents and modifications not described above can be employed without departing from the scope of the present invention, which is defined by the appended claims.
[0202] The present invention has now been described in full, and the following claims further illustrate the invention. Those skilled in the art will recognize or be able to ascertain using no more than routine experimentation many equivalents to the specific methods described herein. These equivalents are within the scope of the invention and the claims. The contents of all references cited in this application, including issued patents and published patent applications, are hereby incorporated by reference in their entirety.
Claims
1. A method for obtaining an additive AI result from a digital file, the method comprising: processing the digital file through at least one first artificial intelligence (AI) classifier and at least one second AI classifier to respectively obtain a first evaluation result and at least one second evaluation result; directing the first evaluation result and the at least one second evaluation result to at least one synthesis processor; creating clusters of AI classifier results and non-AI classifier results in a database; and using at least one of the first evaluation result and the second evaluation result as an instruction to create an output by comparing with at least one data cluster in the database, thereby obtaining an inferential additive AI output result.
2. The method according to claim 1, further comprising measuring the distance from the additive AI result to an example result from the data cluster to obtain an additive AI cluster identifier.
3. The method according to claim 1, wherein the data cluster further comprises a matching written template.
4. The method according to claim 3, further comprising compiling the additive AI cluster identifier and the matching written template to obtain a report.
5. The method according to claim 4, further comprising displaying the report to a user.
6. The method according to claim 1, wherein the at least one second AI classifier is a derivative of the first AI classifier.
7. The method according to claim 6, using at least a portion of the data used to train the first AI classifier to train the second AI classifier.
8. The method according to claim 6, wherein the second AI classifier is related to the first AI classifier.
9. The method according to claim 1, wherein the first AI classifier is a comprehensive classifier.
10. The method according to claim 1, wherein the second AI classifier is a specific classifier.
11. The method according to claim 1, wherein the first AI classifier is a specific classifier.
12. The method according to claim 1, wherein the second AI classifier is a comprehensive classifier.
13. The method according to claim 1, further comprising repeating the steps of directing and comparing for a series of daisy-chained related or derivative AI classifiers.
14. The method according to claim 1, further comprising comparing the first evaluation result with the second evaluation result to compare AI results and test expected performance.
15. The method according to claim 1, further comprising adding the first evaluation result and the second evaluation result to a result database.
16. The method according to claim 1, further comprising using one or more derivative classifiers of the at least one first AI classifier and a first evaluation classifier to create database cluster entries.
17. The method according to claim 16, further comprising using the at least one first AI evaluation classifier and one or more derivative classifiers of the first evaluation classifier to evaluate additional data inputs for comparison with the database cluster entries and return example results from the database.
18. The method according to claim 1 further includes a system or user input that accepts at least one form of data input selected from the following for analysis and storage in a database or for analysis and creation of an output report: pixels, pixel sets, images, sets of grouped images, colors, color arrangements, words, groups of words, partial statements, complete statements, partial paragraphs created from multiple statements together, complete paragraphs, multiple paragraphs, complete templates, articles, papers, or other text-based outputs, measurements, numbers, waveforms, equations, formulas, and mixed data.
19. The method according to claim 1 further includes obtaining the digital file before processing.
20. The method according to claim 1 further includes converting an analog file to the digital file before processing.
21. The method according to claim 1 further includes classifying the digital file by performing at least one of tagging, cropping, editing, and orienting on the digital file before processing.
22. The method according to claim 1 includes clustering AI-derived and non-AI-derived data.
23. The method according to claim 1 further includes applying standard mathematical formulas, rearranging, or weighting the AI results.
24. The method according to claim 1 further includes obtaining and summarizing the evaluation results within the database before utilization.
25. A system programmed to obtain additive AI results by any of the methods according to claim 1, the system comprising: at least one first AI processor; at least one derived AI processor derived from the first AI processor; and an output device.
26. The system according to claim 25 further includes at least one database.
27. The system according to claim 25 further includes a user interface.
Citation Information
Patent Citations
Efficient artificial intelligence analysis of radiographic images with combined predictive modeling
US20210202092A1