System and method for evaluating pet radiological images
By using machine learning models RapidReadNet and AdjustNet, the problems of veterinarians lacking radiology training and difficulty in orienting pet radiology images have been solved, enabling automated processing and accurate classification of pet radiology images, thus improving diagnostic efficiency and accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- MARS INC
- Filing Date
- 2021-12-15
- Publication Date
- 2026-05-19
AI Technical Summary
Veterinarians lack radiology training, making it difficult to effectively utilize image-based diagnostic techniques. Furthermore, pet radiographic images may be misoriented and/or lack lateral markings, leading to diagnostic difficulties.
Employing machine learning models such as RapidReadNet and AdjustNet, this method automates the processing and interpretation of pet radiographic images through automated natural language processing and image classification. This includes image orientation and classification. The machine learning classifier model is trained using mixed training data, and combined with labels generated by natural language processing and human markers, image analysis is performed using convolutional neural networks.
It improves the efficiency of veterinary diagnosis, reduces misdirection and labeling problems, provides reliable clinical results, reduces reliance on radiology specialists, and improves the accuracy and efficiency of image interpretation.
Smart Images

Figure CN116801795B_ABST
Abstract
Description
[0001] Claims of rights
[0002] This application claims the benefit of Provisional Application 63 / 274,482, filed November 1, 2021; Provisional Application 63 / 215,769, filed June 28, 2021; and Provisional Application 63 / 125,912, filed December 15, 2020, pursuant to 35 USC §119, the entire contents of which are incorporated herein by reference for all purposes, as fully set forth herein. Technical Field
[0003] This disclosure generally relates to the use of one or more machine learning models or tools to evaluate radiographic images of pets or animals. Background Technology
[0004] More and more veterinarians are utilizing image-based diagnostic techniques (such as X-rays) to diagnose or identify health problems in animals or pets. However, there are fewer than 1,100 veterinary-trained radiologists worldwide. Therefore, many veterinarians are unable to take advantage of the benefits offered by image-based diagnostic techniques. Even for those veterinarians trained in radiology, reviewing medical images can be time-consuming and cumbersome. Adding to these difficulties, animal or pet radiographic images may be misoriented and / or missing or have incorrect laterality markers. Therefore, a system is needed that can automatically process and interpret pet diagnostic images and return clinically reliable results to veterinarians, whether radiology-trained or not. Summary of the Invention
[0005] In some non-limiting embodiments, this disclosure provides systems and methods for training and using machine learning models to process, interpret, and / or analyze radiographic digital images of animals or pets. Images can be any digital image storage format used for medical condition diagnosis, such as Medical Digital Imaging and Communications (“DICOM”), and other formats for displaying images. In a particular embodiment, radiographic images can be labeled using automated natural language processing (“NLP”) tools: in a computer-implemented method, the NLP tool takes a natural language text summary representation of the radiographic image as input and outputs image labels or tags characterizing the radiographic image. In one embodiment, the natural language text summary of the radiographic image is a radiological report. In one embodiment, the radiographic image and the corresponding NLP-generated labels can be used as training data to train one or more machine learning classifier models configured or programmed to classify radiographic images of animals or pets. In other non-limiting embodiments, veterinary radiologists can manually label various images. In a particular embodiment, one or more machine learning classifier models can be trained using manually labeled training data, such as medical images labeled by veterinary radiologists or another type of human domain-specific expert. As further explained herein, in some embodiments the implemented machine learning model can effectively use a mixture of training data with NLP-generated labeled data and human-generated labeled image training data.
[0006] In one embodiment, this disclosure provides a system and method for the automatic classification of radiographic images of animals or pets. In various embodiments, one or more machine learning models or tools may be used to analyze and / or classify acquired, collected, and / or received images. In some embodiments, the machine learning model may include a neural network, which may be a convolutional neural network (“CNN”). The machine learning model can be used to classify images using various labels, tags, or categories. For example, such classification may indicate healthy tissue or the presence of abnormalities. In one embodiment, an image classified as having an abnormality may be further classified, for example, cardiovascular, pulmonary structures, mediastinal structures, pleural cavity, and / or extrathoracic. In this disclosure, such classification in the classification may be referred to as a subclassification.
[0007] In one embodiment, this disclosure provides techniques for training and using a programmed machine learning model (in some cases, referred to herein as "RapidReadNet") to classify pet radiographic images, wherein RapidReadNet may be an ensemble of separate, calibrated deep neural network Student models, as described further in more detail herein. The term RapidReadNet, and every other similar term or label used in this disclosure, is merely for convenience and brevity to facilitate a concise explanation; other embodiments may implement functionally equivalent tools, systems, or methods without using the term RapidReadNet. In one embodiment, a machine learning neural Teacher model may be trained first using a first human-labeled image training dataset. A larger unlabeled image training dataset containing medical images associated with natural language text summaries may then be labeled using an NLP model. For example, the dataset may contain radiographic reports. Soft pseudo-labels may then be generated on the larger image dataset using the Teacher model. Finally, the soft pseudo-labels may be combined with NLP derived labels to further generate more derived labels, and these derived labels may be used to train one or more machine learning neural Student models. In one embodiment, RapidReadNet may include an ensemble of said Student models.
[0008] In one embodiment, this disclosure provides a system and method for automatically determining the correct anatomical orientation in veterinary radiographic images without relying on DICOM metadata or lateral markers. One disclosed method may include using a trained machine learning model (“AdjustNet”) comprising two sub-models (“RotationNet” and “FlipNet”). In some embodiments, each of RotationNet and FlipNet may be programmed as an ensemble of multiple CNNs. In one embodiment, the RotationNet model may be used to determine whether an image (e.g., an animal or pet radiographic image) is correctly rotated. In one embodiment, the FlipNet model may be used to determine whether an image (e.g., an animal or pet radiographic image) should be flipped. In one embodiment, AdjustNet and / or RotationNet and / or FlipNet may be incorporated into an end-to-end system or pipeline for classifying animal or pet radiographic images, which offers numerous technical advantages compared to reported state-of-the-art systems. The terms AdjustNet, RotationNet, and FlipNet, as well as each other similar term or label in this disclosure, are used merely for convenience and brevity to facilitate a concise explanation; other embodiments may implement functionally equivalent tools, systems, or methods without using the terms AdjustNet, RotationNet, or FlipNet.
[0009] In various embodiments, each of RotationNet and FlipNet can be programmed and / or trained in any of a variety of ways. For example, each model can be a single model or a two-stage model. In non-limiting embodiments, a variety of different weight initialization techniques can be used to develop the model. For example, transfer learning methods can be performed using pre-trained model weights (e.g., on ImageNet), or the model can be randomly initialized and then further pre-trained on augmented data. In non-limiting embodiments, one or more different training pipelines can also be used to develop the model. For example, each of RotationNet and FlipNet can be pre-trained using augmented data and then fine-tuned using real data. In other embodiments, one or both models can be jointly trained using augmented and real data.
[0010] In one embodiment, this disclosure provides an end-to-end system or pipeline for classifying radiographic images of animals or pets. In this context, "end-to-end" can refer to a system or pipeline configured to receive digital image data as input and output classification data or labels. As explained further in more detail herein, the end-to-end system or pipeline may include using AdjustNet to determine the correct anatomical orientation of a target image and using RapidReadNet to classify the target image. In one embodiment, after determining the correct anatomical orientation of the target image (using AdjustNet) and before outputting one or more classifications of the target image (using RapidReadNet), another trained model may be used to verify that the target image corresponds to the correct body part. In one embodiment, the infrastructure pipeline may rely on microservices that can be deployed using software containers, using libraries such as DOCKER from DOCKER Corporation or KUBERNETES from Google, and can be invoked via an application programming interface using a representational state transition (ReSTful API). In one embodiment, an AI Orchestrator container may be programmed to coordinate the execution of inference from different AI models, such as the AdjustNet and RapidReadNet models. This disclosure provides an exemplary novel non-limiting system architecture on which some of the methods or techniques provided in this disclosure may be implemented, but other methods or techniques are also possible.
[0011] The embodiments disclosed herein are merely examples, and the scope of this disclosure is not limited thereto. Certain non-limiting embodiments may include all, some, or none of the components, elements, features, functions, operations, or steps of the embodiments disclosed herein. Embodiments according to the invention are specifically disclosed in the appended claims relating to the method. Dependencies or references in the appended claims are chosen solely for formal reasons. However, claims may also be made regarding any subject matter arising from intentional reference to any prior claim (particularly multiple dependencies), and thus, regardless of the dependency chosen in the appended claims, any combination of the claims and their features may be disclosed and claimed. Claimed subject matter may include not only combinations of features listed in the appended claims but also any other combination of features in the claims, wherein each feature mentioned in a claim may be combined with any other feature or combination of features in the claims. Furthermore, any embodiments and features described or depicted herein may be claimed in individual claims, and / or in any combination with any embodiments or features described or depicted herein, or with any features of the appended claims. Attached Figure Description
[0012] In the attached diagram:
[0013] Figure 1 Radiographic images before and after processing by one or more machine learning models or tools are shown according to certain non-limiting embodiments.
[0014] Figure 2 Labels and annotations are shown for images according to certain non-limiting embodiments.
[0015] Figure 3 An exemplary method for using a machine learning system to evaluate images of animals and / or pets is shown.
[0016] Figure 4 An exemplary computer system or device is shown for facilitating the classification and labeling of images of animals or pets.
[0017] Figure 5 An exemplary workflow for an image orientation task is shown.
[0018] Figure 6 Radiographic images showing examples of incorrectly oriented and correctly oriented images of a feline skull are presented.
[0019] Figure 7 A schematic diagram of an exemplary two-stage model technique with integration, according to certain non-limiting embodiments, is shown.
[0020] Figure 8 An exemplary image of the GradCAM transform used for model decisions regarding image orientation is shown.
[0021] Figure 9 A workflow diagram for an exemplary model deployment is shown.
[0022] Figure 10 Exemplary images for error analysis are shown according to certain non-limiting embodiments.
[0023] Figure 11 An image pool being evaluated for modeling is shown in one embodiment.
[0024] Figure 12 An example of the infrastructure of an X-ray system is shown, in which images can be acquired as part of a clinical workflow.
[0025] Figure 13A This shows the first set of ROC and PR curves found in cardiovascular and pleural cavity studies.
[0026] Figure 13B The second set of ROC and PR curves from cardiovascular and pleural cavity studies is shown.
[0027] Figure 14A This shows the first set of ROC and PR curves found in lung research.
[0028] Figure 14B The second set of ROC and PR curves from lung research is shown.
[0029] Figure 15A The first set of ROC and PR curves found in the mediastinal study are shown.
[0030] Figure 15B The second set of ROC and PR curves found in the mediastinal study is shown.
[0031] Figure 16A The first set of ROC and PR curves found in extrathoracic studies is shown.
[0032] Figure 16B The second set of ROC and PR curves found in extrathoracic studies is shown.
[0033] Figure 17A The third set of ROC and PR curves found in extrathoracic studies is shown.
[0034] Figure 17B The fourth set of ROC and PR curves from extrathoracic studies is shown.
[0035] Figure 18 A visualization of the reconstruction error calculated on a weekly basis is shown.
[0036] Figure 19 The distribution of reconstruction error as a function of the number of tissues represented is shown.
[0037] Figure 20 An exemplary computer implementation or programming method is shown for classifying radiographic images (e.g., radiographic images of animals and / or pets) using machine learning neural models. Detailed Implementation
[0038] The terms used in this specification generally have their ordinary meaning in the art, both in the context of this disclosure and in the specific context in which each term is used. Certain terms are discussed below or elsewhere in the specification to provide additional guidance in describing the compositions and methods of this disclosure and how they are made and used.
[0039] The embodiments are disclosed in sections according to the following outline:
[0040] 1.0 Overview
[0041] 2.0 Machine Learning Techniques for Processing Pet Radiographic Images
[0042] 2.1 Exemplary pet radiographic images used for classification
[0043] 2.2 Labeling of pet radiographic images
[0044] 2.3 Classification of pet radiographic images in one embodiment
[0045] 3.0 AdjustNet: An automated technique for orienting radiographic images in pets.
[0046] 3.1 Input Data and Workflow
[0047] 3.2 Model Development
[0048] 3.3 Model Deployment
[0049] 3.4 User Feedback
[0050] 4.0 End-to-end pet radiology image processing using RapidReadNet
[0051] 4.1 Image dataset used for training RapidReadNet in one embodiment
[0052] 4.2 Neural Model Training Techniques for Image Classification Tasks
[0053] 4.3 Drift Analysis, Experimental Results and Longitudinal Drift Analysis
[0054] 4.4 System Architecture and Method of RapidReadNet in One Implementation
[0055] 5.0 Advantages of certain embodiments
[0056] 5.1 Exemplary Technical Advantages of AdjustNet and RotationNet
[0057] 5.2 Exemplary Technical Advantages of RapidReadNet and the Disclosed End-to-End System for Classifying Pet Radiographic Images in One Embodiment
[0058] 6.0 Implementation Example – Hardware Overview
[0059] ***
[0060] 1.0 Overview
[0061] As used in the specification and appended claims, unless the context clearly specifies otherwise, the singular forms “a,” “an,” and “the” include the plural reference.
[0062] As used herein, the terms “comprises,” “comprising,” or any other variation thereof are intended to cover non-exclusive inclusion, such that a process, method, article, system, or apparatus that comprises a list of elements may include not only those elements but also other elements not expressly listed or inherent to such process, method, article, or apparatus.
[0063] The terms "animal" or "pet" as used in this disclosure refer to domesticated animals, including but not limited to domestic dogs, domestic cats, horses, cattle, ferrets, rabbits, pigs, rats, mice, gerbils, hamsters, goats, etc. Domestic dogs and domestic cats are specific, non-limiting examples of pets. The terms "animal" or "pet" as used in this disclosure may also refer to wild animals, including but not limited to bison, elk, deer, wild deer, ducks, birds, fish, etc.
[0064] As used herein, a “feature” of an image or slice can be determined based on one or more measurable features of that image or slice. For example, a feature could be a blemish in the image, a dark spot, or tissue with various sizes, shapes, or light intensity levels.
[0065] In the detailed description herein, references to "embodiment," "an embodiment," "one embodiment," "in various embodiments," "some embodiments," "a few embodiments," "other embodiments," "some other embodiments," etc., indicate that the embodiment may include a specific feature, structure, or characteristic, but each embodiment may not necessarily include that specific feature, structure, or characteristic. Furthermore, such phrases do not necessarily refer to the same embodiment. Moreover, when a specific feature, structure, or characteristic is described in connection with an embodiment, whether explicitly described or not, it is believed that its influence on such feature, structure, or characteristic in conjunction with other embodiments is within the knowledge of those skilled in the art. After reading this specification, it will be readily apparent to those skilled in the art how this disclosure can be implemented in alternative embodiments.
[0066] As used herein, the term "device" refers to a computing system or mobile device. For example, the term "device" can include a smartphone, tablet, or laptop. Specifically, a computing system may include functions for determining its location, orientation, or orientation, such as a GPS receiver, compass, gyroscope, or accelerometer. Client devices may further include functions for wireless communication, such as Bluetooth communication, near field communication (NFC), or infrared (IR) communication, or communication with a wireless local area network (WLAN) or cellular phone network. Such devices may also include one or more cameras, scanners, touchscreens, microphones, or speakers. Client devices may also execute software applications, such as games, web browsers, or social networking applications. Client devices may include, for example, user devices, smartphones, tablets, laptops, desktop computers, or smartwatches.
[0067] Exemplary processes and embodiments can be operated or performed by a computing system or client device through a mobile application and an associated graphical user interface (“UX” or “GUI”). In some non-limiting embodiments, the computing system or client device may be, for example, a mobile computing system—such as a smartphone, tablet, or laptop. The mobile computing system may include functions for determining its location, orientation, or orientation, such as a GPS receiver, compass, gyroscope, or accelerometer. Such a device may also include functions for wireless communication, such as Bluetooth communication, near-field communication (NFC), or infrared (IR) communication, or communication with wireless local area networks (WLAN), 3G, 4G, LTE, LTE-A, 5G, Internet of Things (IoT), or cellular phone networks. Such a device may also include one or more cameras, scanners, touchscreens, microphones, or speakers. The mobile computing system may also execute software applications, such as games, web browsers, or social networking applications. Through social networking applications, users can connect, communicate, and share information with other users in their social network.
[0068] The terms used in this specification generally have their ordinary meaning in the art, both in the context of this disclosure and in the specific context in which they are used. Certain terms are discussed below or elsewhere in the specification to provide additional guidance in describing the compositions and methods of this disclosure and how they are made and used.
[0069] In one embodiment, this disclosure provides a computer-implemented method comprising: receiving a first labeled training dataset comprising a first plurality of images, each of the first plurality of images being associated with a set of labels; programmatically training a machine learning neural Teacher model on the first labeled training dataset; programmatically applying the machine learning model trained for natural language processing to an unlabeled dataset comprising digital electronic representations of natural language text summaries of a second plurality of images, thereby generating a second labeled training dataset comprising the second plurality of images; programmatically generating a corresponding set of soft pseudo-labels for each of the second plurality of images using the machine learning neural Teacher model; programmatically generating a set of derived labels for each image in the second labeled training dataset using the soft pseudo-labels; training one or more programmatic machine learning neural Student models using the derived labels; receiving a target image; and applying an ensemble of one or more Student models to output one or more classifications of the target image.
[0070] One embodiment further includes using active learning to programmatically update at least one of one or more programmed machine learning neural Student models.
[0071] One embodiment further includes applying noise in one or more machine learning model training steps.
[0072] In one embodiment, the target image is a radiographic image of an animal or pet.
[0073] In one embodiment, each of the first plurality of images and each of the second plurality of images is a radiographic image of an animal or pet.
[0074] In one embodiment, the natural language text summary is a radiology report.
[0075] In one embodiment, the target image is formatted as a Medical Digital Imaging and Communication (“DICOM”) image.
[0076] One embodiment further includes using an infrastructure pipeline that includes microservices deployed using DOCKER containers.
[0077] In one embodiment, at least one of the machine learning neural Student model or the machine learning neural Teacher model is programmed to include an architecture that includes at least one of the open-source libraries known and available at the time of writing: DenseNet-121, ResNet-152, ShuffleNet2, ResNext101, GhostNet, EfficientNet-b5, SeNet-154, Se-ResNext-101, or Inception-v4.
[0078] In one embodiment, at least one of the machine learning neural Student model or the machine learning neural Teacher model is programmed as a convolutional neural network.
[0079] In one embodiment, one of the one or more categories of the target image indicates either healthy tissue or abnormal tissue.
[0080] In one embodiment, one of one or more classifications of the target image indicates abnormal tissue; and the indicated abnormal tissue is further classified as at least one of cardiovascular, pulmonary, mediastinal, pleural, or extrathoracic structures.
[0081] In one embodiment, at least one of one or more categories of the target image is a sub-category.
[0082] One embodiment further includes preprocessing the target image, wherein the preprocessing includes applying a trained machine learning filter model to the target image before one or more classifications of the output target image.
[0083] One embodiment further includes programmatically determining the correct anatomical orientation of the target image before outputting one or more classifications of the target image.
[0084] In one embodiment, determining the correct anatomical orientation of a target image includes executing a trained machine learning model programmed to operate without relying on DICOM metadata associated with the target image or lateral markers associated with the target image.
[0085] In one embodiment, the trained machine learning model is jointly trained on augmented and real data.
[0086] In one embodiment, determining the correct anatomical orientation of the target image includes determining the correct rotation of the target image by executing a first programming model and determining the correct flipping of the target image by executing a second programming model.
[0087] One embodiment further includes programmatically verifying that the target image corresponds to the correct body part after determining the correct anatomical orientation of the target image and before outputting one or more classifications of the target image.
[0088] In one embodiment, verifying that a target image corresponds to the correct body part includes executing a trained machine learning model.
[0089] In various embodiments, this disclosure provides one or more computer-readable non-transitory storage media that, when executed by one or more processors, are operable to perform one or more methods provided by this disclosure.
[0090] In various embodiments, this disclosure provides a system comprising: one or more processors; and one or more computer-readable non-transitory storage media coupled to the one or more processors and including instructions operable when executed by the one or more processors to cause the system to perform one or more methods provided in this disclosure.
[0091] 2.0 Machine Learning Techniques for Processing Radiographic Images of Animals or Pets
[0092] In one embodiment, this disclosure provides an automated technique for classifying radiographic images of animals or pets. One or more digitally stored radiographic images may be in Medical Digital Imaging and Communication (“DICOM”) format. Once an image is received, it can be digitally filtered using a trained machine learning model or tool (e.g., a convolutional neural network model or a transformer-based model) to remove certain features, such as non-chest images. In other examples, the machine learning model or tool may be K-Nearest Neighbors (KNN), Naive Bayes (NB), decision trees or random forests, support vector machines (SVM), deep learning models such as CNNs, region-based CNNs (RCNNs), one-dimensional (1-D) CNNs, recurrent neural networks (RNNs), or any other machine learning model or technique. In other exemplary embodiments, further filtering may be performed to remove the entire image or a portion of the image, such as the chest, pelvis, abdomen, or body. This filtering may be performed based on DICOM body part labels and one or more view locations. The model performing such filtering may be referred to as a “filter model”.
[0093] The resulting machine learning model can be used for various clinical or medical purposes. For example, radiographic images of a pet can be taken by a veterinarian or veterinary assistant. The image can then be processed using a trained machine learning model. During processing, the image can be classified as normal or abnormal. If abnormal, the image can be classified as at least one of cardiovascular, pulmonary structures, mediastinal structures, pleural cavity, or extrapleural. In some non-limiting embodiments, the image can be subclassified. For example, a subclass of pleural cavity could include pleural effusion, pneumothorax, and / or pleural mass. The image can be filtered, segmented, annotated, masked, or labeled, and then displayed to a user using a display device (e.g., the screen of a computing device) along with the determined image category and subclass.
[0094] In one embodiment, the machine learning process and the generated images can be used to provide radiologists with on-demand second opinions, forming the basis for a service that provides veterinary hospitals with immediate assessment of radiological images, and / or improving efficiency and productivity by allowing radiologists to focus on the pet itself rather than the images.
[0095] In some non-limiting embodiments, the machine learning framework may include a convolutional neural network (CNN) component that is trained or has been trained on training data of radiographic images of animals or pets and corresponding ground truth data (e.g., known or defined labels or annotations). The collected training data may, for example, include one or more images captured by a client device. A CNN is an artificial neural network containing one or more convolutional layers and subsampling layers with one or more nodes. One or more layers (including one or more hidden layers) may be stacked to form a CNN architecture. The disclosed CNN can learn to determine image parameters and subsequent classification of radiographic images of animals or pets by accessing a large amount of labeled training data. While in some examples the neural network may train a learned weight for each input-output pair, a CNN may convolve trainable fixed-length kernels or filters along its inputs. In other words, a CNN can learn to recognize small, raw features (low-level) and combine them in complex ways (high-level). In certain embodiments, the CNN may be supervised, semi-supervised, or unsupervised.
[0096] In some non-limiting embodiments, pooling, padding, and / or striding may be used to reduce the size of the CNN output in the dimension in which the convolutions are performed, thereby reducing computational cost and / or the likelihood of overtraining. A strid may describe the stride or number of steps the filter window slides, while padding may include zeroing certain regions of data to buffer the data before or after the strid. In one embodiment, pooling may, for example, include simplifying the information collected by the convolutional layers or any other layers and creating a compressed version of the information contained within those layers.
[0097] In some examples, region-based CNNs (RCNNs) or one-dimensional (1-D) CNNs can be used. RCNNs involve using selective search to identify one or more regions of interest in an image and independently extracting CNN features from each region for classification. The type of RCNN employed in one or more embodiments may include Fast RCNN, Faster RCNN, or Mask RCNN. In other examples, a one-dimensional CNN can process fixed-length time series segments generated using a sliding window. Such a one-dimensional CNN can operate in a many-to-one configuration that utilizes pooling and striding to connect the outputs of the final CNN layers. Fully connected layers can then be used to generate classification predictions at one or more time steps.
[0098] Unlike one-dimensional CNNs that convolve a fixed-length kernel along the input signal, recurrent neural networks (RNNs) process each time step sequentially, so the final output of an RNN layer is a function of each previous time step. In some implementations, a variant of the RNN, known as a Long Short-Term Memory (LSTM) model, can be used. An LSTM may include a storage unit and / or one or more control gates to model temporal dependencies in long sequences. In some examples, the LSTM model can be unidirectional, meaning the model processes the time series in the order it is recorded or received. In another example, if the entire input sequence is available, two parallel LSTM models can be evaluated in opposite directions (temporally forward and backward). The results of the two parallel LSTM models can be concatenated to form a bidirectional LSTM (bi-LSTM), which can model temporal dependencies in both directions.
[0099] In some embodiments, one or more CNN models and one or more LSTM models can be combined. This combined model may include a stack of four non-stepping CNN layers, followed by two LSTM layers and a softmax classifier. The softmax classifier can normalize a probability distribution that includes several probabilities proportional to the exponent of the input. For example, the input signal to the CNN is not padded, so each CNN layer will shorten the time series by several samples, even though the layers are non-stepping. The LSTM layers are unidirectional, so the softmax classification corresponding to the final LSTM output can be used for training and evaluation, as well as for reconstructing the output time series of the sliding window. However, the combined model can also operate in a many-to-one configuration.
[0100] 2.1 Exemplary pet radiographic images used for classification
[0101] Figure 1 The images shown are radiographic images before and after processing by one or more machine learning models or tools, according to certain embodiments. Figure 1 In the example, the preceding image 110 shows an X-ray image of the pet's heart. Therefore, the preceding image 110 has not yet been processed by one or more machine learning models and / or tools. The subsequent image 120 has been classified as cardiovascular and further subclassified as myocardial hypertrophy. This classification is based on one or more features included in the preceding image 110. Specifically, one or more features used for classification in the preceding image 110 may include the size and shape of the pet's heart and / or the relationship of the heart to other body parts (such as the pet's chest cavity or other body parts).
[0102] 2.2 Labeling of pet radiographic images
[0103] Figure 2An example of label 220 is shown, which can be selected by a trained veterinary radiologist or any other veterinary specialist to be applied to input image 210. For example, label 220 may include at least five different classes, and at least 33 different subclasses or further classes, with an ever-expanding range. The first category may be cardiovascular. Subclasses associated with cardiovascular categories may include myocardial hypertrophy, vertebral heart score (VHS), right ventricular enlargement, left ventricular enlargement, right atrial enlargement, left atrial enlargement, aortic enlargement, and aortic pulmonary artery enlargement. The second category may be pulmonary structures. Subclasses associated with pulmonary structures may include interstitial unstructured structures, interstitial nodules, alveoli, bronchi, blood vessels, and pulmonary masses. The third category may be mediastinal structures. Subclasses associated with mediastinal structures may include esophageal dilatation, tracheal collapse, tracheal deviation, lymphadenopathy, and masses. The fourth category may be pleural cavity, which may be associated with subclasses such as pleural effusion, pneumothorax, and pleural masses. The fifth category can be extrathoracic, which can be associated with subcategories such as spinal diseases, cervical tracheal collapse / tracheal relaxation, degenerative joint diseases, gastric dilatation, ascites / loss of detail, intervertebral disc disease, dislocation / subluxation, invasive lesions, hepatomegaly, foreign bodies in the stomach, masses / nodules / lipomas, etc.
[0104] In particular, according to certain non-limiting embodiments, the input image 210 may be annotated and / or labeled. An exemplary method for assigning labels may be to use an automated, natural language processing (“NLP”) model that can input text from one or more relevant radiological reports into the body of the image. In another example, a trained veterinary radiologist may manually apply labels to all extracted images.
[0105] In some non-limiting embodiments, annotated or labeled images can be used to train a machine learning model or tool. In other words, the determined classification can be based on annotations or labels included in the images. The model can be trained using, for example, two discrete steps. In the first step, the machine learning model can be trained using images labeled or annotated by an NLP model, such as a CNN. In the second step, the trained machine learning model can be further trained using extracted images labeled by a trained veterinary radiologist (i.e., an expert). For illustrative and not limiting purposes, the architecture of the machine learning model or tool can be programmed or trained based on at least one of DenseNet-121, ResNet-152, ShuffleNet2, ResNext101, GhostNet, EfficientNet-b5, SeNet-154, Se-ResNext-101, Inception-v4, Visual transformer, SWINtransformer and / or any other known training tool or model.
[0106] 2.3 Classification of pet radiographic images in one embodiment
[0107] Figure 3 An exemplary method for classifying and labeling images using a machine learning system is shown. Figure 3 In the example, method 300 may begin with a first step 310, where an image may be received or acquired at a device. The image may include one or more features.
[0108] In the second step 320, the system can generate a corresponding classification for each of one or more images, wherein the classification is generated by a machine learning model and is associated with one or more features.
[0109] In the third step 330, the system can transmit the corresponding classification and labeled features to the network.
[0110] In the fourth step 340, the system can display the corresponding category on the client device or another computing device on the network.
[0111] 3.0 ADJUSTNET: Automation technology for orienting veterinary radiographic images
[0112] In one embodiment, this disclosure provides a method for automatically determining the correct anatomical orientation in veterinary radiographic images without relying on DICOM metadata or lateral markers. Among other things, this disclosure also provides a deep learning model for determining the correct anatomical orientation in veterinary radiographic images for clinical interpretation, which is independent of DICOM metadata or lateral markers and may include novel real-time deployment capabilities for large-scale remote radiological practices. The disclosed topics can inform a variety of clinical imaging applications, including quality control, patient safety, archive correction, and improving radiologist efficiency. In one embodiment, a model for automatically determining the correct anatomical orientation in veterinary radiographic images may be referred to as AdjustNet. In one embodiment, AdjustNet includes two sub-models. The first sub-model includes an ensemble of three trained machine learning neural models, referred to as RotationNet, for determining the correct rotation of a pet radiographic image. The second sub-model includes an ensemble of three trained machine learning neural models, referred to as FlipNet, for determining whether a pet radiographic image should be flipped. In one embodiment, AdjustNet, RotationNet, and FlipNet may be constructed, trained, and used as detailed in this section of this disclosure.
[0113] Radiographic imaging is crucial for the diagnosis of numerous important medical conditions. Accurate image orientation is essential for optimal clinical interpretation, but errors in digital imaging metadata can lead to incorrect displays, requiring human intervention and hindering efforts to maximize the quality and efficiency of clinical workflows.
[0114] Furthermore, several important workflow considerations in radiological imaging influence radiologists' clinical interpretation, including exposure settings, processing techniques, and anatomical orientation. These considerations are incorporated as textual metadata into the Medical Digital Imaging and Communications (DICOM) common image file format, accompanying pixel data for viewing and radiological interpretation using DICOM viewers. Despite standardized DICOM imaging file conventions, inconsistencies in metadata information, particularly regarding image orientation, are common and lead to inefficiencies in practice during medical interpretation tasks. Manual reorientation of images is often required before interpretation. Additionally, most radiological practices use lateral markers to indicate patient position during radiographic imaging; however, this practice is heterogeneous and error-prone, potentially leading to similar image orientation errors. Therefore, an automated solution for correct DICOM radiographic image orientation, independent of accurate DICOM metadata or lateral markers, could significantly improve radiologists' workflows, reduce interpretation errors, contribute to quality improvement and educational programs, and facilitate data science management of retrospective medical imaging data.
[0115] In one embodiment, this disclosure provides various network architectures for accurate automated radiographic image orientation detection. In one experiment, a convolutional neural network architecture was developed using a dataset containing 50,000 annotated veterinary radiographic images to achieve accurate automated radiographic image orientation detection across 0, 90, 180, and 270 degrees, across anatomical regions and clinical symptom levels, and vertical flipping.
[0116] In one embodiment, a model can be trained for a task involving images in the correct orientation. This model can be a single-stage or two-stage model. In non-limiting embodiments, various weight initialization techniques can be used to develop the model. For example, pre-trained model weights (e.g., on ImageNet) can be used to perform transfer learning, or the model can be randomly initialized and then further pre-trained on augmented data. In non-limiting embodiments, various training pipelines can be used to develop the model. For example, the model can be pre-trained using augmented data and fine-tuned using real data. The model can also be jointly trained using both augmented and real data. In some embodiments, the disclosed subject matter can be used to compute a weighted sum of feature maps of the convolutional layers of a trained neural network.
[0117] Sections 3.1–3.4 of this disclosure, among other things, describe data preparation, data annotation, data augmentation, training, and testing of AdjustNet and its components in various embodiments.
[0118] 3.1 Input Data and Workflow
[0119] In a study approved by the ethics review committee, a dataset of 50,000 checks in DICOM format was obtained. These checks were randomly selected, and the distribution of the number of checks is shown in Table 1.
[0120] Table 1. Exemplary distribution of inspection quantity
[0121]
[0122] All DICOM images were first converted to high-resolution JPEG2000 format, and then to a larger PNG format with a resolution of 512 pixels. The images were resized to 256×256 pixels, and then, as further described in the data augmentation section of this paper, centered cropping was performed using an image crop of 0.8 and random scaling (100%–120%) to minimize lateral markers in the images. The training set consisted of 1550 images; the validation set consisted of 350 images (plus augmentations).
[0123] Figure 5 An exemplary workflow for an image orientation task is shown.
[0124] like Figure 5 As shown, a board-certified veterinary radiologist with 14 years of experience annotated radiographic examinations using an iterative labeling process. After an initial dataset of 251 images was reviewed by human experts, a model was built using this initial dataset (as further described in this paper regarding model training). New examinations were then selected for human review using the trained model, focusing on images incorrectly identified based on human expert review, to provide more data for training the model to focus on difficult examinations. This process was repeated five times; a total of 1550 images were annotated by human experts. Due to the class imbalance in the dataset, images were augmented by rotating and flipping them to the correct orientation. Cropping was also performed because the initial model, similar to previous studies, relied on the location of lateral markers to determine orientation, which was unreliable when analyzing images with original incorrect orientations in the dataset through human expert review, as combined with… Figure 6 Further description.
[0125] Figure 6 Radiographic images showing examples of incorrectly oriented and correctly oriented images of a feline skull are presented.
[0126] In particular, Figure 6 Examples of incorrectly oriented (left) and correctly oriented (right) radiographic images of a feline skull are shown. A lateral marker is highlighted in the image, which is incorrectly applied in this case.
[0127] 3.2 Model Development
[0128] First, a single model was trained for the task of correctly orienting images (rotation and flipping). This multi-class model has 8 output neurons (0-No-Flip, 90-No-Flip, 180-No-Flip, 270-No-Flip, 0-Flip, 90-Flip, 180-Flip, 270-Flip). Table 2 shows the radiographic study counts for each label in the training, validation, and test sets.
[0129] Table 2. Count of radiographic studies for each label in the training, validation and test sets.
[0130]
[0131] Next, a two-stage, step-by-step approach is used to automatically reorient the image by training one model to correctly rotate it and then training a second model to correctly flip it (see [link to article]). Figure 7 ).
[0132] Figure 7 A schematic diagram of an exemplary two-stage model technique with integration, according to certain non-limiting embodiments, is shown.
[0133] In particular, Figure 7 An exemplary two-stage model approach is illustrated, featuring ensembles for each task (rotation, followed by flipping). Each step consists of three different model architectures trained for a given task, with the heuristic approach in practice requiring all three models to agree before a given reorientation of the image. For the RotationNet and FlipNet networks, different CNN architectures (ResNet, Xception, and DenseNet121) were trained and their performance compared. Two different weight initialization techniques were used: (1) performing transfer learning using model weights pre-trained on ImageNet, and (2) the model was randomly initialized and then further pre-trained on augmented data. Two different training pipelines were also used: (1) pre-training the model using augmented data and then fine-tuning it using real data, and (2) jointly training the model using both augmented and real data. In one embodiment, the randomly initialized model, jointly trained on both augmented and real data, outperforms the other methods in accuracy over multiple iterations of the disclosed data acquisition process. Therefore, in one embodiment, this method is used to train the final model.
[0134] In one embodiment, Grad-CAM is used by computing a weighted sum of the feature maps from the last convolutional layer of the trained neural network. The weights are determined by normalizing the sum of gradients of the class labels relative to the weights of the feature maps. These weighted sums are resized to the image size of the input image, converted to RGB, and then superimposed on the original input image (see [link to documentation]). Figure 8 ).
[0135] Figure 8 An exemplary image of the GradCAM transform is shown, which is used for model decisions regarding image orientation.
[0136] like Figure 8 As shown, in one embodiment, the generated images are used to evaluate the model and inform the enhancement process. Figure 8 The GradCAM transformation is illustrated, indicating the pixels in an image used for model decisions regarding image orientation. Importantly, in one embodiment, lateral markers are not included in the pixels. Encouragingly, the number of studies requiring manual transformation after model deployment is reduced by 50% compared to those without model deployment.
[0137] All models were trained on two Tesla V100 graphics processing units (GPUs). These models were designed to minimize cross-entropy loss, using default parameters and 1x10... -3 The Adam optimizer for the learning rate reduces the learning rate by a factor of 0.1 if the model does not improve validation accuracy within approximately two epochs. Since only 10% of the total images are misoriented, high accuracy is required for it to function effectively in a production environment. Therefore, an ensemble strategy is employed where the three networks must agree on rotations or flips.
[0138] 3.3 Model Deployment
[0139] In one embodiment, a two-step model can be centrally deployed using microservices hosted in a DOCKER container and accessible via a ReSTful API, allowing the model to reside permanently in memory for optimal inference speed. The model's output can then be used as input to other AI models deployed throughout the AI pipeline. Each image and its model's predictions can ultimately be archived in a central repository (see...). Figure 9 ).
[0140] Figure 9 A schematic diagram of an exemplary model deployment workflow is shown.
[0141] In one study, to quantify the impact of model post-production for a real-time prospective effect, the production model was removed from the workflow for 24 hours, and a web user interface (UI) log was used to collect instances where consultation with a radiologist was required for manual transformations (e.g., rotation and / or flipping) of radiographic images during the study interpretation period. Data throughout the study were aggregated in overlapping, 24-hour batches using a sliding window that moved every three hours until the end of the study. The metric of interest was the proportion of studies that used one or more manual transformations. Figure 9 The red portion indicates the time period during which the model was turned off. The increased lag due to manual conversion is caused by the time lag between the study being received by the system and the radiologist evaluating the study images.
[0142] 3.4 User Feedback
[0143] In one study, a post-deployment radiologist user survey was developed to determine the user experience of the automated RotationNet through radiologists' interpretations. Four questions with Likert scale answer options were presented: (1) How has the implementation of the automated radiology orientation model (RotationNet) affected your clinical productivity? (2) Would you recommend using RotationNet in your ongoing clinical workflow? (3) Are there any situations where RotationNet should not be used clinically? (4) Please complete if you have any comments / concerns. The survey was sent to 79 radiologists representing a mixed group of teleradiologists, hybrid academic-teleradiologists, and hybrid private clinician-teleradiologists, and 20.2% completed the survey. Regarding the impact on clinical productivity, 88% of respondents rated the impact on clinical productivity as 4 or 5 (“better” or “much better”). Two respondents rated the impact on clinical productivity as 3 – neither good nor bad. No respondents rated the impact of RotationNet on clinical productivity as worse. 94% of respondents said they would recommend using RotationNet in their ongoing clinical workflows; only one respondent was unsure. No respondents indicated they would not recommend RotationNet. When asked if there were any situations where RotationNet should not be used clinically, 69% of respondents answered "no," while others were unsure. Uncertain responses were generally attributed to individuals' lack of knowledge about machine learning. Open-ended comments (the fourth question in the survey) primarily (80% left comments) concerned radiologists' increased sensitivity when RotationNet fails to rotate images correctly or when regular offline maintenance or upgrades are required. One respondent commented on the expectation of cross-sectional imaging that matches their personal preferences. The remaining respondents' comments were either reiterations of previous answers or generally positive assessments of their working environment.
[0144] Tables 3A and 3B show the accuracy of different modeling methods. The accuracy (top) and error (bottom) of the model for a given task are shown. Note that the ensemble model for both rotation (rows) and flip (columns) tasks achieved the highest performance.
[0145] Table 3A. Accuracy of different methods.
[0146] Densenet Resnet Xception integrated Densenet 0.96 0.97 0.95 0.97 Resnet 0.96 0.96 0.96 0.97 Xception 0.96 0.96 0.95 0.97 integrated 0.98 0.98 0.97 0.99
[0147] Table 3B. Accuracy of different modeling methods (balanced data).
[0148] Densent Resnet Xception integrated Densenet 0.92 0.91 0.9 0.91 Resnet 0.91 0.91 0.91 0.91 Xception 0.92 0.9 0.9 0.92 integrated 0.91 0.9 0.92 0.91
[0149] Radiologists typically spend time and cognitive effort processing images to ensure correct orientation before interpretation. In radiology, this involves first determining if the orientation is correct, and then switching, flipping, and rotating them multiple times within and between images. While this work may not be significant for a single study, overall, in the busy practices of many radiologists, this inefficiency of manually adjusting incorrectly oriented images can lead to errors, delayed treatment, and physician burnout, in addition to inconvenience. One study aimed to explore the development of deep learning models to determine the correct anatomical orientation of veterinary radiological images for clinical interpretation and to describe novel real-time deployment experiences in large-scale remote radiology practices. The study found that data augmentation techniques significantly improved all models, with an ensemble of three models (RotationNet) achieving the highest performance (error rate <0.01), outperforming state-of-the-art techniques reported in related work. Furthermore, the successful deployment of RotationNet in a real-world production environment, processing over 300,000 incoming DICOM files for more than 4,600 hospitals in over 24 countries within a month, reduced the number of studies requiring manual intervention from clinical radiologists by 50%. In practice, RotationNet’s automated medical imaging DICOM orientation has achieved state-of-the-art or better performance, optimizing the interpretation workflow of clinical image imaging in large-scale production.
[0150] Figure 10 Exemplary images for error analysis are shown according to certain non-limiting embodiments.
[0151] Error analysis was performed for false positives and false negatives; examples of errors include... Figure 10 As shown. In particular, Figure 10 An exemplary error analysis image of the AdjustNet model is shown. Figure 10 A radial image of the forearm (left side, labeled "L") of a canine is depicted, in which the model incorrectly predicted that the image needed to be rotated 90 degrees, but its orientation in the original state was correct. Figure 10 A radial image of a canine (right, labeled “R”) is also depicted, in which the model incorrectly predicted that the image needed to be flipped 180 degrees, but its orientation in its original state was correct.
[0152] In one embodiment, RotationNet can be used to automatically and retrospectively encode a large number of examinations without DICOM or lateral marking, and it can also be effective for other tasks, such as providing feedback to clinics about incorrect orientation, errors in DICOM metadata, or incorrect lateral marking. In one embodiment, RotationNet can be applied at the point of care as feedback to radiographers for immediate and consistent feedback, rather than after submission for interpretation, which has the potential to improve awareness and baseline functionality, thereby reducing errors (e.g., incorrect marking).
[0153] 4.0 End-to-end pet radiology image processing using RAPIDREADNET
[0154] In one embodiment, this disclosure provides a machine learning neural model called RapidReadNet. In one embodiment, RapidReadNet can be programmed as a multi-label classifier associated with 41 research findings corresponding to various ailments detectable in radiographic images of pets or animals. In one embodiment, RapidReadNet can be constructed, trained, and used as detailed in this section of this disclosure.
[0155] 4.1 An example of an image dataset used to train RAPIDREADNET
[0156] Using large image sets as training datasets to train programmed learning neural models can be beneficial. In various embodiments, unlabeled training data can be labeled by human subject matter experts and / or by existing machine learning models to generate one or more labeled training datasets.
[0157] Figure 11 An image pool evaluated for modeling is shown in one embodiment. In this example, an image pool comprising over 3.9 million veterinary radiographic images from 2007 to 2021 was evaluated. In various embodiments, images may be downsampled or otherwise preprocessed before being used as a training dataset. Figure 11 In the study shown, most of the 3.9 million radiographic images, previously archived as lossy (quality 89) JPEG images, were downsampled to a fixed width of 1024 pixels (px) ("Set 1" in Table 4 below). The remaining radiographic images, mostly provided as (lossless) PNG images, were downsampled to a smaller size (width or height) of 1024 pixels ("Set 2" in Table 4).
[0158] Table 4. Summary of X-ray image data.
[0159]
[0160] Figure 12 The infrastructure of an X-ray system according to one embodiment is shown, where images can be acquired as part of a clinical workflow. The final subset of images is acquired as part of the current clinical workflow. Figure 12 In this example, the submitted DICOM images were downsampled to a fixed height of 512px and then converted to PNG (“Silver Halide” in Table 4). In all cases, the downsampling process preserved the original aspect ratio. In this example, all images were provided as metadata along with a subset of the original DICOM tags. In this example, all images / studies covered real clinical cases received over a 14-year period from various client hospitals and clinics (N>3500), such as... Figure 11 As shown.
[0161] In a particular embodiment, multiple filtering steps may be applied before modeling. In this example, firstly, ImageMagick is used to remove duplicate and low-complexity images. Secondly, a CNN model trained for this purpose is used to filter out imaging artifacts and irrelevant views or body parts. Studies containing more than 10 images are also excluded. Of the 3.9 million veterinary radiographic images, approximately 2.7 million images remain after filtering, representing more than 725,000 different patients.
[0162] In a particular embodiment, the next step in modeling may include annotation and labeling. In this example, the image was annotated with 41 different radiological observations (see Table 5 below).
[0163] Table 5. Radiographic Labels.
[0164]
[0165]
[0166] In this example, for most images (studies prior to 2020), labels were extracted from the corresponding (research-related) radiology reports using an automated, natural language processing (NLP)-based algorithm. In another embodiment, labels could be extracted from radiology reports using a different method. In yet another embodiment, the training data for the initial labels could be generated without using any machine learning models, for example, by using human experts. In this example, the radiology reports compiled all images from a specific study, written by over 2000 different board-certified veterinary radiologists. In this example, images from more recent studies (2020–2021; “Silver Halide” in Table 4) were individually labeled by veterinary radiologists immediately after the study evaluation was completed.
[0167] Various methods can be used to evaluate the accuracy of trained models and inter-annotator variability. In this example, a small number of images (N=615) were randomly selected from the “silver halide” set and labeled by 12 other radiologists. This data was not used for training or validation. One approach to generating baseline truth labels for receiver operating characteristic (ROC) analysis and precision-recall (PR) analysis is to aggregate labels for each image through majority rule voting. In this example, if a majority of the 12 radiologists indicated the presence of a certain finding, that finding was used as the baseline truth label. Point estimates of the false positive rate (FPR) and sensitivity for a particular radiologist were calculated by comparing their labels with the majority rule votes of the other 11 radiologists.
[0168] In this example, the automatic extraction of labels from the "Results" section of the radiology report was accomplished using a modified version of rule-based labeling software known in the art. However, the number of labels was expanded to 41, as shown in Table 5.
[0169] In this example, although the reports contain observations from all images in the study, they never explicitly associate these observations with specific image files. Therefore, in this example, the labels for each study extracted from the reports are initially applied to all images within that study, and then masked using a set of expert-provided rules. Using rules in this way ensures that labels are applied only to images showing the corresponding body part (e.g., the label "cardiac hypertrophy" can be removed from images of the pelvis). This paper discusses the details of the masking process in more detail.
[0170] In this example, the dataset of individually labeled data is (x1, y1), ..., (x n ,y n ) and data marked by markers applicable to veterinary radiology reports Used together. In the examples, each study found... The input can be 0 (negative), 1 (positive), or u (uncertain), and can be determined at the level of the radiological report rather than at the level of a single image. Therefore, label noise may exist on the dataset.
[0171] 4.2 Neural Model Training Techniques for Image Classification Tasks
[0172] In one embodiment, an image dataset annotated by one or more human subject matter experts and a potentially larger dataset annotated using a trained natural language processing (NLP) model can be combined using a distillation approach suitable for multi-label use cases. In one embodiment, the Teacher model θ can be first trained using the human-annotated dataset.t * Training can be performed, and noise can be added during the training process:
[0173]
[0174] Then, the Teacher model can be used to infer the image. Soft pseudo-labels. Noise cannot be used in the inference step:
[0175]
[0176] In one embodiment, the following rule can be used to combine soft pseudo-tags with NLP-derived tags:
[0177]
[0178] In one embodiment, one or more Student models can be trained using derived labels. For example, noise can be added to the Student models to train equal or larger Student models θ. S * :
[0179]
[0180] It is worth noting that in some embodiments, soft pseudo-tags may not be combined with tags generated by NLP.
[0181] In one embodiment, one or more Student models can be updated using an active learning process when new training data is streamed into the system or otherwise obtained. In one embodiment, one of the Student models (or an ensemble of Student models) can be programmed to act as the new Teacher model, and the process can be repeated without the first step. In one embodiment, the active learning process can be programmatically triggered whenever the system receives a large amount of new labeled data.
[0182] In a particular embodiment, the Student and Teacher models can be programmed according to various artificial neural network architectures. For each model, the number of images suitable for memory and used to determine the batch size can be maximized. In this example, this results in batch sizes between 32 and 256. In this example, different image input sizes (ranging from 224×224 to 456×456) are used, and all are trained by reshaping the original images to the input size, zero-padding the images to squares, and then resizing them to maintain the aspect ratio of the original images. Various image augmentation techniques can be performed during training. In this example, the model is pre-trained, trained for a maximum of 30 epochs, and stopped early if the validation loss does not decrease for two consecutive epochs.
[0183] In one embodiment, a piecewise linear transformation can be applied to calibrate the probability of each research finding.
[0184]
[0185] Used for all research findings φ. opt can be set. φ To optimize Youden's J statistic on the independent validation set.
[0186] In one embodiment, the final trained machine learning model used for classifying target radiographic images (e.g., animal and / or pet radiographic images) can be an ensemble of individual, calibrated deep neural network Student models. In one embodiment, averaging the output can lead to better results than voting. In this example, eight best models were used based on a validation set, and then a best subset approach was used to determine the best ensemble. Surprisingly, the best subset is the complete set of eight models; in other words, it includes models that perform poorly compared to the rest of the ensemble. In this example, these poorly performing models still contribute to the overall predictions when included in the ensemble of Student models. The final model containing the ensemble of individual, calibrated deep neural networks can be referred to as RapidReadNet.
[0187] 4.3 Drift Analysis, Experimental Results and Longitudinal Drift Analysis
[0188] To assess whether compensation is needed for differences in images obtained outside of a single example and for industry use, drift analysis can be performed. For example, during system development, images X1,...X... n The study of each image revealed Y1,...Y n It can be considered as originating from the joint distribution P devSampling from (X,Y). In real-world applications, images and research findings can be derived from the distribution P. prod Presented in (X,Y). In veterinary radiology at this example scale, a variety of potential factors could cause differences between these distributions, including breed variation, different radiological equipment, or differences in clinical practice in different regions. Covariate shifts, in other words, variations in marginal distributions, can be considered.
[0189] P dev (Y|X)=P prod (Y∣X)
[0190] P dev (X)≠P prod (X)
[0191] This allows for research and analysis of its impact on model performance. To detect covariate shifts, an autoencoder can be trained using the Alibi-Detect software. In one embodiment, the autoencoder f can be trained during development. A (·,Θ), minimize ∫(Xf) with respect to Θ. A (X∣Θ)) 2 dP dev The trained autoencoder can then be used to reconstruct the image during production and analyze the reconstruction error.
[0192] Figures 13-17 show the ROC and PR analysis results of a set of 615 images labeled by multiple radiologists, comparing model predictions with radiologists' findings on cardiovascular / pleural cavity research. Figure 13A and 13B Studies on the lungs have revealed (Figure 14 and...) Figure 14B Studies on the mediastinum have found that... Figure 15A and 15B ), and findings from extrathoracic studies ( Figure 16A and 16B The labels were compared. Each graph shows the ROC (upper) and PR (lower) curves for each study finding, as well as point estimates of the false positive rate (FPR), precision, and recall (sensitivity) for each radiologist. Studies with fewer than five positive labels were not analyzed. The accuracy of the model is comparable to that of individual radiologists.
[0193] Longitudinal drift analysis was also performed. In this example, an autoencoder was trained using all archived images (see sets 1 and 2 in Table 4) and then applied to images in subsequent studies to examine the L2 norm reconstruction error for each image.
[0194] Figure 18 A visualization of the reconstruction error calculated on a weekly basis is shown.
[0195] In particular, Figure 18 The distribution and quartiles of L2 errors, grouped by week, are shown from November 2020 to June 2021. Figure 18 As shown in the weekly chart, almost no differences were observed between the distributions, indicating that the input data was consistent over the past year and that the model is robust to data from new customers. This lack of perceptual drift can be partly attributed to the high diversity of tissues and animals represented in the training data.
[0196] Figure 19 The distribution of reconstruction error as a function of the multiple tissues represented is shown. Specifically, Figure 19 Data from a small number of organizations (1, 6, ..., 16 organizations) do not appear to represent the overall diversity of the image data well.
[0197] Notably, a positive correlation was observed between data size and model performance on a separate, manually labeled test set. Efficient-Net-b5 was trained on subsets of data of varying sizes, and the resulting models were tested on the same test set. The results are shown in Table 6, indicating the potential for further performance improvements with increasing data scale. Table 6 presents metrics of the model on 30,477 unseen test data points, compared to a benchmark truth from a board-certified radiologist. All data were labeled in a clinical production environment.
[0198] Table 6. Comparison of the model’s metrics on 30,477 unseen test data points with the baseline truth of a board-certified radiologist.
[0199]
[0200] Table 7 (below) shows the ROC results for each study finding. The Number of Positive Cases (Npositive) column lists the number of studies (out of 9311 studies) that had at least one positive label. For studies with fewer than 10 positive cases (Npositive), the area under the receiver operating characteristic curve (AUROC), false positive rate (FPR), and sensitivity were not calculated.
[0201] Table 7. ROC results for each research aspect of the findings.
[0202]
[0203] 4.4 System architecture and method of RAPIDREADNET in one embodiment
[0204] An exemplary infrastructure pipeline can rely on microservices (REST APIs) deployed using DOCKER containers. Each container can deploy ReST API modules using Sebastian Ramirez's FastApi framework, and each container can serve a unique, specialized task. In one embodiment, the production pipeline includes an asynchronous processing approach using a message broker to handle a large number of images (e.g., approximately 15,000) to be processed daily. This could be achieved, for example, by using a NoSQL database that stores each individual incoming request, and the Redis Queuer library as a background processing mechanism to consume each stored request in parallel.
[0205] Predictions from the model can be returned in JSON format and can be stored directly in a MongoDB database for long-term archiving, or stored in Redis JSON storage for short-term archiving. In other embodiments, predictions can be stored in another digital storage, such as a relational database, cloud storage, local hard drive, data lake, or other storage media. In a non-limiting embodiment, using short-term storage (e.g., Redis JSON) allows for aggregation of results at the research level, while also including some contextual information.
[0206] like Figure 12 In the best-case scenario, microservices can be managed by a DOCKER container that executes code organized by functional elements, such as: (1) a message broker: preprocessing, monitoring, and scheduling incoming requests; (2) a model service: an AI coordinator module and a separate model service. Models can be implemented using the PyTorch framework; and (3) a results and feedback loop store: contextualizing model results at the research level and sending the results back.
[0207] The message broker layer can include five DOCKER containers: (i) a Redis queuer module, (ii) a Redis database, (iii) a Redis JSON module, (iv) a Redis queue worker, and (v) a Redis monitoring dashboard. Each incoming image is sent via (i) the Redis queuer module (which temporarily stores the image file and its corresponding research-wide metadata on local disk) and added as an entry to the (ii) Redis database queue. In one embodiment, the Redis queue worker executes in parallel to check for new requests in the Redis database and send them to the AI coordinator. This architecture can serve at least 15,000 images per day.
[0208] The AI coordinator container can be programmed to coordinate the execution of inference from different AI modules described in other sections. In one embodiment, the first AI model from which predictions are collected is AdjustNet (Model 1). In one embodiment, Model 1 examines the orientation of a radiographic image. Then, DxpassNet (Model 2) can verify whether the image corresponds to a body part predicted by the architecture, a task derived from RapidReadNet (a multi-label classifier associated with 41 labels corresponding to various conditions, see Table 6). The results can be contextualized using study-wide metadata provided with the images during their initial upload to the service. In one embodiment, to achieve this aggregation at the study level, records of all study-wide image inferences can be temporarily stored in a Redis JSON module and managed using rule-based expert system tools, such as the C Language Integrated Generative System (CLIPS). In embodiments with rule-based expert system tools, the rules applied to the output of the model and animal metadata can be used to obtain contextual data from radiologist reports. In one embodiment, the Python library can interact with C tools and uses a MongoDB database to store rules, supporting the dynamic nature of the rules by periodically creating new ones. Other embodiments may use different programming languages, codebases, or database types. Furthermore, other embodiments implementing some of the disclosed functionality may include fewer, more, or different AI modules.
[0209] One embodiment may include a feedback storage loop. In one embodiment, all records may (1) be stored in a database in JSON format, such as the mongoDB database described above, (2) using embedded pre- and post-deployment infrastructure, (3) include data describing the number of clinics and radiologists using the exposed system, (4) include a workflow for radiologists to provide label feedback, (5) include a method for adding new labels when acquiring data in a semi-supervised approach, and (6) include a canary performance description / shadow performance description of the process.
[0210] Figure 20 An exemplary computer implementation or programming method is shown for classifying radiographic images (e.g., radiographic images of animals and / or pets) using machine learning neural models.
[0211] Method 2000 can be programmed to begin at step 2002, which includes receiving a first labeled training dataset comprising a first plurality of images, each of the first plurality of images being associated with a set of labels.
[0212] In one embodiment, programmatic control may instruct the execution of step 2004, which includes programmatically training a machine learning neural teacher model on a first labeled training dataset.
[0213] In one embodiment, programmatic control may instruct the execution of step 2006, which includes programmatically applying a machine learning model trained for natural language processing (NLP) to an unlabeled dataset comprising digital electronic representations of natural language text summaries of a second plurality of images, thereby generating a second labeled training dataset comprising the second plurality of images.
[0214] In one embodiment, programmatic control may instruct the execution of step 2008, which includes using a machine learning neural teacher model to programmatically generate a corresponding set of soft pseudo-labels for each of the second plurality of images.
[0215] In one embodiment, programmatic control may instruct the execution of step 2010, which includes using soft pseudo-labels to programmatically generate a set of derived labels for each image in the second labeled training dataset.
[0216] In one embodiment, programmatic control may instruct the execution of step 2012, which includes training one or more programmed machine learning neural Student models using derived labels. In one embodiment, programmatic control may instruct the execution of step 2014, which includes receiving a target image.
[0217] In one embodiment, programmatic control may instruct the execution of step 2016, which includes applying a linear ensemble of one or more Student models to output one or more classifications of the target image.
[0218] In one embodiment, programmatic control may instruct the execution of step 2018, which includes, optionally, using active learning to programmatically update one or more of one or more programmed machine learning neural Student models.
[0219] The embodiments disclosed herein are merely examples, and the scope of this disclosure is not limited thereto. Some non-limiting embodiments may include all, some, or none of the components, elements, features, functions, operations, or steps of the embodiments disclosed herein.
[0220] 5.0 Advantages of certain embodiments
[0221] 5.1 Exemplary Technical Advantages of ROTATIONNET
[0222] In some embodiments, the disclosed data augmentation techniques significantly improve machine learning models used for automatically determining the correct anatomical orientation in veterinary radiographic images. In fact, in one embodiment, the integration of three machine learning models, known as RotationNet, achieved superior performance (e.g., error rate <0.01), outperforming the state-of-the-art techniques reported in related work. Furthermore, successful deployment of the disclosed subject matter can reduce the need for manual intervention by clinical radiologists by at least about 10%, about 20%, about 30%, about 40%, or about 50%. The term “about” or “approximately” refers to an acceptable margin of error for a particular value as determined by a person skilled in the art, which will depend in part on how the value is measured or determined, i.e., the limitations of the measurement system. For example, according to practice in the art, “about” can mean within three or more standard deviations. Alternatively, “about” can mean a range of up to 20% of a given value, preferably up to 10%, more preferably up to 5%, and even more preferably up to 1%. Alternatively, particularly for biological systems or processes, the term may be used to indicate a value within an order of magnitude, preferably within 5 times, and more preferably within 2 times.
[0223] In one example, in large-scale remote radiology practices, approximately 3 million radiographic images are received and interpreted annually. Up to 20% of these images lack correctly encoded orientation information in their Medical Digital Imaging and Communications (DICOM) metadata, or contain errors in lateral marking, leading to inaccurate image orientation during interpretation. In summary, this forces radiologists to expend significant effort reorienting images in the viewer and increases the likelihood of significant errors in downstream clinical decision-making. Because the reporting performance and reliance on lateral marking in conventional systems are insufficient for clinical translation, and because previous work has not demonstrated effectiveness based on deployment data from real-world practice, there is a need for novel automated image orientation methods that are independent of DICOM metadata and have practice-based evidence. In one embodiment, this disclosure provides such a novel system.
[0224] 5.2 Exemplary Technical Advantages of RAPIDREADNET and an End-to-End System for Pet Radiographic Image Classification Disclosed in One Embodiment
[0225] Embodiments of the disclosed technology can allow for the detection of predefined clinical research findings in radiographic images of canines and felines through modeling and methods for data refinement. This modeling and method combines automatic labeling with refinement and demonstrates performance enhancements compared to methods using only automatic labeling. In this disclosure, data scaling and its interaction with different models are evaluated. The time-varying performance of one embodiment is evaluated and compared to time-varying input drift. The deployment process and embedding of embodiments in a larger deep learning-based platform for X-ray image processing are discussed herein. As described herein, using a radiologist student model with applied noise can improve the robustness of X-ray image prediction. Large-scale application of high-performance deep learning diagnostic systems in veterinary care can provide crucial insights that help bridge the gap in translating these promising human and veterinary medical imaging diagnostic technologies into clinical practice.
[0226] This disclosure provides innovations in key areas. In one example, a randomly initialized network can outperform a network pre-trained on ImageNet, and joint training with augmented and real data can outperform a pre-training-fine-tuning pipeline. Pre-training with ImageNet (using rotation and flipping as data augmentation) during training can introduce invariance to image reorientation and limit model development for this task in medical imaging. Furthermore, pre-training on augmented data can bias the network toward certain features indicative of the synthetic orientation of the image, while joint training can act as a regularization technique to avoid these biases. Moreover, the methods described in the context of embodiments including AdjustNet and / or RotationNet and / or RapidReadNet neural models show significant performance advantages. Among other things, this disclosure provides a dedicated network ensemble approach for each task, enabling independent optimization and targeted augmentation for model training. For example, the disclosed techniques have been successfully deployed on radiological images from over 4,600 hospitals in more than 24 countries and have been shown to reduce the need for manual image processing by more than 50%, demonstrating their feasibility and significant positive impact on overall workflow efficiency.
[0227] More generally, developing deep learning models for medical imaging can require significant effort in data preparation, which can be greatly complicated by the frequent mislabeling or incompleteness of DICOM headers in archival data. In one embodiment, this disclosure provides a method for implementing automated image orientation in practice. In other embodiments, the disclosed techniques can be used to provide feedback to clinics regarding misorientation, errors in DICOM metadata, or incorrect lateral labeling. Therefore, various embodiments including AdjustNet and / or RotationNet and / or RapidReadNet can be used for automated retrospective coding of large numbers of examinations in the absence of DICOM or lateral data and can be used for a variety of tasks. In one example, embodiments including AdjustNet and / or RotationNet and / or RapidReadNet can be applied at the point of care as feedback to radiographers for immediate and consistent feedback, rather than after submission for interpretation, which has the potential to improve awareness and baseline functionality, thereby reducing errors (e.g., mislabeling).
[0228] In summary, automated medical imaging DICOM orientation using AdjustNet and / or RotationNet or RapidReadNet can achieve an error rate of less than 0.01 and reduce the need for human expert intervention in image orientation by an average of 50%. The disclosed topics include a novel end-to-end machine learning approach for optimizing radiological image orientation so that all images are always presented in the correct orientation, with significant efficiency gains in large-scale deployments.
[0229] 6.0 Implementation Example – Hardware Overview
[0230] Figure 4 An exemplary computer system 400 for evaluating radiographic images of pets or animals using machine learning tools is illustrated according to some non-limiting embodiments. In some non-limiting embodiments, one or more computer systems 400 perform one or more steps of one or more methods described or illustrated herein. In some other non-limiting embodiments, one or more computer systems 400 provide the functionality described or illustrated herein. In some non-limiting embodiments, software running on one or more computer systems 400 performs one or more steps of one or more methods described or illustrated herein, or provides the functionality described or illustrated herein. Some non-limiting embodiments include one or more portions of one or more computer systems 400. Herein, references to computer systems may include computing devices and vice versa, where appropriate. Furthermore, references to computer systems may include one or more computer systems, where appropriate.
[0231] This disclosure contemplates any suitable number of computer systems 400. This disclosure contemplates computer systems 400 employing any suitable physical form. By way of example and not limitation, computer system 400 may be an embedded computer system, a system-on-a-chip (SOC), a single-board computer system (SBC) (e.g., a computer-level module (COM) or system-level module (SOM)), a desktop computer system, a laptop computer or notebook computer system, an interactive kiosk, a mainframe computer, a computer system network, a mobile phone, a personal digital assistant (PDA), a server, a tablet computer system, an augmented / virtual reality device, or a combination of two or more of these. Where appropriate, computer system 400 may include one or more computer systems 400 that are centralized or distributed, spanning multiple locations, spanning multiple machines, spanning multiple data centers, or residing in the cloud, wherein the cloud may include one or more cloud components in one or more networks. Where appropriate, one or more computer systems 400 may perform one or more steps of one or more methods described or illustrated herein without substantial spatial or temporal limitations. By way of example and not limitation, one or more computer systems 400 may perform one or more steps of one or more methods described or illustrated herein in real time or in batch mode. Where appropriate, one or more computer systems 400 may perform one or more steps of one or more methods described or illustrated herein at different times or in different locations.
[0232] In some non-limiting embodiments, computer system 400 includes processor 402, memory 404, storage 406, input / output (I / O) interface 408, communication interface 410, and bus 412. Although this disclosure describes and illustrates a particular computer system having a particular number of particular components in a particular arrangement, this disclosure contemplates any suitable computer system having any suitable number of any suitable components in any suitable arrangement.
[0233] In some non-limiting embodiments, processor 402 includes hardware for executing instructions, such as instructions constituting a computer program. By way of example and limitation, to execute instructions, processor 402 may retrieve (or fetch) instructions from internal registers, internal caches, memory 404, or memory 406; decode and execute the instructions; and then write one or more results to internal registers, internal caches, memory 404, or memory 406. In some non-limiting embodiments, processor 402 may include one or more internal caches for data, instructions, or addresses. This disclosure contemplates that processor 402 may include any suitable number of suitable internal caches where appropriate. By way of example and not limitation, processor 402 may include one or more instruction caches, one or more data caches, and one or more translation lookup buffers (TLBs). Instructions in the instruction cache may be copies of instructions in memory 404 or memory 406, and the instruction cache may accelerate the retrieval of those instructions by processor 402. The data in the data cache may be a copy of data in memory 404 or memory 406 for operation by instructions executed at processor 402; the result of a previous instruction executed at processor 402 for access to or writing to memory 404 or memory 406 by subsequent instructions executed at processor 402; or other suitable data. The data cache can accelerate read or write operations of processor 402. The TLB can accelerate virtual address translation of processor 402. In some non-limiting embodiments, processor 402 may include one or more internal registers for data, instructions, or addresses. This disclosure contemplates that processor 402 may include any suitable number of suitable internal registers where appropriate. Where appropriate, processor 402 may include one or more arithmetic logic units (ALUs), may be a multi-core processor, or may include one or more processors 402. Although this disclosure describes and illustrates specific processors, this disclosure contemplates any suitable processor.
[0234] In some non-limiting embodiments, memory 404 includes main memory for storing instructions to be executed by processor 402 or data to be operated by processor 402. By way of example and not limitation, computer system 400 may load instructions from memory 406 or another source (e.g., another computer system 400) into memory 404. Processor 402 may then load instructions from memory 404 into internal registers or internal caches. To execute instructions, processor 402 may retrieve and decode instructions from internal registers or internal caches. During or after instruction execution, processor 402 may write one or more results (which may be intermediate or final results) to internal registers or internal caches. Processor 402 may then write one or more of these results to memory 404. In some non-limiting embodiments, processor 402 executes only the instructions in one or more internal registers or internal caches or memory 404 (and not memory 406 or elsewhere), and operates only on the data in one or more internal registers or internal caches or memory 404 (and not memory 406 or elsewhere). One or more memory buses (each of which may include an address bus and a data bus) may couple processor 402 to memory 404. As described below, bus 412 may include one or more memory buses. In some non-limiting embodiments, one or more memory management units (MMUs) exist between processor 402 and memory 404 and facilitate access to memory 404 requested by processor 402. In some other non-limiting embodiments, memory 404 includes random access memory (RAM). Where appropriate, the RAM may be volatile memory. Where appropriate, the RAM may be dynamic RAM (DRAM) or static RAM (SRAM). Furthermore, where appropriate, the RAM may be single-port or multi-port RAM. This disclosure contemplates any suitable RAM. Where appropriate, memory 404 may include one or more memories. Although this disclosure describes and illustrates specific memory components, this disclosure contemplates any suitable memory.
[0235] In some non-limiting embodiments, storage 406 includes a mass storage device for data or instructions. By way of example and not limitation, storage 406 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk drive, magneto-optical disk drive, magnetic tape drive, or Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, storage 406 may include removable or non-removable (or fixed) media. Where appropriate, storage 406 may be internal or external to computer system 400. In some non-limiting embodiments, storage 406 is a non-volatile solid-state memory. In some non-limiting embodiments, storage 406 includes read-only memory (ROM). Where appropriate, the ROM may be a mask-programmable ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically variable ROM (EAROM), or flash memory, or a combination of two or more of these. This disclosure contemplates mass storage 406 in any suitable physical form. Where appropriate, storage 406 may include one or more storage control units that facilitate communication between processor 402 and storage 406. Where appropriate, storage 406 may include one or more storage units 406. Although this disclosure describes and illustrates specific storage units, this disclosure considers any suitable storage unit.
[0236] In some non-limiting embodiments, I / O interface 408 includes hardware, software, or both, providing one or more interfaces for communication between computer system 400 and one or more I / O devices. Where appropriate, computer system 400 may include one or more of these I / O devices. One or more of these I / O devices can enable communication between a person and computer system 400. By way of example and not limitation, I / O devices may include a keyboard, keypad, microphone, monitor, mouse, printer, scanner, speaker, still camera, stylus, tablet computer, touchscreen, trackball, camera, other suitable I / O devices, or combinations of two or more of these. I / O devices may include one or more sensors. This disclosure contemplates any suitable I / O device and any suitable I / O interface 408 for use therewith. Where appropriate, I / O interface 408 may include one or more device or software drivers enabling processor 402 to drive one or more of these I / O devices. Where appropriate, I / O interface 408 may include one or more I / O interfaces 408. Although this disclosure describes and illustrates specific I / O interfaces, this disclosure contemplates any suitable I / O interface.
[0237] In some non-limiting embodiments, the communication interface 410 includes hardware, software, or both that provide one or more interfaces for communication (e.g., packet-based communication) between the computer system 400 and one or more other computer systems 400 or one or more networks. By way of example, and not limitation, the communication interface 410 may include a network interface controller (NIC) or network adapter for communicating with Ethernet or other wired networks, or a wireless NIC (WNIC) or wireless adapter for communicating with wireless networks (e.g., Wi-Fi networks). This disclosure contemplates any suitable network and any suitable communication interface 410 for use therewith. By way of example, and not limitation, the computer system 400 may communicate with one or more portions of an ad hoc network, a personal area network (PAN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), or the Internet, or a combination of two or more of these. One or more portions of one or more of these networks may be wired or wireless. As an example, computer system 400 may communicate with wireless PAN (WPAN) (e.g., Bluetooth WPAN), Wi-Fi network, Wi-Fi Max network, cellular telephone network (e.g., GSM network), or other suitable wireless networks, or combinations thereof. Where appropriate, computer system 400 may include any suitable communication interface 410 for any of these networks. Where appropriate, communication interface 410 may include one or more communication interfaces 410. Although specific communication interfaces are described and illustrated in this disclosure, any suitable communication interface is contemplated herein.
[0238] In some non-limiting embodiments, bus 412 includes hardware, software, or both that couple components of computer system 400 to each other. By way of example and not limitation, bus 412 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infiniband interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCIe) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations thereof. Where appropriate, bus 412 may include one or more buses 412. Although this disclosure describes and illustrates specific buses, any suitable bus or interconnect is contemplated herein.
[0239] Herein, where appropriate, one or more computer-readable non-transitory storage media may include one or more semiconductor-based or other integrated circuits (ICs) (e.g., field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs)), hard disk drives (HDDs), hybrid hard disk drives (HHDs), optical disks, optical disk drives (ODDs), magneto-optical disks, magneto-optical drives, floppy disks, floppy disk drives (FDDs), magnetic tape, solid-state drives (SSDs), RAM drives, secure digital cards or drives, any other suitable computer-readable non-transitory storage media, or any suitable combination of two or more of these. Where appropriate, computer-readable non-transitory storage media may be volatile, non-volatile, or a combination of volatile and non-volatile.
[0240] In this document, "or" is inclusive rather than exclusive, unless otherwise expressly stated or the context otherwise indicates. Therefore, in this document, unless otherwise expressly stated or the context otherwise indicates, "A or B" means "A, B, or both". Furthermore, unless otherwise expressly stated or the context otherwise indicates, "and" is both united and separate. Therefore, in this document, unless otherwise expressly stated or the context otherwise indicates, "A and B" means "A and B, united or separate".
[0241] The scope of this disclosure covers all changes, substitutions, variations, alterations, and modifications to the exemplary embodiments described or illustrated herein that will be understood by those skilled in the art. The scope of this disclosure is not limited to the exemplary embodiments described or illustrated herein. Furthermore, although various embodiments are described and illustrated herein as including specific components, elements, features, functions, operations, or steps, any of these embodiments may include any combination or arrangement of any components, elements, features, functions, operations, or steps described or illustrated anywhere herein that will be understood by those skilled in the art. Furthermore, the device, system, or component of a device or system mentioned in the appended claims being adapted, arranged, enabled, configured, enabled, operable, or operated to perform a specific function includes that device, system, or component, whether or not it or that specific function is activated, turned on, or unlocked, provided that the device, system, or component is so adapted, arranged, enabled, configured, enabled, operable, or operated. Moreover, while this disclosure describes or illustrates some non-limiting embodiments to provide specific advantages, some non-limiting embodiments may not provide these advantages, provide some advantages, or all of these advantages.
[0242] Furthermore, embodiments of the methods presented and described as flowcharts in this disclosure are provided by way of example to provide a more complete understanding of the technology. The disclosed methods are not limited to the operations and logical flows presented herein. Alternative embodiments are contemplated in which the order of various operations is altered, and in which sub-operations described as part of a larger operation are performed independently.
[0243] Although various embodiments have been described for the purposes of this disclosure, such embodiments should not be considered as limiting the teachings of this disclosure to these embodiments. Various changes and modifications can be made to the above elements and operations to obtain results that remain within the scope of the systems and processes described in this disclosure.
[0244] The embodiments disclosed herein are merely examples, and the scope of this disclosure is not limited thereto. Certain non-limiting embodiments may include all, some, or none of the components, elements, features, functions, operations, or steps of the embodiments disclosed above. Embodiments are specifically disclosed in the appended claims for methods, storage media, systems, and computer program products, wherein any feature referred to in one claim class (e.g., method) may also be claimed in another claim class (e.g., system). Dependencies or references in the appended claims are for formal reasons only. However, claims may also be made for any subject matter arising from intentional reference to any prior claim (particularly multiple dependencies), thus any combination of claims and their features may be disclosed and claimed regardless of the dependency chosen in the appended claims. The subject matter that may be claimed includes not only combinations of features listed in the appended claims but also any other combination of features in the claims, wherein individual features referred to in the claims may be combined with any other feature or combination of features in the claims. Furthermore, any embodiments and features described or depicted herein may be claimed in individual claims and / or in any combination with any embodiments or features described or depicted herein or with any features of the appended claims.
[0245] ***
[0246] All patents, patent applications, publications, product descriptions, and agreements referenced in this specification are incorporated herein by reference in their entirety. In the event of a conflict of terms, this disclosure shall prevail.
[0247] While it is obvious that the subject matter described herein is intended to achieve the aforementioned benefits and advantages, the subject matter disclosed herein is not limited in scope to the specific embodiments described herein. It should be understood that modifications, variations, and alterations can be made to the disclosed subject matter without departing from its spirit. Those skilled in the art will recognize or be able to identify many equivalents of the specific embodiments described herein using only conventional experimentation. Such equivalents are intended to be covered by the following claims.
[0248] This document cites various references, all of which are incorporated herein by reference.
Claims
1. A computer-implemented method, comprising: Receive a first labeled training dataset comprising a plurality of digitally stored images, each of the plurality of digitally stored images being associated with a set of labels; Train a machine learning neural teacher model on the first labeled training dataset; A machine learning model trained for Natural Language Processing (NLP) is applied to an unlabeled dataset comprising digital electronic representations of natural language text summaries of a second plurality of images, thereby generating a second labeled training dataset comprising the second plurality of images. Each of the second plurality of images is associated with a label generated by the machine learning model trained for Natural Language Processing (NLP); Using the machine learning neural teacher model, a set of soft pseudo-labels is generated for each of the second plurality of images; Use the soft pseudo-labels to generate a set of derived labels for each image in the second labeled training dataset; Use the derived labels to train one or more programmed machine learning neural student models; Receive the target image; as well as An ensemble of one or more student models is applied to output one or more classifications of the target image.
2. The computer-implemented method according to claim 1 further includes: Active learning is used to programmatically update at least one of the one or more programmed machine learning neural student models.
3. The computer-implemented method according to claim 1 further includes: Noise is applied in one or more training steps of a machine learning model.
4. The computer-implemented method according to claim 1, wherein, The target image is a radiographic image of an animal.
5. The computer-implemented method according to claim 1, wherein, Each of the first plurality of digitally stored images and each of the second plurality of images is a radiographic image of an animal.
6. The computer-implemented method according to claim 1, wherein, The natural language text summary is a radiology report.
7. The computer-implemented method according to claim 1, wherein, The target image is formatted as a Medical Digital Imaging and Communication (DICOM) image.
8. The computer-implemented method according to claim 1, further comprising: Use an infrastructure pipeline, which includes microservices deployed using Docker containers.
9. The computer-implemented method according to claim 1, wherein, At least one of the machine learning neural student model or the machine learning neural teacher model is programmed to include an architecture, the architecture including at least one of DenseNet-121, ResNet-152, ShuffleNet2, ResNext101, GhostNet, EfficientNet-b5, SeNet-154, Se-ResNext-101, or Inception-v4.
10. The computer-implemented method according to claim 1, wherein, At least one of the machine learning neural student model or the machine learning neural teacher model is programmed as a convolutional neural network.
11. The computer-implemented method according to claim 1, wherein, One of the categories of the target image indicates either healthy tissue or abnormal tissue.
12. The computer-implemented method according to claim 11, wherein: One or more classifications of the target image indicate abnormal tissue; and The indicated abnormal tissue is further classified into at least one of the following: cardiovascular, pulmonary, mediastinal, pleural, or extrathoracic.
13. The computer-implemented method according to claim 12, wherein, At least one of the one or more classifications of the target image is a subclassification, the subclassification being selected from at least one of the subclassifications related to the cardiovascular system, the lung structures, the mediastinal structures, the pleural cavity, and the extrathoracic cavity.
14. The computer-implemented method according to claim 1, further comprising: The target image is preprocessed, wherein the preprocessing includes applying a trained machine learning filter model to the target image before outputting one or more classifications of the target image.
15. The computer-implemented method according to claim 1, further comprising: Before outputting one or more classifications of the target image, the correct anatomical orientation of the target image is determined programmatically.
16. The computer-implemented method according to claim 15, wherein, Determining the correct anatomical orientation of the target image includes executing a trained machine learning model programmed to operate without relying on DICOM metadata or lateral markers associated with the target image.
17. The computer-implemented method according to claim 16, wherein, The trained machine learning model is jointly trained on augmented and real data.
18. The computer-implemented method according to claim 15, wherein, Determining the correct anatomical orientation of the target image includes determining the correct rotation of the target image by executing a first programming model, and determining the correct flipping of the target image by executing a second programming model.
19. The computer-implemented method according to claim 15, further comprising: After determining the correct anatomical orientation of the target image, and before outputting one or more classifications of the target image, the target image is programmatically verified to correspond to the correct body part.
20. The computer-implemented method according to claim 19, wherein, Verifying that the target image corresponds to the correct body part includes executing a trained machine learning model.
21. A system for evaluating an image, comprising: One or more processors; as well as One or more computer-readable non-transitory storage media are coupled to the one or more processors and include instructions operable when executed by the one or more processors to cause the system to perform operations including: Receive a first labeled training dataset comprising a plurality of digitally stored images, each of the plurality of digitally stored images being associated with a set of labels; Train a machine learning neural teacher model on the first labeled training dataset; A machine learning model trained for Natural Language Processing (NLP) is applied to an unlabeled dataset comprising digital electronic representations of natural language text summaries of a second plurality of images, thereby generating a second labeled training dataset comprising the second plurality of images. Each of the second plurality of images is associated with a label generated by the machine learning model trained for Natural Language Processing (NLP); Using the machine learning neural teacher model, a set of soft pseudo-labels is generated for each of the second plurality of images; Use the soft pseudo-labels to generate a set of derived labels for each image in the second labeled training dataset; Use the derived labels to train one or more programmed machine learning neural student models; Receive the target image; as well as An ensemble of one or more student models is applied to output one or more classifications of the target image.
22. The system of claim 21, wherein the instructions, when executed, are further operable to cause at least one of the one or more programmed machine learning neural student models to be updated programmatically using active learning.
23. The system of claim 21, wherein the instructions, when executed, are further operable to apply noise in one or more machine learning model training steps.
24. The system according to claim 21, wherein, The target image is a radiographic image of an animal.
25. The system according to claim 21, wherein, Each of the first plurality of digitally stored images and each of the second plurality of images is a radiographic image of an animal.
26. The system according to claim 21, wherein, The natural language text summary is a radiology report.
27. The system according to claim 21, wherein, The target image is formatted as a Medical Digital Imaging and Communication (DICOM) image.
28. The system of claim 21, wherein the instructions are further operable when executed to enable the use of an infrastructure pipeline, the infrastructure pipeline comprising microservices deployed using Docker containers.
29. The system according to claim 21, wherein, At least one of the machine learning neural student model or the machine learning neural teacher model is programmed to include an architecture, the architecture including at least one of DenseNet-121, ResNet-152, ShuffleNet2, ResNext101, GhostNet, EfficientNet-b5, SeNet-154, Se-ResNext-101, or Inception-v4.
30. The system according to claim 21, wherein, At least one of the machine learning neural student model or the machine learning neural teacher model is programmed as a convolutional neural network.
31. The system according to claim 21, wherein, One of the categories of the target image indicates either healthy tissue or abnormal tissue.
32. The system according to claim 31, wherein: One or more classifications of the target image indicate abnormal tissue; and The indicated abnormal tissue is further classified into at least one of the following: cardiovascular, pulmonary, mediastinal, pleural, or extrathoracic.
33. The system according to claim 32, wherein, At least one of the one or more classifications of the target image is a subclassification, the subclassification being selected from at least one of the subclassifications related to the cardiovascular system, the lung structures, the mediastinal structures, the pleural cavity, and the extrathoracic cavity.
34. The system of claim 21, wherein the instructions, when executed, are further operable to preprocess the target image, wherein the preprocessing includes applying a trained machine learning filter model to the target image before outputting one or more classifications of the target image.
35. The system of claim 21, wherein the instructions, when executed, are further operable to programmatically determine the correct anatomical orientation of the target image before outputting one or more classifications of the target image.
36. The system of claim 35, wherein the instructions, when executed, are further operable to determine the correct anatomical orientation of the target image by executing a trained machine learning model, the trained machine learning model being programmed to operate without relying on DICOM metadata associated with the target image or lateral markers associated with the target image.
37. The system according to claim 36, wherein, The trained machine learning model is jointly trained on augmented and real data.
38. The system of claim 35, wherein the instructions, when executed, are further operable to determine the correct rotation of the target image by executing a first programming model and to determine the correct anatomical orientation of the target image by executing a second programming model to determine the correct flipping of the target image.
39. The system of claim 35, wherein the instructions, when executed, are further operable to programmatically verify that the target image corresponds to the correct body part after determining the correct anatomical orientation of the target image and before outputting one or more classifications of the target image.
40. The system of claim 39, wherein the instructions, when executed, are further operable to verify that the target image corresponds to the correct body part by executing a trained machine learning model.