System and method for detecting anomalies in radiological images of pets
A machine learning-based system integrates text and image data to automate the detection of anomalies in pet radiographs, addressing inefficiencies in veterinary diagnostics by generating accurate diagnostic reports.
Patent Information
- Application Number
- JP2025500007
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-07-07
- Filing Date
- 2023-07-07
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2043-07-07
AI Technical Summary
There is a need for a system that can automate the processing and interpretation of diagnostic images of pets to provide clinically reliable results to both radiology-trained and non-radiology-trained veterinarians, addressing the challenges of misoriented or missing laterality markers in radiological images and the time-consuming nature of manual image review.
A system utilizing machine learning models, specifically a contrastive radiology captioning model, to analyze radiographic images of pets, integrating text and image data to generate diagnostic reports, capable of detecting anomalies and providing disease classifications.
The system effectively automates the detection of anomalies in radiographic images, improving efficiency and accuracy in veterinary diagnostics by generating clinically reliable reports, reducing the reliance on manual interpretation and overcoming limitations of unimodal models.
Smart Images

Figure 2025525470000001_ABST
Abstract
Description
Priority
[0001] This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 358,905, filed July 7, 2022, the entire disclosure of which is incorporated herein by reference. [Technical Field]
[0002] The present disclosure relates generally to using one or more machine learning models or tools to evaluate radiological images of pets or animals. [Background technology]
[0003] Veterinarians are increasingly using image-based diagnostic techniques, such as X-rays, to diagnose and identify health problems in animals and pets. However, there are fewer than 1,100 veterinary-trained radiologists worldwide, meaning many veterinarians are unable to take advantage of the benefits offered by image-based diagnostic techniques. Even for radiology-trained veterinarians, reviewing medical images can be a time-consuming and tedious task. Further complicating the above challenges, radiological images of animals and pets can often be misoriented or have missing or incorrect laterality markers. Summary of the Invention [Problem to be solved by the invention]
[0004] Therefore, there is a need for a system that can automate the processing and interpretation of diagnostic images of pets and return clinically reliable results to radiology-trained and non-radiology-trained veterinarians. [Means for solving the problem]
[0005] In certain, but not limiting, embodiments, the present disclosure provides systems and methods for training and using machine learning models to process, interpret, and analyze radiographic images of animals or pets. The images may be in any format used for diagnosing medical conditions, such as image display formats such as DICOM (Digital Imaging and Communications in Medicine). In certain embodiments, the radiographic images may be associated with radiological reports. Traditional radiographic image analysis involves training separate models, such as a text-only model (e.g., a natural language processing model) based on the radiological reports and an image-only model based on the radiographic images. Even at their inception, the text-only and image-only models may be deployed as separate entities that can be used independently of each other (in principle). Unlike these traditional approaches, embodiments disclosed herein allow for training of an integrated text-image model that can be used to directly generate or verify radiological reports based on the radiographic images. In one embodiment, the integrated text-image model can also be programmed to detect anomalies in radiographic images of animals or pets.
[0006] In one embodiment, the present disclosure provides a system and method for automated detection of anomalies in radiographic images of animals or pets. In various embodiments, analysis of captured, collected, or received images and anomaly detection can be performed using one or more machine learning models or tools. In some embodiments, the machine learning models can include one or more neural networks. The neural networks can be, but are not limited to, convolutional neural networks (CNNs). In one embodiment, anomaly detection can indicate, for example, whether a tissue is healthy or abnormal. In one embodiment, tissue classified as abnormal can be further classified, for example, as a cardiovascular anomaly, a pulmonary structural anomaly, a mediastinal structural anomaly, a pleural cavity anomaly, an extrathoracic anomaly, or any combination thereof.
[0007] In some embodiments, the present disclosure provides a method for detecting anomalies in radiographic images of an animal or pet by one or more computing systems, the method including: accessing a plurality of radiographic images of the animal, wherein one or more first radiographic images of the plurality of radiographic images each depict the animal from one or more views and one or more second radiographic images of the plurality of radiographic images each depict one or more body parts of the animal; identifying one or more disease classifications for the animal based on analysis of the plurality of radiographic images with a machine learning model; generating a diagnostic report for the animal based on the machine learning model, the diagnostic report including the one or more disease classifications and a natural language text radiology report; and sending instructions to a user device directing the user to present the diagnostic report.
[0008] In one embodiment, each of the plurality of radiographic images is formatted as a Digital Imaging and Communications in Medicine (DICOM) image.
[0009] In one embodiment, the machine learning model is based on at least one first neural network and at least one second neural network, and the at least one first neural network and the at least one second neural network are coupled to each other.
[0010] In one embodiment, generating a diagnostic report includes accessing a plurality of reference reports, encoding the plurality of reference reports into a feature space, encoding a plurality of radiographic images into the feature space, and determining the diagnostic report based on a similarity search in the feature space.
[0011] In one embodiment, one of the one or more disease categories is indicative of a tissue abnormality.
[0012] In one embodiment, the method further comprises identifying the tissue abnormality as at least one of a cardiovascular abnormality, a pulmonary structural abnormality, a mediastinal structural abnormality, a pleural cavity abnormality, or an extrathoracic abnormality.
[0013] In one embodiment, the method further includes accessing a plurality of training radiography images each associated with a plurality of training radiology reports, and training a machine learning model based on the accessed training radiography images and their respective training radiology reports.
[0014] In one embodiment, the method further includes performing pre-processing on each of the plurality of training radiographic images, the pre-processing including one or more of padding, random dilation, random flip, Gaussian blur, and normalization.
[0015] In one embodiment, the method further includes applying long document coding to each of the plurality of training radiology reports.
[0016] In one embodiment, the method further includes performing pre-processing on each of the plurality of training radiology reports, the pre-processing including one or more of tokenizing, padding, adding classification tokens, and applying an attention mask.
[0017] In one embodiment, the machine learning model comprises an image encoder, a multi-image encoder, a text decoder, and a multi-modal decoder.
[0018] In one embodiment, the method further includes generating, by an image encoder, a feature map based on the plurality of radiographic images; generating, by a multi-image encoder, one or more multi-image keys and values based on the feature map; and generating, by a multi-modal decoder, a natural language text radiology report based on the one or more multi-image keys and values and the initial tokens.
[0019] In one embodiment, the diagnostic report further includes one or more of the plurality of radiographic images.
[0020] In various embodiments, the present disclosure provides one or more computer-readable non-transitory storage media that, when executed by one or more processors, are configured to operate to perform one or more of the methods provided by the present disclosure.
[0021] In various embodiments, the present disclosure provides a system comprising one or more processors and one or more computer-readable non-transitory storage media coupled to one or more of the processors and including instructions, the instructions being configured to be operable, when executed by one or more of the processors, to cause the system to perform one or more of the methods provided by the present disclosure.
[0022] The embodiments disclosed herein are merely examples, and the scope of the disclosure is not limited thereto. Non-limiting specific embodiments are possible that include all or some of the components, elements, features, functions, operations, or steps of the embodiments disclosed herein, as well as those that do not include any of the above. Embodiments of the present invention are disclosed in the appended claims, particularly method claims. Note that any dependency or reference to a prior description in the appended claims is merely selected for formality reasons. Accordingly, subject matter derived by specific reference to any prior claim (especially in the case of multiple dependencies) may be claimed, and therefore, any combination of claims and their features is disclosed and may be claimed regardless of the dependency selected in the appended claims. Claimable subject matter includes not only the combination of features described in the appended claims, but also any other combination of features in the claims, and each feature recited in a claim may be combined with any other feature or combination of features in the claims. Furthermore, any embodiment and feature described or illustrated herein may be claimed as a separate claim, or may be claimed in addition or instead in any combination with any embodiment or feature described or illustrated herein or with any feature set forth in the accompanying claims. [Brief explanation of the drawings]
[0023] [Figure 1] FIG. 1 illustrates an exemplary architecture of a contrastive radiology captioning model. [Figure 2] FIG. 1 illustrates an exemplary method for detecting abnormalities in radiographic images of an animal. [Figure 3] FIG. 1 illustrates an exemplary computer system or device used to assist in detecting anomalies from radiographic images of animals. DETAILED DESCRIPTION OF THE INVENTION
[0024] The terms used herein generally have their ordinary meaning in the art, within the context of this disclosure, and in accordance with the particular context in which they are used. However, certain terms are explained below or elsewhere in this specification to provide further guidance in describing the disclosed compositions and methods and how to make and use them.
[0025] In this specification and claims, the singular forms "a," "an," and "the" are intended to include reference to the corresponding plural forms unless the context clearly dictates otherwise.
[0026] As used herein, the terms "comprise," "including," "having," or other variations thereof, have a non-exclusive inclusion meaning. Thus, when a process, method, article, system, or apparatus "comprises" ("includes" or "has") listed elements, it does not include only the listed elements, but may also include other elements not expressly listed, as well as elements inherent to such process, method, article, or apparatus.
[0027] In this disclosure, the terms "animal" or "pet" refer to domesticated animals, such as, but not limited to, domestic dogs, domestic cats, horses, cows, ferrets, rabbits, pigs, rats, mice, gerbils, hamsters, and goats. Non-limiting examples of pets include domestic dogs and domestic cats, among others. In this disclosure, the terms "animal" or "pet" may also refer to wild animals, such as, but not limited to, bison, elk, deer, ducks, birds, and fish.
[0028] As used herein, a "feature" of an image or slide can be identified based on one or more measurable characteristics of the image or slide. For example, a feature can be a blemish or dark spot in the image, or tissue of various sizes, shapes, light intensity levels, etc.
[0029] In the detailed description herein, when a term such as "an embodiment," "one embodiment," "in various embodiments," "certain embodiment," "some embodiment," "other embodiment," or "certain other embodiment" is used, it is intended to indicate that one or more embodiments described therein may have a particular feature, structure, or characteristic, but not all embodiments necessarily have the particular feature, structure, or characteristic. Furthermore, these terms do not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in a description of one embodiment, it is believed that the same feature, structure, or characteristic also affects other embodiments, whether or not explicitly stated. It will be apparent to those skilled in the art how to implement the present disclosure in alternative embodiments after reading the description herein.
[0030] As used herein, the term "device" refers to a computing system or a mobile device. For example, the term "device" may include a smartphone, a tablet computer, or a laptop computer. In particular, a computing system may include functionality for determining its location, direction, or orientation, such as a GPS receiver, a compass, a gyroscope, or an accelerometer. A client device may further include wireless communication functionality, such as Bluetooth® communication, near-field communication (NFC), or infrared (IR) communication, or communication functionality with a wireless local area network (WLAN) or a cellular telephone network. Such devices may also include one or more cameras, scanners, touchscreens, microphones, or speakers. A client device may also run software applications, such as games, web browsers, and social networking applications. Client devices may include, for example, user equipment, smartphones, tablet computers, laptop computers, desktop computers, or smartwatches.
[0031] The exemplary processes and embodiments may be implemented or performed by a computing system or client device via a mobile application and associated graphical user interface ("UX" or "GUI"). In certain embodiments, the computing system or client device may be, without limitation, a mobile computing system such as a smartphone, tablet computer, or laptop computer. The mobile computing system may include functionality for determining the computing system's location, direction, or orientation, such as a GPS receiver, compass, gyroscope, or accelerometer. Such devices may also include wireless communication capabilities, such as Bluetooth communication, near field communication (NFC), or infrared (IR) communication, or communication capabilities with a wireless local area network (WLAN), 3G, 4G, LTE, LTE-A, 5G, Internet of Things, or cellular telephone network. Such devices may also include one or more cameras, scanners, touchscreens, microphones, or speakers. The mobile computing system may also run software applications, such as games, web browsers, and social networking applications. Social networking applications allow users to connect, communicate, and share information with other users on a social network.
[0032] The terms used herein generally have their ordinary meaning in the art, within the context of this disclosure, and in accordance with the particular context in which they are used. However, certain terms are explained below or elsewhere in this specification to provide further guidance in describing the disclosed compositions and methods and how to make and use them.
[0033] In recent years, semi-supervised multimodal artificial intelligence (AI) models have achieved state-of-the-art results in a variety of downstream tasks. The embodiments disclosed herein leverage the effectiveness of these methods for disease classification and report generation in the field of veterinary radiology. Specifically, a contrastive radiology captioning model is disclosed herein. The contrastive radiology captioning model architecture utilizes contrastive and captioning losses to align X-ray images and reports at both global and local levels. This architecture enables the alignment of multiple X-ray images into a single report when creating a diagnostic report using multiple views and body parts. Experimental results show that for several radiology findings, this architecture achieves significant performance improvements over supervised training methods using other labeling techniques. Ablation experiments are also conducted to demonstrate the importance of each architectural design choice. The text generation capabilities of the contrastive radiology captioning model highlight the potential of radiology report generation using multimodal large-scale language models. The contrastive radiology captioning model is a potentially powerful architecture for training on large unlabeled datasets with multiple image / text pair inputs.
[0034] AI systems using supervised learning methods can be used to assist veterinarians in interpreting x-ray images. However, this approach often relies on the time-consuming and resource-intensive process of manually labeling x-ray images for disease classification. In recent years, semi-supervised multimodal methods have demonstrated significant success, achieving state-of-the-art performance in a variety of downstream tasks. These methods reduce the need for data labeling by utilizing text as ground truth labels. There tends to be a positive correlation between dataset size and model performance. Therefore, semi-supervised methods can be trained on large, unlabeled datasets, improving model performance. This advancement could have significant implications for the field of radiology, as it allows models to be trained using the vast amount of historical reports routinely produced alongside x-ray images.
[0035] Furthermore, these state-of-the-art models have demonstrated the benefits that multimodal approaches bring to the performance of unimodal models. Contrastive approaches align similar text and images by learning a common embedding space between the image and text, achieving zero-shot functionality. Furthermore, optimizing the generative loss of cross-modal alignment has been shown to improve the model's ability to learn fine-grained, local feature representations. Accordingly, embodiments disclosed herein utilize both contrastive and generative approaches to train on radiographic image-text pairs for disease classification and text generation. In certain, but not limited to, embodiments, the present disclosure provides techniques for automated detection of anomalies in radiographic images of animals or pets. One or more radiographic images may be in DICOM (Digital Imaging and Communications in Medicine) format. Upon receiving the radiographic images, anomalies can be identified in the radiographic images by performing image analysis using a trained machine learning model or tool, such as a neural network model. In some embodiments, the approach may employ a vision encoder and a decoupled text unimodal and multimodal decoders.
[0036] However, a particular challenge in model development in the radiology field is that a single patient's report typically references multiple images. This is because multiple x-ray images (e.g., images of different body parts or views) are typically taken during a patient examination. Recent research has highlighted the importance of including relevant images for cross-modal alignment between images and reports. In embodiments disclosed herein, incorporating images from a patient's previous visits has been shown to improve model performance by reducing ambiguity in reports due to the lack of contextual information in the images. Previous research has also suggested that taking multiple image views into account using a CNN-ViT architecture can improve performance in radiology multi-label classification tasks. Therefore, the method disclosed herein similarly utilizes a hybrid CNN-Transformer architecture as a vision encoder to facilitate multi-image embedding representations.
[0037] Some exemplary prior art relevant to the embodiments disclosed herein include RapidRead and StudyFormer. RapidRead is an AI veterinary radiology system that has already been deployed. This system can perform disease classification using an ensemble of CNN models and generate scores using an expert system. The techniques disclosed herein can be compared to this system's model. Meanwhile, the StudyFormer model can generate exam-level embedded representations of patients using a single-image CNN encoder model and a multi-image ViT encoder model. A contrasting radiology captioning model architecture makes some structural modifications to the StudyFormer architecture to use it as a vision encoder.
[0038] Embodiments disclosed herein disclose a contrastive radiology captioning model, which may be based on a self-supervised framework for visual language processing in the radiology field. In certain embodiments, the contrastive radiology captioning model may be based on one or more neural networks. For example, but not limited to, the neural network may be based on a convolutional neural network, a transformer-based network, or an MLP-Mixer. In some embodiments, the architecture of the contrastive radiology captioning model may include a hybrid CNN-ViT vision encoder, a text decoder, and a multimodal decoder.
[0039] In certain, non-limiting embodiments, the contrastive radiology captioning model can be trained by jointly training at least two coupled neural networks: at least one first network for radiographic images and at least one second network for radiology reports. The at least one first network for radiographic images can be considered an image encoder, and the at least one second network for radiology reports can be considered a text encoder. The networks can be based on any suitable architecture, such as Resnet50.
[0040] In some non-limiting embodiments, the joint training of the first and second networks can be based on multiple pairs of radiographic images and radiology reports. The contrastive radiology captioning model can be trained to predict correct pairs of radiographic images and radiology reports from training examples. In some embodiments, the training can include learning a multimodal embedding space by jointly training the first and second networks to maximize the cosine similarity between the radiographic image embeddings and the radiology report embeddings of the correct pairs and minimize the cosine similarity between the embeddings of the incorrect pairs.
[0041] While a neural network may train all of its learned weights for each input / output pair, a CNN may convolve trainable fixed-length kernels or filters along its inputs. That is, a CNN can learn to recognize small primitive features (low-level) and combine them in complex ways (high-level). In particular embodiments, a CNN can be supervised, semi-supervised, or unsupervised.
[0042] In certain non-limiting embodiments, pooling, padding, stride adjustment, or any combination thereof can be used to reduce the output size of a CNN in the dimension in which convolution is performed, thereby reducing computational costs and the likelihood of overtraining. Stride adjustment can refer to the width or number of steps by which a filter window slides, and padding can include buffering data by filling some areas with zeros before or after stride adjustment. In one embodiment, for example, pooling can include simplifying the information collected at any layer, such as a convolutional layer, to create a condensed version of the information contained in that layer.
[0043] In some examples, a region-based CNN (RCNN) or one-dimensional (1D) CNN can be used. RCNN involves using selective search to identify one or more regions of interest within an image and detecting anomalies by extracting CNN features from each region of interest separately. The type of RCNN used in one or more embodiments can include Fast RCNN, Faster RCNN, or Mask RCNN. In other examples, a 1D CNN can process fixed-length time series segments generated by a sliding window. Such a 1D CNN can be implemented in a many-to-one configuration, using pooling and stride adjustment to concatenate outputs at the final CNN layer. A fully connected layer can then be used to generate detections at one or more time steps.
[0044] In some embodiments, one or more CNN models can be combined with one or more LSTM models. The combined model (combined model) can include a stack of four non-stride-tuned CNN layers stacked with two LSTM layers and a softmax classifier. The softmax classifier can normalize a probability distribution containing multiple probabilities proportional to an exponential function of the input. For example, because the input signal to the CNN is not padded, the time series is shortened by a few samples at each CNN layer, even though the CNN layers are not stride-tuned. Because the LSTM layers are unidirectional, the softmax classifier corresponding to the final output of the LSTM can be used in training and evaluation, as well as in reconstructing the output time series from sliding window segments. However, this combined model can operate in a many-to-one configuration.
[0045] FIG. 1 illustrates an exemplary architecture 100 of a contrastive radiology captioning model. In certain, non-limiting embodiments, the contrastive radiology captioning model can be trained using multiple radiography images and their associated radiology reports, i.e., multi-image / text pairs 110. The multi-image / text pairs 110 can include an exam image 112 and its corresponding radiology report 114. The radiography images can be, for example, radiographs, CT scans, etc. The radiology report may contain text descriptions of abnormalities that are long and unstructured compared to tags or labels. For example, a radiology report may contain the description "This dog's heart is enlarged" as a description corresponding to the tag / label "cardiomegaly." Therefore, training a machine learning model based on radiology reports containing long, unstructured text descriptions can be more challenging than traditional training based on tags and labels.
[0046] In a particular embodiment, the vision encoder can be a hybrid CNN transformer based on the StudyFormer architecture. The CNN architecture used can be an Efficient-Net model pre-trained with multi-label classification for multiple x-ray images of a single view. Images of a given exam can first be passed individually through a CNN image encoder 120, resulting in a feature map for each image with 2048×10×10 dimensionality. The feature maps for all images of that exam can then be concatenated to form a feature map 122 with 2048×50×10 dimensionality. This feature map 122 can then be passed through a ViT multi-image encoder 130, which outputs a vector representation of all images of that exam with 501×768 dimensionality. This vector representation can include a CLS embedding representation 140a. In some embodiments, the ViT multi-image encoder 130 may be based on a patch size of 1, depth of 12, attention head of 12, a 2048-dimensional multi-layer perceptron (MLP), and output dimensions of 500x768.
[0047] In some embodiments, both the unimodal text decoder and the multimodal text decoder can be small pre-trained generative pretrained Transformer 2 (GPT2) models. The output dimensionality of the unimodal text decoder 150 after embedding can be [513,768]. This output can include a CLS embedding 140b. The CLS embedding 140 with the unimodal model can be used to compute a contrastive loss.
[0048] The output from the unimodal text decoder 150 can be used as a text query 152 in a multimodal text decoder 160 for a cross-attention mechanism. The output of the ViT multi-image encoder 130 can be used as multi-image keys and values 132. One output 162 of the multimodal text decoder 160 can include a probability distribution over the GPT2 corpus for each position in the text string. Another output 164 of the multimodal text decoder 160 can include a caption loss that can be computed using the tokenized ground truth text labels and the predicted text. The output dimensionality can be [50257, 512].
[0049] In some embodiments, both GPT2 models can be trained using low-rank adaptation of the large-scale language model. The low-rank adaptation of the large-scale language model can utilize low-rank decomposition to learn low-rank matrices for the attention layer. These matrices can represent the changes in weights from the original GPT2 weights to the new task. For example, and not by way of limitation, in this disclosure, the rank of the low-rank adaptation of the large-scale language model is set to 8.
[0050] In certain non-limiting embodiments, preprocessing may be required on radiography images and radiology reports before training a contrastive radiology captioning model. For example, radiography images may be large, such as up to 456x456 pixels. All images may be resized to 300x300. Additionally, a maximum of five images per study may be allowed. In one embodiment, radiography images may be padded to avoid the use of image cropping. For example, but not limited to, studies with fewer than five images may be padded to meet a shape requirement of [5, 3, 300, 300]. Transformations applied to training images may include square padding, random augmentation, random flip, Gaussian blur, etc. All images may be normalized by [0.5, 0.5, 0.5] for both the mean and standard deviation.
[0051] In some embodiments, radiology report preprocessing can be based on long document encoding rather than the commonly used short document encoding. In alternative embodiments, a model for preprocessing radiology reports can be trained based on exam notes. For example, but not by way of limitation, radiology reports can be processed using a GPT2 tokenizer with a maximum token length of 512. To ensure consistency, text with fewer than 512 tokens can be padded with an end-of-sequence (EOS) token. Additionally, classification tokens (CLS) can be added to each tokenized text to facilitate contrastive learning. An attention mask can be applied to the padded tokens.
[0052] Table 1 lists the definitions of the symbols used in each formula in the embodiments disclosed herein.
[0053] [Table 1]
[0054] In certain non-limiting embodiments, a contrastive loss can be used to learn discriminative features by training a model to minimize the distance between similar instances and maximize the distance between dissimilar instances in the embedding space.
[0055] The caption loss allows us to optimize a model to learn to generate a report that describes an input image: Given an image I and its corresponding ground truth caption C = c1c2...c T Given a given set of words, where T is the length of the caption, the caption loss function can be defined as the negative log-likelihood of the correct word sequence.
[0056]
number
[0057] Here, p(c t |I,c 1:t=1 ) is the image I and the word c before the word ct. 1:t=1 Given the correct word c at time step t, t represents the probability of generation.
[0058] And the total loss L combined is defined as follows:
[0059] First, each component S ij r scaled by the temperature parameter θ i and i j We can calculate the similarity matrix S as the dot product of
[0060]
number
[0061] Then, using the label L as the target, we have the similarity matrix S and its transpose matrix S T The average cross entropy loss calculated for the control loss L contrast It can be calculated as:
[0062]
number
[0063] And third, the predicted report pred and true report r true The cross entropy loss calculated for the caption loss L caption It can be calculated as:
[0064]
number
[0065] Finally, the weight w c The contrastive loss and weights w scaled by cap The total loss can be calculated as the sum of the caption losses scaled by
[0066]
number
[0067] In some embodiments, the training data consisted of 3,200,173 image / text pairs, representing 755,263 examinations, while the validation dataset consisted of 50,000 image / text pairs, representing 10,446 examinations. Radiology reports consist of findings and ratings. However, in the embodiments disclosed herein, the model was trained on findings only.
[0068] In some embodiments, the contrastive radiology captioning model was further fine-tuned using a fine-tuning dataset. The training data for the fine-tuning dataset consisted of 594,449 images and labels, representing 145,486 examinations. Meanwhile, the validation dataset consisted of 10,005 images and labels, representing 2,253 examinations. Each image had 41 labels indicating whether or not the image contained a finding. Each label was unique to a single X-ray view. Therefore, the label for a given examination was determined by determining the maximum value of each label across all images in that examination. Therefore, if a finding was present in any of the examination images, the patient was deemed to have a lesion.
[0069] To validate the model's performance, we fine-tuned the vision encoder on a disease classification task. This was achieved by adding a classification layer with 41 outputs to the vision encoder. The model was trained using binary cross-entropy loss. We conducted several experiments using this method, including an ablation experiment, to understand the impact of each part of the design of our contrastive radiology captioning model on classification performance.
[0070] The contrastive radiology captioning model was trained over two weeks on a single GPU. Training continued for 11 epochs and stopped when validation loss performance plateaued. Meanwhile, fine-tuning of the StudyFormer vision encoder with the contrastive radiology captioning model weights took five days on a single GPU. This fine-tuning also stopped when the average precision score stopped improving. This took 50 epochs.
[0071] In certain non-limiting embodiments, the contrastive radiology captioning model can be utilized to detect anomalies in any new radiography images. Furthermore, the contrastive radiology captioning model can predict tags or labels for one or more input radiography images, as well as generate a diagnostic report for such images. For example, rather than predicting "cardiomegaly," the contrastive radiology captioning model can generate a diagnostic report for one or more radiography images that includes a text description such as "This dog's heart is enlarged." In one embodiment, the following steps can be performed to enable the contrastive radiology captioning model to generate a diagnostic report. First, the contrastive radiology captioning model can encode a reference diagnostic report into a feature space. Next, the contrastive radiology captioning model can encode the input radiography image into this shared feature space and perform a similarity search. Furthermore, the contrastive radiology captioning model can select the reference diagnostic report that is closest to the input radiography image as the output diagnostic report.
[0072] In certain embodiments, the contrastive radiology captioning model may not only detect an anomaly but may also specify more detailed information about the detected anomaly. For example, but not limited to, after detecting an anomaly, the contrastive radiology captioning model may further specify a location description of the anomaly, a size description of the anomaly, or a severity description of the anomaly.
[0073] In certain non-limiting embodiments, the contrast radiology captioning model can be used in various clinical or medical applications. For example, a veterinarian or veterinarian assistant can take a radiology image of a pet. The image can then be processed using the contrast radiology captioning model. This processing can result in the image being classified as either "normal" or "abnormal." If the image is classified as "abnormal," it can be classified as at least one of a cardiovascular abnormality, a pulmonary structure abnormality, a mediastinal structure abnormality, a pleural cavity abnormality, or an extrathoracic abnormality. The image can also be further analyzed to identify a location description for the abnormality, a size description for the abnormality, or a severity description for the abnormality. In some non-limiting embodiments, the image can also be divided into subclassifications. For example, subclassifications of pleural cavity abnormalities can include pleural effusion, pneumothorax, and / or pleural mass. Similarly, the image can be further analyzed to identify a location description for these subclassifications, a size description for these subclassifications, or a severity description for these subclassifications. The image may then be displayed to the user along with the anomaly classification and subclassification identified for the image. The image may be displayed on a screen or computing device associated with the user.
[0074] In one embodiment, the contrast radiology captioning model and the resulting images can be used to provide on-demand second opinions to radiologists, form the basis of a service that provides emergency evaluation of radiology images to veterinary clinics, and improve efficiency and productivity by allowing radiologists to focus on the pet rather than the image.
[0075] In certain embodiments disclosed herein, experiments were conducted to verify the effectiveness of the contrastive radiology captioning model. The contrastive radiology captioning model was evaluated by comparing its performance with the current model ensemble employed in the RapidRead system. Performance was measured using ROCAUC and mean precision. The results summarized that the contrastive radiology captioning model had higher ROCAUC for 15 findings and higher mean precision for 10 findings. Table 2 shows the results for findings that outperformed the current ensemble in at least one of the metrics.
[0076] [Table 2]
[0077] To evaluate the impact of multimodal training on StudyFormer's performance, we also fine-tuned a StudyFormer model with Image-Net weights on the same data. We then compared the average precision and ROCAUC scores of this model with those of StudyFormer trained on a control radiology captioning model. The comparison showed significant differences in performance for most findings across both metrics.
[0078] To understand the impact of generative learning on the model's classification performance, we also conducted experiments in which the multimodal text decoder was removed (and thus a contrastive radiology model without a captioner was used). That is, the weights of the image encoder and text decoder were optimized using only the contrastive loss. The results of this experiment showed improved performance for most findings compared to the ImageNet version of StudyFormer, but significantly decreased performance for most findings compared to StudyFormer trained with the contrastive radiology captioning model. A comparison of the average accuracy results for the contrastive radiology captioning model and each ablation experiment is shown in Table 3.
[0079] [Table 3]
[0080] We also tested the text generation capabilities of the contrastive radiology captioning model using unseen test data. First, we generated multi-image embedded representations of X-ray images using the StudyFormer vision encoder. The embedded representations were used as keys and values for a multimodal decoder. We used initial tokens from sentences as initial queries. These keys, values, and queries were then used in the multimodal decoder to autoregressively generate a text string. In this case, this text string represents a diagnostic report. We compared this text with a human-written report. This comparison showed that the model was able to generate an accurate report that closely resembled a human interpreting an image and writing a report about it. However, it also showed that the model generated a significant amount of misinformation through hallucination that was not present in the human-written text. The degree of agreement between the model-generated text and the human-written text varied significantly across tests.
[0081] In an alternative embodiment, a machine learning model configured to detect anomalies in radiographic images of animals or pets can be trained based on a pre-trained Resnet50 architecture instead of the architecture disclosed in FIG. 1 . In the embodiments disclosed herein, further experiments were performed using a trained model based on the Resnet50 architecture. Resnet50 from the OpenAI CLIP library (i.e., a publicly available library) was used to generate image features for 72,105 training images and 10,477 test images. Logistic regression was trained using the training image features and radiology reports, and then tested using the test image features and labels. The ROC-AUC was then calculated for each of the 39 labels. The average ROC-AUC for all 39 measurements for the trained machine learning model was 0.7761819903717145. Meanwhile, the average ROC-AUC for all 39 measurements for OpenAI CLIP (i.e., the state-of-the-art method) was 0.7613616312748511. Table 4 shows a comparison of the ROC-AUC between the pre-trained machine learning model and OpenAI Clip for each of the 39 labels. The comparison results show that the pre-trained machine learning model outperforms conventional techniques.
[0082] [Table 4]
[0083] In embodiments disclosed herein, we investigated the effectiveness of multimodal training methods on the disease classification performance of computer vision models from feline and canine radiographs. The disclosed embodiments demonstrate a contrastive radiology captioning model that utilizes a novel model architecture that utilizes a contrast loss and a captioning loss for training with multiple image / text pairs. In accordance with the disclosed embodiments, performance on multi-label image classification tasks after multimodal alignment was shown to be significantly improved in some respects and comparable to the ensemble of models employed in the current RapidRead system in others.
[0084] Interestingly, training with this architecture yielded the most significant performance improvements for some of the more challenging findings seen with current systems. For example, the average precision score for "ingested material in stomach" was 52% higher using the contrastive radiology captioning model than the current ensemble. This is likely due to the fact that current systems use supervised training methods, with many labels derived using NLP algorithms, which apply rules to extract labels from radiology reports. In contrast, the multimodal approach in this contrastive radiology captioning model allows for the use of the text itself as ground truth labels. This means that it can capture nuances in the text (e.g., syntactic variations in describing lesions) that are often missed by NLP labelers. Therefore, this disclosure suggests that aligning radiology reports with X-ray images may lead to better representation learning for lesions that are difficult to label using alternative methods.
[0085] We also performed ablation experiments to understand the importance of each part of the architecture. First, we show that training a StudyFormer model using ImageNet weights results in significantly worse results than training it with weights trained on a contrasting radiology captioning model. This indicates that the performance improvement also comes from multimodal training and is not solely due to the use of the StudyFormer architecture. Second, we show that omitting the multimodal decoder results in a significant performance degradation compared to StudyFormer trained on a contrasting radiology captioning model. This highlights the importance of the multimodal decoder.
[0086] In the embodiments disclosed herein, the maximum number of images used per test was five. 75% of tests had five or fewer images. Therefore, we chose five as the maximum number of images per test to cover all images in the majority of tests while keeping the requirements and computational cost of the padded images low. However, this meant that 25% of tests had extra images that could not be used. As a result, some images referenced in the test report were inaccessible to the model. This may have affected the model's ability to accurately align the images and reports within the embedding space.
[0087] The embodiments disclosed herein may provide insights into the future development and deployment of deep learning models in the field of radiology. This disclosure demonstrates that a multimodal approach can be used to train a model using multiple X-ray images and their corresponding reports. This allows for the utilization of large unlabeled datasets without the need for other labeling methods, such as NLP algorithms. Specifically, this disclosure demonstrates that this training method can be particularly beneficial in the field of radiology for findings that are difficult to reliably detect using other labeling methods.
[0088] The embodiments disclosed herein also highlight the potential for utilizing large-scale language models to automate the radiology report generation process. This disclosure demonstrates that by simply feeding unseen x-ray images into a contrastive radiology captioning model, it is possible to generate text that closely resembles a human-written diagnostic report.
[0089] That is, this disclosure demonstrates that the contrastive radiology captioning model architecture can be a powerful deep learning model training method compared to supervised learning using other labeling methods. This disclosure also highlights the potential for automated diagnostic report generation by comparing actual reports with reports generated by the contrastive radiology captioning model. Overall, the embodiments disclosed herein demonstrate the potential benefits of using the contrastive radiology captioning model architecture to train models for image classification tasks when using large, unlabeled datasets with multiple image / text pair inputs.
[0090] FIG. 2 illustrates an exemplary method 200 for detecting anomalies in radiographic images of an animal. At step 210, one or more computing systems may access a plurality of radiographic images of an animal. One or more first radiographic images of the plurality of radiographic images each depict the animal from one or more views, and one or more second radiographic images of the plurality of radiographic images each depict one or more body parts of the animal. At step 220, the computing system may identify one or more disease classifications for the animal based on analysis of the plurality of radiographic images by a machine learning model. At step 230, the computing system may generate a diagnostic report for the animal based on the machine learning model. The diagnostic report includes one or more disease classifications and a natural language text radiology report. Then, at step 240, the computing system may send instructions to a user device directing the user device to present the diagnostic report.
[0091] FIG. 3 illustrates an exemplary computer system 300 or device used to assist in the detection of anomalies in radiographic images of animals. In certain, but not limited to, embodiments, one or more computer systems 300 perform one or more steps of one or more methods described and illustrated herein. In other, but not limited to, embodiments, one or more computer systems 300 provide the functionality described and illustrated herein. In certain, but not limited to, embodiments, software executing on one or more computer systems 300 performs one or more steps of one or more methods described and illustrated herein or provides the functionality described and illustrated herein. Some embodiments, but not limited to, include one or more portions of one or more computer systems 300. As used herein, a computer system may encompass a computing device, and where appropriate, a computing device may encompass a computing system. Furthermore, where appropriate, reference to the singular "computer system" preceded by the article "a" may encompass one or more computer systems.
[0092] This disclosure contemplates any suitable number of computer systems 300. This disclosure also contemplates that computer system 300 may take any suitable physical form. For example, and without limitation, computer system 300 may comprise an embedded computer system, a system-on-chip (SOC), a single-board computer system (SBC) (e.g., a computer-on-module (COM) or system-on-module (SOM)), a desktop computer system, a notebook computer system, an interactive kiosk, a mainframe, a mesh computer system, a mobile phone, a personal digital assistant (PDA), a server, a tablet computer system, an augmented / virtual reality device, or a combination of two or more thereof. Also, as desired, computer system 300 may include one or more computer systems 300, may be monolithic or distributed, may be located across multiple locations, across multiple machines, across multiple data centers, or reside in the cloud (which may include one or more cloud components in one or more networks). And, if desired, one or more computer systems 300 may perform one or more steps of one or more methods described and illustrated herein without substantial spatial or temporal limitations. For example, and without limitation, one or more computer systems 300 may perform one or more steps of one or more methods described and illustrated herein in real time or in batch mode. Also, if desired, one or more computer systems 300 may perform one or more steps of one or more methods described and illustrated herein at different times or in different locations.
[0093] In a particular, non-limiting embodiment, computer system 300 includes a processor 302, memory 304, storage 306, input / output (I / O) interface 308, communication interface 310, and bus 312. Note that while the description and illustrations of this disclosure describe particular computer systems as including a particular number of particular components arranged in a particular arrangement, this disclosure contemplates that any suitable computer system may include any suitable number of any suitable components arranged in any suitable arrangement.
[0094] Additionally, in some embodiments, processor 302 includes hardware for executing instructions, such as instructions comprising a computer program. For example, and without limitation, processor 302 may execute instructions by retrieving (fetching) instructions from an internal register, an internal cache, memory 304, or storage device 306, decoding and executing the retrieved instructions, and then writing one or more results to an internal register, an internal cache, memory 304, or storage device 306. For example, and without limitation, processor 302 may include one or more internal caches for data, instructions, or addresses. This disclosure contemplates processor 302 including any suitable number of any suitable internal caches, as appropriate. By way of example and without limitation, processor 302 may include one or more instruction caches, one or more data caches, and one or more translation lookaside buffers (TLBs). Instructions in an instruction cache may be copies of instructions in memory 304 or storage device 306, allowing processor 302 to retrieve these instructions faster. The data in the data cache may be a copy of data in memory 304 or storage device 306 that processor 302 needs to perform an operation by executing instructions, or the results of previously executed instructions by processor 302 that need to be accessed or written to memory 304 or storage device 306 when processor 302 executes instructions, or other suitable data. The data cache may enable faster read and write operations by processor 302. The TLB may enable faster virtual address translation for processor 302. Without limitation, in some embodiments, processor 302 may include one or more internal registers for data, instructions, or addresses. It is to be noted that this disclosure contemplates processor 302 including any suitable number of any suitable internal registers, as desired.Additionally, if desired, processor 302 may include one or more arithmetic logic units (ALUs), may be a multi-core processor, or may include more than one processor 302. Although particular processors are described and illustrated in the present disclosure, any suitable processor is contemplated by this disclosure.
[0095] In some embodiments, memory 304 includes, but is not limited to, a main memory for storing instructions to be executed by processor 302 or data necessary for processor 302 operation. By way of example, but not limitation, computer system 300 may load instructions into memory 304 from storage device 306 or other sources (e.g., other computer systems 300). Processor 302 may then load the instructions from memory 304 into an internal register or cache. When executing an instruction, processor 302 may retrieve and decode the instruction from the internal register or cache. During or after execution of an instruction, processor 302 may write one or more results (intermediate or final) to an internal register or cache. Processor 302 may then write one or more of such results to memory 304. In some embodiments, but not by way of limitation, processor 302 executes only instructions stored in one or more internal registers or caches or memory 304 (rather than instructions stored elsewhere, such as storage device 306), and processor 302 operates only on data stored in one or more internal registers or caches or memory 304 (rather than data stored elsewhere, such as storage device 306). One or more memory buses (each of which may include an address bus and a data bus) may also couple processor 302 to memory 304. As described below, bus 312 may include one or more memory buses. In certain embodiments, but not by way of limitation, one or more memory management units (MMUs) may be provided between processor 302 and memory 304 to facilitate accesses to memory 304 requested by processor 302. In certain other embodiments, but not by way of limitation, memory 304 includes random access memory (RAM). This RAM may be volatile, if desired. Also, depending on the needs, this RAM may be dynamic RAM (DRAM) or static RAM (SRAM).Further, where appropriate, this RAM may be single-ported or multi-ported RAM. Any suitable RAM is contemplated by this disclosure. Where appropriate, memory 304 may comprise one or more memories 304. Although particular memory components are described and illustrated in this disclosure, any suitable memory is contemplated by this disclosure.
[0096] In some embodiments, storage device 306 includes mass storage for storing data or instructions. Examples of storage device 306 include, but are not limited to, a hard disk drive (HDD), a floppy disk drive, flash memory, an optical disk, a magneto-optical disk, magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more thereof. Where appropriate, storage device 306 may include removable or non-removable media. Where appropriate, storage device 306 may be internal or external to computer system 300. While not limited to, in certain embodiments, storage device 306 is non-volatile solid-state memory. While not limited to, in some embodiments, storage device 306 includes read-only memory (ROM). Where appropriate, this ROM may be a pre-programmed mask ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically re-programmable ROM (EAROM), flash memory, or a combination of two or more thereof. It is contemplated by this disclosure that mass storage device 306 may take any suitable physical form. Where desired, storage device 306 may include one or more storage control units that facilitate communication between processor 302 and storage device 306. Where desired, storage device 306 may also consist of one or more storage devices 306. It is noted that although particular storage devices are described and illustrated in this disclosure, any suitable storage device is contemplated by this disclosure.
[0097] In certain embodiments, I / O interface 308 may comprise hardware, software, or both, and may provide one or more interfaces for communication between computer system 300 and one or more I / O devices. Computer system 300 may include one or more of these I / O devices as needed. One or more of these I / O devices may enable communication between a person and computer system 300. For example, but not limited to, the I / O devices may include a keyboard, keypad, microphone, monitor, mouse, printer, scanner, speaker, still camera, stylus pen, tablet, touch panel, trackball, video camera, other suitable I / O device, or a combination of two or more thereof. The I / O devices may include one or more sensors. This disclosure contemplates any suitable I / O device and any I / O interface 308 suitable for that I / O device. If needed, I / O interface 308 may include one or more device or software drivers that enable processor 302 to drive one or more of these I / O devices. Additionally, if desired, I / O interface 308 may consist of one or more I / O interfaces 308. Although particular I / O interfaces are described and illustrated in the present disclosure, any suitable I / O interface is contemplated by this disclosure.
[0098] In some embodiments, but not by way of limitation, communication interface 310 comprises hardware, software, or both, and provides one or more interfaces for communication (e.g., packet-based communication, etc.) between computer system 300 and one or more other computer systems 300 or one or more networks. For example, but not by way of limitation, communication interface 310 may include a network interface controller (NIC) or network adapter for communication with an Ethernet or other wired-based network, or a wireless NIC (WNIC) or wireless adapter for communication with a wireless network, such as a Wi-Fi network. Any suitable network and any communication interface 310 suitable for that network are contemplated by the present disclosure. For example, but not by way of limitation, computer system 300 may communicate with an ad-hoc network, a personal area network (PAN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), or one or more portions of the Internet, or a combination of two or more thereof. One or more portions of one or more of these networks may be wired or wireless. For example, computer system 300 may communicate with a suitable wireless network, such as a wireless personal area network (WPAN) (e.g., a "BLUETOOTH" WPAN), a Wi-Fi network, a Wi-MAX network, or a cellular network (e.g., a Global System for Mobile Communications (GSM) network), or a combination of two or more thereof. If desired, computer system 300 may include any communication interface 310 suitable for any of these networks. If desired, communication interface 310 may consist of one or more communication interfaces 310.It should be noted that although particular communication interfaces are described and illustrated in the present disclosure, any suitable communication interface is contemplated by this disclosure.
[0099] In certain embodiments, bus 312 may include, but is not limited to, hardware, software, or both that couples components of computer system 300 together. Examples of bus 312 include, but are not limited to, a graphics bus such as an Accelerated Graphics Port (AGP), an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport Bus (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a low-pin-count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCIe) bus, a serial advanced technology attachment (SATA) bus, or a Video Electronics Standards Association local (VESA local) bus, or any combination of two or more thereof. Bus 312 may also be comprised of one or more buses 312, if desired. It should be noted that although particular buses are described and illustrated in this disclosure, any suitable bus or interconnect wiring is contemplated by this disclosure.
[0100] As used herein, the one or more computer-readable non-transitory storage media may include any suitable computer-readable non-transitory storage media, such as one or more semiconductor-based or other integrated circuits (ICs) (e.g., field programmable gate arrays (FPGAs) or application-specific ICs (ASICs)), hard disk drives (HDDs), hybrid hard drives (HHDs), optical disks, optical disk drives (ODDs), magneto-optical disks, magneto-optical drives, floppy diskettes, floppy disk drives (FDDs), magnetic tapes, solid-state drives (SSDs), RAM drives, or secure digital cards or drives, or any suitable combination of two or more thereof, as appropriate. Also, as appropriate, the computer-readable non-transitory storage media may be volatile, non-volatile, or a combination of volatile and non-volatile.
[0101] As used herein, "or" indicates inclusiveness rather than exclusivity, unless expressly stated otherwise or the context dictates otherwise. Thus, as used herein, "A or B" means "A, B, or both," unless expressly stated otherwise or the context dictates otherwise. Furthermore, "and" indicates joint and several, unless expressly stated otherwise or the context dictates otherwise. Thus, as used herein, "A and B" means "A and B, jointly or severally," unless expressly stated otherwise or the context dictates otherwise.
[0102] The scope of the present disclosure encompasses all changes, substitutions, variations, alterations, and modifications of the exemplary embodiments described or illustrated herein that would be understood by a person skilled in the art. The scope of the present disclosure is not limited to the exemplary embodiments described or illustrated herein. Furthermore, in describing and illustrating the present disclosure, each embodiment herein is described as including particular components, elements, features, functions, operations, or steps, but any of these embodiments can include any unordered or ordered combination of any components, elements, features, functions, operations, or steps described or illustrated anywhere in the specification that would be understood by a person skilled in the art. Furthermore, when the appended claims describe a device or system, or a device or system component, as being adapted to, arranged to, capable of, configured to, enabled to, operable to, or operative to perform a particular function, such description encompasses the device, system, component, or particular function, to the extent that it is so adapted, arranged, capable of, configured, enabled, operative, or operative, regardless of whether the device, system, component, or particular function is activated, turned on, unlocked, or not. Furthermore, in the description and illustrations of this disclosure, certain non-limiting embodiments are described as offering certain advantages, but certain non-limiting embodiments may offer none, some, or all of these advantages.
[0103] Furthermore, while method embodiments are presented and described in this disclosure as flowcharts, this is done for illustrative purposes to provide a more complete understanding of the present technology. As such, the methods of the present disclosure are not limited to the steps and logical flow shown herein. Alternative embodiments are contemplated in which the order of various steps is changed, or in which substeps described as part of a broader process are performed as separate steps.
[0104] Additionally, while various embodiments have been described for purposes of this disclosure, the teachings of this disclosure should not be considered limited to such embodiments. Various modifications and variations can be made to the elements and steps described above and still obtain results that do not depart from the scope of the systems and processes described in this disclosure.
[0105] The embodiments disclosed herein are merely examples, and the scope of the present disclosure is not limited thereto. Non-limiting specific embodiments are possible that include all or some of the components, elements, features, functions, acts, or steps of the embodiments disclosed in the above description, as well as non-limiting specific embodiments that do not include any of them. The appended claims relate to methods, storage media, systems, and computer program products. While multiple embodiments are disclosed in the appended claims, features recited in one claim category, such as method claims, may also be claimed in other claim categories, such as system claims. The dependent relationships and references to prior disclosure in the appended claims are merely selected for formality reasons. Therefore, subject matter derived by reference to any preceding claim (especially in the case of multiple dependent relationships) may be claimed, and therefore any combination of claims and their features is disclosed and may be claimed, regardless of the dependent relationships selected in the appended claims. Claimable subject matter includes not only combinations of features recited in the appended claims, but also any other combination of features in the claims, and each feature recited in a claim may be combined with any other feature or combination of features in the claim. Furthermore, any embodiment and feature described and illustrated herein may be claimed as a separate claim and may additionally or alternatively be claimed in any combination with any embodiment or feature described and illustrated herein or recited in the appended claims.
[0106] All patents, patent applications, publications, product descriptions, and protocols referenced herein are hereby incorporated by reference in their entirety. In the event of a conflict in terminology, the present disclosure will control.
[0107] While it will be apparent that the subject matter described herein is well-calculated to achieve the benefits and advantages noted above, the specific embodiments described herein do not limit the scope of the subject matter of the present disclosure. It will be understood that the subject matter of the present disclosure is susceptible to modification, variation, and alteration without departing from the spirit thereof. Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many equivalents to the specific embodiments described herein. Such equivalents are intended to be encompassed by the following claims.
[0108] Various references are cited herein, the contents of all of which are incorporated herein by reference. [Explanation of symbols]
[0109] 100 Architecture 110 Multi-Image / Text Pairs 112 Inspection images 114 Radiology Reports 120 CNN Image Encoder 122 feature maps 130 ViT Multi-Image Encoder 132 Multi-Image Key and Value 140 CLS Embedded Representations 150 Unimodal Text Decoder 152 Text Queries 160 Multimodal Text Decoder 162, 164 Multimodal Text Decoder Output 300 Computer Systems 302 processor 304 memory 306 Storage device 308 Input / Output (I / O) Interface 310 Communication Interface 312 Bus
Claims
1. by one or more computing systems, accessing a plurality of radiographic images of an animal, wherein one or more first radiographic images of the plurality of radiographic images respectively depict the animal from one or more views and one or more second radiographic images of the plurality of radiographic images respectively depict one or more body parts of the animal; identifying one or more disease classifications for the animal based on analysis of the plurality of radiographic images by a machine learning model; generating a diagnostic report for the animal based on the machine learning model, the diagnostic report including the one or more disease classifications and a natural language text radiology report; sending instructions to a user device directing the user device to present the diagnostic report; A method comprising:
2. The method of claim 1 , wherein each of the plurality of radiographic images is formatted as a Digital Imaging and Communications in Medicine (DICOM) image.
3. 10. The method of claim 1, wherein the machine learning model is based on at least one first neural network and at least one second neural network, the at least one first neural network and the at least one second neural network being coupled to each other.
4. generating the diagnostic report, accessing a plurality of reference reports; encoding the plurality of reference reports into a feature space; encoding the plurality of radiographic images into the feature space; determining the diagnostic report based on a similarity search in the feature space; The method of claim 1 , comprising:
5. The method of claim 1 , wherein one of the one or more disease classifications is indicative of a tissue abnormality.
6. 6. The method of claim 5, further comprising the step of classifying the tissue abnormality as at least one of a cardiovascular abnormality, a pulmonary structural abnormality, a mediastinal structural abnormality, a pleural cavity abnormality, or an extrathoracic abnormality.
7. accessing a plurality of training radiographic images associated with a plurality of training radiology reports respectively; training the machine learning model based on the accessed training radiographic images and their respective training radiology reports; The method of claim 1 further comprising:
8. further comprising pre-processing each of the plurality of training radiographic images; The method of claim 7 , wherein the preprocessing comprises one or more of padding, random dilation, random flip, Gaussian blur, and normalization.
9. The method of claim 7 , further comprising applying long document coding to each of the plurality of training radiology reports.
10. performing pre-processing on each of the plurality of training radiology reports; The method of claim 7 , wherein the preprocessing comprises one or more of tokenization, padding, adding classification tokens, and applying an attention mask.
11. The method of claim 1 , wherein the machine learning models comprise an image encoder, a multi-image encoder, a text decoder, and a multi-modal decoder.
12. generating, by the image encoder, a feature map based on the plurality of radiographic images; generating, by the multi-image encoder, one or more multi-image keys and values based on the feature map; generating, by the multimodal decoder, a natural language text radiology report based on the one or more multi-image keys and values and initial tokens; The method of claim 11 further comprising:
13. The method of claim 1 , wherein the diagnostic report further includes one or more of the plurality of radiographic images.
14. One or more computer-readable non-transitory storage media comprising software, said software, when executed, accessing a plurality of radiographic images of an animal, wherein one or more first radiographic images of the plurality of radiographic images respectively depict the animal from one or more views and one or more second radiographic images of the plurality of radiographic images respectively depict one or more body parts of the animal; identifying one or more disease classifications for the animal based on analysis of the plurality of radiographic images by a machine learning model; generating a diagnostic report for the animal based on the machine learning model, the diagnostic report including the one or more disease classifications and a natural language text radiology report; sending instructions to a user device directing the user device to present the diagnostic report; A medium configured to be operable to perform the above.
15. 15. The medium of claim 14, wherein each of the plurality of radiographic images is formatted as a Digital Imaging and Communications in Medicine (DICOM) image.
16. 15. The medium of claim 14, wherein the machine learning model is based on at least one first neural network and at least one second neural network, the at least one first neural network and the at least one second neural network being coupled to each other.
17. generating the diagnostic report, accessing a plurality of reference reports; encoding the plurality of reference reports into a feature space; encoding the plurality of radiographic images into the feature space; determining the diagnostic report based on a similarity search in the feature space; The medium of claim 14, comprising:
18. The medium of claim 14 , wherein one of the one or more disease classifications is indicative of a tissue abnormality.
19. 20. The medium of claim 18, wherein the software, upon execution, is further configured to perform the step of identifying the tissue abnormality as at least one of a cardiovascular abnormality, a pulmonary structural abnormality, a mediastinal structural abnormality, a pleural cavity abnormality, or an extrathoracic abnormality.
20. When executed, the software: accessing a plurality of training radiographic images associated with a plurality of training radiology reports respectively; training the machine learning model based on the accessed training radiographic images and their respective training radiology reports.
21. wherein the software, upon execution, is further configured to perform pre-processing on each of the plurality of training radiographic images; 21. The medium of claim 20, wherein the preprocessing includes one or more of padding, random dilation, random flip, Gaussian blur, and normalization.
22. 21. The medium of claim 20, wherein the software, upon execution, is further configured to perform the step of applying long document coding to each of the plurality of training radiology reports.
23. wherein the software, upon execution, is further configured to perform a pre-processing step on each of the plurality of training radiology reports; 21. The medium of claim 20, wherein the preprocessing includes one or more of tokenizing, padding, adding classification tokens, and applying an attention mask.
24. The medium of claim 14 , wherein the machine learning models comprise an image encoder, a multi-image encoder, a text decoder, and a multi-modal decoder.
25. When executed, the software: generating, by the image encoder, a feature map based on the plurality of radiographic images; generating, by the multi-image encoder, one or more multi-image keys and values based on the feature map; generating, by the multimodal decoder, a natural language text radiology report based on the one or more multi-image keys and values and initial tokens; 25. The medium of claim 24, further configured to:
26. The medium of claim 14 , wherein the diagnostic report further includes one or more of the plurality of radiographic images.
27. 1. A system comprising: one or more processors; and a non-transitory memory coupled to the processors and containing instructions executable by the processors, Executing the instructions causes the processor to: accessing a plurality of radiographic images of an animal, wherein one or more first radiographic images of the plurality of radiographic images respectively depict the animal from one or more views and one or more second radiographic images of the plurality of radiographic images respectively depict one or more body parts of the animal; identifying one or more disease classifications for the animal based on analysis of the plurality of radiographic images by a machine learning model; generating a diagnostic report for the animal based on the machine learning model, the diagnostic report including the one or more disease classifications and a natural language text radiology report; sending instructions to a user device directing the user device to present the diagnostic report; 10. A system configured to operate as follows:
28. 28. The system of claim 27, wherein each of the plurality of radiographic images is formatted as a Digital Imaging and Communications in Medicine (DICOM) image.
29. 28. The system of claim 27, wherein the machine learning model is based on at least one first neural network and at least one second neural network, the at least one first neural network and the at least one second neural network being coupled to each other.
30. generating the diagnostic report, accessing a plurality of reference reports; encoding the plurality of reference reports into a feature space; encoding the plurality of radiographic images into the feature space; determining the diagnostic report based on a similarity search in the feature space; 28. The system of claim 27, comprising:
31. 28. The system of claim 27, wherein one of the one or more disease classifications is indicative of a tissue abnormality.
32. 32. The system of claim 31, wherein the processor is further configured to perform the step of identifying the tissue abnormality as at least one of a cardiovascular abnormality, a pulmonary structural abnormality, a mediastinal structural abnormality, a pleural cavity abnormality, or an extrathoracic abnormality by executing the instructions.
33. Execution of the instructions causes the processor to: accessing a plurality of training radiographic images associated with a plurality of training radiology reports respectively; 28. The system of claim 27, further configured to: train the machine learning model based on the accessed training radiographic images and their respective training radiology reports.
34. The processor is configured to be operable to execute the instructions to further perform a pre-processing step on each of the plurality of training radiographic images; 34. The system of claim 33, wherein the preprocessing includes one or more of padding, random dilation, random flip, Gaussian blur, and normalization.
35. 34. The system of claim 33, wherein executing the instructions configures the processor to be further operable to apply long document coding to each of the plurality of training radiology reports.
36. and wherein the processor is configured to be operable to execute the instructions to further perform a pre-processing step on each of the plurality of training radiology reports; 34. The system of claim 33, wherein the preprocessing includes one or more of tokenizing, padding, adding classification tokens, and applying an attention mask.
37. 28. The system of claim 27, wherein the machine learning models comprise an image encoder, a multi-image encoder, a text decoder, and a multi-modal decoder.
38. Execution of the instructions causes the processor to: generating, by the image encoder, a feature map based on the plurality of radiographic images; generating, by the multi-image encoder, one or more multi-image keys and values based on the feature map; generating, by the multimodal decoder, a natural language text radiology report based on the one or more multi-image keys and values and initial tokens; 38. The system of claim 37, further configured to:
39. 28. The system of claim 27, wherein the diagnostic report further includes one or more of the plurality of radiographic images.
Citation Information
Patent Citations
Diagnosis support apparatus, information processing method, diagnosis support system and program
JP2017191469A
System and device for analyzing anatomical image
JP2019195627A
Health management device, method for operating health management device, and program for operating health management device
JP2020166682A
Diagnosis assistance device, diagnosis assistance method, and diagnosis assistance program
JP2021015525A
Diagnosis support device, diagnosis support method, and diagnosis support program
JP2021058270A