System and method for detecting abnormalities from radiographic images of pets
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-07-07
- Publication Date
- 2026-08-14
Smart Images

Figure 0007905518000010 
Figure 0007905518000011 
Figure 0007905518000012
Abstract
Description
Priority
[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 358905 filed on 7 July 2022, and all disclosures of this Provisional Application are incorporated herein by reference. [Technical Field]
[0002] This disclosure generally relates to the use of one or more machine learning models or tools to evaluate radiographic images of pets or animals. [Background technology]
[0003] An increasing number of veterinarians are using image-based diagnostic techniques, such as X-rays, to diagnose and identify health problems in animals and pets. However, there are fewer than 1,100 veterinary radiologists worldwide. Consequently, many veterinarians are unable to take advantage of the benefits offered by image-based diagnostic techniques. Furthermore, even for veterinarians trained in radiology, reviewing medical images can be a time-consuming and cumbersome task. The fact that animal and pet radiographs may be oriented incorrectly or have missing or incorrect lateral markers further complicates these issues. [Overview of the Initiative] [Problems that the invention aims to solve]
[0004] Therefore, there is a need for a system that can automate the processing and interpretation of diagnostic images of pets and provide clinically reliable results to both veterinarians trained in radiology and those without. [Means for solving the problem]
[0005] In certain embodiments, however limited, this disclosure provides systems and methods for training and using machine learning models to process, interpret, and analyze radiographic images of animals or pets. The image format may be any format used for disease diagnosis, such as an image display format like DICOM (Digital Imaging and Communications in Medicine). In certain embodiments, the radiographic images may be associated with a radiology report. Conventional radiographic image analysis has involved training separate models, such as training a text-only model (e.g., a natural language processing model) based on the radiology report and training an image-only model based on the radiographic images. Furthermore, at the time of implementation, the text-only model and the image-only model were sometimes deployed as separate entities that could be used independently (in principle). Unlike these conventional methods, the embodiments disclosed herein allow for the training of a text-image integrated model, which can be used to directly generate and verify radiology reports based on radiographic images. In one embodiment, the text-image integrated model can also be programmed to detect anomalies from radiographic images of animals or pets.
[0006] In one embodiment, this disclosure provides a system and method for automatically detecting anomalies from radiographic images of animals or pets. In various embodiments, the analysis and anomaly detection of captured, collected, and received images can be performed using one or more machine learning models or tools. In some embodiments, the machine learning model may include one or more neural networks. The neural network may be a convolutional neural network (CNN), but is not limited to these. In one embodiment, anomaly detection can indicate, for example, whether a certain tissue is healthy or abnormal. In one embodiment, tissue classified as abnormal may be further classified into, for example, cardiovascular anomalies, lung structure anomalies, mediastinal structure anomalies, pleural cavity anomalies, extrathoracic anomalies, or any combination thereof.
[0007] In some embodiments, the Disclosure provides a method for detecting anomalies from radiographic images of an animal or pet using one or more computing systems. The method includes the steps of: accessing a plurality of radiographic images of an animal, wherein one or more first radiographic images of the plurality of radiographic images depict the animal from one or more views, and one or more second radiographic images of the plurality of radiographic images depict one or more body parts of the animal; identifying one or more disease classifications relating to the animal based on an analysis of the plurality of radiographic images by a machine learning model; generating a diagnostic report relating to the animal, including one or more disease classifications and a radiological report in natural language text, based on the machine learning model; and sending a command to a user device instructing it to present the diagnostic report.
[0008] In one embodiment, each of the multiple radiographic images is formatted as a DICOM (Digital Imaging and Communications in Medicine) image.
[0009] In one embodiment, the machine learning model is based on at least one first neural network and at least one second neural network, wherein the at least one first neural network and at least one second neural network are coupled to each other.
[0010] In one embodiment, the step of generating a diagnostic report includes the steps of accessing a plurality of reference reports, encoding the plurality of reference reports into a feature space, encoding a plurality of radiographic images into a feature space, and determining the diagnostic report based on similarity search in the feature space.
[0011] In one embodiment, one of one or more disease classifications exhibits a tissue abnormality.
[0012] In one embodiment, the method further includes the step of identifying the tissue abnormality as at least one of the following: cardiovascular abnormalities, pulmonary structural abnormalities, mediastinal structural abnormalities, pleural cavity abnormalities, and extrathoracic abnormalities.
[0013] In one embodiment, the method further includes the steps of accessing a plurality of training radiograph images associated with a plurality of training radiology reports, and training a machine learning model based on the accessed training radiograph images and their respective training radiology reports.
[0014] In one embodiment, the method further includes the step of preprocessing each of a plurality of training radiographic images, the preprocessing including one or more of padding, random expansion, random flipping, Gaussian blurring, and normalization.
[0015] In one embodiment, the method further includes the step of applying long document encoding to each of a plurality of training radiology reports.
[0016] In one embodiment, the method further includes preprocessing each of a plurality of training radiology reports, the preprocessing including one or more of tokenization, padding, adding classification tokens, and applying an attention mask.
[0017] In one embodiment, the machine learning model includes an image encoder, a multi-image encoder, a text decoder, and a multimodal decoder.
[0018] In one embodiment, the method further includes generating a feature map based on a plurality of radiological images by an image encoder, generating one or more multi-image keys and values based on the feature map by a multi-image encoder, and generating a radiology report of natural language text based on the one or more multi-image keys and values and a start token by a multimodal decoder.
[0019] In one embodiment, the diagnostic report further includes one or more of the plurality of radiological images.
[0020] In various embodiments, the present disclosure provides one or more computer-readable non-transitory storage media that, when executed by one or more processors, are operably configured to perform one or more of the plurality of methods provided by the present disclosure.
[0021] In various embodiments, the present disclosure provides a system comprising one or more processors and one or more computer-readable non-transitory storage media coupled to one or more of the processors and including instructions that, when executed by one or more of the processors, are operably configured to cause the system to perform one or more of the plurality of methods provided by the present disclosure.
[0022] The embodiments disclosed in this specification are merely illustrative, and the scope of the present disclosure is not limited thereto. Specific embodiments that are not intended to be limiting are possible, including all or some of the components, elements, features, functions, operations, or steps of the embodiments disclosed herein, as well as specific embodiments that do not include any of them. The embodiments according to the present invention are disclosed in the appended claims, particularly in the method claims. Note that the dependencies and references to the preceding descriptions in the appended claims are merely selected for formal reasons. Therefore, since it is possible to claim the subject matter obtained by optionally referring to any preceding claim (especially when there are multiple dependencies), for the claims and their features, any combination is disclosed regardless of the dependencies selected in the appended claims, and it is possible to claim this. The subject matter that can be claimed includes not only the combinations of features described in the appended claims but also any other combinations of features in the claims, and each feature mentioned in the claims can be combined with any other feature or combination of other features in the claims. Furthermore, any embodiment and feature described and illustrated herein can be claimed as a separate claim, and in addition to or instead of this, it can also be claimed in any combination with any embodiment or feature described and illustrated herein or any feature described in the appended claims.
Brief Description of the Drawings
[0023] [Figure 1] Figure showing an exemplary architecture of a contrast radiology medical captioning model [Figure 2] Figure showing an exemplary method for detecting abnormalities from radiographic images of animals [Figure 3] Figure showing an exemplary computer system or device used to assist in detecting abnormalities from radiographic images of animals
Modes for Carrying Out the Invention
[0024] The terms used herein generally have their common meanings in the art, within the context of this disclosure and in the context in which they are used. However, certain terms are defined below or elsewhere in this specification to provide further guidance in describing the structure and methods of this disclosure and how they are manufactured and used.
[0025] In this specification and in the claims, the singular forms "a," "an," and "the" shall also include references to the corresponding plural forms unless it is clearly evident from the context that the plural form is not included.
[0026] In this specification, the terms “comprise,” “comprising,” and “incorporate” or other variations thereof have the meaning of “incorporate non-exclusively.” Thus, when a process, method, article, system, or apparatus “comprises” (“incorporates,” “incorporates”) the enumerated elements, it may include not only the enumerated elements but also other elements not explicitly enumerated, as well as elements specific to such process, method, article, or apparatus.
[0027] In this disclosure, the terms “animal” or “pet” refer to, but are not limited to, domesticated animals such as dogs, cats, horses, cows, ferrets, rabbits, pigs, rats, mice, gerbils, hamsters, and goats. Examples of pets, but not limited to, include dogs and cats. In this disclosure, the terms “animal” or “pet” may also refer to wild animals. Examples of wild animals include, but are not limited to, bison, moose, deer, ducks, birds, and fish.
[0028] In this specification, a “feature” of an image or slide can be identified based on one or more measurable characteristics of the image or slide. For example, a feature may be a blemish or dark spot in the image, or tissue of varying sizes, shapes, and light intensity levels.
[0029] Where the detailed description herein uses terms such as “embodiment,” “one embodiment,” “in various embodiments,” “certain embodiments,” “some embodiments,” “other embodiments,” or “certain other embodiments,” it indicates that one or more embodiments described in that section may have certain features, structures, or characteristics, but not all embodiments necessarily have those features, structures, or characteristics. Furthermore, these expressions do not necessarily refer to the same embodiment. Moreover, where a description of a particular embodiment mentions a particular feature, structure, or characteristic, it is considered to be within the knowledge of a person skilled in the art that this will also affect similar features, structures, or characteristics in other embodiments, whether or not it is explicitly stated so. A person skilled in the art will be able to see how the disclosure should be implemented in alternative embodiments by reading the description herein.
[0030] In this specification, the term "device" refers to a computing system or a mobile device. For example, the term "device" may include a smartphone, a tablet computer, or a notebook computer. More specifically, a computing system may include functions to determine the location, direction, or orientation of the computing system itself, such as a GPS receiver, compass, gyroscope, or accelerometer. A client device may further include wireless communication functions such as Bluetooth® communication, near-field communication (NFC), or infrared (IR) communication, or functions to communicate with a wireless local area network (WLAN) or a cellular telephone network. Such a device may further include one or more cameras, scanners, touchscreens, microphones, or speakers. A client device may also run software applications such as games, web browsers, or social networking applications. Examples of client devices include user devices, smartphones, tablet computers, notebook computers, desktop computers, or smartwatches.
[0031] Exemplary processes and embodiments may be implemented or executed by a computing system or client device via a mobile application and associated graphical user interface ("UX" or "GUI"). In certain embodiments, the computing system or client device may be a mobile computing system such as a smartphone, tablet computer, or notebook computer. This mobile computing system may include a GPS receiver, compass, gyroscope, or accelerometer, or other function to determine the location, orientation, or direction of the computing system itself. Such a device may also include wireless communication functions such as "Bluetooth," near-field communication (NFC), or infrared (IR) communication, or communication functions with a wireless local area network (WLAN), 3G, 4G, LTE, LTE-A, 5G, Internet of Things, or cellular telephone network. Such a device may further include one or more cameras, scanners, touchscreens, microphones, or speakers. The mobile computing system may also run software applications such as games, web browsers, or social networking applications. Social networking applications allow users to connect with, communicate with, and share information with other users on social networks.
[0032] The terms used herein generally have their common meanings in the art, within the context of this disclosure and in the context in which they are used. However, certain terms are defined below or elsewhere in this specification to provide further guidance in describing the structure and methods of this disclosure and how they are manufactured and used.
[0033] In recent years, semi-supervised multimodal artificial intelligence (AI) models have achieved cutting-edge results in various downstream tasks. Embodiments disclosed herein leverage the effectiveness of these methods for disease classification and report generation in the field of veterinary radiology. Specifically, this specification discloses a contrastive radiology captioning model. The architecture of the contrastive radiology captioning model utilizes control loss and captioning loss to align X-ray images and reports at both global and local levels. This architecture allows for the alignment of multiple X-ray images into a single report when creating diagnostic reports using multiple different views or body parts. Experiments have shown that for some radiological findings, this architecture provides a significant performance improvement compared to supervised training methods using other labeling techniques. Ablation experiments were also conducted to demonstrate the importance of each choice in the architecture design. The text generation capabilities of the contrastive radiology captioning model highlight the future potential of radiology report generation using multimodal large-scale language models. A contrastive radiology captioning model can be a powerful architecture for training on large, unlabeled datasets with multiple image / text pair inputs.
[0034] AI systems using supervised learning methods can be used to assist veterinary radiologists in interpreting X-ray images. However, this method has sometimes relied on the time-consuming and resource-intensive process of manually labeling X-ray images with disease classifications. In recent years, however, semi-supervised multimodal methods have achieved state-of-the-art performance across various downstream tasks and have demonstrated great success. This method reduces the need to label data by utilizing text directly as ground truth labels. There tends to be a positive correlation between dataset size and model performance. Therefore, semi-supervised methods can improve model performance by allowing training with large, unlabeled datasets. This advancement has significant implications for the field of radiology, as it enables training models with the vast amount of historical reports that have been routinely created alongside X-ray images.
[0035] Furthermore, these state-of-the-art models have demonstrated the advantages that multimodal methods bring to the performance of unimodal models. The reflective method achieves zero-shot functionality by aligning similar texts and images by learning a common embedding space for images and texts. It has also been shown that optimizing the generation loss of cross-modal alignment improves the model's ability to learn fine-grained local feature representations. Accordingly, embodiments disclosed herein disclose methods for training radiographic image-text pairs for disease classification and text generation, utilizing both reflective and generative methods. In certain embodiments, but not limited to these, the disclosure provides a technique for automatically detecting anomalies from radiographic images of animals or pets. One or more radiographic images may be in DICOM (Digital Imaging and Communications in Medicine) format. Upon receiving radiographic images, anomalies can be identified from the images by performing image analysis using a trained machine learning model or tool, such as a neural network model. In some embodiments, the technique may use a vision encoder and a text unimodal decoder and a text multimodal decoder that are separated from each other.
[0036] However, a particular challenge in model development in the field of radiology is that a single patient's report typically references multiple images. This is because multiple X-ray images (e.g., images of various body parts and views) are usually taken during a patient's examination. Recent research has highlighted the importance of including relevant images in order to achieve cross-modal alignment between images and reports. Embodiments disclosed herein demonstrate that incorporating images from the patient's past visits reduces report ambiguity caused by the lack of contextual information in the images, thereby improving model performance. Previous studies have also suggested that considering multi-image views using a CNN-ViT architecture improves performance in multi-label classification tasks in radiology. Accordingly, the method disclosed herein similarly uses a hybrid CNN-transformer architecture as a vision encoder to facilitate multi-image embedding representations.
[0037] Some exemplary prior art related to embodiments disclosed herein include RapidRead and StudyFormer. RapidRead is an AI veterinary radiology system already in use. This system can perform disease classification using an ensemble of CNN models and generate assessments using an expert system. The method of this disclosure can be compared to the model of this system. On the other hand, the StudyFormer model can generate a patient examination level embedding representation using a single-image CNN encoder model and a multi-image ViT encoder model. The architecture of the contrasting radiology captioning model can be used as a vision encoder by making some structural modifications to the StudyFormer architecture.
[0038] Embodiments disclosed herein disclose a contrastive radiology captioning model that can perform visual language processing in the field of radiology based on a self-supervised framework. In certain embodiments, the contrastive radiology captioning model may be based on one or more neural networks. For example, the neural network may be based on a convolutional neural network, a transformer-based network, or an MLP-Mixer. In some embodiments, the architecture of the contrastive radiology captioning model may comprise a hybrid CNN-ViT vision encoder, a text decoder, and a multimodal decoder.
[0039] In certain embodiments, however limited, the training of a contrastive radiology captioning model can be carried out by training at least two coupled neural networks in a single unit, comprising at least one first network for radiographic images and at least one second network for radiology reports. The at least one first network for radiographic images can be considered an image encoder, and the at least one second network for radiology reports can be considered a text encoder. The networks can be based on any suitable architecture, such as Resnet50.
[0040] In some embodiments, the integrated training of the first and second networks can be performed based on multiple pairs of radiographic image-radiological report pairs. A contrastive radiological captioning model can be trained to predict correct pairs of radiographic image-radiological report pairs from training examples. In some embodiments, this training may include learning a multimodal embedding space by training the first and second networks in conjunction to maximize the cosine similarity of correct pair radiographic image embedding representations and radiological report embedding representations, and minimize the cosine similarity of incorrect pair embedding representations.
[0041] While some neural networks can be trained on all their learned weights for each input / output pair, CNNs can convolve trainable, fixed-length kernels or filters along their input. In other words, CNNs can learn to recognize small primitive features (low-level) and combine them in complex ways (high-level). In certain embodiments, CNNs can be supervised, semi-supervised, or unsupervised.
[0042] In certain embodiments, pooling, padding, stride adjustment, or a combination thereof, can be used to reduce the output size of a CNN in the dimension in which convolutions are performed, thereby reducing computational costs and the likelihood of overtraining. Stride adjustment can represent the width or number of steps to slide the filter window, and padding may include the process of buffering data by filling areas of the data with zeros before or after stride adjustment. In one embodiment, for example, pooling may include the process of simplifying the information collected in any layer, such as a convolutional layer, to create a condensed version of the information contained in that layer.
[0043] In some examples, a region-based CNN (RCNN) or a one-dimensional (1D) CNN can be used. The RCNN includes the step of detecting anomalies by using selective search to identify one or more regions of interest within an image and extracting CNN features individually from each region of interest. Types of RCNNs used in one or more embodiments include Fast RCNN, Faster RCNN, or Mask RCNN. In other examples, a one-dimensional CNN can process fixed-length time-series segments generated by a sliding window. Such a one-dimensional CNN can be run in a many-to-one configuration, concatenating outputs in the final layer of the CNN using pooling and stride adjustment. A fully connected layer can then be used to generate detections in one or more time steps.
[0044] In some embodiments, one or more CNN models can be combined with one or more LSTM models. The combined model (composite model) may include a stack of four non-stride-adjusted CNN layers topped with two LSTM layers and a softmax classifier. The softmax classifier can normalize a probability distribution containing multiple probabilities proportional to the input exponentially. For example, because the input signal to the CNN is not padded, the time series is shortened by a few samples in each CNN layer, even though the CNN layers are not stride-adjusted. Since the LSTM layers are unidirectional, the softmax classification corresponding to the final output of the LSTM can be used not only when reconstructing the output time series from sliding window segments, but also in training and evaluation. However, this composite model can operate in a many-to-one configuration.
[0045] Figure 1 shows an exemplary architecture 100 of a contrastive radiology captioning model. In certain embodiments, the contrastive radiology captioning model can be trained using multiple radiographic images and their associated radiology reports, i.e., multi-image / text pairs 110. The multi-image / text pairs 110 may include examination images 112 and their corresponding radiology reports 114. In addition, radiographic images may be, for example, radiographs, CT scans, etc. Radiology reports may contain longer, unstructured text descriptions of anomalies compared to tags and labels. For example, a radiology report may contain the description "This dog's heart is enlarged" as a description corresponding to the tag / label "cardiomegaly". Therefore, training a machine learning model based on radiology reports containing long, unstructured text descriptions may be more difficult than conventional training based on tags and labels.
[0046] In certain embodiments, the vision encoder may be a hybrid CNN transformer based on the StudyFormer architecture. The CNN architecture used may be an Efficient-Net model pre-trained with multi-label classification on multiple X-ray images of a single view. Images from an examination are first passed individually through the CNN image encoder 120 to obtain a feature map for each image with dimensions of 2048 × 10 × 10. Next, the feature maps of all images from the examination can be concatenated to form a feature map 122 with dimensions of 2048 × 50 × 10. This feature map 122 is then passed through the ViT multi-image encoder 130, which outputs a vector representation of all images from the examination with dimensions of 501 × 768. This vector representation may include a CLS embedding representation 140a. In some embodiments, the ViT multi-image encoder 130 may be based on a patch size of 1, a depth of 12, 12 attention heads, a 2048-dimensional multi-layer perceptron (MLP), and an output dimension of 500 × 768.
[0047] In some embodiments, both the unimodal and multimodal text decoders can be small, pre-trained generative pretrained Transformer 2 (GPT2) models. The output dimension of the embedded unimodal text decoder 150 can be [513,768]. This output can include the CLS embedding representation 140b. The contrast loss can be calculated using the CLS embedding representation 140 from the unimodal model.
[0048] The output from the unimodal text decoder 150 can be used as a text query 152 in the multimodal text decoder 160 for the cross-attention mechanism. The output from the ViT multi-image encoder 130 can be used as a multi-image key and value 132. One output 162 of the multimodal text decoder 160 can include a probability distribution on the GPT2 corpus for each position in the text sequence. Another output 164 of the multimodal text decoder 160 can include a caption loss that can be calculated using tokenized ground truth text labels and predicted text. The dimensionality of this output can be [50257,512].
[0049] In some embodiments, both GPT2 models can be trained using low-rank adaptations of large-scale language models. Low-rank adaptations of large-scale language models can utilize low-rank decompositions to learn low-rank matrices for the attention layer. These matrices can represent the change in weights from the original GPT2 to the new task. In this disclosure, for example, the rank of the low-rank adaptation of the large-scale language model is set to 8.
[0050] In certain embodiments, preprocessing of radiographic images and radiographic reports may be required before training a comparative radiographic captioning model. For example, radiographic images may be large in size, such as up to 456 x 456 pixels. All images can be resized to 300 x 300. The maximum number of images per examination can also be limited to five. In one embodiment, the use of image cropping can be avoided by padding the radiographic images. For example, examinations with fewer than five images can be padded to meet the [5,3,300,300] shape requirement. Transformations that can be applied to training images include square padding, random augmentation, random flip, and Gaussian blur. Normalization of both the mean and standard deviation to [0.5,0.5,0.5] can be applied to all images.
[0051] In some embodiments, the preprocessing of radiology reports can be based on long-document coding rather than the commonly used short-document coding. Alternatively, in alternative embodiments, the model that preprocesses radiology reports can be trained on examination notes. For example, radiology reports can be processed using a GPT2 tokenizer with a maximum token length of 512. To ensure consistency, text with fewer than 512 tokens can be padded with end-of-sequence (EOS) tokens. Furthermore, classification tokens (CLS) can be added to each tokenized text to facilitate comparative learning. An attention mask can be applied to the padded tokens.
[0052] Table 1 is a list of definitions for the symbols used in each formula in the various embodiments disclosed herein.
[0053]
Table 1
[0054] Although not limited thereto, in certain embodiments, in the embedding space, discriminative features can be learned using a contrastive loss by strengthening the model to minimize the distance between similar instances and maximize the distance between dissimilar instances.
[0055] The caption loss can be used to optimize the model to learn to generate a report that describes the input image. Given an image I and its corresponding ground-truth caption C = c1c2... c T (where T is the length of the caption), the caption loss function can be defined as the negative log-likelihood of the correct word sequence.
[0056]
Number
[0057] Here, p(c t |I, c 1:t=1 ) represents the probability of generating the correct word c 1:t=1 at time step t given the image I and the previous word c[[ID=3�]] t of the word ct.
[0058] And the total loss L combined is defined as follows.
[0059] First, a similarity matrix S can be calculated where each component S ij is the dot product of r i and i j scaled by the temperature parameter θ.
[0060]
Number
[0061] Next, using label L as the target, we obtain the similarity matrix S and its transpose matrix S. T The average of the cross-entropy losses calculated for this is the control loss L. contrast It can be calculated as follows.
[0062]
number
[0063] And thirdly, the predicted report r pred and true report true The cross-entropy loss calculated for this is the caption loss L caption It can be calculated as follows.
[0064]
number
[0065] Finally, weight w c scaled contrast loss and weights w cap The sum of the scaled caption losses can be taken as the total loss.
[0066]
number
[0067] In some embodiments, the training data consisted of 3,200,173 image / text pairs, representing 755,263 examinations. The validation dataset, on the other hand, consisted of 50,000 image / text pairs, representing 10,446 examinations. A radiology report consists of findings and evaluations. However, in the embodiments disclosed herein, the model was trained on findings only.
[0068] In some embodiments, the contrastive radiology captioning model was further refined using a fine-tuning dataset. The training data for the fine-tuning dataset consisted of 594,449 images and labels, representing 145,486 examinations. The validation dataset, on the other hand, consisted of 10,005 images and labels representing 2,253 examinations. Each image has 41 labels indicating whether or not there is a finding in the image. Each label is unique to a particular X-ray view. Therefore, the label for a given examination was determined by identifying the maximum value of each label across all images of that examination. Thus, if any of the examination images show a finding, the patient is considered to have a lesion.
[0069] To validate the model's performance, the vision encoder was fine-tuned on a disease classification task. This was achieved by adding a classification layer with 41 outputs to the vision encoder. The model was trained using binary cross-entropy loss. Several experiments, including ablation experiments, were performed using this method to understand the impact of each part of the design of the contrastive radiology captioning model on classification performance.
[0070] The contrastive radiology captioning model was trained on a single GPU over two weeks. Training continued for 11 epochs, stopping when the validation loss performance plateaued. Meanwhile, the weighted StudyFormer vision encoder for the contrastive radiology captioning model was fine-tuned on a single GPU over five days. This fine-tuning was also stopped when the average accuracy score stopped improving. This required 50 epochs.
[0071] In certain embodiments, however limited, a contrast radiology captioning model can be used to detect anomalies from any new radiographic image. Furthermore, the contrast radiology captioning model can not only predict tags and labels for one or more input radiographic images, but also generate diagnostic reports for such images. For example, instead of predicting "cardiomegaly" for one or more radiographic images, the contrast radiology captioning model can generate a diagnostic report that includes textual descriptions such as "This dog's heart is enlarged." In one embodiment, the following steps can be taken to enable the contrast radiology captioning model to generate a diagnostic report. First, the contrast radiology captioning model can encode a reference diagnostic report into a feature space. Next, the contrast radiology captioning model can encode the input radiographic image into this shared feature space and perform a similarity search. Furthermore, the contrast radiology captioning model can select the reference diagnostic report that is closest to the input radiographic image as the output diagnostic report.
[0072] In certain embodiments, however, the comparative radiology captioning model may not only detect anomalies but also identify more detailed information about the detected anomalies. For example, after detecting an anomaly, the comparative radiology captioning model may further identify a description of the anomaly's location, size, or severity.
[0073] In certain embodiments, the comparative radiology captioning model can be used for a variety of clinical or medical applications. For example, a veterinarian or veterinary assistant can take radiographic images of a pet, and then process these images using the comparative radiology captioning model. This processing allows the images to be classified as either "normal" or "abnormal." If an image is classified as "abnormal," it can be further classified into at least one of the following: cardiovascular abnormalities, pulmonary structural abnormalities, mediastinal structural abnormalities, pleural cavity abnormalities, or extrathoracic abnormalities. Further analysis of the image can identify descriptions of the location, size, or severity of the abnormality. In some embodiments, the image can also be subdivided into subcategories. For example, subcategories of pleural cavity abnormalities may include pleural effusion, pneumothorax, and / or pleural masses. Similarly, further analysis of the image can identify descriptions of the location, size, or severity of these subcategories. Subsequently, the image can be displayed to the user along with the identified anomaly classifications and subcategories. The image can then be displayed on the user's associated screen or computing device.
[0074] In one embodiment, by using a comparative radiology captioning model and the images obtained therefrom, it is possible to form the basis for services that provide radiologists with on-demand second opinions, provide veterinary hospitals with emergency assessments of radiological images, and improve efficiency and productivity by allowing radiologists to focus on the pet itself rather than the images.
[0075] In certain embodiments disclosed herein, experiments were conducted to verify the effectiveness of a comparative radiologic captioning model. The comparative radiologic captioning model was evaluated by comparing its performance with the current model ensemble employed in the RapidRead system. Performance was measured using ROCAUC and mean precision as indicators. In summary, the comparative radiologic captioning model showed higher ROCAUC for 15 findings and higher mean precision for 10 findings. Table 2 shows the results for findings in which the comparative model outperformed the current ensemble in at least one indicator.
[0076] [Table 2]
[0077] Furthermore, to evaluate the impact of multimodal training on StudyFormer's performance, a StudyFormer model with Image-Net weights was also fine-tuned using the same data. Next, the mean precision score and ROCAUC score of this model were compared to StudyFormer trained with a contrasting radiology captioning model. The comparison showed significant differences in performance for most findings across all metrics.
[0078] Furthermore, to understand the impact of generative learning on the model's classification performance, we also conducted experiments without the multimodal text decoder (and therefore using a contrasting radiology model without a captioner). In other words, the weights of the image encoder and text decoder were optimized using only the contrasting loss. The results of this experiment showed improved performance for most findings compared to the ImageNet version of StudyFormer, but significantly worse performance for most findings compared to StudyFormer trained with the contrasting radiology captioning model. Table 3 shows a comparison of the average accuracy results of the contrasting radiology captioning model and each ablation experiment.
[0079] [Table 3]
[0080] Furthermore, the text generation capability of a contrasting radiology captioning model was tested using previously unseen test data. First, a multi-image X-ray embedding representation was generated using the StudyFormer vision encoder. This embedding representation was used as the key and value for a multimodal decoder. A sentence-first token was used as the initial query. Then, these keys, values, and queries were used in the multimodal decoder to perform autoregressive generation of a text column. In this case, this text column would represent the diagnostic report. This text was compared to a human-written report. This comparison demonstrated that the model could produce a report that was very close to a human-written report based on image decoding. On the other hand, it was also shown that the model generated a considerable amount of misinformation through hallucination that was not present in the human-written text. The degree of agreement between the model-generated text and the human-written text varied significantly depending on the examination.
[0081] In an alternative embodiment, a machine learning model configured to detect anomalies from radiographic images of animals or pets may be trained based on a pre-trained Resnet50 architecture instead of the architecture disclosed in Figure 1. In the embodiments disclosed herein, further experiments were conducted using a pre-trained model based on the Resnet50 architecture. Image features were generated for 72,105 training images and 10,477 test images using Resnet50 from the OpenAI CLIP library (i.e., a publicly available library). Logistic regression was trained with the features of the training images and radiology reports, and then tested with the features and labels of the test images. The ROC-AUC was then calculated for each of the 39 labels. The mean of all 39 ROC-AUCs for the pre-trained machine learning model was 0.7761819903717145. In contrast, the mean of all 39 ROC-AUCs for OpenAI Clip (i.e., the latest method) was 0.7613616312748511. Table 4 shows a comparison of the ROC-AUC of the trained machine learning model and OpenAI Clip for each of the 39 labels. The comparison results show that the trained machine learning model performs better than the conventional technique.
[0082] [Table 4]
[0083] The embodiments disclosed herein investigated the effectiveness of a multimodal training method for the disease classification performance of computer vision models from cat and dog radiographs. The contrast radiology captioning model disclosed in the embodiments herein utilizes a novel model architecture that uses contrast loss and captioning loss to train with multiple image / text pairs. According to the embodiments disclosed herein, the performance of multi-label image classification tasks after multimodal alignment was shown to be significantly improved in some respects and comparable in other respects to the ensemble of models employed in the current RapidRead system.
[0084] Interestingly, training with this architecture yielded the most significant performance improvements for some of the difficult findings observed in the current system. For example, the average accuracy score for "ingestion in stomach" was 52% higher when using the contrastive radiology captioning model than with the current ensemble. This is likely because the current system uses supervised training, and many labels are derived using NLP algorithms. These algorithms can extract labels from radiology reports by applying rules. In contrast, the multimodal method in this contrastive radiology captioning model allows the text itself to be used as ground truth labels. This means that it can capture nuances of text that are often missed by NLP labelers (e.g., syntactic diversity in describing lesions). Therefore, this disclosure suggests that aligning radiology reports with X-ray images may enable better representation learning for lesions that are difficult to label using alternative methods.
[0085] Furthermore, ablation experiments were conducted to understand the importance of each part of the architecture. First, this disclosure shows that training the StudyFormer model with ImageNet weights results in significantly worse performance compared to training with weights trained on a control radiology captioning model. This indicates that the performance improvement also stems from multimodal training and not solely from the use of the StudyFormer architecture. Next, this disclosure shows that omitting the multimodal decoder results in a significant performance decrease compared to StudyFormer trained on a control radiology captioning model. This highlights the importance of the multimodal decoder.
[0086] In the embodiments disclosed herein, the maximum number of images used in a single inspection was set to five. In 75% of inspections, the number of images was five or less. Therefore, five images were selected as the maximum number of images per inspection to cover all images in most inspections while keeping the requirements and computational costs for padded images low. However, this meant that there were extra images that could not be used in 25% of inspections. As a result, some images referenced in the inspection report were inaccessible to the model. This may have affected the model's ability to accurately align images and reports within the embedding space.
[0087] The embodiments disclosed herein may provide insights into the future development and application of deep learning models in the field of radiology. This disclosure demonstrates that a model can be trained using multiple X-ray images and their corresponding reports using a multimodal method. This makes it possible to utilize large, unlabeled datasets without using other labeling methods such as NLP algorithms. Specifically, this disclosure shows that this training method may be particularly beneficial in the field of radiology for findings that are difficult to reliably detect using other labeling methods.
[0088] Furthermore, the embodiments disclosed herein highlight the potential for automating the radiology report generation process using large-scale language models. This disclosure demonstrates that by simply inputting previously unseen X-ray images into a comparative radiology captioning model, text very similar to that of a human-generated diagnostic report can be generated.
[0089] In other words, this disclosure demonstrates that the contrastive radiology captioning model architecture can be a powerful deep learning model training method compared to supervised learning using other labeling methods. This disclosure also highlights the future potential of automated diagnostic report generation by comparing actual reports with reports generated by the contrastive radiology captioning model. Overall, the embodiments disclosed herein demonstrate the potential merit of training models to perform image classification tasks using the contrastive radiology captioning model architecture when working with large, unlabeled datasets with multiple image / text pairs of inputs.
[0090] Figure 2 shows an exemplary method 200 for detecting anomalies from radiographic images of an animal. In step 210, one or more computing systems can access multiple radiographic images of an animal. One or more first radiographic images of the animal depict the animal from one or more views, and one or more second radiographic images of the animal depict one or more body parts. In step 220, the computing system can identify one or more disease classifications for the animal based on an analysis of the multiple radiographic images by a machine learning model. In step 230, the computing system can generate a diagnostic report for the animal based on the machine learning model. The diagnostic report includes one or more disease classifications and a radiological report in natural language text. Then, in step 240, the computing system can send a command to a user device instructing it to present the diagnostic report.
[0091] Figure 3 shows an exemplary computer system 300 or device used to assist in the detection of anomalies from animal radiographic images. In certain embodiments, one or more computer systems 300 perform one or more steps of one or more methods described and illustrated herein. In other specific embodiments, one or more computer systems 300 provide functions described and illustrated herein. In certain embodiments, software running on one or more computer systems 300 performs one or more steps of one or more methods described and illustrated herein, or provides functions described and illustrated herein. In some embodiments, some embodiments include one or more parts of one or more computer systems 300. In this specification, "computer system" may be a term that encompasses computing devices, and where necessary, "computer device" may be a term that encompasses computing systems. Furthermore, where necessary, one or more computer systems may be encompassed by a singular "computer system" introduced with the article "a".
[0092] This disclosure envisions any suitable number of computer systems 300. Furthermore, this disclosure envisions computer systems 300 taking any suitable physical form. For example, computer systems 300 may comprise, but are not limited to, embedded computer systems, system-on-a-chip (SOC), single-board computer systems (SBC) (e.g., computer-on-a-module (COM) or system-on-a-module (SOM)), desktop computer systems, notebook computer systems, interactive kiosks, mainframes, mesh computer systems, mobile phones, personal digital assistants (PDAs), servers, tablet computer systems, augmented / virtual reality devices, or two or more combinations thereof. Also, as necessary, computer systems 300 may comprise one or more computer systems 300, be integrated or distributed, be located across multiple locations, across multiple machines, across multiple data centers, or reside in a cloud (which may include one or more cloud components in one or more networks). Furthermore, as needed, one or more computer systems 300 can perform one or more steps of one or more methods described and illustrated herein without substantially limiting them spatially or temporally. For example, one or more computer systems 300 can perform one or more steps of one or more methods described and illustrated herein in real time or in batches. Also, as needed, one or more computer systems 300 can perform one or more steps of one or more methods described and illustrated herein at different times or in different locations.
[0093] In certain embodiments, the computer system 300 includes a processor 302, memory 304, storage device 306, input / output (I / O) interface 308, communication interface 310, and bus 312. In the description and illustrations of this disclosure, a particular computer system is described as having a particular number of particular components arranged in a particular configuration, but the intent of this disclosure is that any suitable computer system may have any suitable number of any suitable components arranged in any suitable configuration.
[0094] Furthermore, in some embodiments, the processor 302 includes hardware for executing instructions, such as instructions that constitute a computer program. For example, the processor 302 can execute an instruction by fetching it from an internal register, internal cache, memory 304, or storage device 306, decoding and executing the fetched instruction, and then writing one or more results to an internal register, internal cache, memory 304, or storage device 306. In certain embodiments, the processor 302 may include one or more internal caches for data, instructions, or addresses. In this disclosure, it is intended that the processor 302 may include any appropriate number of appropriate internal caches as appropriate. For example, the processor 302 may include one or more instruction caches, one or more data caches, and one or more translation lookaside buffers (TLBs). Instructions in the instruction cache can be copies of instructions in memory 304 or storage device 306, and the instruction cache can speed up the retrieval of these instructions by the processor 302. The data in the data cache can be copies of data in memory 304 or storage device 306 that are necessary for the processor 302 to execute instructions and perform actions, or appropriate data such as the results of instructions previously executed by the processor 302 that need to be accessed when the processor 302 executes instructions or need to be written to memory 304 or storage device 306. The data cache enables faster read and write operations by the processor 302. The TLB can accelerate the virtual address translation of the processor 302. In some embodiments, the processor 302 may have one or more internal registers for data, instructions, or addresses. In this disclosure, it is intended that the processor 302 may have any appropriate number of appropriate internal registers as needed.Furthermore, the processor 302 may, as needed, comprise one or more arithmetic logic units (ALUs), be a multi-core processor, or include one or more processors 302. Although specific processors are described and illustrated in this disclosure, any suitable processor is intended in this disclosure.
[0095] In some embodiments, the memory 304 includes main memory for storing instructions to be executed by the processor 302 or data necessary for the operation of the processor 302. For example, the computer system 300 may load instructions into the memory 304 from a storage device 306 or another source (e.g., another computer system 300). The processor 302 can then load the instructions from the memory 304 into internal registers or an internal cache. When executing an instruction, the processor 302 can retrieve and decode the instruction from the internal registers or the internal cache. During or after the execution of an instruction, the processor 302 may write one or more results (intermediate or final results) to internal registers or the internal cache. The processor 302 can then write one or more of these results into the memory 304. In some embodiments, the instructions executed by the processor 302 are limited to those in one or more internal registers or internal caches or in memory 304 (rather than instructions in other locations such as the storage device 306), and the data used by the processor 302 for operation is limited to those in one or more internal registers or internal caches or in memory 304 (rather than data in other locations such as the storage device 306). The processor 302 can also be coupled to memory 304 by one or more memory buses (each memory bus may include an address bus and a data bus). As described below, bus 312 may include one or more memory buses. In certain embodiments, one or more memory management units (MMUs) are provided between the processor 302 and memory 304 to assist in accessing memory 304 requested by the processor 302. In certain other embodiments, memory 304 includes random access memory (RAM). This RAM may be volatile memory as needed. Furthermore, this RAM may be dynamic RAM (DRAM) or static RAM (SRAM) as needed.Furthermore, this RAM may be a single-port RAM or a multi-port RAM, as needed. Any suitable RAM is intended in this disclosure. Also, memory 304 may be composed of one or more memory 304s, as needed. Although specific memory components are described and illustrated in this disclosure, any suitable memory is intended in this disclosure.
[0096] In some embodiments, the storage device 306 includes a mass storage device for storing data or instructions. Examples of the storage device 306 include, but are not limited to, a hard disk drive (HDD), a floppy disk drive, flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. If necessary, the storage device 306 may include removable or non-removable media (fixed media). Also, if necessary, the storage device 306 may be built into the computer system 300 or be external. In some embodiments, but are not limited to, the storage device 306 is a non-volatile solid-state memory. In some embodiments, but are not limited to, the storage device 306 includes read-only memory (ROM). If necessary, the ROM may be a pre-programmed mask ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically rewritable ROM (EAROM), flash memory, or a combination of two or more of these. In this disclosure, the large-capacity storage device 306 is intended to take any suitable physical form. If necessary, the storage device 306 may include one or more storage device control units to facilitate communication between the processor 302 and the storage device 306. Furthermore, if necessary, the storage device 306 may be composed of one or more storage devices 306. While specific storage devices are described and illustrated in this disclosure, any suitable storage device is intended.
[0097] In certain embodiments, the I / O interface 308 comprises hardware, software, or both, and provides one or more interfaces for communication between the computer system 300 and one or more I / O devices. The computer system 300 may optionally include one or more of these I / O devices. One or more of these I / O devices can enable communication between a person and the computer system 300. In addition, the I / O devices may include, for example, a keyboard, keypad, microphone, monitor, mouse, printer, scanner, speaker, still camera, stylus pen, tablet device, touch panel, trackball, video camera, other suitable I / O devices, or two or more combinations thereof. The I / O devices may include one or more sensors. Any suitable I / O device and any I / O interface 308 suitable for such I / O device are contemplated in this disclosure. If necessary, the I / O interface 308 may include one or more device or software drivers that enable the processor 302 to drive one or more of these I / O devices. Furthermore, the I / O interface 308 can be configured with one or more I / O interfaces 308 as needed. Although specific I / O interfaces are described and illustrated in this disclosure, any suitable I / O interface is intended in this disclosure.
[0098] In some embodiments, the communication interface 310 comprises hardware, software, or both, and provides one or more interfaces for communication (e.g., packet-based communication) between the computer system 300 and one or more other computer systems 300 or one or more networks. For example, the communication interface 310 may include a network interface controller (NIC) or network adapter for communication with Ethernet or other wired-based networks, or a wireless NIC (WNIC) or wireless adapter for communication with wireless networks such as Wi-Fi networks. Any suitable network and any communication interface 310 suitable for that network are contemplated in this disclosure. For example, the computer system 300 may communicate with one or more parts of an ad-hoc network, a personal area network (PAN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), or the Internet, or two or more combinations thereof. One or more parts of one or more of these networks may be wired or wireless. For example, the computer system 300 can communicate with a suitable wireless network such as a wireless PAN (WPAN) (e.g., a "BLUETOOTH" WPAN), a Wi-Fi network, a Wi-MAX network, or a cellular network (e.g., a Global System for Mobile Communications (GSM) network), or a combination of two or more of these. Furthermore, the computer system 300 may be equipped with any communication interface 310 suitable for any of these networks, if necessary. Also, the communication interface 310 may be configured with one or more communication interfaces 310, if necessary.Although specific communication interfaces are described and illustrated in this disclosure, any suitable communication interface is intended for use in this disclosure.
[0099] In certain embodiments, bus 312 may include hardware, software, or both that connect components of the computer system 300 to one another. Examples of bus 312 may include, but are not limited to, graphics buses such as Accelerated Graphics Port (AGP), EISA (Enhanced Industry Standard Architecture) buses, Front Side Bus (FSB), Hypertransport Bus (HT) interconnect wiring, ISA (Industry Standard Architecture) buses, InfiniBand interconnect wiring, low-pin-count (LPC) buses, memory buses, MCA (Micro Channel Architecture) buses, Peripheral Component Interconnect (PCI) buses, PCI-Express (PCIe) buses, serial advanced technology attachment (SATA) buses, or VESA local (Video Electronics Standards Association local: VLB) buses, or any combination of two or more of these. Furthermore, bus 312 may be composed of one or more buses 312 as needed. Although specific buses are described and illustrated in this disclosure, any suitable bus or interconnect wiring is intended for use in this disclosure.
[0100] In this specification, one or more computer-readable non-temporary storage media may include any suitable computer-readable non-temporary storage media such as one or more semiconductor-based integrated circuits (ICs) (e.g., field-programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs)), hard disk drives (HDDs), hybrid hard drives (HHDs), optical discs, optical disc drives (ODDs), magneto-optical discs, magneto-optical drives, floppy diskettes, floppy disk drives (FDDs), magnetic tapes, solid-state drives (SSDs), RAM drives, or secure digital cards or drives, or any suitable combination of two or more of these as needed. Furthermore, the computer-readable non-temporary storage media may be volatile, non-volatile, or a combination of volatile and non-volatile as needed.
[0101] In this specification, "or" indicates inclusiveness, not exclusivity, unless otherwise explicitly stated or the context indicates otherwise. Therefore, in this specification, "A or B" means "A, B, or both," unless otherwise explicitly stated or the context indicates otherwise. Furthermore, "and" indicates jointness and severalness, unless otherwise explicitly stated or the context indicates otherwise. Therefore, in this specification, "A and B" means "A and B, either jointly or individually," unless otherwise explicitly stated or the context indicates otherwise.
[0102] The scope of this disclosure includes all changes, substitutions, modifications, alterations, and modifications to the exemplary embodiments described or illustrated herein that would be understood by those skilled in the art. The scope of this disclosure is not limited to the exemplary embodiments described or illustrated herein. Furthermore, in the description and illustrations of this disclosure, each embodiment herein is described as including a particular component, element, feature, function, operation, or step, and any embodiment among these embodiments may include any unordered or ordered combination of any component, element, feature, function, operation, or step described or illustrated anywhere in this specification that would be understood by those skilled in the art. Furthermore, where the appended claims state that an apparatus or system, or a component of an apparatus or system, is adapted to, arranged to, capable of, configured to, enabled to, operable to, or operative to perform a particular function, such statement includes, to the extent that the apparatus, system, component, or particular function is adapted to, arranged to, capable of, configured to, enabled to, operable to, or operative to, whether or not it is turned on, activated, or unlocked. Furthermore, while the description and illustrations of this disclosure describe certain embodiments that are not intended to limit, some of these embodiments may not provide any of these advantages, or they may provide some or all of these advantages.
[0103] Furthermore, while embodiments of the method are presented and described in flowcharts in this disclosure, these are illustrative examples intended to provide a more complete understanding of the technology. Therefore, the methods of this disclosure are not limited to the steps and logical flow shown herein. Alternative embodiments are also conceivable, including those with altered order of steps, and those in which lower-level steps described as part of a broader process are performed as independent steps.
[0104] Furthermore, while various embodiments have been described in accordance with the purposes of this disclosure, it should not be assumed that the teachings in this disclosure are limited to such embodiments. Even with various modifications and changes to the elements and processes described above, it is still possible to obtain results that do not deviate from the scope of the systems and processes described in this disclosure.
[0105] The embodiments disclosed herein are illustrative and the scope of this disclosure is not limited thereto. Certain non-limiting embodiments are possible that include all or some of the components, elements, features, functions, operations, or steps of the embodiments disclosed above, as well as certain non-limiting embodiments that do not include any of them. The attached claims relate to methods, storage media, systems, and computer program products, and in particular, multiple embodiments are disclosed in these attached claims. Features mentioned in one claim category, such as a method claim, can similarly be claimed in another claim category, such as a system claim. The dependencies and references to prior art in the attached claims are selected for formal reasons only. Therefore, subject matter obtained by deliberately referencing any prior claim (especially when there are multiple dependencies) can also be claimed. Thus, any combination of claims and their features is disclosed and can be claimed regardless of the dependencies selected in the attached claims. The subject matter that can be claimed includes not only combinations of features described in the appended claims, but also any other combination of features in the claims, and each feature referred to in the claims may be combined with any other feature or combination of features in the claims. Furthermore, any embodiments and features described and illustrated herein may be claimed as separate claims, in addition to or instead of these, in any combination with any embodiments or features described and illustrated herein, or any feature from the features described in the appended claims.
[0106] All patents, patent applications, publications, product descriptions, and protocols referenced herein are incorporated herein by reference in their entirety. In the event of any inconsistency in terminology, this disclosure shall prevail.
[0107] It will be evident that the subject matter described herein is well-calculated to achieve the above-mentioned benefits and advantages; however, the scope of the subject matter of this disclosure is not limited by the specific embodiments described herein. It will be understood that the subject matter of this disclosure can be modified, altered, or changed without departing from its spirit. Those skilled in the art will be able to recognize or confirm many equivalents to the specific embodiments described herein without performing any more than routine experiments. Such equivalents are also intended to be included in the following claims.
[0108] This specification cites various references, all of which are incorporated herein by reference. Preferred embodiments of the present invention are described below in separate sections. Embodiment 1 By one or more computing systems, A step of accessing multiple radiographic images of an animal, wherein one or more first radiographic images among the multiple radiographic images depict the animal from one or more views, and one or more second radiographic images among the multiple radiographic images depict one or more body parts of the animal; A step of identifying one or more disease classifications relating to the animal based on the analysis of the plurality of radiographic images by a machine learning model, A step of generating a diagnostic report for the animal, which includes one or more disease classifications and a radiology report in natural language text, based on the machine learning model; The steps include sending a command to the user device instructing it to present the aforementioned diagnostic report, A method that includes this. Embodiment 2 The method according to Embodiment 1, wherein each of the plurality of radiographic images is formatted as a DICOM (Digital Imaging and Communications in Medicine) image. Embodiment 3 The method according to Embodiment 1, wherein the machine learning model is based on at least one first neural network and at least one second neural network, and the at least one first neural network and the at least one second neural network are coupled to each other. Embodiment 4 The step of generating the aforementioned diagnostic report is, Steps to access multiple reference reports, The steps include encoding the aforementioned multiple reference reports into a feature space, The steps include encoding the plurality of radiographic images into the feature space, The steps include: determining the diagnostic report based on similarity search in the feature space; The method according to Embodiment 1, including the method described above. Embodiment 5 The method according to Embodiment 1, wherein one of the one or more disease classifications described above indicates a tissue abnormality. Embodiment 6 The method according to Embodiment 5, further comprising the step of identifying the abnormality in the tissue as at least one of the following: a cardiovascular abnormality, a pulmonary structural abnormality, a mediastinal structural abnormality, a pleural cavity abnormality, or an extrathoracic abnormality. Embodiment 7 The steps include accessing multiple training radiography images associated with multiple training radiology reports, and The steps include training the machine learning model based on the accessed training radiography images and their respective training radiology reports, The method according to Embodiment 1, further comprising: Embodiment 8 The process further includes the step of pre-processing each of the plurality of training radiographic images, The method according to Embodiment 7, wherein the preprocessing includes one or more of the following: padding, random expansion, random flipping, Gaussian blurring, and normalization. Embodiment 9 The method according to embodiment 7, further comprising the step of applying long document coding to each of the plurality of training radiology reports. Embodiment 10 The method further includes a step of pre-processing each of the aforementioned multiple training radiology reports, The method according to Embodiment 7, wherein the preprocessing includes one or more of tokenization, padding, adding classification tokens, and applying an attention mask. Embodiment 11 The method according to Embodiment 1, wherein the machine learning model comprises an image encoder, a multi-image encoder, a text decoder, and a multimodal decoder. Embodiment 12 The steps include generating a feature map based on the plurality of radiographic images using the image encoder, The steps include generating one or more multi-image keys and values based on the feature map using the multi-image encoder, The steps include generating a radiology report in natural language text using the multimodal decoder based on one or more multi-image keys and values and a sentence-beginning token, The method according to Embodiment 11, further comprising: Embodiment 13 The method according to Embodiment 1, wherein the diagnostic report further includes one or more of the plurality of radiographic images. Embodiment 14 One or more computer-readable non-temporary storage media comprising software, wherein at runtime, the software A step of accessing multiple radiographic images of an animal, wherein one or more first radiographic images among the multiple radiographic images depict the animal from one or more views, and one or more second radiographic images among the multiple radiographic images depict one or more body parts of the animal; A step of identifying one or more disease classifications relating to the animal based on the analysis of the plurality of radiographic images by a machine learning model, A step of generating a diagnostic report for the animal, which includes one or more disease classifications and a radiology report in natural language text, based on the machine learning model; The steps include sending a command to the user device instructing it to present the aforementioned diagnostic report, A medium configured to operate in such a way as to perform the following actions. Embodiment 15 The medium according to Embodiment 14, wherein each of the plurality of radiographic images is formatted as a DICOM (Digital Imaging and Communications in Medicine) image. Embodiment 16 The medium according to Embodiment 14, wherein the machine learning model is based on at least one first neural network and at least one second neural network, and the at least one first neural network and the at least one second neural network are coupled to each other. Embodiment 17 The step of generating the aforementioned diagnostic report is, Steps to access multiple reference reports, The steps include encoding the aforementioned multiple reference reports into a feature space, The steps include encoding the plurality of radiographic images into the feature space, The steps include: determining the diagnostic report based on similarity search in the feature space; The medium described in Embodiment 14, including the medium described in Embodiment 14. Embodiment 18 The medium according to Embodiment 14, wherein one of the one or more disease classifications described above indicates a tissue abnormality. Embodiment 19 The medium according to Embodiment 18, wherein the software is configured to operate at runtime to further identify the abnormality in the tissue as at least one of the following: a cardiovascular abnormality, a pulmonary structure abnormality, a mediastinal structure abnormality, a pleural cavity abnormality, or an extrathoracic abnormality. Embodiment 20 When the software is executed, The steps include accessing multiple training radiography images associated with multiple training radiology reports, and The medium according to Embodiment 14, configured to operate to further perform the step of training the machine learning model based on the accessed training radiographic images and their respective training radiology reports. Embodiment 21 The software is configured to operate at runtime to perform a further step of preprocessing each of the plurality of training radiographic images, The medium according to Embodiment 20, wherein the preprocessing includes one or more of the following: padding, random expansion, random flipping, Gaussian blurring, and normalization. Embodiment 22 The medium according to Embodiment 20, wherein the software is configured to operate at runtime to further perform the step of applying long document encoding to each of the plurality of training radiology reports. Embodiment 23 The software is configured to operate at runtime to perform a further preprocessing step on each of the plurality of training radiology reports, The medium according to Embodiment 20, wherein the preprocessing includes one or more of tokenization, padding, addition of classification tokens, and application of an attention mask. Embodiment 24 The medium according to embodiment 14, wherein the machine learning model comprises an image encoder, a multi-image encoder, a text decoder, and a multimodal decoder. Embodiment 25 When the software is executed, The steps include generating a feature map based on the plurality of radiographic images using the image encoder, The steps include generating one or more multi-image keys and values based on the feature map using the multi-image encoder, The steps include generating a radiology report in natural language text using the multimodal decoder based on one or more multi-image keys and values and a sentence-beginning token, The medium according to embodiment 24, which is configured to be operable to perform the following further. Embodiment 26 The medium according to Embodiment 14, wherein the diagnostic report further includes one or more of the plurality of radiographic images. Embodiment 27 A system comprising one or more processors and non-temporary memory coupled to the processors and containing instructions executable by the processors, By executing the aforementioned instruction, the processor, A step of accessing multiple radiographic images of an animal, wherein one or more first radiographic images among the multiple radiographic images depict the animal from one or more views, and one or more second radiographic images among the multiple radiographic images depict one or more body parts of the animal; A step of identifying one or more disease classifications relating to the animal based on the analysis of the plurality of radiographic images by a machine learning model, A step of generating a diagnostic report for the animal, which includes one or more disease classifications and a radiology report in natural language text, based on the machine learning model; The steps include sending a command to the user device instructing it to present the aforementioned diagnostic report, A system configured to operate in a manner that performs the following actions. Embodiment 28 The system according to Embodiment 27, wherein each of the plurality of radiographic images is formatted as a DICOM (Digital Imaging and Communications in Medicine) image. Embodiment 29 The system according to Embodiment 27, wherein the machine learning model is based on at least one first neural network and at least one second neural network, and the at least one first neural network and the at least one second neural network are coupled to each other. Embodiment 30 The step of generating the aforementioned diagnostic report is, Steps to access multiple reference reports, The steps include encoding the aforementioned multiple reference reports into a feature space, The steps include encoding the plurality of radiographic images into the feature space, The steps include: determining the diagnostic report based on similarity search in the feature space; The system according to embodiment 27, including the system described above. Embodiment 31 The system according to embodiment 27, wherein one of the one or more disease classifications indicates a tissue abnormality. Embodiment 32 The system according to embodiment 31, wherein the processor is configured to operate to perform the step of further identifying the abnormality of the tissue as at least one of the following: a cardiovascular abnormality, a pulmonary structure abnormality, a mediastinal structure abnormality, a pleural cavity abnormality, or an extrathoracic abnormality. Embodiment 33 By executing the aforementioned instruction, the processor, The steps include accessing multiple training radiography images associated with multiple training radiology reports, and The system according to embodiment 27, configured to further perform the step of training the machine learning model based on the accessed training radiography images and their respective training radiology reports. Embodiment 34 By executing the instruction, the processor is configured to operate to perform a further step of preprocessing each of the plurality of training radiographic images. The system according to embodiment 33, wherein the preprocessing includes one or more of the following: padding, random expansion, random flipping, Gaussian blurring, and normalization. Embodiment 35 The system according to embodiment 33, wherein the processor is configured to operate to perform the step of applying long document encoding to each of the plurality of training radiology reports by executing the instruction. Embodiment 36 By executing the aforementioned instruction, the processor is configured to operate to perform a further step of preprocessing each of the plurality of training radiology reports, The system according to embodiment 33, wherein the preprocessing includes one or more of tokenization, padding, adding classification tokens, and applying an attention mask. Embodiment 37 The system according to embodiment 27, wherein the machine learning model comprises an image encoder, a multi-image encoder, a text decoder, and a multimodal decoder. Embodiment 38 By executing the aforementioned instruction, the processor, The steps include generating a feature map based on the plurality of radiographic images using the image encoder, The steps include generating one or more multi-image keys and values based on the feature map using the multi-image encoder, The steps include generating a radiology report in natural language text using the multimodal decoder based on one or more multi-image keys and values and a sentence-beginning token, The system according to embodiment 37, configured to be operable to perform the following further. Embodiment 39 The system according to embodiment 27, wherein the diagnostic report further includes one or more of the plurality of radiographic images. [Explanation of Symbols]
[0109] 100 Architectures 110 Multi-image / Text pairs 112 Examination Images 114 Radiology Report 120 CNN Image Encoders 122 Feature Map 130 ViT multi-image encoder 132 Multi-image Keys and Values 140 CLS embedded representation 150 Unimodal Text Decoders 152 Text Queries 160 Multimodal Text Decoders Output of 162 and 164 multimodal text decoders 300 Computer Systems 302 Processors 304 memory 306 Storage device 308 Input / Output (I / O) Interfaces 310 Communication Interface 312 Bus
Claims
1. By one or more computing systems, A step of accessing multiple radiographic images of an animal, wherein one or more first radiographic images among the multiple radiographic images depict the animal from one or more views, and one or more second radiographic images among the multiple radiographic images depict one or more body parts of the animal; A step of identifying one or more disease classifications relating to the animal based on the analysis of the plurality of radiographic images by a machine learning model, A step of generating a diagnostic report for the animal, which includes one or more disease classifications and a radiology report in natural language text, based on the machine learning model; The steps include sending a command to the user device instructing it to present the aforementioned diagnostic report, Includes, The machine learning model is trained by training at least two coupled neural networks in a single unit, each network comprising at least one first network for radiographic images and at least one second network for radiology reports, the training comprising the step of learning a multimodal embedding space by training the first network and the second network in a single unit to maximize the cosine similarity of correct pairs of radiographic image embedding representations and radiology report embedding representations, and to minimize the cosine similarity of incorrect pairs of embedding representations.
2. The method according to claim 1, wherein each of the plurality of radiographic images is formatted as a DICOM (Digital Imaging and Communications in Medicine) image.
3. The method according to claim 1, wherein the machine learning model is based on at least one first neural network and at least one second neural network, and the at least one first neural network and the at least one second neural network are coupled to each other.
4. The step of generating the aforementioned diagnostic report is, Steps to access multiple reference reports, The steps include encoding the aforementioned multiple reference reports into a feature space, The steps include encoding the plurality of radiographic images into the feature space, The steps include: determining the diagnostic report based on similarity search in the feature space; The method according to claim 1, including the method described in claim 1.
5. The method according to claim 1, wherein one of the one or more disease classifications indicates a tissue abnormality.
6. The method according to claim 5, further comprising the step of identifying the abnormality of the tissue as at least one of the following: a cardiovascular abnormality, a pulmonary structure abnormality, a mediastinal structure abnormality, a pleural cavity abnormality, or an extrathoracic abnormality.
7. The steps include accessing multiple training radiography images associated with multiple training radiology reports, and The steps include training the machine learning model based on the accessed training radiography images and their respective training radiology reports, The method according to claim 1, further comprising:
8. The process further includes the step of pre-processing each of the plurality of training radiographic images, The method according to claim 7, wherein the preprocessing includes one or more of the following: padding, random expansion, random flipping, Gaussian blurring, and normalization.
9. The method according to claim 7, further comprising the step of applying long document encoding to each of the plurality of training radiology reports.
10. The method further includes a step of pre-processing each of the aforementioned multiple training radiology reports, The method according to claim 7, wherein the preprocessing includes one or more of tokenization, padding, adding classification tokens, and applying an attention mask.
11. The method according to claim 1, wherein the machine learning model comprises an image encoder, a multi-image encoder, a text decoder, and a multimodal decoder.
12. The image encoder generates a feature map based on the plurality of radiographic images, which is a feature map generated as an output from a convolutional layer in a convolutional neural network. A step of generating one or more multi-image keys and values based on the feature map using the multi-image encoder, wherein the multi-image keys are generated as the output of a vector representation by encoding the feature map, and the values are text queries output from a unimodal text decoder. The steps include: generating a radiology report in natural language text using the multimodal decoder based on one or more multi-image keys and values and a sentence-beginning token; The method according to claim 11, further comprising:
13. The method according to claim 1, wherein the diagnostic report further includes one or more of the plurality of radiographic images.
14. One or more computer-readable non-temporary storage media comprising software, wherein at runtime, the software A step of accessing multiple radiographic images of an animal, wherein one or more first radiographic images among the multiple radiographic images depict the animal from one or more views, and one or more second radiographic images among the multiple radiographic images depict one or more body parts of the animal; A step of identifying one or more disease classifications relating to the animal based on the analysis of the plurality of radiographic images by a machine learning model, A step of generating a diagnostic report for the animal, which includes one or more disease classifications and a radiology report in natural language text, based on the machine learning model; The steps include sending a command to the user device instructing it to present the aforementioned diagnostic report, It is configured to be operable to perform, The machine learning model is trained by training at least two coupled neural networks in a single unit, each consisting of at least one first network for radiographic images and at least one second network for radiology reports, the training of which includes learning a multimodal embedding space by training the first and second networks in a single unit to maximize the cosine similarity of correct pairs of radiographic image embedding representations and radiology report embedding representations, and to minimize the cosine similarity of incorrect pairs of embedding representations.
15. The medium according to claim 14, wherein each of the plurality of radiographic images is formatted as a DICOM (Digital Imaging and Communications in Medicine) image.
16. The medium according to claim 14, wherein the machine learning model is based on at least one first neural network and at least one second neural network, and the at least one first neural network and the at least one second neural network are coupled to each other.
17. The step of generating the aforementioned diagnostic report is, Steps to access multiple reference reports, The steps include encoding the aforementioned multiple reference reports into a feature space, The steps include encoding the plurality of radiographic images into the feature space, The steps include: determining the diagnostic report based on similarity search in the feature space; The medium according to claim 14, including the following:
18. The medium according to claim 14, wherein one of the one or more disease classifications indicates a tissue abnormality.
19. The medium according to claim 18, wherein the software is configured to perform the step of further identifying the abnormality of the tissue as at least one of the following: a cardiovascular abnormality, a pulmonary structure abnormality, a mediastinal structure abnormality, a pleural cavity abnormality, or an extrathoracic abnormality.
20. When the software is executed, The steps include accessing multiple training radiography images associated with multiple training radiology reports, and The medium according to claim 14, further configured to perform the step of training the machine learning model based on the accessed training radiographic images and their respective training radiology reports.
21. The software is configured to operate at runtime to perform a further step of preprocessing each of the plurality of training radiographic images, The medium according to claim 20, wherein the preprocessing includes one or more of padding, random expansion, random flipping, Gaussian blurring, and normalization.
22. The medium according to claim 20, wherein the software is configured to operate at runtime to further perform the step of applying long document encoding to each of the plurality of training radiology reports.
23. The software is configured to operate at runtime to perform a further preprocessing step on each of the plurality of training radiology reports, The medium according to claim 20, wherein the preprocessing includes one or more of tokenization, padding, addition of classification tokens, and application of an attention mask.
24. The medium according to claim 14, wherein the machine learning model comprises an image encoder, a multi-image encoder, a text decoder, and a multimodal decoder.
25. When the software is executed, The image encoder generates a feature map based on the plurality of radiographic images, which is a feature map generated as an output from a convolutional layer in a convolutional neural network. A step of generating one or more multi-image keys and values based on the feature map using the multi-image encoder, wherein the multi-image keys are generated as the output of a vector representation by encoding the feature map, and the values are text queries output from a unimodal text decoder. The steps include: generating a radiology report in natural language text using the multimodal decoder based on one or more multi-image keys and values and a sentence-beginning token; The medium according to claim 24, configured to be operable to perform the following further.
26. The medium according to claim 14, wherein the diagnostic report further includes one or more of the plurality of radiographic images.
27. A system comprising one or more processors and non-temporary memory coupled to the processors and containing instructions executable by the processors, By executing the aforementioned instruction, the processor, A step of accessing multiple radiographic images of an animal, wherein one or more first radiographic images among the multiple radiographic images depict the animal from one or more views, and one or more second radiographic images among the multiple radiographic images depict one or more body parts of the animal; A step of identifying one or more disease classifications relating to the animal based on the analysis of the plurality of radiographic images by a machine learning model, A step of generating a diagnostic report for the animal, which includes one or more disease classifications and a radiology report in natural language text, based on the machine learning model; The steps include sending a command to the user device instructing it to present the aforementioned diagnostic report, It is configured to be operable to perform, The machine learning model is trained by training at least two coupled neural networks in a single unit, each consisting of at least one first network for radiographic images and at least one second network for radiology reports, the training of which includes learning a multimodal embedding space by training the first network and the second network in a single unit to maximize the cosine similarity of correct pairs of radiographic image embedding representations and radiology report embedding representations, and to minimize the cosine similarity of incorrect pairs of embedding representations.
28. The system according to claim 27, wherein each of the plurality of radiographic images is formatted as a DICOM (Digital Imaging and Communications in Medicine) image.
29. The system according to claim 27, wherein the machine learning model is based on at least one first neural network and at least one second neural network, and the at least one first neural network and the at least one second neural network are coupled to each other.
30. The step of generating the aforementioned diagnostic report is, Steps to access multiple reference reports, The steps include encoding the aforementioned multiple reference reports into a feature space, The steps include encoding the plurality of radiographic images into the feature space, The steps include: determining the diagnostic report based on similarity search in the feature space; The system according to claim 27, including the system described in claim 27.
31. The system according to claim 27, wherein one of the one or more disease classifications indicates a tissue abnormality.
32. The system according to claim 31, wherein the processor is configured to perform the step of further identifying the abnormality of the tissue as at least one of a cardiovascular abnormality, a pulmonary structure abnormality, a mediastinal structure abnormality, a pleural cavity abnormality, or an extrathoracic abnormality, by executing the instruction.
33. By executing the aforementioned instruction, the processor, The steps include accessing multiple training radiography images associated with multiple training radiology reports, and The system according to claim 27, further configured to perform the step of training the machine learning model based on the accessed training radiographic images and their respective training radiology reports.
34. By executing the instruction, the processor is configured to operate to perform a further step of preprocessing each of the plurality of training radiographic images. The system according to claim 33, wherein the preprocessing includes one or more of the following: padding, random expansion, random flipping, Gaussian blurring, and normalization.
35. The system according to claim 33, wherein the processor is configured to operate to perform the step of applying long document encoding to each of the plurality of training radiology reports by executing the instruction.
36. By executing the aforementioned instruction, the processor is configured to operate to perform a further step of preprocessing each of the plurality of training radiology reports, The system according to claim 33, wherein the preprocessing includes one or more of tokenization, padding, adding classification tokens, and applying an attention mask.
37. The system according to claim 27, wherein the machine learning model comprises an image encoder, a multi-image encoder, a text decoder, and a multimodal decoder.
38. By executing the aforementioned instruction, the processor, The image encoder generates a feature map based on the plurality of radiographic images, which is a feature map generated as an output from a convolutional layer in a convolutional neural network. A step of generating one or more multi-image keys and values based on the feature map using the multi-image encoder, wherein the multi-image keys are generated as the output of a vector representation by encoding the feature map, and the values are text queries output from a unimodal text decoder. The steps include: generating a radiology report in natural language text using the multimodal decoder based on one or more multi-image keys and values and a sentence-beginning token; The system according to claim 37, configured to be operable to perform the following further.
39. The system according to claim 27, wherein the diagnostic report further includes one or more of the plurality of radiographic images.
Citation Information
Patent Citations
Diagnosis support apparatus, information processing method, diagnosis support system and program
JP2017191469A
System and device for analyzing anatomical image
JP2019195627A
Health management device, method for operating health management device, and program for operating health management device
JP2020166682A
Diagnosis assistance device, diagnosis assistance method, and diagnosis assistance program
JP2021015525A
Diagnosis support device, diagnosis support method, and diagnosis support program
JP2021058270A