System, method, and apparatus for analyzing x-ray images using multi-view classification

By integrating a convolutional neural network with an attention-based vision transformer, the system improves multi-label classification of X-ray images, addressing the limitations of single-view models and enhancing diagnostic accuracy in veterinary medicine.

JP2026507663APending Publication Date: 2026-03-04MARS INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-22
Publication Date
2026-03-04

AI Technical Summary

Technical Problem

Existing AI models for medical image analysis, particularly in veterinary medicine, often rely on single-view X-ray images, which can be less sensitive in detecting diseases like structured interstitial lung disease, and traditional multi-view classification architectures have limitations in performance.

Method used

A system that integrates information from multiple views using a convolutional neural network to extract feature maps, followed by an attention mechanism implemented with a vision transformer, enabling improved multi-label classification of X-ray images.

Benefits of technology

The proposed system outperforms single-view and traditional multi-view models by effectively utilizing multiple views, achieving enhanced diagnostic accuracy and adaptability across various use cases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026507663000001_ABST
    Figure 2026507663000001_ABST
Patent Text Reader

Abstract

In one embodiment, the method includes evaluating a number of X-ray images associated with an animal, the X-ray images corresponding to each of multiple views; generating, using a first machine learning model, a number of feature maps for each of the number of X-ray images; and determining, using a second machine learning model including an attention mechanism, a diagnostic label indicative of a disease associated with the animal based on the number of feature maps.
Need to check novelty before this filing date? Find Prior Art

Description

Description of Related Applications

[0001] This application claims the benefit of priority to U.S. Provisional Patent Application No. 63 / 486,569, filed February 23, 2023, and U.S. Provisional Patent Application No. 63 / 488,746, filed March 6, 2023, the contents of each of which are incorporated herein in their entirety by reference and to which priority is claimed. [Technical Field]

[0002] Embodiments described in this disclosure relate to radiological studies of animals. For example, some non-limiting embodiments relate to the analysis of X-ray images to aid in the detection of disease in animals. [Background technology]

[0003] Accurate and efficient classification of medical images, such as X-ray images, can be crucial in disease diagnosis and treatment. For example, chest X-ray images are widely used in medical diagnosis. Over the years, artificial intelligence (AI) models have been developed to assist in the interpretation of these images. Many AI tools for medical imaging rely on models that process a single view of an X-ray image. However, multiple views are commonly taken when a patient visits a radiologist. In veterinary medicine, there is strong evidence that multiview chest radiology examinations can yield results that are more sensitive to the likelihood of structured interstitial lung disease, including metastatic disease.

[0004] Machine learning (ML) is a field of study dedicated to understanding and building methods for "learning," or using data to improve performance at a set of tasks. It is considered a subset of artificial intelligence. Machine learning algorithms build models based on example data, known as training data, to make predictions or decisions without being explicitly programmed. Machine learning algorithms are used in a wide range of fields, including medicine, email filtering, speech recognition, and computer vision, where traditional algorithms are difficult or impossible to perform the required tasks. Summary of the Invention

[0005] The objects and advantages of the disclosed subject matter will be set forth in and become apparent from the following description, as well as be learned by practice of the disclosed subject matter. Additional advantages of the disclosed subject matter will be realized and attained by the methods and systems particularly pointed out in the description and claims hereof, as well as from the accompanying drawings.

[0006] To achieve these and other advantages, and in accordance with the objects (as embodied and broadly described) of the disclosed subject matter, the disclosed subject matter provides systems, methods, and apparatus that can be used to collect, receive, and / or analyze data. For example, one non-limiting embodiment can be used to analyze x-ray images of animals.

[0007] In one non-limiting embodiment, the present disclosure describes a method for integrating information from multiple views to improve the performance of X-ray image classification. This approach can be based on the use of a convolutional neural network to extract feature maps from each view, followed by an attention mechanism implemented using a vision transformer. The presently disclosed subject matter demonstrates the effectiveness of this approach through experiments using a dataset of 363,000 X-ray images. The experimental results show that the resulting model can perform multi-label classification on 41 labels, outperforming both single-view models and traditional multi-view classification architectures.

[0008] In certain non-limiting embodiments, the one or more computing systems can access a plurality of X-ray images associated with the animal. In some non-limiting embodiments, each of the plurality of X-ray images can correspond to a plurality of views. The computing system can generate a plurality of feature maps for each of the plurality of X-ray images using a first machine learning model. Further, the computing system can determine one or more diagnostic labels using a second machine learning model based on the plurality of feature maps. In some embodiments, the second machine learning model can include an attention mechanism. In one feature, the one or more diagnostic labels can be indicative of one or more diseases associated with the animal.

[0009] In certain non-limiting embodiments, one or more non-transitory computer-readable storage media embodying software are operable, when executed, to access a plurality of X-ray images associated with an animal. In some non-limiting embodiments, the plurality of X-ray images can each correspond to a plurality of views. The software is further operable, when executed, to generate, with a first machine learning model, a plurality of feature maps for each of the plurality of X-ray images. The software is further operable, when executed, to determine, with a second machine learning model, one or more diagnostic labels based on the plurality of feature maps. In some non-limiting embodiments, the second machine learning model can include an attention mechanism. In one feature, the one or more diagnostic labels can be indicative of one or more diseases associated with the animal.

[0010] In certain non-limiting embodiments, the system includes one or more processors and a non-transitory memory coupled to the processors and storing instructions executable by the processors. The processor, upon executing the instructions, is operable to access a plurality of X-ray images associated with an animal. In some non-limiting embodiments, each of the plurality of X-ray images can correspond to a plurality of views. The processor, upon executing the instructions, is further operable to generate, via a first machine learning model, a plurality of feature maps for each of the plurality of X-ray images. The processor, upon executing the instructions, is further operable to determine, via a second machine learning model, one or more diagnostic labels based on the plurality of feature maps. In some non-limiting embodiments, the second machine learning model can include an attention mechanism. In one feature, the one or more diagnostic labels can be indicative of one or more diseases associated with the animal.

[0011] Furthermore, the disclosed embodiments of the method, non-transitory computer-readable storage medium, and system may have further non-limiting features as described below.

[0012] In one non-limiting embodiment, the multiple x-ray images can depict one or more organs of the animal.

[0013] In certain non-limiting embodiments, the first machine learning model can be based on one or more convolutional neural networks. In certain non-limiting embodiments, the second machine learning model can be based on one or more vision transformers. In some embodiments, the second machine learning model can be based on one or more loss functions including one or more of binary cross-entropy or focal loss.

[0014] In one non-limiting embodiment, the computing system can concatenate the feature maps to form a tensor, which can then be input to a second machine learning model.

[0015] In one feature, the one or more diagnostic labels can include one or more of a pulmonary mass, a pulmonary interstitial nodule, or a mediastinal mass effect.

[0016] In one non-limiting embodiment, the computing system can generate one or more attention maps for a plurality of x-ray images, and in some embodiments, each of the one or more attention maps can highlight focal regions indicative of disease.

[0017] It is to be understood that both the foregoing general description and the following detailed description are exemplary and are intended to provide further explanation of the disclosed subject matter as claimed. [Brief explanation of the drawings]

[0018] The foregoing and other objects, features, and advantages of the present disclosure will become apparent from the following description of embodiments as illustrated in the accompanying drawings, in which reference characters refer to the same parts throughout the various views. The drawings are not necessarily to scale, emphasis instead being placed upon illustrating the principles of the present disclosure. [Figure 1] Diagram showing the example architecture of the model [Figure 2] Illustrative samples of X-ray images used in this disclosure [Figure 3] Graph showing example training and validation losses during ViT-only training [Figure 4] Graph showing an example comparison of ROC curves for lung nodule labeling for different models [Figure 5] An example attention map of a lung disease-specific model for a patient with a lung mass. [Figure 6] An example attention map focused on the chest [Figure 7]Another example attention map [Figure 8] Another example attention map [Figure 9] 1 is a flow diagram illustrating an exemplary method for multi-view classification of X-ray images. [Figure 10] FIG. 1 illustrates an exemplary computer system. DETAILED DESCRIPTION OF THE INVENTION

[0019] The present disclosure is described in more detail below with reference to the accompanying drawings, which form a part hereof, and which show, by way of illustration, certain exemplary embodiments. However, because subject matter can be embodied in a variety of different forms, it is intended that the subject matter as claimed or recited therein not be construed as limited to any exemplary embodiment set forth herein. The exemplary embodiments are provided for illustrative purposes only. Likewise, the broadest reasonable scope of claimed or recited subject matter is intended. Among other things, for example, subject matter may be embodied as methods, devices, components, and / or systems. Thus, embodiments can take the form of, for example, hardware, software, firmware, or any combination thereof (other than software itself). Accordingly, the following detailed description is not intended to be construed in a limiting sense.

[0020] The present disclosure provides a system, method, and / or device that can accept a variable number of X-ray images as input and perform multi-label classification using a model. The model's architecture can combine a convolutional neural network with an attention mechanism implemented using a vision transformer. This approach enables the model to extract relevant features from each view and effectively combine them to improve classification performance. The disclosed subject matter demonstrates the effectiveness of the model through experiments using a dataset of 363,000 veterinary X-ray images, showing that it outperforms both single-view models and traditional multi-view classification architectures. Furthermore, the model can accept a variable number of views at any position, making it highly adaptable to various use cases.

[0021] In the detailed description herein, references to "an embodiment," "one embodiment," "one non-limiting embodiment," "in various embodiments," etc., indicate that the described embodiment may include a particular feature, structure, or characteristic, but that not all embodiments necessarily include that particular feature, structure, or characteristic. Moreover, such phrases do not necessarily refer to the same embodiment. Furthermore, when a particular feature, structure, or characteristic is described in connection with one embodiment, it is believed to be within the knowledge of one skilled in the art to affect such feature, structure, or characteristic in connection with other embodiments, whether or not explicitly stated. After reading this description, it will become apparent to one skilled in the relevant art how to implement the present disclosure in alternative embodiments.

[0022] Generally, terminology can be understood, at least in part, from its use in context. For example, terms such as "and," "or," or "and / or" as used herein can include a variety of meanings that may depend, at least in part, on the context in which the terms are used. Typically, when "or" is used to relate a list such as A, B, or C, it is intended to refer to both A, B, and C when used in an inclusive sense, and A, B, or C when used in an exclusive sense. Moreover, the term "one or more" as used herein can be used to describe any feature, structure, or characteristic in a singular sense, or can be used to describe a combination of features, structures, or characteristics in a plural sense, depending, at least in part, on the context. Similarly, terms such as "a," "an," and "the" can refer to either the singular or the plural, depending, at least in part, on the context. Moreover, the term "based on" can be understood as not necessarily intended to refer to an exclusive set of factors, and again, depending, at least in part, on the context, it can allow for the presence of additional factors not necessarily explicitly recited. As used herein, the words "may" and "can" are used in a permissive sense (i.e., meaning having the possibility) rather than a mandatory sense (i.e., meaning required). Similarly, the words "include," "including," and "includes" mean including, but not limited to.

[0023] As used herein, the terms "comprises," "comprising," or similar terms are intended to be non-exclusive inclusive, and thus a process, method, article, or apparatus that includes a list of elements does not include only those elements, but may include other elements not expressly listed or inherent in the process, method, article, or apparatus.

[0024] The term "animal" or "veterinary" as used in accordance with this disclosure may refer to domestic animals including domestic dogs, domestic cats, horses, cows, ferrets, rabbits, pigs, rats, mice, gerbils, hamsters, goats, etc. The term "animal" or "veterinary" as used in accordance with this disclosure may also refer to wild animals including, but not limited to, bison, elk, deer, venison, ducks, poultry, fish, etc.

[0025] Certain non-limiting embodiments are described below with reference to block diagrams and operational diagrams of methods, processes, devices, and apparatus. It is understood that each block of the block diagrams or operational diagrams, and combinations of blocks in the block diagrams or operational diagrams, can be implemented by analog or digital hardware and computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an ASIC, or other programmable data processing device to modify its functionality as detailed herein, such that the instructions, when executed by the processor of the computer or other programmable data processing device, perform the functions / operations specified in the block diagram or operational block(s). In some alternative implementations, the functions / operations noted in the blocks may occur in a different order than that noted in the operational diagrams. For example, two blocks shown in succession may in fact be executed substantially concurrently, or the blocks may be executed in the reverse order, depending on the functions / operations involved.

[0026] These computer program instructions can be provided to a processor of a general purpose computer; a special purpose computer; an ASIC; or other programmable digital data processing apparatus to modify its functions for a specific purpose, such that the instructions, when executed by the processor of the computer or other programmable data processing apparatus, perform the functions / acts specified in the block diagrams or operational blocks, thereby transforming those functions in accordance with the embodiments herein.

[0027] In some non-limiting embodiments, a computer-readable medium (or computer-readable storage medium / media) stores computer data in machine-readable form, which may include computer program code (or computer-executable instructions) executable by a computer. By way of example and without limitation, a computer-readable medium may include a computer-readable storage medium for tangible or fixed storage of data or a communication medium for transitory interpretation of signals containing code. As used herein, a computer-readable storage medium refers to physical or tangible storage (as opposed to signals) and includes, but is not limited to, volatile and non-volatile, removable and non-removable media implemented in any method or technology for tangibly storing information, such as computer-readable instructions, data structures, program modules, or other data. Computer-readable storage media include, but are not limited to, RAM, ROM, EPROM, EEPROM, flash memory, or other solid-state memory technology, CD-ROM, DVD, or other optical storage media, magnetic cassettes, magnetic tape, magnetic disk storage, or other magnetic storage devices, or any other physical or material medium that can be used to tangibly store desired information, data, or instructions that can be accessed by a computer or processor.

[0028] The disclosed subject matter is in the field of multi-view classification using a concatenation method and feature map expansion combined with a vision transformer-based attention mechanism. This model can significantly improve classification performance compared to existing models, such as MVCNN-based models.

[0029] In recent years, the problem of multi-view classification has attracted considerable attention in the machine learning community. A common approach to address this problem is the multi-view convolutional neural network (MVCNN) architecture, which first processes each view independently using a convolutional neural network (CNN), followed by a view pooling step and final classification using either a multilayer perceptron or another CNN. Several studies have demonstrated the effectiveness of multi-view classification techniques, particularly in the context of X-ray images. The disclosed method advances these existing approaches by introducing an attention-based dynamic mechanism for integrating information from different views.

[0030] Transformers have rapidly gained popularity in the field of natural language processing and have emerged as a promising architecture for a variety of tasks. Transformers have been adapted to computer vision tasks and have subsequently achieved state-of-the-art performance in image classification. In the Vision Transformer (ViT) model, an image can be decomposed into a set of patches, which can then be converted into tokens and processed with a Transformer. The disclosed method leverages the effectiveness of Vision Transformers by incorporating them as a key component of an attention-based dynamic multi-view classifier for X-ray images.

[0031] The presently disclosed subject matter compares its approach to classification models trained on large-scale medical image data.

[0032] The currently disclosed subject matter provides a method for multi-label classification of X-ray images. This technique can use an adaptive vision transformer (ViT) that takes as input features extracted from multiple views of an X-ray image. The features can be extracted using a convolutional neural network (CNN) that can be pre-trained on a single-view image. The same CNN can be applied to each view to extract features, which can then be concatenated and expanded to form a square matrix. This square matrix can be input to ViT to generate the final multi-label classification result.

[0033] 1 illustrates an example architecture 100 of the model. In particular embodiments, there may be a variable number of inputs 110 to the architecture 100. The variable number of inputs may include multiple input x-ray views 110a, 110b, 110c, etc. By way of example, and not limitation, each of the input x-ray views may have dimensions of 3x320x320.

[0034] In particular embodiments, each of the input X-ray views 110 may be processed by a CNN feature extraction module 120. By way of example, and not limitation, the CNN used for feature extraction may be based on a Densenet 121 architecture (i.e., a traditional CNN architecture) and may have pre-trained weights from a model trained on single-view multi-label classification. The output of the CNN feature extraction module 120 may include feature maps for multiple views, e.g., a feature map 130a for view 1, a feature map 130b for view 2, and a feature map 130c for view 3. By way of example, and not limitation, the output feature map 130 of the convolution portion may have dimensions 10x10x1024.

[0035] In particular embodiments, feature maps 130 may be processed by processing, expanding, and concatenating 140. Data expansion can be used to generate additional feature maps, if necessary, to form a rectangular matrix 150d. Feature maps 130 may be concatenated to form a square with width W, where W is the square root of the total number of feature maps. As shown in FIG. 1 , square 150a may correspond to a feature map for view 1, square 150b may correspond to a feature map for view 2, and square 150c may correspond to a feature map for view 3. In the present disclosure, the model can be trained for W=2, 3, and 4 and can accommodate up to 16 views (W=4).

[0036] The input to the Vision Transformer (ViT) 160 can be a matrix of size (W x 10, W x 10, 1024), which may require adaptation of the Vision Transformer 160. A multi-layer perceptron (MLP) with a patch size of 1, depth of 6, attention heads of 16, and dimensions of 2048 can be used. In certain embodiments, the Vision Transformer 160 can utilize a self-attention mechanism to model contextual relationships associated with the input. The Vision Transformer 160 can predict multi-label classifications 170.

[0037] Certain embodiments used example data to train machine learning models and evaluate their performance. The example data consisted of 390,850 x-ray images taken from 98,660 veterinary visits. These images were annotated by radiologists for over 41 diseases in a multi-label fashion and included feedback from the use of radiology learning tools. Additionally, a dataset of 800 images with high-quality annotations was used. These annotations were created by 12 radiologists collaborating on each image. Figure 2 shows an example sample of x-ray images used in this disclosure.

[0038] To ensure the results are fair, the X-ray images can be divided into three datasets: training, validation, and testing. This takes into account the fact that a CNN can be pre-trained on this dataset. On average, there can be approximately 3.96 images per study, which can sometimes range from 1 to more than 10 images. A total of 390,850 images can correspond to 98,660 studies. This dataset can be divided as follows: The training set can consist of 363,820 images, or 92,434 studies. The validation set can consist of 13,515 images, or 3,113 studies. The testing set can consist of 13,481 images, or 3,113 studies.

[0039] While an initial label can be for each x-ray image, in this disclosure, a label for the study may be required. Different views of a study may often have different labels, since a disease may be visible in some views but not in others. For a given label, its value may be the maximum value of that label among all views of the study. Formally, L i,j,k Let be the value of the i-th label in the k-th view of the j-th study, and let LS i,j is a research-level label, and thus:

[0040]

number

[0041] This may reflect the fact that a patient is considered to have a disease if the disease is detected in at least one view of the study.

[0042] Preprocessing operations can be consistent with those used in CNN preprocessing pipelines. When switching to a different CNN, corresponding preprocessing can be applied. First, images can be converted to tensors and resized to 3x320x320 (channels x height x width), and then normalized. Normalization can be performed using the ImageNet mean and standard deviation, which are [0.485, 0.456, 0.406] for the mean and [0.229, 0.224, 0.225] for the standard deviation, respectively. During training, data augmentation can be performed using horizontal and vertical flipping, as well as various data augmentation methods.

[0043] The CNN used can be pre-trained for the specific application in this disclosure, and the first training phase can focus only on the vision transformer by freezing the CNN weights. This phase can require over 50 epochs and can be done in several phases by saving the weights and optimizer. The entire network, including the CNN, can then be trained for 10 epochs until the validation loss stops decreasing. Most of the training can be done using a GPU, which would take approximately 100 hours.

[0044] Several loss functions were evaluated, including BCE (binary cross entropy) and focal loss. Focal loss showed promise due to the presence of imbalanced classes, but its results were not superior to BCE. In this disclosure, BCE was selected as the loss function. Figure 3 shows example training and validation losses during ViT-only training.

[0045] A model can be specifically trained for lung-related diseases, aiming to improve the performance of these labels. This model can classify three diseases: pulmonary mass, pulmonary interstitial nodule, or mediastinal mass effect.

[0046] For comparison, an alternative architecture based on multi-view CNN (MVCNN) can also be implemented, in which the vision transformer can be replaced with a CNN. The results of the model outperformed this architecture, as shown in the results described below.

[0047] The performance of the original CNN, the multi-view CNN (MVCNN)-based architecture, the model, and a lung disease-specific model was compared by evaluating the ROC-AUC scores for different diseases. The ROC-AUC scores of the models are presented in Table 1. For the single-view CNN model, the maximum score across all views was considered as the output of the CNN for each label.

[0048] [Table 1]

[0049] Figure 4 shows an example comparison of ROC curves for lung nodule labeling for different models, e.g., lung-dedicated StudyFormer 410, single-view-based CNN 420, and StudyFormer 430. The curves for the lung-dedicated model lie above the curve for the single-view CNN.

[0050] Figures 5-8 present visualizations of the Vision Transformer (ViT) attention map. The Vision Transformer used is dedicated to lung disease, and attention maps are shown for patients showing the positive label "lung mass." The input to the Vision Transformer can be a concatenated feature map, and an X-ray image is shown mapped to the augmented feature map after the same concatenation and transformation.

[0051] Figure 5 shows an example attention map of a lung disease-specific model for a patient with a lung mass. Input X-ray images 510a, 510b, 510c, and 510d represent lung anatomy images taken from different perspectives. The attention map is visualized in images 520a, 520b, 520c, and 520d, respectively.

[0052] Figure 6 shows an example attention map focused on the thorax. Input X-ray images 610a, 610b, 610c, and 610d represent the thorax body part taken from different perspectives. The bottom left image (i.e., image 610c) was created from the first image 610a after data augmentation because only three views were initially available for the study. The attention maps are visualized in images 620a, 620b, 620c, and 620d, respectively. As expected, attention is concentrated on the thorax where the lungs are located.

[0053] 7 shows another example attention map. Input X-ray images 710a, 710b, 710c, and 710d represent the lung body part taken from different viewpoints. The attention map is visualized in images 720a, 720b, 720c, and 720d, respectively. The third image 720c and the fourth image 720d do not show the lung region, which would indicate that the Vision Transformer is not paying attention to these images, as expected.

[0054] Figure 8 shows another example attention map. Input X-ray images 810a, 810b, 810c, and 810d represent the chest body part taken from different perspectives. The attention maps are visualized in images 820a, 820b, 820c, and 820d, respectively. In this study, the backgrounds of the images are different. This does not interfere with the focus of the Vision Transformer, which remains focused on the chest region. This demonstrates the robustness of the Vision Transformer. The attention maps show that, as expected, the Vision Transformer focuses on the chest region, where the lungs are located. The results also demonstrate that the Vision Transformer continues to focus on the chest region even when the backgrounds in the X-ray images are different. This highlights the robustness of the Vision Transformer.

[0055] These experimental results demonstrate that the model architecture disclosed herein can outperform single-view based architectures in terms of classification performance. The results also show that the model disclosed herein, with its unique architecture that allows a variable number of views to be input in a random order, can perform better than networks applied in MVCNN under the conditions and hyperparameters tested.

[0056] Visualization of the attention map of the Vision Transformer (ViT) can support the conclusion that attention focuses on the correct regions of the image, highlighting the robustness of the network. This finding may be particularly relevant to medical imaging, where focusing on specific regions of the image can be important for accurate diagnosis.

[0057] Furthermore, we can observe that using networks specific to a particular group of labels, such as in the case of a lung disease-specific network, would result in improved performance. This conclusion highlights the possibility of building domain-specific networks to achieve better results.

[0058] In summary, the disclosed subject matter can contribute to the advancement of multi-view networks and highlight the potential of using region-specific networks for medical image classification.

[0059] FIG. 9 illustrates an example method 900 for multi-view classification of X-ray images. The method begins at step 910, where a computing system may access a plurality of X-ray images associated with an animal, where each of the plurality of X-ray images corresponds to a plurality of views. In step 920, the computing system may generate a plurality of feature maps for each of the plurality of X-ray images using a first machine learning model. In step 930, the computing system may determine one or more diagnostic labels using a second machine learning model based on the plurality of feature maps, where the second machine learning model includes an attention mechanism, and the one or more diagnostic labels indicate one or more diseases associated with the animal. Particular embodiments may repeat one or more steps of the method of FIG. 9 as necessary. Although this disclosure describes and illustrates certain steps of the method of FIG. 9 occurring in a particular order, this disclosure contemplates any suitable steps of the method of FIG. 9 occurring in any suitable order. Additionally, although this disclosure describes and illustrates an example method for multiview classification of x-ray images that includes particular steps of the method of Figure 9, this disclosure contemplates any suitable method for multiview classification of x-ray images that includes any suitable steps, including all, some, or none of the steps of the method of Figure 9, as appropriate. Additionally, although this disclosure describes and illustrates particular components, devices, or systems that perform particular steps of the method of Figure 9, this disclosure contemplates any suitable combination of any suitable components, devices, or systems that perform any suitable steps of the method of Figure 9.

[0060] For purposes of this specification, the terms "user," "subscriber," "consumer," or "customer" should be understood to refer to a user of an application or applications described herein and / or a consumer of data provided by a data provider. By way of example and without limitation, the terms "user" or "subscriber" may refer to an individual who receives data over the Internet or data provided by a service provider in a browser session, or may refer to an automated software application that receives the data and stores or processes it.

[0061] Those skilled in the art will appreciate that the disclosed method and system can be implemented in a variety of ways and, therefore, are not limited by the illustrative embodiments and examples set forth above. In other words, functional elements can be performed by single or multiple components, various combinations of hardware and software or firmware, and individual functions can be distributed among software applications at either the client level or the server level, or both. In this regard, several features of different embodiments described herein can be combined into a single or multiple embodiments, and alternative embodiments can have fewer or more than all of the features described herein.

[0062] Functionality may be distributed, in whole or in part, across multiple components in ways now known or that will become known in the future. Thus, numerous software / hardware / firmware combinations are possible for implementing the functions, features, interfaces, and preferences described herein. Moreover, the scope of the present disclosure covers conventionally known methods for implementing the described features, functions, and interfaces, as well as variations and modifications that may be made to the hardware, software, or firmware components described herein, both now and in the future, as would be understood by one of ordinary skill in the art.

[0063] Furthermore, the method embodiments presented and described as flow charts in this disclosure are provided as examples to provide a more complete understanding of the technology. The disclosed methods are not limited to the operations and logic flow presented herein. Alternative embodiments are contemplated in which the order of various operations is changed and sub-operations described as part of a larger operation are performed independently.

[0064] Although various embodiments have been described for purposes of this disclosure, such embodiments should not be construed as limiting the teachings of this disclosure to those embodiments. Various changes and modifications to the elements and operations described above can be made to obtain results that remain within the scope of the systems and processes described in this disclosure.

[0065] While the disclosed subject matter has been described herein with respect to certain preferred embodiments, those skilled in the art will recognize that various changes and modifications can be made to the disclosed subject matter without departing from its scope. Furthermore, although individual features of one non-limiting embodiment of the disclosed subject matter may be discussed herein or shown in the drawings of one non-limiting embodiment and not shown in other embodiments, it should be apparent that individual features of one non-limiting embodiment may be combined with one or more features of another embodiment or features from multiple embodiments.

[0066] FIG. 10 illustrates an exemplary computer system 1000. In particular embodiments, one or more computer systems 1000 perform one or more steps of one or more methods described or illustrated herein. In particular embodiments, one or more computer systems 1000 provide functionality described or illustrated herein. In particular embodiments, software running on one or more computer systems 1000 performs one or more steps of one or more methods described or illustrated herein or provides functionality described or illustrated herein. Particular embodiments include one or more portions of one or more computer systems 1000. Herein, reference to a computer system may encompass a computing device, and vice versa, where appropriate. Furthermore, reference to a computer system may encompass one or more computer systems, where appropriate.

[0067] The present disclosure contemplates any suitable number of computer systems 1000. The present disclosure contemplates computer system 1000 taking any suitable physical form. By way of example and without limitation, computer system 1000 may be an embedded computer system, a system-on-chip (SOC), a single-board computer system (SBC) (e.g., a computer-on-module (COM) or system-on-module (SOM)), a desktop computer system, a laptop or notebook computer system, an interactive kiosk, a mainframe, a mesh of computer systems, a mobile phone, a personal digital assistant (PDA), a server, a tablet computer system, or a combination of two or more thereof. Where appropriate, computer system 1000 may include one or more computer systems 1000; may be single or distributed; may span multiple locations; may span multiple machines; may span multiple data centers; or may reside in a cloud, which may include one or more cloud components in one or more networks. Where appropriate, one or more computer systems 1000 may perform one or more steps of one or more methods described or illustrated herein without substantial spatial or temporal limitations. By way of example, and without limitation, one or more computer systems 1000 may perform one or more steps of one or more methods described or illustrated herein in real time or batch mode. One or more computer systems 1000 may, where appropriate, perform one or more steps of one or more methods described or illustrated herein at different times or in different locations.

[0068] In particular embodiments, computer system 1000 includes a processor 1002, memory 1004, storage 1006, an input / output (I / O) interface 1008, a communication interface 1010, and a bus 1012. Although this disclosure describes and illustrates particular computer systems having particular numbers of particular components in particular configurations, this disclosure contemplates any suitable computer system having any suitable number of any suitable components in any suitable configuration.

[0069] In particular embodiments, processor 1002 includes hardware for executing instructions, such as instructions that make up a computer program. By way of example and without limitation, to execute instructions, processor 1002 may retrieve (or fetch) instructions from an internal register, an internal cache, memory 1004, or storage 1006, decode and execute them, and then write one or more results to an internal register, an internal cache, memory 1004, or storage 1006. In particular embodiments, processor 1002 may include one or more internal caches for data, instructions, or addresses. This disclosure contemplates processor 1002 including any suitable number of any suitable internal caches, where appropriate. By way of example and without limitation, processor 1002 may include one or more instruction caches, one or more data caches, and one or more translation lookaside buffers (TLBs). Instructions in an instruction cache may be copies of instructions in memory 1004 or storage 1006, and the instruction cache may speed up retrieval of those instructions by processor 1002. The data in the data cache may be a copy of data in memory 1004 or storage 1006 on which instructions executing on processor 1002 operate; results of previous instructions executed by processor 1002 for access by subsequent instructions executing on processor 1002 or for writing to memory 1004 or storage 1006; or other suitable data. A data cache may speed up read or write operations by processor 1002. A TLB may speed up virtual address translation for processor 1002. In particular embodiments, processor 1002 may include one or more internal registers for data, instructions, or addresses. This disclosure contemplates processor 1002 including any suitable number of any suitable internal registers, where appropriate. Where appropriate, processor 1002 may include one or more arithmetic logic units (ALUs); may be a multi-core processor; or may include one or more processors 1002.Although this disclosure describes and illustrates a particular processor, this disclosure contemplates any suitable processor.

[0070] In particular embodiments, memory 1004 includes main memory for storing instructions to be executed by processor 1002 or data to be manipulated by processor 1002. By way of example and without limitation, computer system 1000 may load instructions into memory 1004 from storage 1006 or another source (e.g., another computer system 1000). Processor 1002 may then load the instructions from memory 1004 into an internal register or cache. To execute the instructions, processor 1002 may retrieve the instructions from the internal register or cache and decode them. During or after execution of an instruction, processor 1002 may write one or more results (which may be intermediate or final results) to an internal register or cache. Processor 1002 may then write one or more of those results to memory 1004. In particular embodiments, processor 1002 executes only instructions stored in one or more internal registers or internal caches, or memory 1004 (rather than storage 1006 or elsewhere), and manipulates only data stored in one or more internal registers or internal caches, or memory 1004 (rather than storage 1006 or elsewhere). One or more memory buses (each of which may include an address bus and a data bus) may connect processor 1002 and memory 1004. Bus 1012 may include one or more memory buses, as described below. In particular embodiments, one or more memory management units (MMUs) are disposed between processor 1002 and memory 1004 to facilitate accesses to memory 1004 requested by processor 1002. In particular embodiments, memory 1004 includes random access memory (RAM). This RAM may be volatile memory, where appropriate. This RAM may be dynamic RAM (DRAM) or static RAM (SRAM), where appropriate. Furthermore, this RAM may be single-ported RAM or multi-ported RAM, where appropriate. This disclosure contemplates any suitable RAM.Memory 1004 may include, where appropriate, one or more memories 1004. Although this disclosure describes and illustrates particular memory, this disclosure contemplates any suitable memory.

[0071] In particular embodiments, storage 1006 includes mass storage for data or instructions. By way of example, and without limitation, storage 1006 may include a hard disk drive (HDD), a floppy disk drive, flash memory, an optical disk, a magneto-optical disk, magnetic tape, or a universal serial bus (USB) drive, or a combination of two or more thereof. Storage 1006 may include removable or non-removable (fixed) media, where appropriate. Storage 1006 may be located internal or external to computer system 1000, where appropriate. In particular embodiments, storage 1006 is non-volatile solid-state memory. In particular embodiments, storage 1006 includes read-only memory (ROM). Where appropriate, this ROM may be mask-programmed ROM, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), electrically alterable ROM (EAROM), flash memory, or a combination of two or more thereof. This disclosure contemplates mass storage 1006 taking any suitable physical form. Storage 1006 may include, where appropriate, one or more storage control units that facilitate communications between processor 1002 and storage 1006. Where appropriate, storage 1006 may include one or more storages 1006. Although this disclosure describes and illustrates particular storage, this disclosure contemplates any suitable storage.

[0072] In particular embodiments, I / O interface 1008 includes hardware, software, or both that provide one or more interfaces for communication between computer system 1000 and one or more I / O devices. Computer system 1000 may include one or more of these I / O devices, where appropriate. One or more of these I / O devices may enable communication between a person and computer system 1000. By way of example and without limitation, an I / O device may include a keyboard, keypad, microphone, monitor, mouse, printer, scanner, speaker, still camera, stylus, tablet, touch screen, trackball, video camera, another suitable I / O device, or a combination of two or more thereof. An I / O device may include one or more sensors. This disclosure contemplates any suitable I / O device and any suitable I / O interface 1008 thereto. Where appropriate, I / O interface 1008 may include one or more device drivers or software drivers that enable processor 1002 to drive one or more of these I / O devices. I / O interface 1008 may, where appropriate, include one or more I / O interfaces 1008. Although this disclosure describes and illustrates particular I / O interfaces, this disclosure contemplates any suitable I / O interface.

[0073] In particular embodiments, communication interface 1010 includes hardware, software, or both that provide one or more interfaces for communications (e.g., packet-based communications) between computer system 1000 and one or more other computer systems 1000 or one or more networks. By way of example and without limitation, communication interface 1010 may include a network interface controller (NIC) or network adapter for communications with an Ethernet or other wired network, or a wireless NIC (WNIC) or wireless adapter for communications with a wireless network, such as a Wi-Fi network. This disclosure contemplates any suitable network and any suitable corresponding communication interface 1010. By way of example and without limitation, computer system 1000 may communicate with an ad-hoc network, a personal area network (PAN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), or a portion of the Internet, or a combination of two or more thereof. One or more portions of one or more of these networks may be wired or wireless. By way of example, computer system 1000 may communicate with a wireless PAN (WPAN) (e.g., a BLUETOOTH® WPAN, etc.), a Wi-Fi network, a Wi-MAX network, a cellular network (e.g., a Global System for Mobile Communications (GSM) network, etc.), or other suitable wireless network, or a combination of two or more thereof. Computer system 1000 may include any suitable communication interface 1010 corresponding to any of these networks, where appropriate. Communication interface 1010 may include one or more communication interfaces 1010, where appropriate. Although this disclosure describes and illustrates a particular communication interface, this disclosure contemplates any suitable communication interface.

[0074] In particular embodiments, bus 1012 includes hardware, software, or both for interconnecting components of computer system 1000. By way of example, and without limitation, bus 1012 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an INFINIBAND interconnect, a Low Pin Count (LPC) bus, a memory bus, a MicroChannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCIe) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (WLB) bus, or another suitable bus, or a combination of two or more thereof. Bus 1012 may include one or more buses 1012, where appropriate. Although this disclosure describes and illustrates a particular bus, this disclosure contemplates any suitable bus or interconnect.

[0075] Here, a non-transitory computer-readable storage medium or media may include one or more semiconductor-based or other integrated circuits (ICs) (e.g., field programmable gate arrays (FPGAs) or application-specific integrated circuits (ASICs)), hard disk drives (HDDs), hybrid hard drives (HHDs), optical disks, optical disk drives (ODDs), magneto-optical disks, magneto-optical drives, floppy disks, floppy disk drives (FDDs), magnetic tapes, solid-state drives (SSDs), RAM drives, SECURE DIGITAL cards or drives, any other suitable non-transitory computer-readable storage media, or a suitable combination of two or more of these. A non-transitory computer-readable storage medium may be volatile, non-volatile, or, where appropriate, a combination of volatile and non-volatile. [Explanation of symbols]

[0076] 110, 110a, 110b, 110c Input X-ray view 120, 120a, 120b, 120c CNN feature extraction module 130, 130a, 130b, 130c Output feature maps 150a, 150b, 150c, 150d square matrix 160 Vision Transformer 170 Multi-label Classification

Claims

1. evaluating a plurality of three-dimensional (3D) x-ray images associated with the animal, the plurality of 3D x-ray images corresponding to a respective one of a plurality of views; generating a plurality of 3D feature maps for each of the plurality of 3D X-ray images by a first machine learning model including one or more convolutional neural networks, the plurality of 3D feature maps corresponding to the plurality of views, respectively; generating one or more additional 3D feature maps from the plurality of 3D feature maps based on one or more data augmentation mechanisms; generating a joint 3D feature map based on the plurality of 3D feature maps and the one or more additional 3D feature maps; inputting the integrated 3D feature map into a second machine learning model including one or more adaptive vision transformers configured to process 3D inputs; modeling, by the second machine learning model, contextual relationships associated with the integrated 3D feature map based on a self-attention mechanism associated with the second machine learning model; determining, by the second machine learning model, one or more diagnostic labels based on the integrated 3D feature map and the contextual relationships associated with the integrated 3D feature map, the one or more diagnostic labels indicative of one or more diseases associated with the animal; and generating one or more attention maps for the plurality of 3D X-ray images, each of the one or more attention maps highlighting focal regions representative of the one or more diseases associated with the one or more diagnostic labels; 10. A method according to one or more computing systems, comprising:

2. The method of claim 1 , wherein the plurality of 3D X-ray images depict one or more organs of the animal.

3. 3. The method of claim 1 or 2, wherein the one or more diagnostic labels include one or more of a pulmonary mass, a pulmonary interstitial nodule, or a mediastinal mass effect.

4. evaluating a plurality of x-ray images relating to the animal, the x-ray images corresponding to a respective plurality of views; generating a plurality of feature maps for each of the plurality of X-ray images using a first machine learning model; and determining, by a second machine learning model including an attention mechanism, one or more diagnostic labels indicative of one or more diseases associated with the animal based on the plurality of feature maps; 10. A method according to one or more computing systems, comprising:

5. The method of claim 4 , wherein the plurality of X-ray images depict one or more organs of the animal.

6. The method of claim 4 or 5, wherein the first machine learning model is based on one or more convolutional neural networks.

7. The method of claim 4 , wherein the second machine learning model is based on one or more vision transformers.

8. generating one or more additional feature maps from the plurality of feature maps based on data augmentation; 8. The method of claim 4, further comprising:

9. concatenating the feature maps to form a tensor; and inputting the tensor into the second machine learning model; 9. The method of claim 4, further comprising:

10. 10. The method of claim 4, wherein the second machine learning model is based on one or more loss functions including one or more of binary cross-entropy or focal loss.

11. 11. The method of any one of claims 4 to 10, wherein the one or more diagnostic labels include one or more of a pulmonary mass, a pulmonary interstitial nodule, or a mediastinal mass effect.

12. generating one or more attention maps for the plurality of X-ray images, each of the one or more attention maps highlighting focal regions indicative of disease; 12. The method of claim 4, further comprising:

13. One or more non-transitory computer-readable storage media embodying software, the software, when executed, accessing a plurality of x-ray images associated with the animal, the x-ray images corresponding to a respective one of a plurality of views; generating a plurality of feature maps for each of the plurality of X-ray images using a first machine learning model; determining, by a second machine learning model including an attention mechanism, one or more diagnostic labels indicative of one or more diseases associated with the animal based on the plurality of feature maps; So viable, medium.

14. The medium of claim 13 , wherein the plurality of X-ray images depict one or more organs of the animal.

15. The medium of claim 13 or 14, wherein the first machine learning model is based on one or more convolutional neural networks.

16. The medium of claim 13 , wherein the second machine learning model is based on one or more vision transformers.

17. The software, when executed, concatenating the feature maps to form a tensor; inputting the tensor into the second machine learning model; 17. The medium of any one of claims 13 to 16, further operable to:

18. 18. The medium of claim 13, wherein the second machine learning model is based on one or more loss functions including one or more of binary cross-entropy or focal loss.

19. 19. The medium of any one of claims 13 to 18, wherein the one or more diagnostic labels include one or more of a pulmonary mass, a pulmonary interstitial nodule, or a mediastinal mass effect.

20. The software, when executed, generating one or more attention maps for the plurality of X-ray images, each of which highlights a focal region indicative of disease; 20. The medium of any one of claims 13 to 19, further operable to:

21. In a system including one or more processors and a non-transitory memory coupled to the processors that stores instructions executable by the processors, the processors, when executing the instructions, accessing a plurality of x-ray images associated with the animal, the x-ray images corresponding to a respective one of a plurality of views; generating a plurality of feature maps for each of the plurality of X-ray images using a first machine learning model; determining, by a second machine learning model including an attention mechanism, one or more diagnostic labels indicative of one or more diseases associated with the animal based on the plurality of feature maps; A system capable of operating as follows.

22. 22. The system of claim 21, wherein the plurality of X-ray images depict one or more organs of the animal.

23. 23. The system of claim 21 or 22, wherein the first machine learning model is based on one or more convolutional neural networks.

24. 24. The system of claim 21, wherein the second machine learning model is based on one or more vision transformers.

25. When the processor executes the instructions, concatenating the feature maps to form a tensor; inputting the tensor into the second machine learning model; 25. The system of any one of claims 21 to 24, further operable to:

26. 26. The system of any one of claims 21 to 25, wherein the second machine learning model is based on one or more loss functions including one or more of binary cross-entropy or focal loss.

27. 27. The system of any one of claims 21 to 26, wherein the one or more diagnostic labels include one or more of a pulmonary mass, a pulmonary interstitial nodule, or a mediastinal mass effect.

28. When the processor executes the instructions, generating one or more attention maps for the plurality of X-ray images, each of which highlights a focal region indicative of disease; 28. The system of any one of claims 21 to 27, further operable to: