Classification Based on Annotation Information

The end-to-end deep learning framework addresses the limitations of conventional AI by integrating region-of-interest masks and boundary boxes to enhance the accuracy and efficiency of digital image classification and analysis in medical imaging.

CN112262395BActive Publication Date: 2025-07-08GE PRECISION HEALTHCARE LLC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN201980038768.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-09-27
Filing Date
2019-06-28
Publication Date
2025-07-08
Estimated Expiration
2040-01-24

AI Technical Summary

Technical Problem

Conventional artificial intelligence (AI) techniques struggle with the accuracy and efficiency of digital image classification and analysis, often requiring labor-intensive processes like pixel annotation and lacking effective utilization of region-of-interest information.

Method used

A novel end-to-end deep learning framework that incorporates region-of-interest masks, image-level labels, and boundary boxes to train convolutional neural networks, using multiple loss functions to improve classification and localization accuracy in medical image analysis.

Benefits of technology

Enhances the accuracy and efficiency of digital image classification and analysis by leveraging richer annotation information, improving the performance of machine learning models in processing medical imaging data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112262395B_ABST
    Figure CN112262395B_ABST
Patent Text Reader

Abstract

The present invention provides systems and techniques for classification based on annotation information. In one example, a system trains a convolutional neural network based on training data and a plurality of images. The plurality of images are associated with a plurality of masks, a plurality of image-level labels, and / or bounding boxes. The system also generates a first loss function based on the plurality of masks, a second loss function based on the plurality of image-level labels, and a third loss function based on the bounding boxes. Additionally, the system generates a fourth loss function based on the first loss function, the second loss function, and the third loss function, wherein the fourth loss function is iteratively backpropagated to tune the parameters of the convolutional neural network. The system also predicts a classification label of an input image based on the convolutional neural network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure generally relates to artificial intelligence. Background Art

[0002] Artificial intelligence (AI) can be used for the classification and / or analysis of digital images. For example, AI can be used for image recognition. In certain technical applications, AI can be used to enhance imaging analysis. In one example, a deep neural network based on regions of interest can be employed to locate features in a digital image. However, it is generally difficult to achieve the accuracy and / or efficiency of the classification and / or analysis of digital images using conventional manual techniques. In addition, conventional manual techniques for the classification and / or analysis of digital images typically require labor-intensive processes, such as pixel annotation, voxel-level annotation, etc. Therefore, the conventional manual techniques for the classification and / or analysis of digital images can be improved. Summary of the Invention

[0003] The following presents a simplified summary of the present specification in order to provide a basic understanding of some aspects of the present specification. This summary is not an exhaustive overview of the present specification. It is neither intended to identify the key or important elements of the present specification nor to describe any scope of any specific implementation of the present specification or any scope of any claims. Its sole purpose is to present some concepts of the present specification in a simplified form as a prelude to the more detailed description presented later.

[0004] According to one embodiment, a system includes a training component, a first loss function component, a second loss function component, a third loss function component, a fourth loss function component, and a classification component. The training component trains a convolutional neural network based on training data and a plurality of images. The training data is associated with a plurality of patients from at least one imaging device. The plurality of images is associated with a plurality of masks from a plurality of objects, a plurality of image-level labels of the plurality of images, and / or bounding boxes that link regions of interest to class labels. The first loss function component generates a first loss function based on the plurality of masks. The second loss function component generates a second loss function based on the plurality of image-level labels of the plurality of images. The third loss function component generates a third loss function based on the bounding boxes that link regions of interest to class labels. The fourth loss function component generates a fourth loss function based on the first loss function, the second loss function, and the third loss function, wherein the fourth loss function is iteratively backpropagated to tune the parameters of the convolutional neural network. The classification component predicts a classification label of an input image based on the convolutional neural network.

[0005] According to another embodiment, a method is provided. The method includes receiving, from at least one imaging device, a plurality of images associated with a plurality of patients. The method further includes receiving, from a plurality of objects, a plurality of masks, wherein each image includes at least one mask associating an object of interest with a corresponding class label, at least one image-level label for the image, and / or a bounding box linking the object of interest to the corresponding class label. Additionally, the method includes training a convolutional neural network based on the plurality of images, the plurality of masks, the bounding boxes, and / or the at least one image-level label, wherein the convolutional neural network includes a pre-trained classifier network that outputs a convolutional feature map and a classification / localization network that outputs a corresponding localization map. The method further includes generating a first loss function based on the plurality of masks. The method also includes generating a second loss function based on the at least one image-level label of the image. The method further includes generating a third loss function based on the bounding box linking the object of interest to the corresponding class label. The method also includes generating a fourth loss function based on the first loss function, the second loss function, and the third loss function. Additionally, the method includes iteratively backpropagating the fourth loss function to tune the parameters of the convolutional neural network. The method further includes predicting a classification label of an input image based on the convolutional neural network.

[0006] According to yet another embodiment, a computer-readable storage device is provided. The computer-readable storage device includes instructions that, in response to execution, cause a system including a processor to perform operations, the operations including receiving, from at least one imaging device, a plurality of images associated with a plurality of patients. The processor further performs operations including receiving, from a plurality of objects, a plurality of masks, wherein each image includes at least one mask associating an object of interest with a corresponding class label, at least one image-level label for the image, and / or a bounding box linking the object of interest to the corresponding class label. The processor further performs operations including training a convolutional neural network based on the plurality of images, the plurality of masks, the bounding boxes, and / or the at least one image-level label, wherein the convolutional neural network includes a pre-trained classifier network that outputs a convolutional feature map and a classification / localization network that outputs a corresponding localization map. Additionally, the processor performs operations including generating a first loss function based on the plurality of masks. Additionally, the processor performs operations including generating a second loss function based on the at least one image-level label of the image. Additionally, the processor performs operations including generating a third loss function based on the bounding box linking the object of interest to the corresponding class label. Additionally, the processor performs operations including generating a fourth loss function based on the first loss function, the second loss function, and the third loss function. The processor further performs operations including iteratively backpropagating the fourth loss function to tune the parameters of the convolutional neural network. The processor further performs operations including predicting a classification label of an input image based on the convolutional neural network.

[0007] The following detailed description and the accompanying drawings set forth certain illustrative aspects of the present specification. However, these aspects merely indicate some of the various ways in which the principles of the present specification may be employed. When considered in conjunction with the accompanying drawings, other advantages and novel features of the present specification will become apparent from the following detailed description of the specification. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Many aspects, specific implementations, objects, and advantages of the present invention will become apparent when the following detailed description is considered in conjunction with the accompanying drawings, in which like reference numerals represent like components throughout, and in which:

[0009] Figure 1 A high-level block diagram of an exemplary machine learning component in accordance with various aspects and specific implementations described herein is shown;

[0010] Figure 2 A high-level block diagram of another exemplary machine learning component in accordance with various aspects and specific implementations described herein is shown;

[0011] Figure 3 A system in accordance with various aspects and specific implementations described herein is shown, the system including an exemplary machine learning component and an exemplary medical imaging diagnostic process;

[0012] Figure 4 Another exemplary system associated with a segmentation-classification network in accordance with various aspects and specific implementations described herein is shown;

[0013] Figure 5 Another exemplary system associated with a segmentation-classification network implementing a loss function in accordance with various aspects and specific implementations described herein is shown;

[0014] Figure 6 An exemplary loss function in accordance with various aspects and specific implementations described herein is shown;

[0015] Figure 7 An exemplary system in accordance with various aspects and specific implementations described herein for generating a loss function using masks, bounding boxes, and / or labels is shown;

[0016] Figure 8 Another exemplary multi-dimensional visualization in accordance with various aspects and specific implementations described herein is shown;

[0017] Figure 9 A flowchart depicting another exemplary method for classification and / or localization based on annotation information in accordance with various aspects and specific implementations described herein is shown;

[0018] Figure 10 is a schematic block diagram showing a suitable operating environment; and

[0019] Figure 11 is a schematic block diagram of a sample computing environment. Detailed implementation manners

[0020] Aspects of the present disclosure will now be described with reference to the accompanying drawings, in which like reference numerals are always used to denote like elements. In the following description, for purposes of explanation, numerous specific details are set forth in order to provide a thorough understanding of one or more aspects. However, it should be understood that certain aspects of the present disclosure may be practiced without these specific details, or with other methods, components, materials, etc. In other instances, well-known structures and devices are shown in block diagram form to facilitate the description of one or more aspects.

[0021] The present invention provides systems and techniques for providing classification and / or localization based on annotation information. For example, a novel end-to-end deep learning framework is disclosed herein to automatically detect and / or localize diseases in medical images, for example, given mask annotations related to regions of interest. The classification and localization network can be a fully convolutional neural network and can output image-level labels and localization maps during inference. Thus, compared with conventional classification that only uses image-level labels, the classification and / or localization accuracy when using mask information can be improved. In one embodiment, one or more image-level labels, bounding boxes, and / or masks related to one or more regions of interest of an image can be employed, for example, to improve the performance of a classifier. For example, weighted losses from mask annotations, ground truth weak labels (e.g., image-level labels), and / or bounding boxes can be backpropagated through the deep learning framework to, for example, backpropagate classification losses and / or segmentation losses and also improve the localization results. In addition, by employing the novel end-to-end deep learning framework described herein, the detection and / or localization of one or more features associated with image data (e.g., the detection and / or localization of one or more conditions of a patient associated with medical imaging data) can be improved. In addition, the accuracy and / or efficiency of the classification and / or analysis of image data (e.g., medical imaging data) can be improved. Additionally, the effectiveness of a machine learning model for the classification and / or analysis of image data (e.g., medical imaging data) can be improved, the performance of one or more processors executing the machine learning model for the classification and / or analysis of image data (e.g., medical imaging data) can be improved, and / or the efficiency of one or more processors executing the machine learning model for the classification and / or analysis of image data (e.g., medical imaging data) can be improved.

[0022] First, refer to Figure 1, shows an exemplary system 100 for classification and / or localization based on annotation information according to one aspect of the present disclosure. System 100 can be adopted by various systems, such as but not limited to, medical device systems, medical imaging systems, medical diagnostic systems, medical systems, medical modeling systems, enterprise imaging solution systems, advanced diagnostic tool systems, simulation systems, image management platform systems, care delivery management systems, artificial intelligence systems, machine learning systems, neural network systems, modeling systems, aviation systems, power systems, distributed power systems, energy management systems, thermal management systems, transportation systems, oil and gas systems, mechanical systems, machine systems, equipment systems, cloud-based systems, heating systems, HVAC systems, medical systems, automotive systems, aircraft systems, watercraft systems, water filtration systems, cooling systems, pump systems, engine systems, prediction systems, machine design systems, etc. In one example, system 100 can be associated with a classification system to facilitate the visualization and / or interpretation of medical imaging data. Additionally, system 100 and / or components of system 100 can be used to solve problems that are highly technical in nature (e.g., related to processing digital data, related to processing medical imaging data, related to medical modeling, related to medical imaging, related to artificial intelligence, etc.) using hardware and / or software, and these problems are not abstract and cannot be performed by humans as a series of mental acts.

[0023] System 100 may include a machine learning component 102, which may include a training component 104, a loss function component 106, and a classification component 108. In one embodiment, the loss function component 106 may include a first loss function component 109, a second loss function component 111, a third loss function component 113, and a fourth loss function component 115. Aspects of the systems, devices, or processes explained in this disclosure may constitute machine-executable components embodied within a machine (e.g., embodied in one or more computer-readable media associated with one or more machines). When executed by one or more machines (e.g., computers, computing devices, virtual machines, etc.), such components may cause the machines to perform the operations. System 100 (e.g., the machine learning component 102) may include a memory 112 for storing computer-executable components and instructions. System 100 (e.g., the machine learning component 102) may also include a processor 110 to facilitate the operation of system 100 (e.g., the machine learning component 102) on the instructions (e.g., computer-executable components and instructions).

[0024] The machine learning component 102 (e.g., the training component 104) may receive medical imaging data (e.g., Figure 1the medical imaging data shown). The medical imaging data may be associated with multiple patients. Additionally, the medical imaging data may be a set of images (e.g., a set of medical images). The medical imaging data may be two-dimensional medical imaging data and / or three-dimensional medical imaging data generated by one or more medical imaging devices. For example, the medical imaging data may be an electromagnetic radiation image captured via a set of sensors (e.g., a set of sensors associated with a medical imaging device). In some embodiments, the medical imaging data may be a series of electromagnetic radiation images captured via a set of sensors (e.g., a set of sensors associated with a medical imaging device) during a time interval. The medical imaging data may be received directly from one or more medical imaging devices. Alternatively, the medical imaging data may be stored in one or more databases that receive and / or store medical imaging data associated with one or more medical imaging devices. The medical imaging device may be, for example, an x-ray device, a computed tomography (CT) device, another type of medical imaging device, etc. In addition or alternatively, the machine learning component 102 (e.g., the training component 104) may receive label data (e.g., Figure 1 the label data shown). For example, the medical imaging data may be associated with label data that includes multiple image-level labels for multiple images. The multiple image-level labels included in the label data may be, for example, multiple weak ground truth labels for the multiple images. In addition or alternatively, the machine learning component 102 (e.g., the training component 104) may receive bounding box data (e.g., Figure 1 the bounding box data shown). For example, the medical imaging data may be associated with bounding box data that includes one or more bounding boxes that link one or more regions of interest in the image to class labels. The bounding boxes included in the bounding box data may identify where an object is located in the image and / or may link the region of interest associated with the object to a class label. The bounding boxes included in the bounding box data may include, for example, a set of coordinates (e.g., top-left coordinates, top-right coordinates, bottom-left coordinates, bottom-right coordinates, etc.) that provide a location (e.g., a region) for the region of interest. Additionally, the bounding boxes included in the bounding box data may, in addition or alternatively, include a height value and / or a width value of the location (e.g., the region) associated with the region of interest. In addition or alternatively, the machine learning component 102 (e.g., the training component 104) may receive mask data (e.g., Figure 1The mask data shown). In one embodiment, the mask data can be a set of masks from multiple objects. For example, each medical image from medical imaging data can be associated with one or more masks. For example, a mask can include one or more weights of one or more regions of interest in the image (e.g., in medical imaging data). In one example, a mask can include a set of pixels that define the location of a region of interest using binary filtering. In one embodiment, medical imaging data and / or mask data can be used as training data to, for example, train a convolutional neural network. In certain embodiments, medical imaging data and / or mask data can be stored in a database that receives and / or stores training data associated with at least one imaging device. In certain embodiments, medical imaging data can be associated with a set of weights from a pre-trained model.

[0025] In one embodiment, the training component 104 can train a convolutional neural network based on medical imaging data (e.g., multiple images) and / or mask data. For example, the training component 104 can execute the training phase of a machine learning process to, for example, train the neural network model of a convolutional neural network. The convolutional neural network can include a decoder composed of at least one upsampling layer and / or at least one convolutional layer. Additionally, in some embodiments, the convolutional neural network can include a pre-trained classifier network that outputs a convolutional feature map. In addition or alternatively, in some embodiments, the convolutional neural network can include a classification / localization network that outputs a corresponding localization map. In some embodiments, the convolutional neural network can be a spring network of convolutional layers. For example, the convolutional neural network can perform multiple sequential and / or parallel downsampling and upsampling on the medical imaging data associated with the convolutional layers of the convolutional neural network. In one example, the convolutional neural network can perform a first convolutional layer process associated with sequential downsampling of the medical imaging data and a second convolutional layer process associated with sequential upsampling of the medical imaging data. The spring network of convolutional layers can include a first convolutional layer process associated with sequential downsampling and a second convolutional layer process associated with sequential upsampling. The spring network of convolutional layers associated with the convolutional neural network can change the convolutional layer filters, similar to the function of a spring. For example, the convolutional neural network can analyze the medical imaging data based on a first convolutional layer filter including a first size, a second convolutional layer filter including a second size different from the first size, and a third convolutional layer filter including the first size associated with the first convolutional layer filter. In some embodiments, the training component 104 can train a convolutional neural network based on medical imaging data and / or mask data (e.g., training data) to determine whether a first class exists in the medical imaging data. In addition or alternatively, the training component 104 can train a convolutional neural network based on medical imaging data and / or mask data (e.g., training data) to form at least a portion of the convolutional neural network associated with a neural network architecture. The neural network architecture can be, for example, a binary neural network architecture that performs machine learning associated with one or more binary classifications of medical imaging data.

[0026] The loss function component 106 can generate a loss function based on a plurality of masks associated with medical imaging data, a plurality of image-level labels, and / or one or more bounding boxes. The loss function can be, for example, a loss function of a convolutional neural network. In some embodiments, the loss function component 106 can employ a decoder to generate a localization map. For example, the loss function component 106 can perform a decoding process associated with upsampling and / or one or more convolutional neural network layers to generate a localization map. The localization map can include, for example, information representing probability scores of one or more regions of the medical imaging data. In one embodiment, the localization map can include a visualization of probability scores of one or more regions of the medical imaging data. In some embodiments, the decoder can be a set of decoders. In one aspect, the decoder can be a set of decoders that perform different decoding processes associated with upsampling and / or one or more convolutional neural network layers. For example, the decoder can include a first decoder that performs a first decoding process associated with upsampling and / or one or more convolutional neural network layers, a second decoder that performs a second decoding process associated with upsampling and / or one or more convolutional neural network layers, a third decoder that performs a third decoding process associated with upsampling and / or one or more convolutional neural network layers, and so on. In another aspect, the number of decoders included in the set of decoders can be determined during the training of the convolutional neural network.

[0027] The first loss function component 109 may generate a first loss function based on multiple masks associated with medical imaging data. For example, the first loss function component 109 may generate a first loss function based on the probabilities of the classes associated with the multiple masks. In one example, the first loss function component 109 may generate a first loss function based on the probabilities associated with the classification outputs from a convolutional neural network and multiple masks. The second loss function component 111 may generate a second loss function based on multiple image-level labels associated with medical imaging data (e.g., multiple images). For example, the second loss function component 111 may generate a second loss function based on the probabilities of the classes associated with the multiple image-level labels. In one example, the second loss function component 111 may generate a second loss function based on the probabilities associated with the classification outputs from a convolutional neural network and multiple image-level labels. The third loss function component 113 may generate a third loss function based on one or more bounding boxes. The one or more bounding boxes may link one or more regions of interest to one or more class labels. In one example, the third loss function component 113 may generate a third loss function based on the bounding boxes that link the regions of interest in the image to class labels. The fourth loss function component 115 may generate a fourth loss function based on the first loss function, the second loss function, and / or the third loss function. For example, the fourth loss function component 115 may apply a first weight to the first loss function, may apply a second weight to the second loss function, and / or may apply a third weight to the third loss function. Additionally, the fourth loss function component 115 may combine the first loss function, the second loss function, and / or the third loss function (e.g., the fourth loss function component 115 may add the first loss function, the second loss function, and / or the third loss function). In one example, the second weight may be different from the first weight and / or the third weight. In another example, the second weight may correspond to the first weight and / or the third weight. In one aspect, the fourth loss function may be iteratively backpropagated to tune one or more parameters of the convolutional neural network. For example, the convolutional neural network may be modified based on the fourth loss function to improve the classification output from the convolutional neural network. In certain embodiments, the machine learning component 102 (e.g., the loss function component 106) may generate loss function data including the first loss function, the second loss function, the third loss function, and / or the fourth loss function generated by the loss function component 106. For example, the loss function data may include the first loss function associated with multiple masks, the second loss function associated with multiple image-level labels, the third loss function associated with bounding boxes, and / or the fourth loss function that can be used to tune one or more parameters of the convolutional neural network.

[0028] The classification component 108 may predict the classification label of the input image based on the convolutional neural network. In one embodiment, the classification component 108 may generate classification data that may include the classification label of the input image (e.g.,Figure 1 the classification data shown). In addition or alternatively, the classification component 108 may generate localization data that may include a localization map of the input image (e.g., Figure 1 the localization data shown). The localization map of the input image may include information representing, for example, probability scores of one or more regions of the input image. In one embodiment, the localization map of the input image may include a visualization representing probability scores of one or more regions of the input image. In addition or alternatively, the classification component 108 may generate predicted bounding box data (e.g., Figure 1 the predicted bounding box data shown). The predicted bounding box data may be a bounding box of the input image that links one or more regions of interest in the input image to one or more class labels associated with the input image. In one aspect, the predicted bounding box of the input image may provide a location for a region of interest in the input image. In addition or alternatively, the classification component 108 may generate predicted mask data (e.g., Figure 1The predicted mask data shown). The predicted mask data may include a set of masks for the input image. For example, the predicted mask data may include one or more weights for one or more regions of interest in the input image. In one example, the predicted mask data may include a set of pixels that define the location of one or more regions of interest in the input image using, for example, binary filtering. The convolutional neural network employed by the classification component 108 may be of the form of a convolutional neural network tuned based on a fourth loss function. The input image may be, for example, a medical image. The input image may be a two-dimensional image (e.g., a two-dimensional medical image) and / or a three-dimensional image (e.g., a three-dimensional medical image) generated by one or more medical imaging devices. For example, the input image may be a two-dimensional image (e.g., a two-dimensional medical image) and / or a three-dimensional image (e.g., a three-dimensional medical image) generated by an x-ray device, a CT device, another type of medical imaging device, etc. In one example, the input image may be an electromagnetic radiation image captured via a set of sensors (e.g., a set of sensors associated with a medical imaging device). In certain embodiments, the input image may be a series of electromagnetic radiation images captured via a set of sensors (e.g., a set of sensors associated with a medical imaging device) during a time interval. The input image may be received directly from one or more medical imaging devices. Alternatively, the input image may be stored in one or more databases that receive and / or store input images associated with one or more medical imaging devices. In one aspect, the convolutional neural network may include a classification / localization network that outputs a corresponding localization map based on a convolutional feature map. In another aspect, the dimensions of the bounding boxes from one or more bounding boxes may match the dimensions of the convolutional feature map from the convolutional feature map. In another aspect, the dimensions of the masks in the plurality of masks may match the dimensions of the convolutional feature map in the convolutional feature map. In addition or alternatively, based on a mask pooling process, the dimensions of the masks in the plurality of masks may match the dimensions of the convolutional feature map from the convolutional feature map.

[0029] In certain embodiments, the classification component 108 may extract information indicative of relevance, reasoning, and / or expression from an input image based on a convolutional neural network (e.g., a form of convolutional neural tuned based on a fourth loss function). The classification component 108 may generate a learned imaging output based on the execution of at least one machine learning model associated with the convolutional neural network (e.g., a form of convolutional neural tuned based on a fourth loss function). In one aspect, the classification component 108 may generate a learned imaging output. The learned imaging output generated by the classification component 108 may include, for example, learning, relevance, reasoning, and / or expression associated with the input image. In one aspect, the classification component 108 may perform learning explicitly or implicitly with respect to the input image using a convolutional neural network (e.g., a form of convolutional neural tuned based on a fourth loss function). The classification component 108 may also employ an automatic classification system and / or an automatic classification process to facilitate the analysis of the input image. For example, the classification component 108 may employ probability- and / or statistics-based analysis (e.g., taking into account analysis utility and cost) to learn and / or generate reasoning with respect to the input image. The classification component 108 may employ, for example, a support vector machine (SVM) classifier to learn and / or generate reasoning about the imaging data. In addition or alternatively, the classification component 108 may employ other classification techniques associated with Bayesian networks, decision trees, and / or probabilistic classification models. The classifier employed by the classification component 108 may be explicitly trained (e.g., via general training data) and implicitly trained (e.g., via receiving external information). For example, with respect to an SVM, the SVM may be configured via a learning or training phase within a classifier constructor and a feature selection module. A classifier may be a function that maps an input attribute vector x = (x1, x2, x3, x4, xn) to a confidence that the input belongs to a class - i.e., f(x) = confidence(class).

[0030] It should be appreciated that the technical features of the machine learning component 102 are highly technical in nature and are not abstract ideas. The processing threads of the machine learning component 102 for processing and / or analyzing medical imaging data, determining abnormal medical imaging data, etc. cannot be performed by a human (e.g., the processing threads exceed the mental capabilities of a single human). For example, the amount of medical imaging data processed by the machine learning component 102 within a particular time period, the speed of processing the medical imaging data, and / or the data type of the medical imaging data processed may be greater, faster, and different, respectively, compared to the amount, speed, and data type that a single human mind can process within the same time period. In addition, the medical imaging data processed by the machine learning component 102 may be one or more medical images generated by a sensor of a medical imaging device. In addition, the machine learning component 102 may be fully operational (e.g., fully powered on, fully executing, etc.) for performing one or more other functions while also processing the medical imaging data.

[0031] Now refer toFigure 2 , which shows a non - limiting specific implementation of the system 200 according to various aspects and specific implementations of the present disclosure. For the sake of brevity, the repeated description of similar elements employed in other implementation schemes described herein is omitted.

[0032] The system 200 includes a machine learning component 102. The machine learning component 102 may include a training component 104, a loss function component 106, a classification component 108, a visualization component 202, a processor 110, and / or a memory 112. In one implementation, the loss function component 106 may include a first loss function component 109, a second loss function component 111, a third loss function component 113, and / or a fourth loss function component 115. The visualization component 202 may generate a multi - dimensional visualization associated with the classification label of the input image classified by the classification component 108. In addition or alternatively, the visualization component 202 may generate a multi - dimensional visualization associated with the localization information of the input image classified by the classification component 108. For example, the visualization component 202 may generate a human - interpretable visualization of the classification label of the input image and / or the localization information of the input image. In addition or alternatively, the visualization component 202 may generate a human - interpretable visualization of the input image and / or medical imaging data. In one implementation, the visualization component 202 may generate deep - learning data based on the classification and / or localization of a portion of the anatomical region associated with the input image. The deep - learning data may include, for example, the classification and / or location of one or more diseases located in the input image. In certain implementations, the deep - learning data may include probability data indicating the probability of one or more diseases located in the input image. The probability data may be, for example, a probability array of data values of one or more diseases located in the input image. In addition or alternatively, the visualization component 202 may generate a multi - dimensional visualization associated with the classification and / or localization of a portion of the anatomical region associated with the input image.

[0033] Multidimensional visualization can be a graphical representation of an input image that shows the classification and / or location with respect to a patient's body of one or more diseases. The visualization component 202 can also generate a display of a multidimensional visualization of a diagnosis provided by a medical imaging diagnostic process. For example, the visualization component 202 can present a 2D visualization of a portion of an anatomical region on a user interface associated with a display of a user device such as, but not limited to, a computing device, a computer, a desktop computer, a laptop computer, a monitor device, a smart device, a smart phone, a mobile device, a handheld device, a tablet, a portable computing device, or another type of user device associated with a display. In one aspect, the multidimensional visualization can include deep learning data. In another aspect, the deep learning data can also be presented as one or more dynamic visual elements on a 3D model. In one implementation, the visualization component 202 can change at least a portion of the visual characteristics (e.g., color, size, hue, shading, etc.) of the deep learning data associated with the multidimensional visualization based on the classification and / or localization of a portion of the anatomical region. For example, based on the results of deep learning and / or medical imaging diagnosis, the classification and / or localization of a portion of the anatomical region can be presented with different visual characteristics (e.g., color, size, hue, or shading, etc.). In another aspect, the visualization component 202 can allow a user to zoom in or out with respect to the deep learning data associated with the multidimensional visualization. For example, the visualization component 202 can allow a user to zoom in or out with respect to the classification and / or location of one or more diseases identified in an anatomical region of a patient's body. In this way, a user can view, analyze, and / or interact with the deep learning data associated with the multidimensional visualization of the input image.

[0034] Now referring to Figure 3 , a non-limiting specific implementation of a system 300 in accordance with various aspects and specific implementations of the present disclosure is shown. For the sake of brevity, repeated descriptions of similar elements employed in other implementations described herein are omitted.

[0035] System 300 includes a machine learning component 102 and a medical imaging diagnostic process 302. The machine learning component 102 can provide classification data and / or localization data to the medical imaging diagnostic process 302. The classification data and / or localization data can include one or more classification and / or localization information associated with the input image. In one aspect, the classification data and / or localization data can be generated by a classification component 108. Additionally or alternatively, in some embodiments, the machine learning component 102 can provide predicted bounding box data and / or predicted mask data to the medical imaging diagnostic process 302. In one aspect, the medical imaging diagnostic process 302 can perform deep learning to facilitate classification and / or localization of one or more diseases associated with the input image and / or medical imaging data. In another aspect, the medical imaging diagnostic process 302 can perform deep learning based on a convolutional neural network that receives the input image and / or medical imaging data. Diseases classified and / or localized by the medical imaging diagnostic process 302 can include, for example, lung diseases, heart diseases, tissue diseases, bone diseases, tumors, cancers, tuberculosis, heart hypertrophy, pulmonary hypoinflation, pulmonary opacity, hypertension, spinal degenerative diseases, calcinosis, or other types of diseases associated with an anatomical region of a patient's body. In one aspect, the medical imaging diagnostic process 302 can determine a prediction of a disease associated with the input image and / or medical imaging data. For example, the medical imaging diagnostic process 302 can determine a probability score of a disease associated with the input image and / or medical imaging data (e.g., a first percentage value representing the likelihood of a negative prognosis of the disease and a second value representing the likelihood of a positive prognosis of the disease).

[0036] Now referring to Figure 4 , a non-limiting specific implementation of a system 400 according to various aspects and specific implementations of the present disclosure is shown. For the sake of brevity, repeated descriptions of similar elements employed in other embodiments described herein are omitted.

[0037] System 400 may be a classification - localization network. In one implementation, system 400 may represent a machine learning process and / or another process performed by machine learning components 102 (e.g., training component 104, loss function component 106, classification component 108, and / or visualization component 202). Image 402 (e.g., input image) may be processed by convolutional neural network 404. Image 402 may be, for example, a medical image. For example, image 402 may be a two - dimensional image (e.g., two - dimensional medical image) and / or a three - dimensional image (e.g., three - dimensional medical image) generated by one or more medical imaging devices. In one example, image 402 may be a two - dimensional image (e.g., two - dimensional medical image) and / or a three - dimensional image (e.g., three - dimensional medical image) generated by an x - ray device, a CT device, another type of medical imaging device, etc. In another example, image 402 may be an electromagnetic radiation image captured via a set of sensors (e.g., a set of sensors associated with a medical imaging device). In certain implementations, image 402 may be a series of electromagnetic radiation images captured via a set of sensors (e.g., a set of sensors associated with a medical imaging device) during a time interval. Image 402 may be received directly from one or more medical imaging devices. Alternatively, image 402 may be stored in one or more databases that receive and / or store images 402 associated with one or more medical imaging devices. In one implementation, image 402 may be an input image analyzed by machine learning components 102 (e.g., an input image classified by classification component 108).

[0038] The convolutional neural network 404 can output a convolutional feature map 406, which can be adopted by a convolutional neural network 408 (e.g., a classification and localization network) that creates a scoring map 410. In one aspect, the convolutional neural network 404 can encode an image 402 into the convolutional feature map 406. In one implementation, the convolutional neural network 404 can be a spring network of convolutional layers. For example, the convolutional neural network can perform multiple sequential and / or parallel downsamplings and upsamplings on the image 402 associated with the convolutional layers of the convolutional neural network 404 to generate the convolutional feature map 406. In one example, the convolutional neural network 404 can perform a first convolutional layer process associated with the sequential downsampling of the image 402 and a second convolutional layer process associated with the sequential upsampling of the image 402 to generate the convolutional feature map 406. The spring network of convolutional layers can include a first convolutional layer process associated with sequential downsampling and a second convolutional layer process associated with sequential upsampling. The spring network of convolutional layers associated with the convolutional neural network can change the convolutional layer filters, similar to the function of a spring. For example, the convolutional neural network 404 can analyze the image 402 based on a first convolutional layer filter including a first size, a second convolutional layer filter including a second size different from the first size, and a third convolutional layer filter including the first size associated with the first convolutional layer filter to generate the convolutional feature map 406. The convolutional feature map 406 can be, for example, data representing the output of the convolutional layer filter applied to the previous convolutional layer. For example, the first convolutional feature map from the convolutional feature map 406 can include first data representing the output of the first convolutional layer filter applied to the previous convolutional layer, the second convolutional feature map from the convolutional feature map 406 can include second data representing the output of the second convolutional layer filter applied to the previous convolutional layer, the third convolutional feature map from the convolutional feature map 406 can include third data representing the output of the third convolutional layer filter applied to the previous convolutional layer, and so on. In another implementation, the convolutional neural network 408 can be a 1x1 convolutional layer that generates a scoring map 410 based on the convolutional feature map 406. The scoring map 410 can include predicted scores for the classes associated with the regions of interest of the image 402.

[0039] In one aspect, during the training of the convolutional neural network 404, the mask 416 of the image 402 can be matched in size to the convolutional feature map 406 via mask pooling 418. For example, mask pooling 418 can compare the mask 516 to a downsampled mask 420 (e.g., a predicted mask). The size of the downsampled mask 420 can correspond to the size of the mask 416, for example. In one example, during the training of the convolutional neural network 404, the mask 416 can be a mask of the region of interest of the image 402 that matches the size of at least one convolutional feature map from the convolutional feature map 406. Additionally, mask pooling 418 can perform reasonable mask pooling to compare the mask 416 (e.g., a predicted mask) to a downsampled mask 420 of the same size (e.g., a downsampled ground truth mask). In one implementation, the class label of the image 402 can be implicit and can be determined based on the mask 416. For example, mask elements above a defined threshold associated with the mask 416 can signal the presence of a class. For testing, the scoring map 410 can provide a predicted classification label with a localization map 422. The localization map 422 can include information representing probability scores of one or more regions of the image 402, for example. In certain implementations, the localization map 422 can include a visualization of the probability scores of one or more regions of the image 402.

[0040] The system 400 can also include a decoder 411. The decoder 411 can include an upsampling 412 and / or a convolutional neural network layer 414. In one aspect, the decoder 411 can be implemented as a repeatable segmentation network, where the upsampling 412 and the convolutional neural network layer 414 can be blocks repeated a certain number of times. In another aspect, the decoder 411 can generate the localization map 422. For example, the decoder 411 can perform a decoding process associated with the upsampling 412 and / or the convolutional neural network layer 414 to generate the localization map 422. The decoder 411 can provide improved localization results associated with the image 402. In one implementation, during the training of the convolutional neural network 404, multiple decoder blocks associated with the decoder 411 can be considered hyperparameters. In another implementation, the upsampling 412 can perform bilinear interpolation to upsample the scoring map 412 to a specific size. In another implementation, the convolutional neural network layer 414 can be configured as an identification network that includes a set of filters, a batch normalization process, and / or a set of rectified linear units to generate a set of predictions for the localization map 422. The decoder 411 can also provide a smoother and more accurate heat map in the final classification and / or localization results of the image 402. In another aspect, the system 400 can provide improved performance of the classifier based on the mask 416 related to the region of interest and / or the image-level label of the image 402.

[0041] System 400 may also include global pooling 424, prediction labels 426, and / or image-level labels 428 to facilitate improving classification accuracy given weak and more abundant annotation information. Global pooling 424 may perform a global pooling process (e.g., global average pooling process) associated with the score map 410. For example, global pooling 424 may modify the dimensions of the score map 410 (e.g., reduce dimensions or increase dimensions). Prediction labels 426 may be generated based on the score map 410 and the image-level labels 428. For example, image-level labels 428 and global pooling 424 of the score map 410 may be employed to generate prediction labels 426. Image-level labels 428 may be a set of labels for a set of images, where each image is annotated with a label. The label may be a description associated with the image (e.g., a textual description of a disease, etc.). For example, an image associated with the image-level labels 428 may be labeled with a specific disease included in the image. Prediction labels 426 may include one or more predicted categories for the score map 410. For example, prediction labels 426 may be a set of predicted class labels for the score map 410. In one embodiment, system 400 may also include bounding boxes 430. Bounding boxes 430 may include one or more bounding boxes. For example, bounding boxes 430 may be one or more bounding boxes that may link one or more regions of interest to one or more class labels associated with the prediction labels 426 and / or the image-level labels 428. In one example, bounding boxes 430 may link regions of interest in the image 402 to class labels associated with the prediction labels 426 and / or the image-level labels 428. In one aspect, the dimensions of the bounding boxes 430 may match the dimensions of the convolutional feature map 406. For example, bounding boxes 430 may provide positions for regions of interest in the image 402, and the dimensions of the bounding boxes 430 may match the dimensions of at least one convolutional feature map from the convolutional feature map 406. In one embodiment, bounding boxes 430 may be used to generate bounding box predictions 432. The bounding box predictions 432 may be, for example, predicted bounding boxes. For example, bounding boxes 430 may be ground truth bounding boxes, and bounding box predictions 432 may be predicted bounding boxes. In one aspect, bounding box predictions 432 may provide predictions for the bounding boxes 430 using, for example, object detection techniques. In certain embodiments, bounding box predictions 432 may be generated based on a convolutional neural network 434. The convolutional neural network 434 may be, for example, a 1x1 convolutional layer to facilitate object detection within the image 402.

[0042] Now referring Figure 5 , a non-limiting specific implementation of a system 500 in accordance with various aspects and specific implementations of the present disclosure is shown. For the sake of brevity, repeated descriptions of similar elements employed in other implementations described herein are omitted.

[0043] System 500 can be a classification - localization network including a loss function 502. In one embodiment, system 500 can represent a machine - learning process and / or another process executed by machine - learning components 102 (e.g., training component 104, loss - function component 106, first loss - function component 109, second loss - function component 111, third loss - function component 113, fourth loss - function component 115, classification component 108, and / or visualization component 202). System 500 can include an image 402, a convolutional neural network 404, a convolutional feature map 406, a convolutional neural network 408, a scoring map 410, and a decoder 411 including an upsampling 412 and a convolutional - neural - network layer 414. System 500 can also include a mask 416, a mask pooling 418, a downsampled mask 420, a localization map 422, a global pooling 424, a predicted label 426, an image - level label 428, a bounding box 430, a bounding - box prediction 432, a convolutional neural network 434, and a loss function 502. The loss function 502 can be a loss function created during the training of the convolutional neural network 404 based on the downsampled mask 420 (e.g., the downsampled ground - truth mask) and the mask 416 (e.g., the predicted mask). Additionally or alternatively, the loss function 502 can be created based on the predicted label 426 and / or the image - level label 428. Additionally or alternatively, the loss function 502 can be created based on the bounding box 430 and / or the bounding - box prediction 432. In one embodiment, the loss function 502 can correspond to the fourth loss function generated by the fourth loss - function component 115. The loss function 502 can be represented, for example, by the following equation:

[0044] LOss = λ mask ηLoss mask +λ bbx (1 - η)γLoss bbx +λ label (1 - η)(1 - γ)Loss label

[0045] In the case where Loss mask can correspond to a first loss function associated with multiple masks, Loss label can correspond to a second loss function associated with multiple image - level labels, and Loss bbx can correspond to a third loss function associated with one or more bounding boxes. Additionally, λ mask can be the first weight of the first loss function Loss mask λ label can be the second weight of the second loss function Loss label and λ bbx can be the third weight of the third loss function Loss bbxThe third weight. Additionally, η can be a variable indicating whether a mask exists, and Υ can be a variable indicating whether a bounding box exists. For example, the variable η can have a value equal to 0 or 1 depending on the existence of a mask, and the variable Υ can have a value equal to 0 or 1 depending on the existence of a bounding box. In one embodiment, the first loss function Loss mask can be equal to:

[0046]

[0047] where is the probability that image i is positive for class k with respect to the total area in image i and / or the area covered by the mask. Additionally, the second loss function Loss label can be equal to:

[0048]

[0049] where is the probability that image i is positive for class k with respect to the total area in image i and / or the area covered by the image-level label. Additionally, the third loss function Loss bbx can be equal to:

[0050]

[0051] where is the probability that image i is positive for class k with respect to the total area in image i and / or the region of interest covered by the bounding box. Loss mask can correspond to the first loss function generated by the first loss function component 109, Loss labels can correspond to the second loss function generated by the second loss function component 111, Loss bbx can correspond to the third loss function generated by the third loss function component 113, and Loss can correspond to the fourth loss function generated by the fourth loss function component 115. In one embodiment, Loss (e.g., the fourth loss function) can be equal to w1*Loss labels +w2*Loss mask +w3*Loss bbx , where w1 is the first weight, w2 is the second weight, and w3 is the third weight. Additionally, y k can be the k-th output from the convolutional neural network 404, which represents whether image i is positive for class k, where x iis the i-th image. In one embodiment, the loss function 502 can be generated based on the downsampling mask 420, the localization map 422, the predicted label 426, and / or the image-level label 428. For example, the loss function 502 can be generated based on the first probability of the class associated with the downsampling mask 420 and the localization map 422. In addition or alternatively, the loss function 502 can be generated based on the second probability of the class associated with the predicted label 426, the image-level label 428, and / or the localization map 422. In addition or alternatively, the loss function 502 can be generated based on the third probability class of the class associated with the bounding box 430 and / or the bounding box prediction 432. Further, the loss function 502 can be provided to the convolutional neural network layer 414. Additionally, the loss function 502 can be backpropagated from the convolutional neural network layer 414 to the convolutional neural network 404. For example, the loss function 502 can be backpropagated through the system 500, starting from the convolutional neural network layer 414 and ending at the convolutional neural network 404. In one embodiment, the loss function 502 can be backpropagated through the localization map 422, the convolutional neural network layer 414, the upsampling 412, the scoring map 410, the convolutional neural network 408, the convolutional feature map 406, and / or the convolutional neural network 404. In addition or alternatively, the loss function 502 can be backpropagated through the predicted label 426, the global pooling 424, the scoring map 410, the convolutional neural network 408, the convolutional feature map 406, and / or the convolutional neural network 404. In this way, the weighted loss associated with the image-level label 428 and / or the downsampling mask 420 can be backpropagated to the classification loss and / or the segmentation loss associated with the convolutional neural network 404. In one aspect, the loss function 502 can tune one or more parameters of the convolutional neural network 404. For example, the convolutional neural network 404 can be modified based on the loss function 502 to improve the classification and / or localization results associated with the localization map 422.

[0052] Now referring to Figure 6 , a non-limiting example of the loss function 502 in accordance with various aspects and embodiments of the present disclosure is shown. For the sake of brevity, repeated descriptions of similar elements employed in other embodiments described herein are omitted.

[0053] As described above, the loss function 502 can be represented by the following equation:

[0054] Loss = λ mask ηLoss mask + λ bbx (1 - η)γLoss bbx + λ label (1 - η)(1 - γ)Loss label

[0055] For example, the loss function 502 can be generated based on a first probability of a category associated with the downsampled mask 420 and / or the localization map 422. In addition or alternatively, the loss function 502 can be generated based on a second probability of a category associated with the predicted label 426, the image-level label 428, and / or the localization map 422. In addition or alternatively, the loss function 502 can be generated based on a third probability of a category associated with the bounding box 430 and / or the bounding box prediction 432. By employing the loss function 502 and / or the annotation information (e.g., the mask 416 and / or the downsampled mask 420), the classification accuracy can be improved. The system 400 and / or the system 500 can also output an improved localization map (e.g., a more accurate localization map). For example, the loss function 502 and / or the annotation information (e.g., the mask 416 and / or the downsampled mask 420) can be used to provide improved localization information associated with the localization map 422.

[0056] In a non-limiting implementation using the system 400 and / or the system 500, the experiments conducted on the dataset can consist of medical and non-medical condition X-ray images extracted from a database. The medical conditions can include, for example, lung diseases, heart diseases, tissue diseases, bone diseases, tumors, cancers, tuberculosis, heart hypertrophy, pulmonary hypoinflation, pulmonary opacity, hypertension, spinal degenerative diseases, calcinosis, pneumothorax, or other types of medical conditions associated with an anatomical region of the patient's body. The medical condition masks can be annotated by a radiologist, for example. A total of 1806 images can be divided into 1444 images for training (e.g., 80% of the images), 180 images for validation (e.g., 10% of the images), and 182 images for testing (10% of the images), as shown in Table I below. The experimental results are shown in Table II below. The test accuracy of the system 400 and / or the system 500 is 0.923, and the AUC is 0.979, and the dice coefficient is 0.5, which is better than the traditional classification network trained only with image-level labels.

[0057] Table I - Description of the Medical Condition Dataset

[0058] Dataset Training (80%) Validation (10%) Testing (10%) Medical Condition 722 90 91 Non - Medical Condition 722 90 91 Total 1444 180 182

[0059] Table II - Experimental Results

[0060]

[0061] As can be seen from the experimental results in Table II, by providing richer annotation information (e.g., masks), the classification accuracy can be improved, and the convolutional neural network can also output an improved localization map (e.g., a more accurate localization map). This can be achieved by the same underlying prediction model for both tasks. Due to the optional convolutional neural network framework associated with system 400 and / or system 500, the repeatable segmentation network associated with system 400 and / or system 500, and the tunable mask size associated with system 400 and / or system 500, system 400 and / or system 500 can also be flexible and generalized to other applications. In this way, system 400 and / or system 500 can jointly model classification and / or localization. In addition, system 400 and / or system 500 can apply classification and / or localization to disease detection (e.g., medical condition detection, etc.) in medical imaging data (e.g., X-ray images) and / or other digital images.

[0062] Reference is now made to Figure 7 , which shows a non-limiting specific implementation of system 700 in accordance with various aspects and specific implementations of the present disclosure. For the sake of brevity, repeated descriptions of similar elements employed in other embodiments described herein are omitted.

[0063] System 700 includes a convolutional neural network 702, a classification detection segmentation network 704, and a loss function 706. The convolutional neural network 702 can be a deep artificial neural network associated with machine learning. In one embodiment, the convolutional neural network 702 can encode an image 701 into a set of convolutional feature maps. The image 701 can be, for example, the input image of the convolutional neural network 702. In one embodiment, the image 701 can be a medical image. For example, the image 701 can be a two-dimensional image (e.g., a two-dimensional medical image) and / or a three-dimensional image (e.g., a three-dimensional medical image) generated by one or more medical imaging devices. In one example, the image 701 can be a two-dimensional image (e.g., a two-dimensional medical image) and / or a three-dimensional image (e.g., a three-dimensional medical image) generated by an x-ray device, a CT device, another type of medical imaging device, etc. In another example, the image 701 can be an electromagnetic radiation image captured via a set of sensors (e.g., a set of sensors associated with a medical imaging device). In certain embodiments, the image 701 can be a series of electromagnetic radiation images captured via a set of sensors (e.g., a set of sensors associated with a medical imaging device) during a time interval. The image 701 can be received directly from one or more medical imaging devices. Alternatively, the image 701 can be stored in one or more databases that receive and / or store images 701 associated with one or more medical imaging devices. In one embodiment, the image 701 can correspond to the image 402 and / or the medical imaging data received by the machine learning component 102.

[0064] In certain embodiments, the convolutional neural network 702 can be a spring network of convolutional layers. For example, the convolutional neural network 702 can perform multiple sequential and / or parallel downsampling and upsampling on an image 701 associated with the convolutional layers of the convolutional neural network 702 to generate, for example, a set of convolutional feature maps. In one example, the convolutional neural network 702 can perform a first convolutional layer process associated with sequential downsampling of the image 701 and a second convolutional layer process associated with sequential upsampling of the image 701 to generate, for example, a set of convolutional feature maps. The spring network of convolutional layers can include a first convolutional layer process associated with sequential downsampling and a second convolutional layer process associated with sequential upsampling. The spring network of convolutional layers associated with a convolutional neural network can change the convolutional layer filters, similar to the function of a spring. For example, the convolutional neural network 702 can analyze the image 701 based on a first convolutional layer filter including a first size, a second convolutional layer filter including a second size different from the first size, and a third convolutional layer filter including the first size associated with the first convolutional layer filter to generate, for example, a set of convolutional feature maps. In certain embodiments, the convolutional neural network 702 can correspond to the convolutional neural network 404.

[0065] The classification detection segmentation network 704 can be used to classify and / or localize one or more regions of interest associated with the image 701. For example, the classification detection segmentation network 704 can receive a set of convolutional feature maps generated by the convolutional neural network 702 to facilitate the classification and / or localization of one or more regions of interest associated with the image 701. In one example, the classification detection segmentation network 704 can generate a score map associated with the classification and / or localization of one or more regions of interest associated with the image 701. The score map can provide, for example, a predicted classification label with a localization map. In certain embodiments, the classification detection segmentation network 704 can correspond to the convolutional neural network 408, the score map 410, the decoder 411 (e.g., upsampling 412 and / or convolutional neural network layer 414), the global pooling 424, and / or the predicted label 426. The loss function 706 can be a loss function created based on the mask 708, the bounding box 714, and / or the label 720. The mask 708 can be generated, for example, during the training of the convolutional neural network 702. The mask 708 can include one or more weights of one or more regions of interest in the image 701 (e.g., the image 701 provided to the convolutional neural network 702). In one example, the mask 708 can include a set of pixels that use binary filtering to define the positions of the regions of interest in the image 701 (e.g., the image 701 provided to the convolutional neural network 702). In one embodiment, the mask 708 can correspond to the mask 416. The bounding box 714 can link one or more regions of interest to one or more class labels associated with the label 720. In one example, the bounding box 714 can link the region of interest in the image 701 (e.g., the image 701 provided to the convolutional neural network 702) to the class label associated with the label 720. In one aspect, the size of the bounding box 714 can match the convolutional feature maps generated by the convolutional neural network 702. For example, the bounding box 714 can provide the position of the region of interest in the image 701 (e.g., the image 701 provided to the convolutional neural network 702), and the size of the bounding box 714 can match the size of at least one convolutional feature map generated by the convolutional neural network 702. In one example, the bounding box 714 can be a predicted bounding box. In another example, the bounding box 714 can be a ground truth bounding box. In one embodiment, the bounding box 714 can correspond to the bounding box 430 and / or the bounding box prediction 432. The label 720 can be, for example, an image-level label. In one aspect, the label 720 can be a description associated with the image 701 (e.g., the image 701 provided to the convolutional neural network 702) (e.g., a text description of a disease, etc.). For example, the label 720 can label at least a portion of the image 701 (e.g., the image 701 provided to the convolutional neural network 702) with a specific disease included in the image 701.In some embodiments, the label 720 may label one or more regions of interest in the image 701 (e.g., the image 701 provided to the convolutional neural network 702). For example, the label 720 may label (e.g., provide a description of) the region of interest associated with the bounding box 714. In one embodiment, the label 720 may correspond to the image-level label 428 and / or the prediction label 426. In one embodiment, the loss function 706 may correspond to the fourth loss function generated by the fourth loss function component 115. For example, the loss function 706 may correspond to the loss function 502. In some embodiments, the mask 708 may be processed by dilation 710 and / or pooling 712 before being received by the classification detection segmentation network 704. The dilation 710 may be, for example, a dilated pooling process. For example, the dilation 710 may be a convolution applied to the mask 708 having a set of defined gaps. The pooling 712 may be, for example, a mask pooling process that compares the mask 708 with a downsampled version of the mask 708 (e.g., the predicted mask). In some embodiments, the bounding box 714 may be processed by the padding mask 716 and / or pooling 718 before being received by the classification detection segmentation network 704. The padding mask 716 may be, for example, a process of padding at least a portion of the bounding box 714 with the mask. The pooling 712 may be, for example, a mask pooling process that compares the bounding box 714 with the mask 708 and / or a downsampled version of the mask 708 (e.g., the predicted mask).

[0066] Figure 8 An exemplary multi-dimensional visualization 800 and an exemplary input image 801 in accordance with various aspects and embodiments described herein are shown. In Figure 8In the illustrated embodiment, the multi-dimensional visualization 800 may display a medical imaging diagnosis for a patient, for example. For example, the multi-dimensional visualization 800 may display one or more classifications and / or one or more localizations of one or more medical conditions identified in the imaging data (e.g., input image 801). However, it should be understood that the multi-dimensional visualization 800 may be associated with another type of classification and / or localization of one or more features located in the imaging data. In one aspect, the multi-dimensional visualization 800 may include localization data 802 for medical imaging diagnosis. The localization data 802 may be a predicted location of a medical condition associated with the input image and / or medical imaging data processed by the machine learning component 102. Based on the information provided by the machine learning component 102, the visual characteristics (e.g., color, size, hue, shade, etc.) of the localization data 802 may be dynamic. For example, a first portion of the localization data 802 may include a first visual characteristic, a second portion of the localization data 802 may include a second visual characteristic, a third portion of the localization data 802 may include a third visual characteristic, and so on. In one embodiment, the display environment associated with the multi-dimensional visualization 800 may include a heat bar 804. The heat bar 804 may include a set of colors corresponding to different values of the localization data 802. For example, a first color (e.g., red) in the heat bar 804 may correspond to a first value of the localization data 802, a second color (e.g., green) in the heat bar 804 may correspond to a second value of the localization data 802, a third color (e.g., blue) in the heat bar 804 may correspond to a third value of the localization data 802, and so on.

[0067] Figure 9 A method and / or flowchart in accordance with the disclosed subject matter is shown. For simplicity of illustration, the method is described and shown as a series of acts. It should be understood and appreciated that the present invention is not limited by the acts and / or order of acts shown, e.g., acts may occur in various orders and / or concurrently, and have other acts not presented and described herein. Further, not all acts shown may be required to implement the method in accordance with the disclosed subject matter. In addition, those skilled in the art will understand and appreciate that the method may alternatively be represented as a series of interrelated states via a state diagram or events. Additionally, it should be further appreciated that the methods disclosed herein and throughout the specification are capable of being stored on an article of manufacture to facilitate the transfer and conveyance of such methods to a computer. As used herein, the term "article of manufacture" is intended to encompass a computer program accessible from any computer-readable device or storage medium.

[0068] See Figure 9, which shows a non-limiting specific implementation of a method 900 for classification and / or localization based on annotation information according to an aspect of the present subject innovation. At 902, a plurality of images associated with a plurality of patients are received from at least one imaging device (e.g., by the training component 104). The plurality of images may be associated with a plurality of patients. Additionally, the plurality of images may be a set of medical images. The plurality of images may be two-dimensional images and / or three-dimensional images generated by one or more medical imaging devices. For example, the plurality of images may be electromagnetic radiation images captured via a set of sensors (e.g., a set of sensors associated with a medical imaging device). In certain embodiments, the plurality of images may be a series of electromagnetic radiation images captured via a set of sensors (e.g., a set of sensors associated with a medical imaging device) during a time interval. The plurality of images may be received directly from one or more medical imaging devices. Alternatively, the plurality of images may be stored in one or more databases that receive and / or store the plurality of images associated with one or more medical imaging devices. The medical imaging device may be, for example, an x-ray device, a CT device, another type of medical imaging device, etc. In one embodiment, each image from the plurality of images may be associated with one or more masks.

[0069] At 904, a plurality of masks are received from a plurality of objects (e.g., by the training component 104), where each image includes at least one mask associating an object of interest with a corresponding class label, at least one image-level label for the image, and / or a bounding box linking the object of interest to the corresponding class label. The mask can be a filter for masking one or more regions in an image (e.g., an image from a plurality of images). For example, the mask can include one or more weights for one or more regions of interest in the image (e.g., an image from a plurality of images). In one example, the mask can include a set of pixels that define the location of the region of interest using binary filtering. The at least one image-level label can be a set of labels for a set of images, where each image is annotated with a label. The label can be a description associated with the image (e.g., a text description of a disease, etc.). For example, an image associated with at least one image-level label can be labeled with a specific disease included in the image. In one embodiment, at least one image-level label can be used to generate a prediction label associated with a score map. The prediction label can include one or more predicted classes of the score map. For example, the prediction label can be a set of predicted class labels for the score map. The bounding box can link one or more regions of interest to one or more class labels associated with at least one image-level label. In one example, the bounding box can link a region of interest in at least one image from a plurality of images to a class label associated with at least one image-level label. In one aspect, the size of the bounding box can match the convolutional feature map generated by the convolutional neural network. In another aspect, the bounding box can provide the location of the region of interest in the image from a plurality of images. In one example, the bounding box can be a predicted bounding box. In another example, the bounding box can be a ground truth bounding box.

[0070] At 906, a convolutional neural network is trained based on the plurality of images, the plurality of masks, the bounding box, and / or the at least one image-level label (e.g., by the training component 104), where the convolutional neural network includes a decoder, which is composed of at least one upsampling layer and at least one convolutional layer, a pre-trained classifier network that outputs a convolutional feature map, and / or a classification / localization network that outputs a corresponding localization map. The decoder can be implemented as a repeatable segmentation network, where at least one upsampling layer and / or at least one convolutional neural network layer can be a block that is repeated a certain number of times.

[0071] At 908, a first loss function is generated based on multiple masks (e.g., by the first loss function component 109). In one aspect, the first loss function can be generated by using a decoder to generate a localization map. In certain embodiments, the number of decoders associated with the decoder can be determined during the training of the convolutional neural network. In another aspect, the first loss function can be generated based on the probabilities of the classes associated with the multiple masks. In one embodiment, the first loss function can be generated based on a downsampled mask (e.g., a downsampled ground truth mask) and another mask (e.g., a predicted mask) during the training of the convolutional neural network. In another embodiment, the first loss function can be generated based on the downsampled mask and / or the localization map. For example, the first loss function can be generated based on the probabilities of the classes associated with the downsampled mask and / or the mask.

[0072] At 910, a second loss function is generated based on at least one image-level label associated with multiple images (e.g., by the second loss function component 111). In one aspect, the second loss function can be generated by using a decoder to generate a localization map. In certain embodiments, the number of decoders associated with the decoder can be determined during the training of the convolutional neural network. In another aspect, the second loss function can be generated based on the probabilities of the classes associated with the at least one image-level label. In one embodiment, the second loss function can be generated based on the image-level label, the predicted label, and / or the localization map. For example, the second loss function can be generated based on the probabilities of the classes associated with the image-level label, the predicted label, and / or the localization map.

[0073] At 912, a third loss function is generated based on a bounding box that links an object of interest to a corresponding class label (e.g., by the third loss function component 113). The object of interest can be, for example, a region of interest in an image. In one aspect, the third loss function can be generated by using a decoder to generate a localization map. In certain embodiments, the number of decoders associated with the decoder can be determined during the training of the convolutional neural network. In another aspect, the third loss function can be generated based on the probabilities of the classes associated with the bounding box.

[0074] At 914, a fourth loss function is generated based on the first loss function, the second loss function, and the third loss function (e.g., by the fourth loss function component 115). For example, a first weight can be applied to the first loss function, a second weight can be applied to the second loss function, and a third weight can be applied to the third loss function. Additionally, the first loss function, the second loss function, and the third loss function can be combined (e.g., the first loss function, the second loss function, and the third loss function can be added). In one example, the second weight can be different from the first weight and / or the third weight. In another example, the second weight can correspond to the first weight and / or the third weight.

[0075] At 916, the fourth loss function is iteratively backpropagated (e.g., by the third loss function component 113) to tune the parameters of the convolutional neural network based on the training data. For example, the fourth loss function can be provided to at least one convolutional neural network layer of the decoder. Additionally, the fourth loss function can be backpropagated from at least one convolutional neural network layer to the convolutional neural network to modify one or more portions of the convolutional neural network.

[0076] At 918, a classification label, a localization map, and / or a bounding box of the input image are predicted based on the convolutional neural network (e.g., by the classification component 108). The convolutional neural network for predicting the classification label can be a form of convolutional neural tuned based on the fourth loss function. The image can be, for example, a medical image. The input image can be a two-dimensional image (e.g., a two-dimensional medical image) and / or a three-dimensional image (e.g., a three-dimensional medical image) generated by one or more medical imaging devices. For example, the input image can be a two-dimensional image (e.g., a two-dimensional medical image) and / or a three-dimensional image (e.g., a three-dimensional medical image) generated by an x-ray device, a CT device, another type of medical imaging device, etc. In one example, the input image can be an electromagnetic radiation image captured via a set of sensors (e.g., a set of sensors associated with a medical imaging device). In some embodiments, the input image can be a series of electromagnetic radiation images captured via a set of sensors (e.g., a set of sensors associated with a medical imaging device) during a time interval. The input image can be received directly from one or more medical imaging devices. Alternatively, the input image can be stored in one or more databases that receive and / or store input images associated with one or more medical imaging devices. The localization map of the input image can include, for example, information representing probability scores of one or more regions of the input image. In one embodiment, the localization map of the input image can include a visualization representing probability scores of one or more regions of the input image. The bounding box of the input image can be a predicted bounding box that links one or more regions of interest in the input image to one or more class labels associated with the input image. In one aspect, the bounding box of the input image can provide a location for the region of interest in the input image. In some embodiments, method 900 can further include matching the size of the mask from a plurality of masks to the size of the convolutional feature map from the convolutional feature map. In some embodiments, method 900 can further include matching the size of the mask from a plurality of masks to the size of the convolutional feature map from the convolutional feature map based on a mask pooling process. In some embodiments, method 900 can further include generating a multi-dimensional visualization associated with the classification label of the input image. In some embodiments, the decoder can generate the localization map. For example, the decoder can perform a decoding process associated with at least one upsampling layer and / or at least one convolutional neural network layer to generate the localization map.

[0077] The foregoing systems and / or devices have been described with respect to interactions between several components. It should be appreciated that such systems and components can include those components or sub-components specified herein, some of the specified components or sub-components, and / or additional components. Sub-components can also be implemented as components that are communicatively coupled to other components other than those included within a parent component. Still further, one or more components and / or sub-components can be combined into a single component that provides an aggregated function. Components can also interact with one or more other components that are not specifically described herein for the sake of brevity but are known to those of ordinary skill in the art.

[0078] To provide context for various aspects of the disclosed subject matter, Figure 10 and Figure 11 the following discussion is intended to provide a brief general description of a suitable environment in which the various aspects of the disclosed subject matter can be implemented.

[0079] Referring Figure 10 , a suitable environment 1000 for implementing various aspects of the present disclosure includes a computer 1012. The computer 1012 includes a processing unit 1014, a system memory 1016, and a system bus 1018. The system bus 1018 couples system components, including but not limited to the system memory 1016, to the processing unit 1014. The processing unit 1014 can be any of a variety of available processors. Dual microprocessors and other multi-processor architectures can also be used as the processing unit 1014.

[0080] The system bus 1018 can be any of a variety of types of bus structures, including a memory bus or memory controller, a peripheral bus or external bus, and / or a local bus using various available bus architectures, including but not limited to Industry Standard Architecture (ISA), Micro Channel Architecture (MSA), Extended ISA (EISA), Intelligent Drive Electronics (IDE), VESA Local Bus (VLB), Peripheral Component Interconnect (PCI), Card Bus, Universal Serial Bus (USB), Advanced Graphics Port (AGP), Personal Computer Memory Card International Association Bus (PCMCIA), FireWire (IEEE 1394), and Small Computer System Interface (SCSI).

[0081] The system memory 1016 includes volatile memory 1020 and non-volatile memory 1022. The basic input / output system (BIOS) (basic routines that transfer information between elements contained within the computer 1012, such as during startup) is stored in the non-volatile memory 1022. By way of example and not limitation, the non-volatile memory 1022 can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), flash memory, or non-volatile random access memory (RAM) (e.g., ferroelectric RAM (FeRAM)). The volatile memory 1020 includes random access memory (RAM), which acts as an external cache memory. By way of example and not limitation, the RAM can be provided in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), direct Rambus RAM (DRRAM), direct Rambus dynamic RAM (DRDRAM), and Rambus dynamic RAM.

[0082] The computer 1012 also includes removable / non-removable, volatile / non-volatile computer storage media. Figure 10 For example, a disk storage device 1024 is shown. The disk storage device 1024 includes, but is not limited to, devices such as disk drives, floppy disk drives, tape drives, Jaz drives, Zip drives, LS-100 drives, flash memory cards, or memory sticks. The disk storage device 1024 can also include the storage medium alone or in combination with other storage media, which include, but are not limited to, optical disk drives, such as compact disk ROM devices (CD-ROM), CD recordable drives (CD-R drives), CD rewritable drives (CD-RW drives), or digital versatile disk ROM drives (DVD-ROM). To facilitate connection of the disk storage device 1024 to the system bus 1018, a removable / non-removable interface, such as interface 1026, is typically used.

[0083] Figure 10Software acting as an intermediary between the user and the basic computer resources described in the suitable operating environment 1000 is also depicted. For example, such software includes the operating system 1028. The operating system 1028, which can be stored on the disk storage device 1024, is used to control and allocate the resources of the computer system 1012. The system applications 1030 utilize the operating system 1028 for the management of resources through, for example, program modules 1032 and program data 1034 stored in the system memory 1016 or on the disk storage device 1024. It should be recognized that the present disclosure can be implemented with various operating systems or combinations of operating systems.

[0084] The user inputs commands or information into the computer 1012 through the input device 1036. The input device 1036 includes, but is not limited to, pointing devices such as a mouse, trackball, stylus, touchpad, keyboard, microphone, joystick, gamepad, satellite antenna, scanner, TV tuner card, digital camera, digital video camera, web camera, etc. These and other input devices are connected to the processing unit 1014 via the interface port 1038 through the system bus 1018. The interface port 1038 includes, for example, serial ports, parallel ports, game ports, and Universal Serial Bus (USB). Some of the output devices 1040 use the same type of ports as the input device 1036. Thus, for example, a USB port can be used to provide input to the computer 1012 and output information from the computer 1012 to the output device 1040. An output adapter 1042 is provided to illustrate the presence of some output devices 1040 such as monitors, speakers, and printers, as well as other output devices 1040 that require special adapters. By way of example and not limitation, the output adapter 1042 includes video and sound cards that provide a connection means between the output device 1040 and the system bus 1018. It should be noted that other devices and / or systems of devices provide both input capabilities and output capabilities, such as the remote computer 1044.

[0085] The computer 1012 can operate in a networked environment using a logical connection to one or more remote computers, such as remote computer 1044. The remote computer 1044 can be a personal computer, server, router, network PC, workstation, microprocessor-based device, peer device, or other common network node, etc., and generally includes many or all of the elements described relative to the computer 1012. For simplicity purposes, only the memory storage device 1046 is shown for the remote computer 1044. The remote computer 1044 is logically connected to the computer 1012 via a network interface 1048 and then physically connected via a communication connection 1050. The network interface 1048 encompasses wired and / or wireless communication networks, such as local area networks (LANs), wide area networks (WANs), cellular networks, etc. LAN technologies include Fiber Distributed Data Interface (FDDI), Copper Distributed Data Interface (CDDI), Ethernet, Token Ring, etc. WAN technologies include but are not limited to point-to-point links, circuit-switched networks such as Integrated Services Digital Network (ISDN) and its variants thereon, packet-switched networks, and Digital Subscriber Line (DSL).

[0086] The communication connection 1050 refers to the hardware / software for connecting the network interface 1048 to the bus 1018. Although the communication connection 1050 is shown within the computer 1012 for clarity, this communication connection can also be external to the computer 1012. For illustrative purposes only, the hardware / software required to connect to the network interface 1048 includes internal and external technologies, such as modems, including conventional telephone-grade modems, cable modems, and DSL modems, ISDN adapters, and Ethernet cards.

[0087] Figure 11 It is a schematic block diagram of a sample computing environment 1100 with which the subject matter of the present disclosure can interact. The system 1100 includes one or more clients 1110. The one or more clients 1110 can be hardware and / or software (e.g., threads, processes, computing devices). The system 1100 also includes one or more servers 1130. Thus, in addition to other models, the system 1100 can correspond to a two-tier client-server model or a multi-tier model (e.g., client, middle-tier server, data server). The server 1130 can also be hardware and / or software (e.g., threads, processes, computing devices). For example, the server 1130 can host threads to perform conversions by adopting the present disclosure. One possible communication between the client 1110 and the server 1130 can be in the form of data packets transmitted between two or more computer processes.

[0088] System 1100 includes a communication framework 1150 that can be used to facilitate communication between a client 1110 and a server 1130. The client 1110 is operatively connected to one or more client data repositories 1120, which can be used to store information local to the client 1110. Similarly, the server 1130 is operatively connected to one or more server data repositories 1140, which can be used to store information local to the server 1130.

[0089] It should be noted that various aspects or features of the present disclosure can be utilized in substantially any radio telecommunications or radio technology, such as, for example, Wi-Fi; Bluetooth; Worldwide Interoperability for Microwave Access (WiMAX); Enhanced General Packet Radio Service (Enhanced GPRS); 3rd Generation Partnership Project (3GPP) Long Term Evolution (LTE); 3rd Generation Partnership Project 2 (3GPP2) Ultra Mobile Broadband (UMB); 3GPP Universal Mobile Telecommunications System (UMTS); High Speed Packet Access (HSPA); High Speed Downlink Packet Access (HSDPA); High Speed Uplink Packet Access (HSUPA); GSM (Global System for Mobile Communications) EDGE (Enhanced Data Rates for GSM Evolution) Radio Access Network (GERAN); UMTS Terrestrial Radio Access Network (UTRAN); LTE Advanced (LTE-A); etc. Additionally, some or all of the aspects described herein can be utilized in traditional telecommunications technologies (such as, for example, GSM). Furthermore, mobile as well as non-mobile networks (such as, for example, the Internet, data service networks such as Internet Protocol Television (IPTV), etc.) can utilize the aspects or features described herein.

[0090] Although the subject matter has been described above in the general context of computer-executable instructions of a computer program running on one and / or more computers, those skilled in the art will recognize that the present disclosure can also or may be implemented in conjunction with other program modules. Generally, program modules include routines, programs, components, data structures, etc. that perform particular tasks and / or implement particular abstract data types. Additionally, those skilled in the art should recognize that the methods of the present invention can be practiced with other computer system configurations, including single-processor or multi-processor computer systems, small computing devices, mainframe computers, and personal computers, handheld computing devices (such as, for example, PDAs, telephones), microprocessor-based or programmable consumer or industrial electronic products, etc. The illustrated aspects can also be practiced in a distributed computing environment where tasks are performed by remote processing devices linked through a communication network. However, some (if not all) aspects of the present disclosure can be practiced on a stand-alone computer. In a distributed computing environment, program modules can be located in local and remote memory storage devices.

[0091] As used in this application, the terms "component", "system", "platform", "interface", etc. may refer to and / or may include computer-related entities or entities related to an operating machine with one or more specific functions. Entities disclosed herein may be hardware, a combination of hardware and software, software, or software in execution. For example, a component may be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an executing thread, a program, and / or a computer. By way of example, an application running on a server and the server can both be components. One or more components may reside within a process and / or an executing thread, and a component may be located on one computer and / or distributed between two or more computers.

[0092] In another example, corresponding components may execute according to various computer-readable media on which various data structures are stored. Components may communicate, such as via signals having one or more data packets (e.g., data from one component that interacts with another component in a local system, a distributed system, and / or via signals across a network (such as the Internet) with other systems), via local and / or remote processes. As another example, a component may be a device having a specific function provided by a mechanical part operated by an electrical or electronic circuit, where the electrical or electronic circuit is operated by a software or firmware application executed by a processor. In this case, the processor may be internal or external to the device and may execute at least a portion of the software or firmware application. As yet another example, a component may be a device that provides a specific function through electronic components rather than mechanical parts, where the electronic components may include a processor or other means for executing software or firmware that at least partially imparts functionality to the electronic components. In one aspect, a component may be emulated, for example, via a virtual machine within a cloud computing system.

[0093] Furthermore, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or". That is, unless otherwise specified or clear from the context, "X employs A or B" is intended to mean any natural inclusive permutation. That is, if X employs A; X employs B; or X employs both A and B, then "X employs A or B" is satisfied in any of the foregoing instances. Additionally, the articles "a" and "an" as used in this specification and the drawings are generally to be construed to mean "one or more" unless otherwise specified or clear from the context to be the singular form.

[0094] As used herein, the terms "example" and / or "exemplary" are used to mean serving as an example, instance, or illustration. To avoid doubt, the subject matter disclosed herein is not limited by such examples. In addition, any aspect or design described herein as "example" and / or "exemplary" should not necessarily be construed as more preferred or advantageous than other aspects or designs, nor does it imply the exclusion of equivalent exemplary structures and techniques known to those of ordinary skill in the art.

[0095] The various aspects or features described herein can be implemented as a method, apparatus, system, or article of manufacture using standard programming or engineering techniques. In addition, the various aspects or features disclosed in this disclosure can be implemented by program modules that implement at least one or more of the methods disclosed herein, the program modules being stored in a memory and executed by at least a processor. Other combinations of hardware and software, or hardware and firmware, can implement or carry out the aspects described herein, including the disclosed methods. As used herein, the term "article of manufacture" can cover a computer program accessible from any computer-readable device, carrier, or storage medium. For example, computer-readable storage media can include, but are not limited to, magnetic storage devices (e.g., hard disks, floppy disks, magnetic strips...), optical disks (e.g., compact disks (CDs), digital versatile disks (DVDs), Blu-ray disks (BDs)...), smart cards, and flash memory devices (e.g., cards, sticks, key drives...), etc.

[0096] As employed in this specification, the term "processor" can generally refer to any computing processing unit or device, including but not limited to a single-core processor; a single processor with software multithreading execution capabilities; a multi-core processor; a multi-core processor with software multithreading execution capabilities; a multi-core processor with hardware multithreading technology; a parallel platform; and a parallel platform with distributed shared memory. Additionally, a processor can refer to an integrated circuit, an application-specific integrated circuit (ASIC), a digital signal processor (DSP), a field-programmable gate array (FPGA), a programmable logic controller (PLC), a complex programmable logic device (CPLD), discrete gate or transistor logic components, discrete hardware components, or any combination thereof designed to perform the functions described herein. Additionally, a processor can utilize nanoscale architectures (such as but not limited to molecule- and quantum dot-based transistors, switches, and gates) in order to optimize space usage or enhance the performance of user equipment. A processor can also be implemented as a combination of computing processing units.

[0097] In the present disclosure, terms such as "store", "storage device", "data storage", "data storage device", "database", and substantially any other information storage component related to the operation and function of a component are used to refer to a "memory component", an entity embodied in a "memory", or a component including a memory. It should be recognized that the memory and / or memory components described herein can be volatile memory or non-volatile memory, or can include both volatile and non-volatile memory.

[0098] By way of illustration and not limitation, non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable ROM (EEPROM), flash memory, or non-volatile random access memory (RAM) (e.g., ferroelectric RAM (FeRAM)). For example, volatile memory can include RAM, which can serve as an external cache memory. By way of illustration and not limitation, RAM can be provided in various forms, such as synchronous RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), direct Rambus RAM (DRRAM), direct Rambus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM). Additionally, the disclosed memory components of the systems or methods herein are intended to include, but are not limited to, including these and any other suitable types of memory.

[0099] It should be recognized and understood that components described with respect to a particular system or method can include the same or similar functionality as corresponding components (e.g., separately named components or similarly named components) described with respect to other systems or methods disclosed herein.

[0100] The foregoing includes examples of systems and methods that provide the advantages of the present disclosure. Of course, it is not possible to describe every conceivable combination of components or methods for the purpose of describing the present disclosure, but one of ordinary skill in the art can recognize that many additional combinations and permutations of the present disclosure are possible. In addition, to the extent that the terms "comprising", "having", "owning", etc. are used in the detailed description, claims, appendices, and drawings, such terms are intended to be inclusive in a manner similar to the term "comprising" as interpreted when used as a transitional word in a claim.

Claims

1. A machine learning system, the machine learning system comprising: A memory that stores computer-executable components; A processor that executes the computer-executable components stored in the memory, wherein the computer-executable components include: A training component that trains a convolutional neural network based on training data and a plurality of images, wherein the training data is associated with a plurality of patients from at least one imaging device, the plurality of images are three-dimensional images, and wherein the plurality of images are associated with a plurality of masks from a plurality of objects, or a plurality of image-level labels of the plurality of images, or bounding boxes that link regions of interest to class labels; A first loss function component that generates a first loss function based on the plurality of masks; A second loss function component that generates a second loss function based on the plurality of image-level labels of the plurality of images; A third loss function component that generates a third loss function based on the bounding boxes that link regions of interest to the class labels; A fourth loss function component that generates a fourth loss function based on the first loss function, the second loss function, and the third loss function, wherein the fourth loss function is iteratively backpropagated to tune the parameters of the convolutional neural network; and A classification component that predicts a classification label of an input image based on the convolutional neural network, wherein the convolutional neural network includes a pre-trained classifier network that outputs a convolutional feature map, and wherein the size of the mask from the plurality of masks is matched to the size of the convolutional feature map from the convolutional feature map based on a mask pooling process.

2. The machine learning system according to claim 1, wherein the convolutional neural network includes a classification / localization network that outputs a corresponding localization map based on the convolutional feature map.

3. The machine learning system according to claim 1, wherein the size of the bounding box is matched to the size of the convolutional feature map from the convolutional feature map.

4. The machine learning system according to claim 1, wherein the first loss function component generates the first loss function based on the probability of the class associated with the plurality of masks.

5. The machine learning system according to claim 1, wherein the second loss function component generates the second loss function based on the probability of the class associated with the plurality of image-level labels.

6. The machine learning system according to claim 1, wherein the fourth loss function component applies a first weight to the first loss function, a second weight to the second loss function, and a third weight to the third loss function.

7. The machine learning system according to claim 1, wherein the computer-executable components further include: A visualization component that generates a multi-dimensional visualization associated with the classification label of the input image.

8. A method, the method comprising using a processor operatively coupled to a memory to execute computer-executable components to perform the following actions: Receiving a plurality of images associated with a plurality of patients from at least one imaging device, wherein the plurality of images are three-dimensional images; Receiving a plurality of masks from a plurality of objects, wherein each image includes at least one mask associating an object of interest with a corresponding class label, or at least one image-level label of the image, or a bounding box linking the object of interest to the corresponding class label; Training a convolutional neural network based on the plurality of images, the plurality of masks, the bounding box, and / or the at least one image-level label, wherein the convolutional neural network includes a pre-trained classifier network that outputs a convolutional feature map and a classification / localization network that outputs a corresponding localization map; Matching the size of the masks from the plurality of masks with the size of the convolutional feature map from the convolutional feature map based on a mask pooling process; Generating a first loss function based on the plurality of masks; Generating a second loss function based on the at least one image-level label of the image; Generating a third loss function based on the bounding box linking the object of interest to the corresponding class label; Generating a fourth loss function based on the first loss function, the second loss function, and the third loss function; Iteratively backpropagating the fourth loss function to tune the parameters of the convolutional neural network; and Predicting a classification label of an input image based on the convolutional neural network.

9. The method according to claim 8, further comprising matching the size of the bounding box with the size of the convolutional feature map from the convolutional feature map.

10. The method according to claim 8, wherein generating the first loss function includes generating the first loss function based on the probability of the class associated with the plurality of masks.

11. The method according to claim 8, wherein generating the second loss function includes generating the second loss function based on the probability of the class associated with the at least one image-level label.

12. The method according to claim 8, wherein generating the fourth loss function includes applying a first weight to the first loss function, applying a second weight to the second loss function, and applying a third weight to the third loss function.

13. The method according to claim 8, further comprising generating a multi-dimensional visualization associated with the classification label of the input image.

14. A computer-readable storage device comprising instructions that, when executed, cause a system including a processor to perform operations, the operations including: Receiving a plurality of images associated with a plurality of patients from at least one imaging device, wherein the plurality of images are three-dimensional images; Receiving a plurality of masks from a plurality of objects, wherein each image includes at least one mask associating an object of interest with a corresponding class label, or at least one image-level label of the image, or a bounding box linking the object of interest to the corresponding class label; Train a convolutional neural network based on the plurality of images, the plurality of masks, the bounding boxes, and / or the at least one image-level label, wherein the convolutional neural network includes a pre-trained classifier network that outputs a convolutional feature map and a classification / localization network that outputs a corresponding localization map; Match the size of the masks from the plurality of masks to the size of the convolutional feature map from the convolutional feature map based on a mask pooling process; Generate a first loss function based on the plurality of masks; Generate a second loss function based on the at least one image-level label of the image; Generate a third loss function based on the bounding boxes that link the object of interest to the corresponding class labels; Generate a fourth loss function based on the first loss function, the second loss function, and the third loss function; Iteratively backpropagate the fourth loss function to tune the parameters of the convolutional neural network; and Predict a classification label of an input image based on the convolutional neural network.

15. The computer-readable storage device according to claim 14, wherein generating the first loss function includes generating the first loss function based on the probability of the class associated with the plurality of masks.

16. The computer-readable storage device according to claim 14, wherein generating the second loss function includes generating the second loss function based on the probability of the class associated with the at least one image-level label.

17. The computer-readable storage device according to claim 14, wherein the operations further include generating a multi-dimensional visualization associated with the classification label of the input image.