Multimodal ai models for identification and severity scoring of regions of interest

Multimodal and multi-instance AI models address the limitations of current medical imaging analysis by integrating data from multiple imaging modalities and timepoints, enhancing disease assessment and treatment planning through consistent ROI detection and severity scoring.

WO2025149900A1PCT designated stage expired Publication Date: 2025-07-17JANSSEN RESEARCH & DEVELOPMENT LLC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2025/050161
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-11
Filing Date
2025-01-07
Publication Date
2025-07-17

AI Technical Summary

Technical Problem

Current AI solutions for medical imaging analysis fail to accurately and efficiently incorporate ROI-based features across different imaging modalities and timepoints, lacking the ability to standardize scores and assess disease progression due to modality-specific detection techniques.

Method used

Implementing multimodal and multi-instance AI models that concatenate image data from various imaging modalities and timepoints to identify and assess regions of interest, using trained machine learning models for disease detection and severity scoring.

Benefits of technology

Enhances the accuracy and efficiency of disease assessment by integrating ROI detection with disease assessment and treatment planning, providing consistent scoring across modalities and timepoints, and enabling automated disease progression analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025050161_17072025_PF_FP_ABST
    Figure IB2025050161_17072025_PF_FP_ABST
Patent Text Reader

Abstract

Various embodiments of the present disclosure disclose multimodal and / or multi- instance models for the identification and assessments of medical conditions, and methods thereof. An example method comprises: receiving, from a plurality of imaging modalities, a 5 plurality of respective modality-specific image data commonly representing a patient anatomy comprising regions of interest (ROI); identifying portions of the modality-specific image data representing the ROI; generating, for each ROI, an ROI-specific image dataset comprising the respective portions; concatenating the respective portions; and applying, to a trained machine learning model, the concatenated ROI image data for each ROI to determine one or more 0 attributes of the medical condition for one or more of the plurality of ROI.
Need to check novelty before this filing date? Find Prior Art

Description

MULTIMODAL Al MODELS FOR IDENTIFICATION AND SEVERITY SCORING OF REGIONS OF INTERESTRELATED APPLICATIONS

[0001] The present application claims priority to U.S. provisional patent application serial number 63 / 620,102 filed January 11, 2023, the entire content of which is incorporated herein by reference and relied upon.FIELD

[0002] The present application relates to identifying and assessing diseased regions, and more specifically to artificial intelligence (Al) models for identifying and assessing diseased regions of interest.BACKGROUND

[0003] The diagnosis and treatment for many diseases and medical conditions benefit from the proper analysis of medical imaging for detection and monitoring. Such diseases and medical conditions include different types of arthritis, pneumonia, or COVID- 19. For example, psoriatic arthritis (PsA) and rheumatoid arthritis (RA) involve chronic inflammation of joints for which early detection and appropriate treatment are instrumental in preventing permanent joint damage and long-term disability. Structural damage in PsA and RA is usually evaluated using various imaging modalities, such as X-ray, MRI, and ultrasound. Scoring systems for evaluating structural damage for PsA and RA from radiographs include the Sharp scoring method and one of its modifications, the van DerHeijde-Sharpe (vdH-S) scoring method. Such scoring methods measure erosions in bone morphology, which reflect bone destruction, and joint space narrowing (JSN), which reflect cartilage loss. However, such scoring mechanisms also necessitate careful analysis of images, which are often obtained from various disparate medical imaging modalities, and careful consideration of patient-specific information. For example, for PsA and RA, image readers are required to score various joints of a patient anatomy from medical images of a patient. Similarly, tumor analysis (e.g., for non-small cell lung tumors) may typically involve manually assessing changes in severity from images of the tumor at different time points. Such analysis involves not only a correct detection of tumor regions, but an accurate and precise determination of the severity of the tumor or the progression or regression of the tumor over time.

[0004] For the diagnosis and treatment of various diseases, clinicians (e.g., radiologists) often tasked with manually identifying regions of interest (ROI) from medical images, often from different medical imaging modalities and / or timepoints, manually reviewing the ROIs for granular assessment of a given disease, and manually scoring the ROIs based on their disease severity. This can be time consuming, laborious, subjective, and expensive. Furthermore, medical images received from different imaging modalities, even if the medical images pertain to the same patient anatomy, present new challenges as there is a need to standardize scores and analysis for an ROI, irrespective of the imaging modality. Additionally, the extraction of differences in severity across longitudinal timepoints is crucial to assess disease progression, change in disease severity, and response to treatment. Yet, few systems and methods exist for effectively automating longitudinal image analysis. Although there have been some technological development to utilizing artificial intelligence (Al) for medical imaging analysis, current Al solutions fail to incorporate: (1) ROI-based features across different imaging modalities into machine learning models accurately, reliable, and efficiently, and (2) the implicit dependency between images received: (a) from different imaging modalities, and (b) across different timepoints.

[0005] Various embodiments of the present disclosure address one or more of the abovedescribed issues.SUMMARY

[0006] In various embodiments, multimodal and / or multi-instance Al models are disclosed for the identification and disease assessment of regions of interest (ROI).

[0007] In some embodiments, a method of evaluating a medical condition using multimodal image data is disclosed. A computing device having a processor receives, from a plurality of imaging modalities, a plurality of respective modality-specific image data commonly representing a patient anatomy comprising a plurality of regions of interest (ROI). For each of the plurality of ROI, a portion of each of the modality-specific image data as representing the ROI is identified. For each ROI of the plurality of ROI, an ROI-specific image dataset is generated comprising the respective portions representing the ROI from the plurality of modality-specific image data. For each ROI, the respective portions in the respective ROI- specific image dataset are concatenated to generate a concatenated ROI image data for the ROI. Furthermore, the concatenated ROI image data for each ROI is applied to a trained machinelearning model to determine one or more attributes of the medical condition for one or more of the plurality of ROI.

[0008] In some embodiments, a system for evaluating a medical condition using multimodal image data is disclosed. The system comprises: a processor; and memory storing instructions that, when executed by the processor, cause the processor to perform the aforementioned method.

[0009] In some embodiments, a method of evaluating a medical condition using longitudinal image data is disclosed. A computing device having a processor receives, from at least one imaging modality, a first image data of a patient anatomy at a first time point and a second image data of the patient anatomy at a second time point. A region of interest (ROI) that is common to both the first image data and the second image data is identified by applying the first image data and the second image data to a first trained machine learning model. Portions of the first image data and the second image data corresponding to the ROI may comprise a first ROI-specific image data and a second ROI-specific image data, respectively. The first ROI-specific image data and the second ROI-specific image data may be applied to a trained multi -instance machine learning model to generate one or more attributes of the medical condition.

[0010] In some embodiments, a system for evaluating a medical condition using longitudinal image data is disclosed. The system comprises: a processor; and memory storing instructions that, when executed by the processor, cause the processor to perform the aforementioned method.

[0011] In some embodiments, a method of evaluating a medical condition using multimodal longitudinal image data is disclosed. A computing device having a processor may receive, from each of a plurality of imaging modalities, a plurality of modality-specific image data commonly representing a region of interest (ROI) of a patient anatomy at different time points, including a first time point and a second time point. The computing device may generate, from the plurality of modality-specific image data received from each of the plurality of imaging modalities, a first set of modality-specific image data received from different imaging modalities and representing the ROI at the first time point, and a second set of modality-specific image data received from the different imaging modalities and representing the ROI at the second time point. The method may further include concatenating: the first set of modality-specific image data to form a concatenated first image data, and the second set ofmodality-specific image data to form a concatenated second image data. Furthermore, the concatenated first image data and the concatenated second image data may be applied to a trained multi -instance machine learning model to generate one or more attributes of the medical condition.

[0012] In some embodiments, a system for evaluating a medical condition using multimodal longitudinal image data is disclosed. The system comprises: a processor; and memory storing instructions that, when executed by the processor, cause the processor to perform the aforementioned method.

[0013] In some embodiments, non-transitory computer-readable media for use on a computer device is disclosed. The non-transitory computer-readable media may contain computer-executable programming instructions that may cause processors to perform one or more methods described herein.

[0014] Additional features and advantages of the disclosed method and apparatus are described in, and will be apparent from, the following Detailed Description and the Figures. The features and advantages described herein are not all-inclusive and, in particular, many additional features and advantages will be apparent to one of ordinary skill in the art in view of the figures and description. Moreover, it should be noted that the language used in the specification has been principally selected for readability and instructional purposes, and not to limit the scope of the inventive subject matter.BRIEF DESCRIPTION OF THE DRAWINGS

[0015] For a more complete understanding of the principles disclosed herein, and the advantages thereof, reference is now made to the following descriptions taken in conjunction with the accompanying drawings, in which:

[0016] FIG. 1 is a block diagram of a computer system and network environment in accordance with various embodiments.

[0017] FIG. 2 is a flowchart illustrating an example method of assessing a medical condition using multimodal image data, according to various embodiments.

[0018] FIG. 3 is a diagram illustrating example processes for multimodal data input, according to various embodiments.

[0019] FIG. 4 is a diagram illustrating an example process for assessing regions of interest (ROI) for a severity of arthritis, according to various embodiments.

[0020] FIG. 5 is a diagram illustrating an example process for predicting a drug response for treating tumor in regions of interest (ROI) of a lung, according to various embodiments.

[0021] FIG. 6 is a flowchart illustrating an example method of assessing a change in a medical condition using multi-instance learning, according to various embodiments.

[0022] FIG. 7 is a diagram illustrating an example process for assessing regions of interest (ROI) for a change in severity of arthritis using multi-instance learning, according to various embodiments.

[0023] FIG. 8 is a diagram illustrating an example process for assessing regions of interest (ROI) for a change in severity of non-small cell lung cancer using multi-instance learning, according to various embodiments.

[0024] FIG. 9 is a flowchart illustrating an example method of assessing a medical condition using multimodal image data and multi-instance learning, according to various embodiments.

[0025] FIG. 10 is a diagram illustrating an example experiment where Al-based models were trained and deployed to identify ROI and assess attributes for arthritis, according to various embodiments.

[0026] FIG. 11 is a graph illustrating the performance of multimodal Al models for the identification and severity scoring of regions of interest, according to various embodiments.

[0027] It is to be understood that the figures are not necessarily drawn to scale, nor are the objects in the figures necessarily drawn to scale in relationship to one another. The figures are depictions that are intended to bring clarity and understanding to various embodiments of apparatuses, systems, and methods disclosed herein. Wherever possible, the same reference numbers will be used throughout the drawings to refer to the same or like parts. Moreover, it should be appreciated that the drawings are not intended to limit the scope of the present teachings in any way.DETAILED DESCRIPTIONOverview

[0028] The diagnosis and treatment of many diseases require the ability to efficiently, reliably, and accurately interpret medical imaging data. Such diseases include different typesof arthritis, pneumonia, or COVID-19. For example, as previously discussed, the detection and monitoring of psoriatic arthritis (PsA) and rheumatoid arthritis (RA) involve readers interpreting radiographs of a patient anatomy involving joints and scoring various regions of interest using, for example, the Sharp scoring method or the van Der Heijde-Sharpe (vdH-S) scoring method. Such scoring systems measure erosions in bone morphology and joint space narrowing (JSN). Similarly, tumor analysis (e.g., for non-small cell lung tumors) may typically involve manually assessing severity or changes in severity from images of the tumors. Such analysis involves not only detection of tumor regions, but an accurate and precise determination of the severity of the tumor, or the progression or regression of the tumor over time. Conventionally, such assessments often occurs without consideration of patient-specific data related to demographics, treatment assignments, participant clinical characteristics. Such scoring also typically occurs without consideration to the background of the medical image being examined, such as the sequence of the image or the medical imaging modality from which it was acquired.

[0029] Therefore, current methodologies and systems by which clinicians (e.g., radiologists) examine medical images, identify regions of interest (ROI), and assess for disease severity, progression, or regression can be time consuming, laborious, subjective, and expensive. Furthermore, scoring and assessing ROIs based on medical images received from different imaging modalities, even if the medical images pertain to the same patient anatomy, present new challenges because it may be difficult to maintain a consistency of scores for the same ROI across different imaging modalities. Additionally, the extraction of differences in severity across longitudinal timepoints is crucial to assess disease progression, change in disease severity, and / or response to treatment. There is thus a desire and need to automate and enhance processes for using medical images to diagnose, monitor, and treat diseases in a manner that considers the background of the patient and the source of the medical images.

[0030] Although there have been some technological development to utilizing artificial intelligence (Al) for medical imaging analysis, current Al solutions do not incorporate: (1) ROI-based features across different imaging modalities for accurate, reliable, and efficient disease assessments, and (2) the implicit dependency between images that are received: (a) from different imaging modalities, and (b) across different timepoints. For example, ROI detection techniques are generally modality-specific, which has prevented the assessment of disease based on images acquired across modalities. Furthermore, conventional Al-based models for disease assessment are limited in the types of data on which the model is trained,or which can be considered in disease assessment. Thus, there is also a desire and need for more holistic Al-based models for disease assessment that spans across imaging modalities and imaging timepoints, while considering non-imaging, patient-specific data for more accurate, reliable, and efficient disease assessments.

[0031] The present disclosure describes systems and methods for identifying and assessing disease that and address one or more of the above described issues. For example, various embodiments of the present disclosure describe Al-based systems and methods for fast, robust, and reliable identification of ROIs from various imaging modalities. In some aspects, such embodiments incorporate 2-dimensional or 3 -dimensional concatenation to leverage multimodal and longitudinal image inputs. As another example, various embodiments of the present disclosure describe Al-based systems and methods for severity scoring for a wide range of applications. In some aspects, such systems and methods may receive imaging data from the different modalities directly as input, thus integrating ROI detection with disease assessment and treatment planning, as end to end systems and methods. Furthermore, such systems and methods allow for single timepoint based disease assessments (e.g., via a convolutional neural network model) or and / or longitudinal disease assessments (e.g., via a multi-instance learning (MIL) models).

[0032] With respect to disease assessments, various embodiments of the present disclosure incorporate single timepoint-based scoring (e.g., of severity and other attributes of a disease) as well as dependency between images or timepoints for assessments based on the change in a disease severity or other attribute (e.g., for assessing a progression or regression of a disease). Various Al-based models described herein may also include intermediate labels (e.g., ROI- based labels), which may be based on intermediate models (e.g., for ROI detection) into the subsequent models for the disease assessments (e.g., severity scoring). Various embodiments of the present disclosure provide much needed improvement to disease treatment since conventional techniques for the analysis of medica imaging data are laborious, time intensive, and often inconsistent. For example, based on disease assessments based on various Al-based models described herein, some embodiments of the present disclosure describe making recommendations for the enrollment or non-enrollment of a patient (e.g., for a medical treatment plan) or a specified treatment response.

[0033] These and other aspects and implementations are discussed in detail below. The foregoing information and the following detailed description include illustrative examples of various aspects and implementations, and provide an overview or framework for understandingthe nature and character of the claimed aspects and implementations. The drawings provide illustration and a further understanding of the various aspects and implementations, and are incorporated in and constitute a part of this specification.Example Descriptions of Terms

[0034] As used herein the specification, “a” or “an” may mean one or more. As used herein in the claim(s), when used in conjunction with the word “comprising,” the words “a” or “an” may mean one or more than one. Some embodiments of the disclosure may consist of or consist essentially of one or more elements, method steps, and / or methods of the disclosure. It is contemplated that any method or composition described herein can be implemented with respect to any other method or composition described herein and that different embodiments may be combined.

[0035] As used herein, “substantially” means sufficient to work for the intended purpose. The term “substantially” thus allows for minor, insignificant variations from an absolute or perfect state, dimension, measurement, result, or the like such as would be expected by a person of ordinary skill in the field but that do not appreciably affect overall performance. When used with respect to numerical values or parameters or characteristics that can be expressed as numerical values, “substantially” means within ten percent.

[0036] The use of the term “or” in the claims is used to mean “and / or” unless explicitly indicated to refer to alternatives only or the alternatives are mutually exclusive, although the disclosure supports a definition that refers to only alternatives and “and / or.” For example, “x, y, and / or z” can refer to “x” alone, “y” alone, “z” alone, “x, y, and z,” “(x and y) or z,” “x or (y and z),” or “x or y or z.” It is specifically contemplated that x, y, or z may be specifically excluded from an embodiment. As used herein “another” may mean at least a second or more.

[0037] The term “ones” means more than one.

[0038] As used herein, the term “plurality” can be 2, 3, 4, 5, 6, 7, 8, 9, 10, or more.

[0039] As used herein, the term “set of’ means one or more. For example, a set of items includes one or more items.

[0040] As used herein, the phrase “at least one of,” when used with a list of items, means different combinations of one or more of the listed items may be used and only one of the items in the list may be needed. The item may be a particular object, thing, step, operation, process, or category. In other words, “at least one of’ means any combination of items or number ofitems may be used from the list, but not all of the items in the list may be required. For example, without limitation, “at least one of item A, item B, or item C” means item A; item A and item B; item B; item A, item B, and item C; item B and item C; or item A and C. In some cases, “at least one of item A, item B, or item C” means, but is not limited to, two of item A, one of item B, and ten of item C; four of item B and seven of item C; or some other suitable combination.

[0041] As used herein, the term “about” refers to include the usual error range for the respective value readily known. Reference to “about” a value or parameter herein includes (and describes) embodiments that are directed to that value or parameter per se. For example, description referring to “about X” includes description of “X”. In some embodiments, “about” may refer to ±15%, ±10%, ±5%, or ±1% as understood by a person of skill in the art.

[0042] While the present teachings are described in conjunction with various embodiments, it is not intended that the present teachings be limited to such various embodiments. On the contrary, the present teachings encompass various alternatives, modifications, and equivalents, as will be appreciated by those of skill in the art.

[0043] In describing the various embodiments, the specification may have presented a method and / or process as a particular sequence of steps . However, to the extent that the method or process does not rely on the particular order of steps set forth herein, the method or process should not be limited to the particular sequence of steps described, and one skilled in the art can readily appreciate that the sequences may be varied and still remain within the spirit and scope of the various embodiments.Network Environment for the Multimodal Identification and Assessment of Disease

[0044] FIG. 1 is a block diagram of a system and network environment 100 (“environment” 100) in accordance with various embodiments. As shown in FIG. 1, the environment 100 may include a patient 102, one or more medical imaging modalities 104 (also referred to herein as “imaging modalities”), one or more image data 106, one or more computing systems 108, and a communication network 130. In some embodiments, the environment 100 may further include external systems in communication with the computing systems 108, with examples of such external systems including but not limited to one or more electronic health record systems 132 and pharmacy systems 136.

[0045] The patient 102 is an illustrative example of a patient, subject, or individual possessing an anatomy (“patient anatomy”) for which disease identification and / or assessment is desired or sought. In some aspects, patient 102 shown in FIG. 1 may represent a plurality ofpatients from which reference images were captured to generate reference image data to label and use as training data.

[0046] The environment 100 may further include one or more medical imaging modalities 104 for acquiring images of a patient anatomy of the patient 102. The patient anatomy may be of a diseased region of the patient or a part of body of the patient being examined for the detection or identification of a disease (if any). As used herein, a medical imaging modality may refer to a medical imaging device or technique distinguishable from another medical imaging modality based on the mechanism by which signals that reflect either anatomical structures or physiological events in a patient are acquired or imaged. Examples of medical imaging modalities may include but are not limited to a computed tomography (CT) modality, a magnetic resonance imaging (MRI) modality, a functional MRI (fMRI) modality, scintigraphy modality, a positron emissions tomography (PET) modality, a single-photon emission computed tomography (SPECT) modality, an ultrasound modality, a functional nearinfrared spectroscopy modality, a near-infrared spectroscopy modality, a magnetic particle imaging (MPI) modality, an elastography modality, an electrical impedance tomography modality, a photoacoustic imaging modality, or any combination thereof

[0047] Each medical imaging modality 104 may generate image data 106 that may be received by one or more computing systems 108. As used herein, “image data” may refer to the resulting data when captured images and / or visual representations of slices of a patient anatomy acquired by the medical imaging modality is converted into a format suitable for use in relevant software. For example, an image data 106 may include a 2D or 3D bitmap of greyscale intensities containing a voxel or pixel grid. In some aspects, an image data may correspond to an entirety or a portion of an image or slice of visual representation of a patient anatomy captured by the time. Thus, the image data may correspond to an image or a subset of an image. One or more image data 106 generated by the same medical imaging modality 104 may be referred to herein as “modality-specific image data” for example, to group the one or more image data as originating from the generation by the same medical imaging modality 104 for ease of explanation. As will be described herein, multiple image data received respectively from multiple medical imaging modalities 104 may nevertheless be based on images corresponding to the same patient anatomy. If the said same patient anatomy also has a region of interest (ROI), it is contemplated that each of the multiple image data may have portions based on respective portions of the respective images corresponding to the same ROI in the same patient anatomy. The respective portions of each of the multiple image data, eachcorresponding to the same ROI, may be referred to herein as “ROI-specific image data” for ease of explanation. The image data generated from the one or more medical imaging modalities 104 may be transmitted or otherwise sent to the one or more computing systems 108 via a wired or wireless network (e.g., such as communication network 130).

[0048] The one or more computing systems 108 may be used to train and apply machine learning models to identify and assess diseases from multimodal and / or multi-instance medical imaging data. The one or more computing systems 108 may include and / or comprise a general computing device or a special purpose computing device (e.g., with hardware configured to facilitate numerous iterative processes comprising large data sets). For simplicity, the one or more computing systems 108 as used herein may refer to any one of or a subset of the one or more computing systems 108. In some aspects, while one computing systems or one set of computing systems of the one or more computing systems 108 may be configured to train the machine learning models, another or another set of computing systems of the one or more computing systems 108 may be configured to apply the machine learning model to image data for a target patient (e.g., for which disease is assessed). In another aspects, the training and application may be performed by the same computing system or same set of computing systems. In some embodiments, the one or more computing systems may include a computing system that specializes or focuses in ROI detection across one or more images (e.g., from multiple imaging modalities or across different timepoints) and another computing system that specializes or focuses in disease assessment and / or treatment planning.

[0049] In some embodiments, an example computing system of the one or more computing systems 108 may comprise one or more of the components shown in FIG. 1, such as one or more processors 110, memory 112, a concatenation module 114, a network interface 116, a treatment recommendation engine 118, an ROI identification module 118, and an attribute assessment module 124. The one or more processors 110 may comprise any one or more types of digital circuit configured to perform operations on a data stream, including functions described in the present disclosure. In some aspects, the one or more processors 110 may include special purpose processors such as an image processor, etc. Also or alternatively, the one or more processors 110 may include a high performance processor having functionalities (e.g., processor speed, core count, etc.) configured to read and execute on large data sets (e.g., millions of voxels or pixels from a large number of reference image data). Also or alternatively, the one or more processors 110 may comprise a general purpose processor. The memory 112 may comprise any type of long term, short term, volatile, nonvolatile, or other memory and isnot to be limited to any particular type of memory or number of memories, or type of media upon which memory is stored. The memory 112 may store instructions that, when executed by the processor 110, can cause the one or more computing systems 108 to perform one or more methods discussed herein.

[0050] The concatenation module 114 may comprise a software, program, module, and / or plug-in that, when executed by a processor, such as but not limited to the one or more processors 110, causes the processor to concatenate (e.g., stitch together, augment, average, etc.) two or more image data. Such two or more image data may correspond to image data from different imaging modalities and / or from different timepoints. Non-limiting embodiments of such concatenation techniques of the present disclosure are described below in relation to FIG. 2.

[0051] The network interface 116 (e.g., a wired interface (e.g., electrical, RF (via coax), optical interface (via fiber)), a wireless interface, a modem, etc.) may allow the one or more computing systems 108 to communicate with other systems (e.g., electronic health records systems 132) over the communication network 130.

[0052] The treatment recommendation engine 118 may comprise a software, program, module, and / or plug-in that, when executed by a processor, such as but not limited to the one or more processors 110, causes the processor to generate treatment recommendations based on the disease assessments, as described herein. For example, the treatment may generate a treatment recommendation based on one or more attributes of a detected disease and / or ROI satisfying a predetermined threshold, or a change in the one or more attributes satisfying a predetermined threshold. Also or alternatively, the treatment recommendation engine 118 may allow the one or more computing systems 108 to establish communications with an external system (e.g., a pharmacy computing system, a specialized clinic or treatment center, etc.) to implement a treatment recommendation.

[0053] The ROI identification module 119 may comprise a software, program, module, and / or plug-in that, when executed by a processor, such as but not limited to the one or more processors 110, causes the processor to identify a region of interest (ROI) of an image or slice of a visual representation of a patient anatomy captured by the one or more medical imaging modalities 104. The ROI identification module 119 may use the image data 106 generated from a respective medical imaging modality 104 for the identification. For example, the ROI identification module 119 may input a received image data 106 of a patient anatomy into one or more ROI machine learning (ML) models 120 to generate an identification of one or moreregions of interest of the patient anatomy. The ROI ML models 120 may be trained, using training data, to receive an image data of a patient anatomy, as input, and generate an identification of a ROI, as output. The training data for ROI identification (referred to herein as ROI training data 122) may include a plurality of reference image data corresponding a plurality of reference images of a patient anatomy, with an identification or labeling of known ROI in all or at least a portion of the reference images. In some aspects, the identification may be a bounding box around an identified ROI. In some aspects, the identification may comprise an identification of elements (e.g., a 2D bitmap, 3D bitmap, pixmap, voxmap, pixels, voxels, etc.) of an image data corresponding to the identified ROI. In some embodiments, each ROI ML model may be trained to identify a separate type of ROI. For example, the type of ROI may include but is not limited to, a type of disease, a type of sub-anatomical region (e.g., joint, vascular region, airway, organ, etc.), a tissue, a layer, a cell, or a type of diseased region. Also or alternatively, each ROI ML model may be trained to identify ROIs from image data received from a separate imaging modality or group of imaging modalities 104. For example, one ROI ML model may be trained to identify ROI from MRI image data, whereas another ROI ML model may be trained to identify ROI from CT image data. In some embodiments, as will be described herein, one or more ROI ML models 120 may be trained to identify ROIs from image data generated from concatenating a plurality of image data. The plurality of image data being concatenated may be received, for example, from different imaging modality 104 and / or across different timepoints.

[0054] The attribute assessment module 124 may comprise a software, program, module, and / or plug-in that, when executed by a processor, such as but not limited to the one or more processors 110, causes the processor to assess one or more attributes of a disease from an identified ROI. The one or more attributes may include but are not limited to the one or more attribute comprises at least one of: an identification of a disease or medical condition (e.g., an identify of a disease or category of the disease), a presence or absence of the disease or medical condition; a measurement of a severity of the medical condition; a measurement of a change in a severity of the medical condition, etc. The attribute assessment module 124 may input an image data of an ROI (ROI-specific image data) into one or more attribute machine learning (ML) models 126 to generate an assessment of the one or more attributes of a disease for the ROI in the image data. The assessment may comprise, stem from, or may be based on an output feature vector output by the attribute ML model 126. For example, the output feature vector may indicate one or more values for the one or more attributes, such as a score for a severity,or a binary value (e.g., truth or false) to indicate a presence or absence of a disease or medical condition. The attribute ML models 126 may be trained, using training data, to receive an ROI- specific image data of an ROI, as input, and generate an output feature vector of values for an attribute, as output. The training data for the attribute assessment (referred to herein as attribute training data 128) may include a plurality of reference ROI-specific image data corresponding a plurality of reference images of an ROI, with an identification or labeling of known or validated assessments of one or more attributes for a disease or medical condition in the ROI. Furthermore, the ROI-specific image data may comprise a concatenated ROI-specific image data, where the concatenated ROI-specific image data results from a concatenation of image data received from multiple imaging modalities 104 and / or from different timepoints.

[0055] As previously discussed, the one or more computing systems may receive and / or transmit information (e.g., image data, treatment recommendations, patient-specific data, etc.,), or otherwise communicate with one or more components of the environment 100 over a communication network 130. The communication network 130 may comprise any wired or wireless network for communicating data or signals between systems located proximately, locally, or remotely.

[0056] In some embodiments, examples of external systems with which the one or more computing systems 108 may transit or receive information include, but are not limited to one or more electronic health record systems 730 and pharmacy information systems 136. An electronic health record system 132 may that facilitate the transfer of patient-specific data 134 and the storage of such data in a patient-specific electronic health record (EHR). The patientspecific data 134 may be used to train one or more ML models (e.g., attribute ML models 126), feed into ML models (e.g., as part of an input feature vector) and / or otherwise link image data to a patient. In some aspects, the electronic health records systems 730 may further encrypt various aspects of the patient-specific data, for example, to comply with various regulations (e.g., HIPAA).

[0057] A pharmacy system 136 may comprise any one or more computing devices, systems, or servers associated with a pharmacy. The one or more computing systems 108 may communicate with the pharmacy system 136, for example, to provide treatment recommendations (e.g., a prescription order), based on a disease assessment.Example Embodiments For Multimodal and / or Multi-Instance Models for Identifying And Assessing A Medical Condition

[0058] FIG. 2 is a flowchart illustrating an example method 200 of assessing a medical condition using multimodal image data, according to various embodiments. As illustrated, method 200 includes a number of enumerated steps, but aspects of method 200 may include additional steps before, after, and in between the enumerated steps. In some embodiments, one or more enumerated blocks, or the steps or processes involved in the enumerated blocks, may be omitted or performed in a different order. Method 200 may be performed by one or more computing systems (e.g., such as but not limited to the one or more computing systems 108). For example, method 200 may be performed by one or more processors (such as, but not limited to, one or more processors 110) based on computer-executable or machine readable instructions stored in a memory (such as, but not limited to, memory 112) of the one or more computing systems. In some aspects, the training of machine learning models involved in method 200 (e.g., the machine learning model at block 210, any machine learning models relied on for ROI identification at block 204, etc.) may be performed by a computing system separate or distinct from the computing system applying the trained machine learning model (e.g., at blocks 204 or 210), for example, to conserve computer resources and / or bandwidth, especially where training involves large datasets that cannot practically be performed by the same computing system as the application.

[0059] At block 202, one or more computing systems may receive, from a plurality of imaging modalities (such as, but not limited to, imaging modalities 104), a plurality of respective modality-specific image data (such as, but not limited to, modality-specific image data 106) commonly representing a patient anatomy comprising a plurality of regions of interest (ROI). Examples of imaging modalities may include but are not limited to: an X-Ray modality, a computed tomography (CT) modality, a magnetic resonance imaging (MRI) modality, a functional MRI (fMRI) modality, a scintigraphy modality, a positron emissions tomography (PET) modality, a single-photon emission computed tomography (SPECT) modality, an ultrasound modality, a functional near-infrared spectroscopy modality, a nearinfrared spectroscopy modality, a magnetic particle imaging (MPI) modality, an elastography modality, an electrical impedance tomography modality, or a photoacoustic imaging modality. Thus, a plurality of imaging modalities may include at least two or more imaging modalities from the aforementioned list.

[0060] At block 204, one or more computing systems may identify, for each of the plurality of ROI, a portion of each of the modality-specific image data representing the ROI. In some embodiments, for each of the plurality of ROI, the portion of each of the modality-specificimage data representing the ROI is identified by applying each modality-specific image data to a second trained machine learning model (such as, but not limited to one or more ROI models 120). For example, the ROI identification module 119 may input each of the modality-specific image data into one or more ROI ML models 120 to output a portion of the modality-specific image data representing the ROI. The process of identifying said portion representing the ROI may be referred to herein as identifying the ROI for ease of explanation. In some embodiments, each ROI ML model may be trained to identify a separate type of ROI. For example, the type of ROI may include but is not limited to, a type of disease, a type of sub-anatomical region (e.g., joint, vascular region, airway, organ, etc.), a tissue, a layer, a cell, or a type of diseased region. Also or alternatively, each ROI ML model may be trained to identify ROIs from image data received from a separate imaging modality or group of imaging modalities 104. For example, one ROI ML model may be trained to identify ROI from MRI image data, whereas another ROI ML model may be trained to identify ROI from CT image data. In some aspects, the second trained machine learning model may be a trained convolutional neural network. The ROI may be identified via a bounding box.

[0061] At block 206, one or more computing systems may generate, for each ROI of the plurality of ROI, an ROI-specific image dataset comprising the respective portions of each modality-specific image data representing the ROI. As previously discussed, the ROI-specific image dataset for a given ROI may comprise portions of images corresponding to the same ROI in the same patient anatomy, even if the images may each be generated by different imaging modalities.

[0062] At block 208, the one or more computing systems may concatenate, for each ROI, the respective portions in the respective ROI-specific image dataset to generate a concatenated ROI image data for the ROI. For example, for each ROI, the concatenation module 114 may concatenate image feature parameters for each of the respective portions in the respective ROI-specific image dataset to generate the concatenated ROI image data for the ROI. The concatenated ROI image data can be a 2-dimensional or a 3-dimensional image data. In some aspects, the concatenation into a 3-dimensional image data can result from the concatenation of 2-dimensional image data. For example, at least one modality-specific image data of the plurality of modality-specific image data can be a 2-dimensional image data. In some aspects, the concatenated ROI image data is a 2-dimensional image data, and the respective portions of the respective ROI-specific image dataset that are concatenated to form the concatenated ROI image data can also comprise 2-dimensional image data. Exampleembodiments of the concatenation process applied in block 208 is discussed in relation to FIG. 3, as will be described below.

[0063] At block 210, the one or more computing systems may apply, to a trained machine learning model, the concatenated ROI image data for each ROI to determine one or more attributes of the medical condition for one or more of the plurality of ROIs. As discussed, the one or more attributes may include but are not limited to, for example, an identification of a disease or medical condition (e.g., an identify of a disease or category of the disease), a presence or absence of the disease or medical condition; a measurement of a severity of the medical condition; a measurement of a change in a severity of the medical condition, etc. For example, the attribute assessment module 124 may be configured to apply, as input, the concatenated ROI image data for each ROI to one or more attribute ML models 126 to output the one or more attributes. In some aspects, each attribute ML model 126 may be trained to receive, as input, a concatenated ROI image data for a specific ROI, and to provide, as output, the one or more attributes of the medical condition for that specific ROI. Also or alternatively, each attribute ML model 126 may be trained to receive, as input, a concatenated ROI image data, and to provide, as output, a specific attribute of the medical condition.

[0064] FIG. 3 is a diagram illustrating example processes for multimodal data input, according to various embodiments. As shown in FIG. 3, various sets of image data 306 (“image datasets” 306) pertaining to regions of interest (ROI) may be respectively extracted from various imaging modalities 304 (such as, but not limited to, imaging modalities 104.). As the various imaging modalities 104 may be different (e.g., a CT versus an MRI), the image datasets extracted from one modality may be characteristically different from the image dataset extracted from another modality due to the nature of signals used to generate the images (e.g., magnetic resonance, X-rays, etc.), even if both image datasets are based on visual representations of the same set of regions of interest (ROI). The present disclosure describes techniques to be able to utilize features from each of the imaging modalities for disease assessment, such that various imaging modalities are able to contribute information in aggregate. For example, as shown in FIG. 3, image data extracted from each of the different imaging modalities 304 but pertaining to the same ROI may be concatenated via 3D concatenation 308 or 2D concatenation 310. Under 3D concatenation 308, two or more image data may undergo image fusion (e.g., via image registration, image transformation, and image combination), whereby common image features (e.g., landmarks of the ROI) may be used to co-register the image data (e.g., along a coordinate system), and simple or weightedconcatenation may occur along a z-axis. Also or alternatively, image feature parameters may be extracted from the respective ROIs to generate a three-dimensional tensor. Under 2D concatenation 310, two or more image data may be stitched together. For example, the two or more image data may be co-registered (e.g., along a coordinate system), transformed, and then combined, such that image feature parameters for each of the two or more image data are arranged along a 2 dimensional coordinate axis.

[0065] Also or alternatively, two or more image data need not be concatenated, and may instead be inputted into a machine learning model (e.g., as individual inputs or as a nonconcatenated input of a set of image data). For example, as shown in FIG. 3, multimodal image data (e.g., two or more image data from two or more imaging modalities, respectively) may be left as independent instances 312 that can be applied to a multi-instance learning model. For example, if such independent instances 312 correspond to image data from different time points, the multi-instance learning model may be used to assess disease progression, disease regression, and / or an predicted outcome or time of a predicted outcome.

[0066] FIG. 4 is a diagram illustrating an example process for assessing regions of interest (ROI) for a severity of arthritis, according to various embodiments. As shown in FIG. 4, the example process involves the capture of image data of a patient anatomy that is a patient hand. However, the patient hand has several regions of interest (ROI) corresponding to joints in the patient hand. In the example shown in FIG. 4, the imaging modality used is an X-ray, thus yielding an X-ray image data. However, other imaging modalities or combinations of imaging modalities are possible. As discussed, for example, in relation to FIG. 3, the present disclosure describes techniques for the concatenation of image data obtained via multiple imaging modalities. FIG. 4 further shows that a plurality of ROI-specific image data 404 can be extracted from the original image data 402, each ROI-specific image data representing a respective one of the several ROI in the patient hand. In some embodiments, the ROI can be identified and / or extracted from the original image data, using existing ROI identification models, such as YOLOv8. In other embodiments, the ROI identification may involve concatenation, and may thus leverage the techniques described in the present disclosure. The extracted ROI-specific image data, at least one for each of the plurality of ROI in the patient hand, may each be applied to a machine learning model (“attribute ML model” 406).

[0067] The attribute ML model 406 may be trained to determine one or more disease attributes for the ROI. The attribute ML model 406 is a non-limiting example of the previously discussed one or more attribute ML models 126 of FIG. 1. The attribute ML model 406 maybe trained to receive, as input, an ROI-specific image data of an ROI, and output one or more disease attributes for the ROI. In the example shown in FIG. 4, the attribute determined by the attribute ML model 406 is a van Der Heijde-Sharpe (vdH-S) score for the joint of the patient hand. The vdH-S score for each joint (i.e., the ROI in the example) may be based on one or more of an erosion in bone morphology and / or a joint space narrowing (JSN) of the joint. In some aspects, the attribute may comprise a score indicating the JSN for the joint, or an erosion for the joint.

[0068] The attribute ML model 406 may be trained using a training dataset comprising a plurality of reference image data (e.g., radiographs) representing ROI (e.g., joints). In this dataset, labels are placed (e.g., based on known vdH-S scores) for each radiograph for supervised learning. It is contemplated that in some embodiments, labels may be placed only for a portion of the reference image data. In such embodiments, if an image data pertaining to an ROI without a label may be associated with image data pertaining to the same ROI with a label, for example, based on the two image data pertaining to the same patient at the same timepoint but being generated from different imaging modalities. In some embodiments, for example, where an attribute is a prediction to a forecasted event (e.g., when a medical condition reaches a certain severity state or a binary event), the label may be a forecasted time to the event based on the image data of the ROI. In some embodiments, the training of the attribute ML model 406 may involve the aforementioned training dataset for the training the attribute ML model 406 and one or both of a validation dataset (e.g., a subset of the training dataset) for a validation of the trained attribute ML 406, and a testing dataset (e.g., another subset of the training dataset) for a testing of the trained attribute ML model 406. Additional details and embodiments for the training, validation, testing, and application of the attribute ML model 406 is further described in relation to FIG. 10.

[0069] As shown in FIG. 4, the trained attribute ML model 506 may be used to output attributes (e.g., vdH-S scores) per ROI 408 for each ROI of the extracted ROI-specific image data 404. In the example shown in FIG. 4, where the ROI represents joints in the patient anatomy of a hand, the attributes may comprise vdH-S scores indicating the JSN for the joint, or an erosion for the joint. In some aspects, the attribute may comprise a composite score based on the aforementioned vdH-S scores or other scores.

[0070] Based on the attributes (e.g., the vdH-S scores), a treatment may be assessed, recommended, or otherwise implemented (such processes collectively referred to herein as “treatment recommendation” 410) for the patient for which the diseases is being assessed. Forexample, if the vdH-S score is found to satisfy a predetermined threshold, a recommendation may be generated (e.g., by the treatment recommendation engine 118 of the one or more computing systems 108) to enroll the patient in a treatment program. Also or alternatively, a prescription may be generated (e.g., a medication or therapy) based on the attribute satisfying a predetermined threshold. In some aspects, the prescription may be transmitted (e.g., by the one or more computing systems 108) to an external system or server for the implementation of the recommended treatment (e.g., to a pharmacy system 136). In some embodiments, the treatment recommendation may rely on patient demographic data, prior medical treatment data, co-morbidity, and other patient-specific data, which may be received, from electronic health record systems 132.

[0071] FIG. 5 is a diagram illustrating an example process for predicting a drug response for treating tumor in regions of interest (ROI) of a lung, according to various embodiments. Although the patient anatomy in the example shown in FIG. 5 is a whole body, the regions of interest are lungs and tumor in the lungs, and the imaging modality used is computed tomography (CT) for a whole body CT scan, it is contemplated that similar methods and systems as described in relation to FIG. 5 and in the present disclosure can be applied for other patient anatomies, regions of interests, and / or imaging modalities, including a combination of varied imagining modalities. As shown in FIG. 5, the example process involves the capture of image data 502A-502B via a whole body CT scan. However, the region of interest (ROI) in this example is a patient lung 506A. In some embodiments (e.g., as shown in the bottom process), the region of interest (ROI) is at least one tumor 506B in the lung. In the example shown in FIG. 5, the imaging modality used is a CT scan, thus yielding an whole body CT scan image data. However, other imaging modalities or combinations of imaging modalities are possible. Also or alternatively, the imaging modalities may be used to capture image data for a portion of the whole body (e.g., the lung, the chest, the upper body, etc.), with ROI identification in subsequent steps for a portion of that portion (e.g., lung, a tumor in the lung, etc.). The ROI may be detected from the patient anatomy via ROI ML models 504A-504B (such as, but not limited to the previously described ROI ML models 120). The ROI ML models 504A-504B that may be trained to receive, as input, image data of a patient anatomy (in this example, a whole body CT scan), and provide, as output, a region of interest. In this example, the ROI ML model 504 A may be trained to output a bounding box around a lung and the ROI ML model 504B may be trained to output a bounding box around at least one tumor in a lung. The training may involve annotating the set of reference image data of lungs, with a boundingbox around the respective ROI in at least a portion of the set of reference image data (e.g., for supervised learning).

[0072] The extracted ROI-specific image data, in this example the ROI-specific image data 506A corresponding to the lung and ROI-specific image data 506B corresponding to the at least one tumor, may be applied to attribute ML models, in this example attribute ML models 508A and 508B, respectively. The attribute ML models 508A-508B, such as but not including the previously discussed attribute ML model 126 of FIG. 1, may be trained to determine one or more disease attributes for the ROI. In this example, the disease attribute is a predicted drug response for the ROI 510A-510B. Thus attribute ML models 506A-506B may be trained to receive, as input, an ROI-specific image data of an ROI, and output a drug response for the ROI 510A-510B. For example, the drug response may be a classification of whether the ROI would be a drug responder, a non-drug responder, or any intermediary classification between the aforementioned classifications. Also or alternatively, the drug response may be a probability of the ROI responding to the drug. In some embodiments, one or more parameters of the drug (e.g., potency, viability, concentration, amount, etc.) may be an input parameter applied to the respective attribute ML models 508A-508B. In such embodiments, the respective attribute ML models 508A-508B may be trained (e.g., using reference datasets) to receive, as input parameters, the ROI-specific image data and the one or more parameters of the drug, and output an indicator of the drug response 510A-510B.

[0073] FIG. 6 is a flowchart illustrating an example method of assessing a change in a medical condition using multi -instance learning, according to various embodiments. As illustrated, method 600 includes a number of enumerated steps, but aspects of method 600 may include additional steps before, after, and in between the enumerated steps. In some embodiments, one or more enumerated blocks, or the steps or processes involved in the enumerated blocks, may be omitted or performed in a different order. Method 600, which may be performed by one or more computing systems (e.g., such as but not limited to the one or more computing systems 108). For example, method 600 may be performed by one or more processors (such as, but not limited to, one or more processors 110) based on computerexecutable or machine readable instructions stored in a memory (such as, but not limited to, memory 112) of the one or more computing systems. In some aspects, the training of one or more machine learning models involved in method 600 (e.g., the first machine learning model at 604, the multi-instance machine learning model at block 606) may be performed by a computing system separate or distinct from the computing system applying the trained machinelearning model (e.g., at blocks 604 and 606), for example, to conserve computer resources and / or bandwidth, especially where training involves large datasets that cannot practically be performed by the same computing system as the application.

[0074] At block 602, the one or more computing systems may receive, from at least one imaging modality (such as, but not limited to, at least one imaging modality 104), multiple image data across different time points. For ease of explanation, the multiple image data across different time points may be understood to include, at least, a first image data of a patient anatomy at a first time point and a second image data of the patient anatomy at a second time point.

[0075] At block 604, the one or more computing systems may identify, by applying the first image data and the second image data to a first trained machine learning model, a region of interest (ROI) that is common to both the first image data and the second image data. The ROI model 120 is but one non-limiting example of the first trained machine learning model applied in block 604. Portions of the first image data and the second image data corresponding to the ROI may include a first ROI-specific image data and a second ROI-specific image data, respectively. For example, the ROI identification module 119 may input each of the first images data and the second image data into one or more ROI ML models 120 to output a portion of each of the first image data and the second image data as representing the ROI. In some embodiments, each ROI ML model may be trained to identify a separate type of ROI. For example, the type of ROI may include but is not limited to, a type of disease, a type of sub- anatomical region (e.g., joint, vascular region, airway, organ, etc.), a tissue, a layer, a cell, or a type of diseased region. For example, as will be described in relation to FIG. 7, a patient anatomy such as a hand may have a plurality of joints, with each type of joint comprising an ROI. In some aspects, the first trained machine learning model (e.g., ROI ML model 120) may be trained to identify a specific ROI (e.g., a specific joint) for each of the plurality of joints in the hand. In other aspects, the first trained machine learning model (e.g., ROI ML model 120) may be trained to identify a joint generally, which may be used to identify the plurality of ROI (e.g., joints) in the hand. In some embodiments, the first trained machine learning model may comprise a trained convolutional neural network.

[0076] At block 606, the one or more computing systems may apply, to a trained multiinstance machine learning model, the first ROI-specific image data and the second ROI- specific image data to generate one or more attributes of the medical condition. The one or more attributes may comprise, for example, an indicia of a progression of the medicalcondition; or an indicia of a regression of the medical condition. Also or alternatively, the one or more attributes may include a prediction of a severity of the medical condition at a third time point. The third time occurs after the first time point and the second time point. For example, the attribute assessment module 124 may apply, as input, the first ROI-specific image data and the second ROI-specific image data to the one or more attribute ML models 126 to output the one or more attributes of the medical condition. In some aspects, a time relation between the first and second timepoints may also be input into the one or more attribute ML models. In some aspects, each attribute ML model 126 may be trained to receive, as input, ROI-specific image data across different timepoints for a specific ROI, and to provide, as output, the one or more attributes of the medical condition for that specific ROI. Also or alternatively, each attribute ML model 126 may be trained to receive, as input, ROI-specific image data across timepoints, and to provide, as output, a specific attribute of the medical condition. In some embodiments, one or more the attribute ML models 126 may be trained to also receive, as input, patient-specific data such as, but not limited to, patient demographic information, patient medical history, patient physiological information, patient genetic information, and the like, and to provide, as output, the one or more attributes for the medical condition.

[0077] FIG. 7 is a diagram illustrating an example process for assessing regions of interest (ROI) for a change in severity of arthritis using multi-instance learning, according to various embodiments. Although the patient anatomy in the example shown in FIG. 7 is a patient hand, the regions of interest are joints in the hand, and the imaging modality used is an X-ray, it is contemplated that similar methods and systems as described in relation to FIG. 7 and in the present disclosure can be applied for other patient anatomies, regions of interests, and imaging modalities, including a combination of varied imagining modalities. For example, the present disclosure describes techniques for the concatenations of image data from multiple imaging modalities for enhanced accuracy and reliability of disease assessments.

[0078] As shown in FIG. 7, the example process involves the capture, at a plurality of timepoints, of X-ray image data of the patient hand 702 having the plurality of joints as regions of interest (ROI). For example, the original image data captured via X-ray may include a first image data of the patient hand at a first time point and a second image data of the patient hand at a second time point.

[0079] FIG. 7 further shows that a plurality of ROI-specific image data 704 can be extracted from the original image data of the patient hand 702 captured at the plurality of timepoints. Each ROI-specific image data may represent a respective one of the several ROI inthe patient hand at a specific timepoint of the plurality of time points. However, given the original image data was captured at a plurality of timepoints including, for example, a first timepoint and a second timepoint, the plurality of ROI-specific image data may be divided into subsets of ROI-specific image data for given timepoints, such as a first ROI-specific image data corresponding to a given ROI at a first timepoint and a second ROI-specific image data corresponding to the same ROI but at the second timepoint, for each of the plurality of ROI. In some embodiments, the ROI can be identified and / or extracted from the original image data, using ROI identification models described herein.

[0080] As shown in FIG. 7, time dependencies between the timepoints of the plurality of ROI-specific image data may be determined (step 706). For example, for any given ROI, the time duration between the capture of image corresponding to one set of ROI-specific image data (first ROI-specific image data) and another set of ROI-specific image data (second ROI- specific image data) may be determined. For example, the one or more computing systems 108 may rely on metadata from received image data to determine time dependencies between the various image data to determine and / or classify first ROI-specific image data and the second ROI-specific image data. The time duration may be useful for analysis of a medical condition over time, as timepoints close to one another may not necessarily reveal meaningful longitudinal data. As shown in FIG. 7, and for ease of explanation, the timepoint corresponding to capture of the original image data for the first ROI-specific image data is labeled “z” and the timepoint corresponding to the capture of the original image data for the second ROI-specific image data is labeled ‘y .”

[0081] The extracted first and second ROI-specific image data, as well as the time dependencies, may be collectively applied as inputs to an attribute ML model to determine, via the output from the attribute ML model, the one or more attributes for each of the ROIs. In some embodiments, the inputs may further comprise relevant patient demographic and other patient-specific data (e.g., prior medical conditions, patient physiology, etc.). In the example shown in FIG. 7, the attribute ML model is a multi-instance model 708, and the attribute being determined is a change in a severity for each of the joints. The multi-instance model 708 may be trained using a training dataset comprising a plurality of reference ROI-specific image data the ROI from different timepoints, along with time dependencies between such timepoints, with known attributes of a disease or medical condition concerning the ROI. Examples of attributes determined from the multi -instance model may include but are not limited to: a progression of a medical condition, a regression of a medical condition, a predicted outcomeof a medical condition, or a time to a specified outcome for the medical condition. In the example shown in FIG. 7, the change in the severity of the arthritis at a joint (e.g., a change in a vdH-S score) may thus be a regression or a progression.

[0082] The process shown in FIG. 7 may further include a treatment recommendation 710 based on the determined one or more attributes for the patient for which the diseases is being assessed. For example, if the change in severity is found to satisfy a predetermined threshold (e.g. it is a significant regression), a recommendation may be generated (e.g., by the treatment recommendation engine 118 of the one or more computing systems 108) to enroll the patient in a treatment program. Also or alternatively, a prescription may be generated (e.g., a medication or therapy) based on the attribute satisfying a predetermined threshold. In some aspects, the prescription may be transmitted (e.g., by the one or more computing systems 108) to an external system or server for the implementation of the recommended treatment (e.g., to a pharmacy system 136). In some embodiments, the treatment recommendation may rely on patient demographic data, prior medical treatment data, co-morbidity, and other patientspecific data, which may be received, from electronic health record systems 132.

[0083] FIG. 8 is a diagram illustrating an example process for assessing regions of interest (ROI) for a change in severity of non-small cell lung cancer using multi-instance learning, according to various embodiments. Although the patient anatomy in the example shown in FIG. 8 is a patient lung, the regions of interest are tumors in the lung, and the imaging modality used is computed tomography (CT), it is contemplated that similar methods and systems as described in relation to FIG. 8 and in the present disclosure can be applied for other patient anatomies, regions of interests, and imaging modalities, including a combination of varied imaging modalities. For example, the present disclosure describes techniques for the concatenations of image data from multiple imaging modalities for enhanced accuracy and reliability of disease assessments.

[0084] As shown in FIG. 8, the example process involves the capture of CT image data, in at least two timepoints, of the patient lung having at least one tumor as a region of interest (ROI). For example, the original image data captured via CT may include a first image data 802A of the patient lung at a first timepoint ( / ) and a second image data 802B of the patient lung at a second time point ( / ).

[0085] FIG. 8 further shows that ROI-specific image data 806A-806B (the ROI in this example being a non-small cell lung tumor) can be extracted from the original image data802A-802B of the patient lung captured at the respective timepoints, i and j, via tumor detection models 804A-804B, respectively. The tumor-detection models 804A-804B may comprise one or more ROI ML models 120, as previously discussed, that may be trained to receive, as input, image data of a patient anatomy (in this example, a lung), and provide, as output, a region of interest (in this example, a non-small cell lung tumor). The training may involve annotating the set of reference image data of lungs, with a bounding box around the non-small cell lung tumor in at least a portion of the set of reference image data (e.g., for supervised learning).

[0086] As shown in FIG. 8, there may be a plurality of ROI-specific image data 808 resulting from original image data captured for each timepoint, i and j . That is, the original image data 802A-802B may be enough to provide for the extraction of two or more ROI- specific image data 806A at timepoint i and / or two or more ROI-specific image data 806B at timepoint j. Thus, the plurality of ROI-specific image data 808 may be divided into subsets of ROI-specific image data for given timepoints, such as a first ROI-specific image data corresponding to a given ROI at a first timepoint, z, and a second ROI-specific image data corresponding to the same ROI but at the second timepoint, / , for each of the plurality of ROI.

[0087] The extracted first and second ROI-specific image data may be collectively applied as inputs to an attribute ML model to determine, via the output from the attribute ML model, the one or more attributes for each of the ROIs. In the example shown in FIG. 8, the one or more attributes being determined is a change (e.g., a delta) in severity of the at least one non- small cell lung tumor 812, between at least timepoints z and j. Furthermore, the attribute ML model shown in FIG. 8 is a multi -instance model 810. The multi -instance model 810 may be trained using a training dataset comprising a plurality of reference ROI-specific image data of tumors (e.g., lung tumors) from different timepoints with known attributes (e.g., severity) of the tumor concerning the ROI. An example dataset and training process is described further below in relation to FIG. 10. Although the example shown in FIG. 8 concerns the attribute of a change (e.g., progression and / or regression) of a severity of a medical condition (e.g., tumor) 812, other examples of attributes determined from the multi-instance model may include a predicted outcome of a medical condition, or a time to a specified outcome for the medical condition.

[0088] In some embodiments, time dependencies between the timepoints of the plurality of ROI-specific image data may be determined and also considered (e.g., as an input feature parameter) in the attribute ML model. For example, for any given ROI, the time duration between the capture of images corresponding to one set of ROI-specific image data (first ROI-specific image data) and another set of ROI-specific image data (second ROI-specific image data) may be determined. In at least one instance, the one or more computing systems 108 may rely on metadata from received image data to determine time dependencies between the various image data to determine and / or classify first ROI-specific image data and the second ROI- specific image data. The time duration may be useful for analysis of the tumor over time, as timepoints close to one another may not necessarily reveal meaningful longitudinal data. In some embodiments, the inputs may further comprise relevant patient demographic and other patient-specific data (e.g., prior medical conditions, patient physiology, etc.) that may be relevant to the tumor. Such additional inputs may be applied to the multi -instance model 810 (e.g., as part of the input feature vector) in order to determine the change in severity of the tumor 812. In embodiments, the multi-instance model 810 may be trained based on training data that further comprises these additional inputs.

[0089] Based on the change in severity of the non-small cell lung tumor, a treatment may be recommended and / or altered. For example, if the change in severity is found to satisfy a predetermined threshold (e.g. it is a significant regression), a recommendation may be generated (e.g., by the treatment recommendation engine 118 of the one or more computing systems 108) to enroll the patient in a treatment program. Also or alternatively, a prescription may be generated (e.g., a medication or therapy) based on the attribute satisfying a predetermined threshold. In some aspects, the prescription may be transmitted (e.g., by the one or more computing systems 108) to an external system or server for the implementation of the recommended treatment (e.g., to a pharmacy system 136). In some embodiments, the treatment recommendation may rely on patient demographic data, prior medical treatment data, comorbidity, and other patient-specific data, which may be received, from electronic health record systems 132.

[0090] FIG. 9 is a flowchart illustrating an example method 900 of assessing a medical condition using multimodal image data and multi-instance learning, according to various embodiments. As illustrated, method 900 includes a number of enumerated steps, but aspects of method 900 may include additional steps before, after, and in between the enumerated steps. In some embodiments, one or more enumerated blocks, or the steps or processes involved in the enumerated blocks, may be omitted or performed in a different order. Method 900, which may be performed by one or more computing systems (e.g., such as but not limited to the one or more computing systems 108). For example, method 900 may be performed by one or more processors (such as, but not limited to, one or more processors 110) based on computer-executable or machine readable instructions stored in a memory (such as, but not limited to, memory 112) of the one or more computing systems. In some aspects, the training of machine learning models involved in method 700 (e.g., multi-instance machine learning model at block 908) may be performed by a computing system separate or distinct from the computing system applying the trained machine learning model (e.g., at block 908), for example, to conserve computer resources and / or bandwidth, especially where training involves large datasets that cannot practically be performed by the same computing system as the application.

[0091] At block 902, the one or more computing systems may receive from each of a plurality of imaging modalities (such as, but not limited to, imaging modalities 104), a plurality of modality-specific image data (such as, but not limited to, modality-specific image data 106) commonly representing a region of interest (ROI) of a patient anatomy at different time points. For ease of explanation, the different time points includes at least a first time point and a second time point. In some embodiments, the received image data may not necessarily represent only the ROI of the patient anatomy. In such embodiments, the one or more computing systems may identify the ROI from the image data, for example, using processes described in blocks 204 and 504 of FIGS. 2 and 5, respectively. The identified ROI-specific portions from the plurality of modality-specific image data may thus be used in subsequent steps of method 900.

[0092] At block 904, the one or more computing systems may generate, from the plurality of modality-specific image data received from each of the plurality of imaging modalities, a first set of modality-specific image data received from different imaging modalities and representing the ROI at the first time point, and a second set of modality-specific image data received from the different imaging modalities and representing the ROI at the second time point. In some embodiments, the second time point may occur after at least a predetermined duration after the first time point. Such predetermined duration may be useful for analysis of a medical condition over time, as timepoints close to one another may not necessarily reveal meaningful longitudinal data. In some embodiments, the one or more computing systems may rely on metadata within the received plurality of modality-specific image data to determine time dependencies between the various modality-specific image data to determine and / or classify first image data and the second image data.

[0093] At block 906, the one or more computing systems may concatenate: the first set of modality-specific image data to form a concatenated first image data, and the second set of modality-specific image data to form a concatenated second image data. For example, the concatenation module 114 may concatenate image feature parameters for each of the first setof modality-specific image data to form the concatenated first image data, and concatenate image feature parameters for each of the second set of modality-specific image data to form a concatenated second image data. Each of the concatenated first image data or the concatenated second image data can be a 2-dimensional or a 3-dimensional image data. In some aspects, the concatenation into a 3 -dimensional image data can result from the concatenation of two or more 2-dimensional image data. For example, at least one modality-specific image data of the first set and / or at least one modality-specific image data of the second set may comprise 2- dimensional image data. In some aspects, each of the concatenated first image data and / or the concatenated second image data can be 2-dimensional image data. For example, the modalityspecific image data of each of the first set and the second set are 2-dimensional image data. Example embodiments of the concatenation process applied in block 906 are discussed herein in relation to FIG. 3.

[0094] At block 908, the one or more computing systems may apply, to a trained multiinstance machine learning model, the concatenated first image data and the concatenated second image data to generate one or more attributes of the medical condition. Non-limiting embodiments of the multi-instance machine learning model are described herein, for example, in relation to FIG. 7.Example Experiment for The Training And Application of AI-Based Models For ROI Identification and Disease Assessment

[0095] FIG. 10 is a diagram illustrating an example experiment where Al-based models were trained and used to identify ROI and assess attributes for arthritis, according to various embodiments. Although the example experiment illustrated in FIG. 10 concerns attributes for arthritis, it is contemplated that similar processes and techniques described in the example experiment may be used for assessing attributes for other medical conditions, such as COVID- 19, pneumonia, and lung cancer.

[0096] Specifically, the Al-driven processes relied on machine learning models 1004A- 1004B trained to receive, as input, radiographs (DICOM) of the hands and feet, respectively, and identified joints of interest (e.g., according to vdH-S). The Al-driven processes further relied on machine learning models 1006A-1006B and 1008A-1008B to evaluate the identified joints for specific attributes. Specifically, ML models 1006A-1006B were used to assess erosion (according to vdH-S) in the hands and feet, respectively, whereas ML models 1008A- 1008B were used to assess JSN in the hands and feet, respectively. However, it is contemplatedthat similar datasets, techniques, and / or processes for the training of machine learning models as will be described below may be used to train machine learning models to evaluate other regions of interest of other patient anatomies for the assessment of attributes for other medical conditions (e.g., COVID-10, pneumonia, lung cancer, etc.) In this example experiment, the assessments outputted by these models was specifically a patient vdH-S score 1010 for the erosion and the JSN for each of the identified joints. In some embodiments, assessments outputted by similarly development models may indicate, for example, a drug response for a region of interest (e.g., lung tumor), a severity or a change in severity for another patient anatomy or region of interest (e.g., non-small cell lung tumor, pneumonia, COVID-19), a forecasted or predicted event, a time to a forecasted or predicted event, and the like. Details regarding the training and application of the above described ML models will be further described below.

[0097] As shown in FIG. 10, the example experiment involved ML models for joint identification 1004A-1004B (such as, but not limited to, the previously discussed ROI ML model 120), and ML models for vdH-S scoring of the joint 1006A-1006B and 1008A-1008B (such as, but not limited to, the previously discussed attribute ML model 126.

[0098] The ML models for joint identification 1004A-1004B were trained to receive, as input, radiographic images of hands or feet. It is contemplated that such images may be received in a standardized format (e.g., Digital Imaging and Communications in Medicine (DICOM) format). It is contemplated that in embodiments involving the identification of other regions of interest from other patient anatomies (e.g., lung, lung tumor, etc.), machine learning models may be trained to receive, as input, radiographic images of patient anatomies that include the regions of interest. Furthermore, the ML models 1004A-1004B in the experiment shown in FIG. 10 used an object detection convolutional neural network (CNN) based on YOLOv8 to identify joints from the image inputs. The CNN was trained to identify all joints in the hands and feet as specified in the vdH-S scoring. It is contemplated that in other embodiments, other object detection CNNs may be used to identify other ROI from image inputs, by training the CNN (e.g., by annotating the image data via a bounding box surrounding the ROI). In the experiment, unique CNNs were used for joint identification from radiographs of the hands and feet, respectively (i.e., ML model 1004A for joint identification in the hands and ML model 1004B for joint identification in the feet). The outputs of the joint identification were a region of the inputted radiographic image corresponding to each joint specified in the vdH-S scoring system. It is contemplated that for ROI ML models to detect other ROI (e.g.,lung, lung tumor, etc.), the outputs for the ROI detection would be a region of the inputted radiographic image corresponding to the ROI (lung, lung tumor, etc.) for which an assessment of an attribute (e.g., a disease progression, a disease regression, a disease severity, change in the severity, a drug response, a forecasted event, a time to the forecasted event, etc.) is desired.

[0099] The training of the foregoing CNNs for the joint detection models 1004A-1004B as well as the training for attribute ML models 1006A-1006B and 1008A-1008B relied on training multiple datasets from several clinical trials as set forth below:

[0100] The inclusion criteria for each dataset were generally consistent and targeted an enrollment population with moderate to severe psoriatic arthritis. It is contemplated that the combined datasets may also be used for the training of the ROI ML models and attribute ML models for an assessment of one or more attributes of other medical conditions, concerning other patient anatomies, and / or other ROIs.

[0101] In order to label the training data (e.g., for supervised learning), bounding boxes that bounded the joint areas of interest were manually annotated to serve as ground truth for the training of the CNNs for the ROI ML models 1004A-1004B for joint identification. The performance of the joint detection CNNs was subsequently evaluated, and the predicted joint areas of the joint detection CNNs were found to align with the ground-truth joint areas. As noted above, after extracting the ROIs from the image data, the experiment proceed to generate vdH-S scores for erosion and JSN, using ML models for vdH-S scoring for JSN 1008A-1008B and erosion 1006A-1006B (as non-limiting examples of the one or more attribute ML models 126). Such ML models 1006A-1006B and 1008A-1008B, and the training thereof, are non-limiting examples of the previously described one or more attribute ML models 126. The attribute ML models 1006A-1006B and 1008A-1008B also used two separate CNNs for the vdH-S scoring of the hands and feet. Ground-truth vdH-S scores were also provided for the training data for the attribute ML models 1006A-1006B and 1008A-1008B based on the scores obtained from the combined datasets.

[0102] It is contemplated that in other embodiments (e.g., concerning other patient anatomies, ROIs, medical conditions, or attributes), the training data may also be labeled accordingly, bounding boxes that bound the ROIs appropriate for such embodiments (e.g., lung tumor, lungs, etc.) can also be manually annotated to serve as ground truth for the training of the CNNs for the ROI ML models 1004A-1004B. The performance of the ROI detection CNNs can be subsequently evaluated. After extracting the ROIs from the image data, one or more desired attributes (e.g., a drug response for an ROI, a change in severity of a disease, etc.) may be determined using one or more respective attribute ML models (such as, but not limited to attribute models 126). In some aspects, known attributes (e.g., known drug response probabilities, known changes in severity, etc.) may be provided as ground-truth for the training data for the attribute ML models for training.

[0103] In the example experiment illustrated in FIG. 10, and as would be applicable for other embodiments concerning other patient anatomies, ROIs, medical conditions, or attributes, the combined datasets (e.g., each comprising radiographs and patient-specific data) were each divided into three separate groups corresponding to different phases for the training of the ROI ML model: ML model training, ML model validation, and ML model testing. During ML model training, the ROI models (in this experiment, the j oint detection models) and the attribute ML models (in this experiment, the vdH-S scoring models) were trained using the combined datasets along with the aforementioned corresponding ground truth labels provided to such datasets. The weights and biases of the neural networks were trained and / or fitted using the combined datasets.

[0104] During ML model validation, a validation dataset was formed from a sample of the training dataset to provide an unbiased evaluation of the ML model after each training step of the ML model (e.g., after each “epoch”). During ML model testing, an ML model test dataset was formed from a reserved sample of the training dataset to provide unbiased evaluation of the ML model after performance was established on the validation dataset.

[0105] The training datasets, validation dataset, and testing dataset were each further split based on images of the patient anatomy (the hands and feet in the experiment shown in FIG. 10-), with an approximately equal number of images of the particular patient anatomy (e.g., an equal number of hands and feet images for the experiment illustrated in FIG. 10) being available for each participant. All images for the ML model training, ML model validation, and ML model testing were stratified by participant such that images from the same participant were contained within the same split.

[0106] FIG. 11 includes graphs 1100A and 1100B illustrating the performance of ML models for the identification and severity scoring of regions of interest, according to various embodiments. Specifically, graph 1100A shows how well other readers were able to correctly determine the vdH-S score for joint space narrowing and erosion in the joints from the dataset (based on a comparison of the readers’ scoring with ground truth scores obtained from clinical studies). Graph 1100B shows how well the attribute ML models 1006A-1006B and 1008A- 1008B described above performed in correctly determining the vdH-S score for joint space narrowing and erosion in the joints from the dataset (based on a comparison of the readers’ scoring with the ground truth scores obtained from clinical studies). As shown from these graphs 1100A and 1100B, and the foregoing description, the Al-driven systems, methods, and ML models for identifying and assessing medical conditions (in this experiment, arthritis), appears to align well, if not outperform, manual and / or conventional techniques of identifying and assessing medical conditions.

[0107] It should be appreciated that the methodologies described herein, flow charts, diagrams, and accompanying disclosure can be implemented using the one or more computer systems 108, or a computing system thereof, as a standalone device or on a distributed network of shared computer processing resources such as a cloud computing network.

[0108] The methodologies described herein may be implemented by various means depending upon the application. For example, these methodologies may be implemented in hardware, firmware, software, or any combination thereof. For a hardware implementation, the processing unit may be implemented within one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, electronic devices, other electronic units designed to perform the functions described herein, or a combination thereof.

[0109] The methods of the present teachings may be implemented as firmware and / or a software program and applications written in conventional programming languages such as C, C++, Python, etc. If implemented as firmware and / or software, the embodiments described herein can be implemented on a non-transitory computer-readable medium in which a program is stored for causing a computer to perform the methods described above. It should be understood that the various engines described herein can be provided on a computer system, such as computing system 108, whereby processor 110 would execute the analyses and determinations provided by these engines, subject to instructions provided by any one of, or a combination of, the memory and user input.

[0110] While the present teachings are described in conjunction with various embodiments, it is not intended that the present teachings be limited to such various embodiments. On the contrary, the present teachings encompass various alternatives, modifications, and equivalents, as will be appreciated by those of skill in the art.[oni] In describing the various embodiments, the specification may have presented a method and / or process as a particular sequence of steps. However, to the extent that the method or process does not rely on the particular order of steps set forth herein, the method or process should not be limited to the particular sequence of steps described, and one skilled in the art can readily appreciate that the sequences may be varied and still remain within the spirit and scope of the various embodiments.Additional Recited EmbodimentsEmbodiment 1 : A method of evaluating a medical condition using multimodal image data, the method comprising: receiving, by a computing device having a processor, from a plurality of imaging modalities, a plurality of respective modality-specific image data commonly representing a patient anatomy comprising a plurality of regions of interest (ROI); identifying, for each of the plurality of ROI, a portion of each of the modality-specific image data as representing the ROI; generating, for each ROI of the plurality of ROI, an ROI-specific image dataset comprising the respective portions representing the ROI from the plurality of modality-specific image data; concatenating, for each ROI, the respective portions in the respective ROI-specific image dataset to generate a concatenated ROI image data for the ROI; and applying, to a trained machine learning model, the concatenated ROI image data for eachROI to determine one or more attributes of the medical condition for one or more of the plurality of ROI.Embodiment 2: The method of embodiment 1, wherein the one or more attribute comprises at least one of: an identification of the medical condition; a presence of the medical condition; or a measurement of a severity of the medical condition.Embodiment 3 : The method of embodiment 1 or 2, wherein the concatenated ROI image data is a 3 -dimensional image data, wherein at least one modality-specific image data of the plurality of modality-specific image data is a 2-dimensional image data.Embodiment 4: The method of any one of embodiments 1-3, wherein the concatenatedROI image data is a 2-dimensional image data, wherein the respective portions of the respective ROI-specific image dataset that are concatenated to form the concatenated ROI image data are 2-dimensional image data.Embodiment 5: The method of any one of embodiments 1-4, wherein concatenating the respective portions of the respective ROI-specific image dataset comprises aggregating image feature parameters for each of a plurality of image features common to the respective portions of the respective ROI-specific image dataset.Embodiment 6: The method of any one of embodiments 1-5, wherein concatenating the respective portions of the respective ROI-specific image dataset comprises averaging image feature parameters for each of a plurality of image features common to the respective portions of the respective ROI-specific image dataset.Embodiment 7: The method of any one of embodiments 1-6, wherein the plurality of imaging modalities comprises at least two or more imaging modalities selected from a group consisting of: an X-Ray modality, a computed tomography (CT) modality, a magnetic resonance imaging (MRI) modality, a functional MRI (fMRI) modality, a scintigraphy modality, a positron emissions tomography (PET) modality, a single-photon emission computed tomography (SPECT) modality, an ultrasound modality, a functional near-infrared spectroscopy modality, a near-infrared spectroscopy modality, a magnetic particle imaging(MPI) modality, an elastography modality, an electrical impedance tomography modality, and a photoacoustic imaging modality.Embodiment 8: The method of any one of embodiments 1-7, wherein, for each of the plurality of ROI, the portion of each of the modality-specific image data representing the ROI is identified by applying each modality-specific image data to a second trained machine learning model.Embodiment 9: The method of embodiment 8, wherein the second trained machine learning model is a trained convolutional neural network.Embodiment 10: The method of any one of embodiments 1-9, wherein the medical condition is an arthritis, wherein the plurality of ROI in the patient anatomy correspond to a plurality of joints, the method further comprising: training, using a reference image dataset, a machine learning model to generate the trained machine learning model, wherein the reference image dataset comprises a plurality of reference image data, each reference image data labeled with one or more known attributes.Embodiment 11: The method of embodiment 10, wherein the one or more attributes of the medical condition and the known attributes comprises a van Der Heijde-Sharpe (vdH-S) score for each of the plurality of joints.Embodiment 12: The method of any one of embodiments 1-11, wherein the medical condition is COVID-19, wherein the patient anatomy is a lung, the method further comprising: training, using a reference image dataset, a machine learning model to generate the trained machine learning model, wherein the reference image dataset comprises a plurality of reference image data, each reference image data labeled with one or more known attributes.Embodiment 13: The method of any one of embodiments 1-12, wherein the medical condition is pneumonia, wherein the patient anatomy is a lung, the method further comprising: training, using a reference image dataset, a machine learning model to generate the trained machine learning model, wherein the reference image dataset comprises a plurality of reference image data, each reference image data labeled with one or more known attributes.Embodiment 14: A system for evaluating a medical condition using multimodal image data, the system comprising: a processor; and memory storing instructions that, when executed by the processor, cause the processor to: receive, from a plurality of imaging modalities, a plurality of respective modality-specific image data commonly representing a patient anatomy comprising a plurality of regions of interest (ROI); identify, for each of the plurality of ROI, a portion of each of the modality-specific image data as representing the ROI; generate, for each ROI of the plurality of ROI, an ROI-specific image dataset comprising the respective portions representing the ROI from the plurality of modality-specific image data; concatenate, for each ROI, the respective portions in the respective ROI-specific image dataset to generate a concatenated ROI image data for the ROI; and apply, to a trained machine learning model, the concatenated ROI image data for each ROI to determine one or more attributes of the medical condition for one or more of the plurality of ROI.Embodiment 15: The system of embodiment 14, wherein the one or more attribute comprises at least one of: an identification of the medical condition; a presence of the medical condition; or a measurement of a severity of the medical condition.Embodiment 16: The system of embodiment 14 or 15, wherein the concatenated ROI image data is a 3-dimensional image data, wherein at least one modality-specific image data of the plurality of modality-specific image data is a 2-dimensional image data.Embodiment 17 : The system of any one of embodiments 14-16, wherein the concatenatedROI image data is a 2-dimensional image data, wherein the respective portions of the respective ROI-specific image dataset that are concatenated to form the concatenated ROI image data are 2 -dimensional image data.Embodiment 18: The system of any one of embodiments 14-17, wherein the instructions, when executed, cause the processor to concatenate the respective portions of the respective ROI-specific image dataset by aggregating image feature parameters for each of a plurality of image features common to the respective portions of the respective ROI-specific image dataset.Embodiment 19: The system of any one of embodiments 14-18, wherein the instructions, when executed, cause the processor to concatenate the respective portions of the respectiveROI-specific image dataset by averaging image feature parameters for each of a plurality of image features common to the respective portions of the respective ROI-specific image dataset.Embodiment 20: The system of any one of embodiments 14-19, wherein the plurality of imaging modalities comprises at least two or more imaging modalities selected from a group consisting of: an X-Ray modality, a computed tomography (CT) modality, a magnetic resonance imaging (MRI) modality, a functional MRI (fMRI) modality, a scintigraphy modality, a positron emissions tomography (PET) modality, a single-photon emission computed tomography (SPECT) modality, an ultrasound modality, a functional near-infrared spectroscopy modality, a near-infrared spectroscopy modality, a magnetic particle imaging (MPI) modality, an elastography modality, an electrical impedance tomography modality, and a photoacoustic imaging modality.Embodiment 21 : The system of any one of embodiments 14-20, wherein, for each of the plurality of ROI, the portion of each of the modality-specific image data representing the ROI is identified by applying each modality-specific image data to a second trained machine learning model.Embodiment 22: The system of embodiment 21, wherein the second trained machine learning model is a trained convolutional neural network.Embodiment 23 : The system of any one of embodiments 14-22, wherein the medical condition is an arthritis, wherein the plurality of ROI in the patient anatomy correspond to a plurality of joints, wherein the instructions, when executed, further cause the processor to: train, using a reference image dataset, a machine learning model to generate the trained machine learning model, wherein the reference image dataset comprises a plurality of reference image data, each reference image data labeled with one or more known attributes.Embodiment 24: The system of embodiment 23, wherein the one or more attributes of the medical condition and the known attributes comprises a van Der Heijde-Sharpe (vdH-S) score for each of the plurality of joints.Embodiment 25 : The system of any one of embodiments 14-24, wherein the medical condition is COVID-19, wherein the patient anatomy is a lung, wherein the instructions, whenexecuted, further cause the processor to: train, using a reference image dataset, a machine learning model to generate the trained machine learning model, wherein the reference image dataset comprises a plurality of reference image data, each reference image data labeled with one or more known attributes.Embodiment 26: The system of any one of embodiments 14-25, wherein the medical condition is pneumonia, wherein the patient anatomy is a lung, wherein the instructions, when executed, further cause the processor to: train, using a reference image dataset, a machine learning model to generate the trained machine learning model, wherein the reference image dataset comprises a plurality of reference image data, each reference image data labeled with one or more known attributes.Embodiment 27 : A method of evaluating a medical condition using longitudinal image data, the method comprising: receiving, by a computing device having a processor, from at least one imaging modality, a first image data of a patient anatomy at a first time point and a second image data of the patient anatomy at a second time point; identifying, by applying the first image data and the second image data to a first trained machine learning model, a region of interest (ROI) that is common to both the first image data and the second image data, wherein portions of the first image data and the second image data corresponding to the ROI comprise a first ROI-specific image data and a second ROI-specific image data, respectively; and applying, to a trained multi -instance machine learning model, the first ROI-specific image data and the second ROI-specific image data to generate one or more attributes of the medical condition.Embodiment 28: The method of embodiment 27, wherein the one or more attributes comprises at least one of: an indicia of a progression of the medical condition; or an indicia of a regression of the medical condition.Embodiment 29: The method of embodiment 27 or 28, wherein the one or more attributes comprises a prediction of a severity of the medical condition at a third time point, wherein the third time occurs after the first time point and the second time point.Embodiment 30: The method of any one of embodiments 27-29, wherein the at least one imaging modality comprises at least one of: an X-Ray modality, a computed tomography (CT)modality, a magnetic resonance imaging (MRI) modality, a functional MRI (fMRI) modality, a scintigraphy modality, a positron emissions tomography (PET) modality, a single-photon emission computed tomography (SPECT) modality, an ultrasound modality, a functional nearinfrared spectroscopy modality, a near-infrared spectroscopy modality, a magnetic particle imaging (MPI) modality, an elastography modality, an electrical impedance tomography modality, or a photoacoustic imaging modality.Embodiment 31 : The method of any one of embodiments 27-30, wherein the first trained machine learning model is a trained convolutional neural network.Embodiment 32: The method of any one of embodiments 27-31, wherein the medical condition is an arthritis, wherein the ROI corresponds to a joint, the method further comprising: training, using a reference image dataset, a multi-instance learning model to generate the trained multi-instance learning model, wherein the reference image dataset comprises a plurality of subsets of reference image data, each reference image data within any given subset corresponding to a different time point, each subset labeled with one or more known attributes.Embodiment 33: The method of embodiment 32, wherein the one or more attributes of the medical condition include a change in a van Der Heijde-Sharpe (vdH-S) score for the joint.Embodiment 34: The method of any one of embodiments 27-33, wherein the medical condition is COVID-19, wherein the patient anatomy is a lung, the method further comprising: training, using a reference image dataset, a multi-instance learning model to generate the trained multi-instance learning model, wherein the reference image dataset comprises a plurality of subsets of reference image data, each reference image data within any given subset corresponding to a different time point, each subset labeled with one or more known attributes.Embodiment 35: The method of any one of embodiments 27-34, wherein the medical condition is pneumonia, wherein the patient anatomy is a lung, the method further comprising: training, using a reference image dataset, a multi-instance learning model to generate the trained multi-instance learning model, wherein the reference image dataset comprises a plurality of subsets of reference image data, each reference image data within any given subset corresponding to a different time point, each subset labeled with one or more known attributes.Embodiment 36: The method of any one of embodiments 27-35, wherein the medical condition is non-small cell lung cancer, wherein the patient anatomy is a lung, the method further comprising: training, using a reference image dataset, a multi-instance learning model to generate the trained multi-instance learning model, wherein the reference image dataset comprises a plurality of subsets of reference image data, each reference image data within any given subset corresponding to a different time point, each subset labeled with one or more known attributes.Embodiment 37: A system for evaluating a medical condition using longitudinal image data, the system comprising: a processor; and memory storing instructions that, when executed by the processor, cause the processor to: receive, from at least one imaging modality, a first image data of a patient anatomy at a first time point and a second image data of the patient anatomy at a second time point; identify, by applying the first image data and the second image data to a first trained machine learning model, a region of interest (ROI) that is common to both the first image data and the second image data, wherein portions of the first image data and the second image data corresponding to the ROI comprise a first ROI-specific image data and a second ROI-specific image data, respectively; and apply, to a trained multi-instance machine learning model, the first ROI-specific image data and the second ROI-specific image data to generate one or more attributes of the medical condition.Embodiment 38: The system of embodiment 37, wherein the one or more attributes comprises at least one of: an indicia of a progression of the medical condition; or an indicia of a regression of the medical condition.Embodiment 39: The system of embodiment 37 or 38, wherein the one or more attributes comprises a prediction of a severity of the medical condition at a third time point, wherein the third time occurs after the first time point and the second time point.Embodiment 40: The system of any one of embodiments 37-39, wherein the at least one imaging modality comprises at least one of: an X-Ray modality, a computed tomography (CT) modality, a magnetic resonance imaging (MRI) modality, a functional MRI (fMRI) modality, a scintigraphy modality, a positron emissions tomography (PET) modality, a single-photon emission computed tomography (SPECT) modality, an ultrasound modality, a functional nearinfrared spectroscopy modality, a near-infrared spectroscopy modality, a magnetic particleimaging (MPI) modality, an elastography modality, an electrical impedance tomography modality, or a photoacoustic imaging modality.Embodiment 41 : The system of any one of embodiments 37-40, wherein the first trained machine learning model is a trained convolutional neural network.Embodiment 42: The system of any one of embodiments 37-41, wherein the medical condition is an arthritis, wherein the ROI corresponds to a joint, wherein the instructions, when executed, further cause the processor to: train, using a reference image dataset, a multi -instance learning model to generate the trained multi-instance learning model, wherein the reference image dataset comprises a plurality of subsets of reference image data, each reference image data within any given subset corresponding to a different time point, each subset labeled with one or more known attributes.Embodiment 43 : The system of embodiment 42, wherein the one or more attributes of the medical condition include a change in a van Der Heijde-Sharpe (vdH-S) score for the joint.Embodiment 44: The system of any one of embodiments 37-43, wherein the medical condition is COVID-19, wherein the patient anatomy is a lung, wherein the instructions, when executed, further cause the processor to: train, using a reference image dataset, a multi -instance learning model to generate the trained multi-instance learning model, wherein the reference image dataset comprises a plurality of subsets of reference image data, each reference image data within any given subset corresponding to a different time point, each subset labeled with one or more known attributes.Embodiment 45 : The system of any one of embodiments 37-44, wherein the medical condition is pneumonia, wherein the patient anatomy is a lung, wherein the instructions, when executed, further cause the processor to: train, using a reference image dataset, a multi -instance learning model to generate the trained multi-instance learning model, wherein the reference image dataset comprises a plurality of subsets of reference image data, each reference image data within any given subset corresponding to a different time point, each subset labeled with one or more known attributes.Embodiment 46: The system of any one of embodiments 37-45, wherein the medical condition is non-small cell lung cancer, wherein the patient anatomy is a lung, wherein the instructions, when executed, further cause the processor to: train, using a reference image dataset, a multi-instance learning model to generate the trained multi-instance learning model, wherein the reference image dataset comprises a plurality of subsets of reference image data, each reference image data within any given subset corresponding to a different time point, each subset labeled with one or more known attributes.Embodiment 47: A method of evaluating a medical condition using multimodal longitudinal image data, the method comprising: receiving, by a computing device having a processor, from each of a plurality of imaging modalities, a plurality of modality-specific image data commonly representing a region of interest (ROI) of a patient anatomy at different time points, including a first time point and a second time point; generating, from the plurality of modality-specific image data received from each of the plurality of imaging modalities, a first set of modality-specific image data received from different imaging modalities and representing the ROI at the first time point, and a second set of modality-specific image data received from the different imaging modalities and representing the ROI at the second time point; concatenating: the first set of modality-specific image data to form a concatenated first image data, and the second set of modality-specific image data to form a concatenated second image data; and applying, to a trained multi -instance machine learning model, the concatenated first image data and the concatenated second image data to generate one or more attributes of the medical condition.Embodiment 48: The method of embodiment 47, wherein the one or more attributes comprises at least one of: an indicia of a progression of the medical condition; or an indicia of a regression of the medical condition.Embodiment 49: The method of embodiment 47 or 48, wherein the one or more attributes comprises a prediction of a severity of the medical condition at a third time point, wherein the third time occurs after the first time point and the second time point.Embodiment 50: The method of any one of embodiments 47-49, wherein the concatenated first image data and the concatenated second image data are each 3 -dimensionalimage data, wherein at least one modality-specific image data of the first set and at least one modality-specific image data of the second set are 2-dimensional image data.Embodiment 51 : The method of any one of embodiments 47-50, wherein the concatenated first image data and the concatenated second image data are each 2-dimensional image data, wherein the modality-specific image data of each of the first set and the second set are 2-dimensional image data.Embodiment 52: The method of any one of embodiments 47-51, wherein concatenating the first set of modality-specific image data comprises aggregating image feature parameters for each of a plurality of image features common to the modality-specific image data of the first set, and wherein concatenating the second set of modality-specific image data comprises aggregating image feature parameters for each of a plurality of image features common to the modality-specific image data of the second set.Embodiment 53: The method of any one of embodiments 47-52, wherein concatenating the first set of modality-specific image data comprises averaging image feature parameters for each of a plurality of image features common to the modality-specific image data of the first set, and wherein concatenating the second set of modality-specific image data comprises averaging image feature parameters for each of a plurality of image features common to the modality-specific image data of the second set.Embodiment 54: The method of any one of embodiments 47-53, wherein the plurality of imaging modalities comprises at least two or more imaging modalities selected from a group consisting of: an X-Ray modality, a computed tomography (CT) modality, a magnetic resonance imaging (MRI) modality, a functional MRI (fMRI) modality, a scintigraphy modality, a positron emissions tomography (PET) modality, a single-photon emission computed tomography (SPECT) modality, an ultrasound modality, a functional near-infrared spectroscopy modality, a near-infrared spectroscopy modality, a magnetic particle imaging (MPI) modality, an elastography modality, an electrical impedance tomography modality, and a photoacoustic imaging modality.Embodiment 55: The method of any one of embodiments 47-54, wherein the medical condition is an arthritis, wherein the ROI corresponds to a joint, the method further comprising:training, using a reference image dataset, a multi-instance learning model to generate the trained multi-instance learning model, wherein the reference image dataset comprises a plurality of subsets of reference image data, each reference image data within any given subset corresponding to a different time point, each subset labeled with one or more known attributes.Embodiment 56: The method of embodiment 55, wherein the one or more attributes of the medical condition include a change in a van Der Heijde-Sharpe (vdH-S) score for the joint.Embodiment 57: The method of any one of embodiments 47-56, wherein the medical condition is COVID-19, wherein the patient anatomy is a lung, the method further comprising: training, using a reference image dataset, a multi-instance learning model to generate the trained multi-instance learning model, wherein the reference image dataset comprises a plurality of subsets of reference image data, each reference image data within any given subset corresponding to a different time point, each subset labeled with one or more known attributes.Embodiment 58: The method of any one of embodiments 47-57, wherein the medical condition is pneumonia, wherein the patient anatomy is a lung, the method further comprising: training, using a reference image dataset, a multi-instance learning model to generate the trained multi-instance learning model, wherein the reference image dataset comprises a plurality of subsets of reference image data, each reference image data within any given subset corresponding to a different time point, each subset labeled with one or more known attributes.Embodiment 59: The method of any one of embodiments 47-58, wherein the medical condition is non-small cell lung cancer, wherein the patient anatomy is a lung, the method further comprising: training, using a reference image dataset, a multi-instance learning model to generate the trained multi-instance learning model, wherein the reference image dataset comprises a plurality of subsets of reference image data, each reference image data within any given subset corresponding to a different time point, each subset labeled with one or more known attributes.Embodiment 60: A system for evaluating a medical condition using multimodal longitudinal image data, the system comprising: a processor; and memory storing instructions that, when executed by the processor, cause the processor to: receive, from each of a plurality of imaging modalities, a plurality of modality-specific image data commonly representing aregion of interest (ROI) of a patient anatomy at different time points, including a first time point and a second time point; generate, from the plurality of modality-specific image data received from each of the plurality of imaging modalities, a first set of modality-specific image data received from different imaging modalities and representing the ROI at the first time point, and a second set of modality-specific image data received from the different imaging modalities and representing the ROI at the second time point; concatenate : the first set of modality-specific image data to form a concatenated first image data, and the second set of modality-specific image data to form a concatenated second image data; and apply, to a trained multi-instance machine learning model, the concatenated first image data and the concatenated second image data to generate one or more attributes of the medical condition.Embodiment 61 : The system of embodiment 60, wherein the one or more attributes comprises at least one of: an indicia of a progression of the medical condition; or an indicia of a regression of the medical condition.Embodiment 62: The system of embodiment 60 or 61, wherein the one or more attributes comprises a prediction of a severity of the medical condition at a third time point, wherein the third time occurs after the first time point and the second time point.Embodiment 63 : The system of any one of embodiments 60-62, wherein the concatenated first image data and the concatenated second image data are each 3 -dimensional image data, wherein at least one modality-specific image data of the first set and at least one modalityspecific image data of the second set are 2-dimensional image data.Embodiment 64: The system of any one of embodiments 60-63, wherein the concatenated first image data and the concatenated second image data are each 2-dimensional image data, wherein the modality-specific image data of each of the first set and the second set are 2- dimensional image data.Embodiment 65 : The system of any one of embodiments 60-64, wherein concatenating the first set of modality-specific image data comprises aggregating image feature parameters for each of a plurality of image features common to the modality-specific image data of the first set, and wherein concatenating the second set of modality-specific image data comprisesaggregating image feature parameters for each of a plurality of image features common to the modality-specific image data of the second set.Embodiment 66: The system of any one of embodiments 60-65, wherein the instructions, when executed, cause the processor to concatenate the first set of modality-specific image data by averaging image feature parameters for each of a plurality of image features common to the modality-specific image data of the first set, and wherein the instructions, when executed, cause the processor to concatenate the second set of modality-specific image data by averaging image feature parameters for each of a plurality of image features common to the modality-specific image data of the second set.Embodiment 67 : The system of any one of embodiments 60-66, wherein the plurality of imaging modalities comprises at least two or more imaging modalities selected from a group consisting of: an X-Ray modality, a computed tomography (CT) modality, a magnetic resonance imaging (MRI) modality, a functional MRI (fMRI) modality, a scintigraphy modality, a positron emissions tomography (PET) modality, a single-photon emission computed tomography (SPECT) modality, an ultrasound modality, a functional near-infrared spectroscopy modality, a near-infrared spectroscopy modality, a magnetic particle imaging (MPI) modality, an elastography modality, an electrical impedance tomography modality, and a photoacoustic imaging modality.Embodiment 68: The system of any one of embodiments 60-67, wherein the medical condition is an arthritis, wherein the ROI corresponds to a joint, wherein the instructions, when executed, further cause the processor to: train, using a reference image dataset, a multi -instance learning model to generate the trained multi-instance learning model, wherein the reference image dataset comprises a plurality of subsets of reference image data, each reference image data within any given subset corresponding to a different time point, each subset labeled with one or more known attributes.Embodiment 69: The system of embodiment 68, wherein the one or more attributes of the medical condition include a change in a van Der Heijde-Sharpe (vdH-S) score for the joint.Embodiment 70: The system of any one of embodiments 60-69, wherein the medical condition is COVID-19, wherein the patient anatomy is a lung, wherein the instructions, whenexecuted, further cause the processor to: train, using a reference image dataset, a multi -instance learning model to generate the trained multi-instance learning model, wherein the reference image dataset comprises a plurality of subsets of reference image data, each reference image data within any given subset corresponding to a different time point, each subset labeled with one or more known attributes.Embodiment 71 : The system of any one of embodiments 60-70, wherein the medical condition is pneumonia, wherein the patient anatomy is a lung, wherein the instructions, when executed, further cause the processor to: train, using a reference image dataset, a multi -instance learning model to generate the trained multi-instance learning model, wherein the reference image dataset comprises a plurality of subsets of reference image data, each reference image data within any given subset corresponding to a different time point, each subset labeled with one or more known attributes.Embodiment 72: The system of any one of embodiments 60-71, wherein the medical condition is non-small cell lung cancer, wherein the patient anatomy is a lung, wherein the instructions, when executed, further cause the processor to: train, using a reference image dataset, a multi-instance learning model to generate the trained multi-instance learning model, wherein the reference image dataset comprises a plurality of subsets of reference image data, each reference image data within any given subset corresponding to a different time point, each subset labeled with one or more known attributes.Embodiment 73: A non-transitory computer-readable medium storing computer instructions that, when executed by a computer, cause the computer to perform the method of any one of embodiments 1-13.Embodiment 74: A non-transitory computer-readable medium storing computer instructions that, when executed by a computer, cause the computer to perform the method of any one of embodiments 27-36.Embodiment 75: A non-transitory computer-readable medium storing computer instructions that, when executed by a computer, cause the computer to perform the method of any one of embodiments 47-59.

Claims

CLAIMSWhat is claimed is:

1. A method of evaluating a medical condition using multimodal image data, the method comprising: receiving, by a computing device having a processor, from a plurality of imaging modalities, a plurality of respective modality-specific image data commonly representing a patient anatomy comprising a plurality of regions of interest (ROI); identifying, for each of the plurality of ROI, a portion of each of the modality-specific image data as representing the ROI; generating, for each ROI of the plurality of ROI, an ROI-specific image dataset comprising the respective portions representing the ROI from the plurality of modality-specific image data; concatenating, for each ROI, the respective portions in the respective ROI-specific image dataset to generate a concatenated ROI image data for the ROI; and applying, to a trained machine learning model, the concatenated ROI image data for each ROI to determine one or more attributes of the medical condition for one or more of the plurality of ROI.

2. The method of claim 1, wherein the one or more attribute comprises at least one of: an identification of the medical condition; a presence of the medical condition; or a measurement of a severity of the medical condition.

3. The method of claim 1 or 2, wherein the concatenated ROI image data is a 3- dimensional image data, wherein at least one modality-specific image data of the plurality of modality-specific image data is a 2-dimensional image data.

4. The method of any one of claims 1-3, wherein the concatenated ROI image data is a 2-dimensional image data, wherein the respective portions of the respective ROI-specific image dataset that are concatenated to form the concatenated ROI image data are 2-dimensional image data.

5. The method of any one of claims 1-4, wherein concatenating the respective portions of the respective ROI-specific image dataset comprises aggregating image feature parameters for each of a plurality of image features common to the respective portions of the respective ROI-specific image dataset.

6. The method of any one of claims 1-5, wherein concatenating the respective portions of the respective ROI-specific image dataset comprises averaging image feature parameters for each of a plurality of image features common to the respective portions of the respective ROI-specific image dataset.

7. The method of any one of claims 1-6, wherein the plurality of imaging modalities comprises at least two or more imaging modalities selected from a group consisting of: an X-Ray modality, a computed tomography (CT) modality, a magnetic resonance imaging (MRI) modality, a functional MRI (fMRI) modality, a scintigraphy modality, a positron emissions tomography (PET) modality, a single-photon emission computed tomography (SPECT) modality, an ultrasound modality, a functional near-infrared spectroscopy modality, a near-infrared spectroscopy modality, a magnetic particle imaging (MPI) modality, an elastography modality, an electrical impedance tomography modality, and a photoacoustic imaging modality.

8. The method of any one of claims 1-7, wherein, for each of the plurality of ROI, the portion of each of the modality-specific image data representing the ROI is identified by applying each modality-specific image data to a second trained machine learning model.

9. The method of claim 8, wherein the second trained machine learning model is a trained convolutional neural network.

10. The method of any one of claims 1-9, wherein the medical condition is an arthritis, wherein the plurality of ROI in the patient anatomy correspond to a plurality of joints, the method further comprising: training, using a reference image dataset, a machine learning model to generate the trained machine learning model, wherein the reference image dataset comprises a plurality of reference image data, each reference image data labeled with one or more known attributes.

11. The method of claim 10, wherein the one or more attributes of the medical condition and the known attributes comprises a van Der Heijde-Sharpe (vdH-S) score for each of the plurality of joints.

12. The method of any one of claims 1-11, wherein the medical condition is COVID-19, wherein the patient anatomy is a lung, the method further comprising: training, using a reference image dataset, a machine learning model to generate the trained machine learning model, wherein the reference image dataset comprises a plurality of reference image data, each reference image data labeled with one or more known attributes.

13. The method of any one of claims 1-12, wherein the medical condition is pneumonia, wherein the patient anatomy is a lung, the method further comprising: training, using a reference image dataset, a machine learning model to generate the trained machine learning model, wherein the reference image dataset comprises a plurality of reference image data, each reference image data labeled with one or more known attributes.

14. A system for evaluating a medical condition using multimodal image data, the system comprising: a processor; and memory storing instructions that, when executed by the processor, cause the processor to: receive, from a plurality of imaging modalities, a plurality of respective modalityspecific image data commonly representing a patient anatomy comprising a plurality of regions of interest (ROI); identify, for each of the plurality of ROI, a portion of each of the modality-specific image data as representing the ROI;generate, for each ROI of the plurality of ROI, an ROI-specific image dataset comprising the respective portions representing the ROI from the plurality of modality-specific image data; concatenate, for each ROI, the respective portions in the respective ROI-specific image dataset to generate a concatenated ROI image data for the ROI; and apply, to a trained machine learning model, the concatenated ROI image data for each ROI to determine one or more attributes of the medical condition for one or more of the plurality of ROI.

15. The system of claim 14, wherein the one or more attribute comprises at least one of: an identification of the medical condition; a presence of the medical condition; or a measurement of a severity of the medical condition.

16. The system of claim 14 or 15, wherein the concatenated ROI image data is a 3- dimensional image data, wherein at least one modality-specific image data of the plurality of modality-specific image data is a 2-dimensional image data.

17. The system of any one of claims 14-16, wherein the concatenated ROI image data is a 2-dimensional image data, wherein the respective portions of the respective ROI- specific image dataset that are concatenated to form the concatenated ROI image data are 2- dimensional image data.

18. The system of any one of claims 14-17, wherein the instructions, when executed, cause the processor to concatenate the respective portions of the respective ROI- specific image dataset by aggregating image feature parameters for each of a plurality of image features common to the respective portions of the respective ROI-specific image dataset.

19. The system of any one of claims 14-18, wherein the instructions, when executed, cause the processor to concatenate the respective portions of the respective ROI- specific image dataset by averaging image feature parameters for each of a plurality of image features common to the respective portions of the respective ROI-specific image dataset.

20. The system of any one of claims 14-19, wherein the plurality of imaging modalities comprises at least two or more imaging modalities selected from a group consisting of: an X-Ray modality, a computed tomography (CT) modality, a magnetic resonance imaging (MRI) modality, a functional MRI (fMRI) modality, a scintigraphy modality, a positron emissions tomography (PET) modality, a single-photon emission computed tomography (SPECT) modality, an ultrasound modality, a functional near-infrared spectroscopy modality, a near-infrared spectroscopy modality, a magnetic particle imaging (MPI) modality, an elastography modality, an electrical impedance tomography modality, and a photoacoustic imaging modality.

21. The system of any one of claims 14-20, wherein, for each of the plurality of ROI, the portion of each of the modality-specific image data representing the ROI is identified by applying each modality-specific image data to a second trained machine learning model.

22. The system of claim 21, wherein the second trained machine learning model is a trained convolutional neural network.

23. The system of any one of claims 14-22, wherein the medical condition is an arthritis, wherein the plurality of ROI in the patient anatomy correspond to a plurality of joints, wherein the instructions, when executed, further cause the processor to: train, using a reference image dataset, a machine learning model to generate the trained machine learning model, wherein the reference image dataset comprises a plurality of reference image data, each reference image data labeled with one or more known attributes.

24. The system of claim 23, wherein the one or more attributes of the medical condition and the known attributes comprises a van Der Heijde-Sharpe (vdH-S) score for each of the plurality of joints.

25. The system of any one of claims 14-24, wherein the medical condition is COVID-19, wherein the patient anatomy is a lung, wherein the instructions, when executed, further cause the processor to: train, using a reference image dataset, a machine learning model to generate the trained machine learning model, wherein the reference image dataset comprises a plurality of reference image data, each reference image data labeled with one or more known attributes.

26. The system of any one of claims 14-25, wherein the medical condition is pneumonia, wherein the patient anatomy is a lung, wherein the instructions, when executed, further cause the processor to: train, using a reference image dataset, a machine learning model to generate the trained machine learning model, wherein the reference image dataset comprises a plurality of reference image data, each reference image data labeled with one or more known attributes.

27. A method of evaluating a medical condition using longitudinal image data, the method comprising: receiving, by a computing device having a processor, from at least one imaging modality, a first image data of a patient anatomy at a first time point and a second image data of the patient anatomy at a second time point; identifying, by applying the first image data and the second image data to a first trained machine learning model, a region of interest (ROI) that is common to both the first image data and the second image data, wherein portions of the first image data and the second image data corresponding to the ROI comprise a first ROI-specific image data and a second ROI-specific image data, respectively; and applying, to a trained multi-instance machine learning model, the first ROI-specific image data and the second ROI-specific image data to generate one or more attributes of the medical condition.

28. The method of claim 27, wherein the one or more attributes comprises at least one of: an indicia of a progression of the medical condition; or an indicia of a regression of the medical condition.

29. The method of claim 27 or 28, wherein the one or more attributes comprises a prediction of a severity of the medical condition at a third time point, wherein the third time occurs after the first time point and the second time point.

30. The method of any one of claims 27-29, wherein the at least one imaging modality comprises at least one of: an X-Ray modality, a computed tomography (CT) modality, a magnetic resonance imaging (MRI) modality, a functional MRI (fMRI) modality, a scintigraphy modality, a positron emissions tomography (PET) modality, a single-photon emission computed tomography (SPECT) modality, an ultrasound modality, a functional near-infrared spectroscopy modality, a near-infrared spectroscopy modality, a magnetic particle imaging (MPI) modality, an elastography modality, an electrical impedance tomography modality, or a photoacoustic imaging modality.

31. The method of any one of claims 27-30, wherein the first trained machine learning model is a trained convolutional neural network.

32. The method of any one of claims 27-31, wherein the medical condition is an arthritis, wherein the ROI corresponds to a joint, the method further comprising: training, using a reference image dataset, a multi-instance learning model to generate the trained multi-instance learning model, wherein the reference image dataset comprises a plurality of subsets of reference image data, each reference image data within any given subset corresponding to a different time point, each subset labeled with one or more known attributes.

33. The method of claim 32, wherein the one or more attributes of the medical condition include a change in a van Der Heijde-Sharpe (vdH-S) score for the joint.

34. The method of any one of claims 27-33, wherein the medical condition is COVID-19, wherein the patient anatomy is a lung, the method further comprising: training, using a reference image dataset, a multi-instance learning model to generate the trained multi-instance learning model, wherein the reference image dataset comprises a plurality of subsets of reference image data, each reference image data within any given subset corresponding to a different time point, each subset labeled with one or more known attributes.

35. The method of any one of claims 27-34, wherein the medical condition is pneumonia, wherein the patient anatomy is a lung, the method further comprising: training, using a reference image dataset, a multi-instance learning model to generate the trained multi-instance learning model, wherein the reference image dataset comprises a plurality of subsets of reference image data, each reference image data within any given subset corresponding to a different time point, each subset labeled with one or more known attributes.

36. The method of any one of claims 27-35, wherein the medical condition is nonsmall cell lung cancer, wherein the patient anatomy is a lung, the method further comprising: training, using a reference image dataset, a multi-instance learning model to generate the trained multi-instance learning model, wherein the reference image dataset comprises a plurality of subsets of reference image data, each reference image data within any given subset corresponding to a different time point, each subset labeled with one or more known attributes.

37. A system for evaluating a medical condition using longitudinal image data, the system comprising: a processor; and memory storing instructions that, when executed by the processor, cause the processor to: receive, from at least one imaging modality, a first image data of a patient anatomy at a first time point and a second image data of the patient anatomy at a second time point; identify, by applying the first image data and the second image data to a first trained machine learning model, a region of interest (RO I) that is common to both the first image data and the second image data, wherein portions of the first image data and the second image data corresponding to the ROI comprise a first ROI-specific image data and a second ROI-specific image data, respectively; andapply, to a trained multi -instance machine learning model, the first ROI-specific image data and the second ROI-specific image data to generate one or more attributes of the medical condition.

38. The system of claim 37, wherein the one or more attributes comprises at least one of: an indicia of a progression of the medical condition; or an indicia of a regression of the medical condition.

39. The system of claim 37 or 38, wherein the one or more attributes comprises a prediction of a severity of the medical condition at a third time point, wherein the third time occurs after the first time point and the second time point.

40. The system of any one of claims 37-39, wherein the at least one imaging modality comprises at least one of: an X-Ray modality, a computed tomography (CT) modality, a magnetic resonance imaging (MRI) modality, a functional MRI (fMRI) modality, a scintigraphy modality, a positron emissions tomography (PET) modality, a single-photon emission computed tomography (SPECT) modality, an ultrasound modality, a functional near-infrared spectroscopy modality, a near-infrared spectroscopy modality, a magnetic particle imaging (MPI) modality, an elastography modality, an electrical impedance tomography modality, or a photoacoustic imaging modality.

41. The system of any one of claims 37-40, wherein the first trained machine learning model is a trained convolutional neural network.

42. The system of any one of claims 37-41, wherein the medical condition is an arthritis, wherein the ROI corresponds to a joint, wherein the instructions, when executed, further cause the processor to: train, using a reference image dataset, a multi-instance learning model to generate the trained multi-instance learning model, wherein the reference image dataset comprises a plurality of subsets of reference image data, each reference image data within any given subset corresponding to a different time point, each subset labeled with one or more known attributes.

43. The system of claim 42, wherein the one or more attributes of the medical condition include a change in a van Der Heijde-Sharpe (vdH-S) score for the joint.

44. The system of any one of claims 37-43, wherein the medical condition is COVID-19, wherein the patient anatomy is a lung, wherein the instructions, when executed, further cause the processor to: train, using a reference image dataset, a multi-instance learning model to generate the trained multi-instance learning model, wherein the reference image dataset comprises a plurality of subsets of reference image data, each reference image data within any given subset corresponding to a different time point, each subset labeled with one or more known attributes.

45. The system of any one of claims 37-44, wherein the medical condition is pneumonia, wherein the patient anatomy is a lung, wherein the instructions, when executed, further cause the processor to: train, using a reference image dataset, a multi-instance learning model to generate the trained multi-instance learning model, wherein the reference image dataset comprises a plurality of subsets of reference image data, each reference image data within any given subset corresponding to a different time point, each subset labeled with one or more known attributes.

46. The system of any one of claims 37-45, wherein the medical condition is nonsmall cell lung cancer, wherein the patient anatomy is a lung, wherein the instructions, when executed, further cause the processor to: train, using a reference image dataset, a multi-instance learning model to generate the trained multi-instance learning model, wherein the reference image dataset comprises a plurality of subsets of reference image data, each reference image data within any given subset corresponding to a different time point, each subset labeled with one or more known attributes.

47. A method of evaluating a medical condition using multimodal longitudinal image data, the method comprising: receiving, by a computing device having a processor, from each of a plurality of imaging modalities, a plurality of modality-specific image data commonly representing a region of interest (ROI) of a patient anatomy at different time points, including a first time point and a second time point; generating, from the plurality of modality-specific image data received from each of the plurality of imaging modalities, a first set of modality-specific image data received from different imaging modalities and representing the ROI at the first time point, and a second set of modality-specific image data received from the different imaging modalities and representing the ROI at the second time point; concatenating: the first set of modality-specific image data to form a concatenated first image data, and the second set of modality-specific image data to form a concatenated second image data; and applying, to a trained multi-instance machine learning model, the concatenated first image data and the concatenated second image data to generate one or more attributes of the medical condition.

48. The method of claim 47, wherein the one or more attributes comprises at least one of: an indicia of a progression of the medical condition; or an indicia of a regression of the medical condition.

49. The method of claim 47 or 48, wherein the one or more attributes comprises a prediction of a severity of the medical condition at a third time point, wherein the third time occurs after the first time point and the second time point.

50. The method of any one of claims 47-49, wherein the concatenated first image data and the concatenated second image data are each 3 -dimensional image data, wherein at least one modality-specific image data of the first set and at least one modality-specific image data of the second set are 2-dimensional image data.

51. The method of any one of claims 47-50, wherein the concatenated first image data and the concatenated second image data are each 2-dimensional image data, wherein the modality-specific image data of each of the first set and the second set are 2-dimensional image data.

52. The method of any one of claims 47-51, wherein concatenating the first set of modality-specific image data comprises aggregating image feature parameters for each of a plurality of image features common to the modality-specific image data of the first set, and wherein concatenating the second set of modality-specific image data comprises aggregating image feature parameters for each of a plurality of image features common to the modality-specific image data of the second set.

53. The method of any one of claims 47-52, wherein concatenating the first set of modality-specific image data comprises averaging image feature parameters for each of a plurality of image features common to the modalityspecific image data of the first set, and wherein concatenating the second set of modality-specific image data comprises averaging image feature parameters for each of a plurality of image features common to the modality-specific image data of the second set.

54. The method of any one of claims 47-53, wherein the plurality of imaging modalities comprises at least two or more imaging modalities selected from a group consisting of: an X-Ray modality, a computed tomography (CT) modality, a magnetic resonance imaging (MRI) modality, a functional MRI (fMRI) modality, a scintigraphy modality, a positron emissions tomography (PET) modality, a single-photon emission computed tomography (SPECT) modality, an ultrasound modality, a functional near-infrared spectroscopy modality, a near-infrared spectroscopy modality, a magnetic particle imaging (MPI) modality,an elastography modality, an electrical impedance tomography modality, and a photoacoustic imaging modality.

55. The method of any one of claims 47-54, wherein the medical condition is an arthritis, wherein the ROI corresponds to a joint, the method further comprising: training, using a reference image dataset, a multi-instance learning model to generate the trained multi-instance learning model, wherein the reference image dataset comprises a plurality of subsets of reference image data, each reference image data within any given subset corresponding to a different time point, each subset labeled with one or more known attributes.

56. The method of claim 55, wherein the one or more attributes of the medical condition include a change in a van Der Heijde-Sharpe (vdH-S) score for the joint.

57. The method of any one of claims 47-56, wherein the medical condition is COVID-19, wherein the patient anatomy is a lung, the method further comprising: training, using a reference image dataset, a multi-instance learning model to generate the trained multi-instance learning model, wherein the reference image dataset comprises a plurality of subsets of reference image data, each reference image data within any given subset corresponding to a different time point, each subset labeled with one or more known attributes.

58. The method of any one of claims 47-57, wherein the medical condition is pneumonia, wherein the patient anatomy is a lung, the method further comprising: training, using a reference image dataset, a multi-instance learning model to generate the trained multi-instance learning model, wherein the reference image dataset comprises a plurality of subsets of reference image data, each reference image data within any given subset corresponding to a different time point, each subset labeled with one or more known attributes.

59. The method of any one of claims 47-58, wherein the medical condition is nonsmall cell lung cancer, wherein the patient anatomy is a lung, the method further comprising: training, using a reference image dataset, a multi-instance learning model to generate the trained multi-instance learning model, wherein the reference image dataset comprises a plurality of subsets of reference image data, each reference image data within any given subset corresponding to a different time point, each subset labeled with one or more known attributes.

60. A system for evaluating a medical condition using multimodal longitudinal image data, the system comprising: a processor; and memory storing instructions that, when executed by the processor, cause the processor to: receive, from each of a plurality of imaging modalities, a plurality of modalityspecific image data commonly representing a region of interest (ROI) of a patient anatomy at different time points, including a first time point and a second time point; generate, from the plurality of modality-specific image data received from each of the plurality of imaging modalities, a first set of modality-specific image data received from different imaging modalities and representing the ROI at the first time point, and a second set of modality-specific image data received from the different imaging modalities and representing the ROI at the second time point; concatenate: the first set of modality-specific image data to form a concatenated first image data, and the second set of modality-specific image data to form a concatenated second image data; and apply, to a trained multi-instance machine learning model, the concatenated first image data and the concatenated second image data to generate one or more attributes of the medical condition.

61. The system of claim 60, wherein the one or more attributes comprises at least one of: an indicia of a progression of the medical condition; or an indicia of a regression of the medical condition.

62. The system of claim 60 or 61, wherein the one or more attributes comprises a prediction of a severity of the medical condition at a third time point, wherein the third time occurs after the first time point and the second time point.

63. The system of any one of claims 60-62, wherein the concatenated first image data and the concatenated second image data are each 3 -dimensional image data, wherein at least one modality-specific image data of the first set and at least one modality-specific image data of the second set are 2-dimensional image data.

64. The system of any one of claims 60-63, wherein the concatenated first image data and the concatenated second image data are each 2-dimensional image data, wherein the modality-specific image data of each of the first set and the second set are 2-dimensional image data.

65. The system of any one of claims 60-64, wherein concatenating the first set of modality-specific image data comprises aggregating image feature parameters for each of a plurality of image features common to the modality-specific image data of the first set, and wherein concatenating the second set of modality-specific image data comprises aggregating image feature parameters for each of a plurality of image features common to the modality-specific image data of the second set.

66. The system of any one of claims 60-65, wherein the instructions, when executed, cause the processor to concatenate the first set of modality-specific image data by averaging image feature parameters for each of a plurality of image features common to the modality-specific image data of the first set, and wherein the instructions, when executed, cause the processor to concatenate the second set of modality-specific image data by averaging image feature parameters for each of a plurality of image features common to the modality-specific image data of the second set.

67. The system of any one of claims 60-66, wherein the plurality of imaging modalities comprises at least two or more imaging modalities selected from a group consisting of: an X-Ray modality, a computed tomography (CT) modality, a magnetic resonance imaging (MRI) modality, a functional MRI (fMRI) modality, a scintigraphy modality, a positron emissions tomography (PET) modality, a single-photon emission computed tomography (SPECT) modality, an ultrasound modality, a functional near-infrared spectroscopy modality, a near-infrared spectroscopy modality,a magnetic particle imaging (MPI) modality, an elastography modality, an electrical impedance tomography modality, and a photoacoustic imaging modality.

68. The system of any one of claims 60-67, wherein the medical condition is an arthritis, wherein the ROI corresponds to a joint, wherein the instructions, when executed, further cause the processor to: train, using a reference image dataset, a multi-instance learning model to generate the trained multi-instance learning model, wherein the reference image dataset comprises a plurality of subsets of reference image data, each reference image data within any given subset corresponding to a different time point, each subset labeled with one or more known attributes.

69. The system of claim 68, wherein the one or more attributes of the medical condition include a change in a van Der Heijde-Sharpe (vdH-S) score for the joint.

70. The system of any one of claims 60-69, wherein the medical condition is COVID-19, wherein the patient anatomy is a lung, wherein the instructions, when executed, further cause the processor to: train, using a reference image dataset, a multi-instance learning model to generate the trained multi-instance learning model, wherein the reference image dataset comprises a plurality of subsets of reference image data, each reference image data within any given subset corresponding to a different time point, each subset labeled with one or more known attributes.

71. The system of any one of claims 60-70, wherein the medical condition is pneumonia, wherein the patient anatomy is a lung, wherein the instructions, when executed, further cause the processor to: train, using a reference image dataset, a multi-instance learning model to generate the trained multi-instance learning model, wherein the reference image dataset comprises a plurality of subsets of reference image data, each reference image data within any given subset corresponding to a different time point, each subset labeled with one or more known attributes.

72. The system of any one of claims 60-71, wherein the medical condition is nonsmall cell lung cancer, wherein the patient anatomy is a lung, wherein the instructions, when executed, further cause the processor to:train, using a reference image dataset, a multi-instance learning model to generate the trained multi-instance learning model, wherein the reference image dataset comprises a plurality of subsets of reference image data, each reference image data within any given subset corresponding to a different time point, each subset labeled with one or more known attributes.

73. A non-transitory computer-readable medium storing computer instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1-13.

74. A non-transitory computer-readable medium storing computer instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 27-36.

75. A non-transitory computer-readable medium storing computer instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 47-59.

Citation Information

Patent Citations

  • Port mega foundation structure construction method for automation contruction of mega grade port logistics

    KR1020230017573A

  • Anatomy Aware Articulated Registration for Image Segmentation

    US20150023575A1

  • Multi-Modality Image Fusion for 3D Printing of Organ Morphology and Physiology

    US20170217102A1

  • Systems and methods to determine disease progression from artificial intelligence detection output

    US20200211694A1

  • System and Method for Interpretation of Multiple Medical Images Using Deep Learning

    US20220254023A1