Imaging based on a set of medical imaging modalities

A machine learning method fuses and transforms medical imaging data from multiple modalities into a unified representation, addressing ergonomic challenges and improving medical evaluation by integrating diverse imaging information without additional imaging needs.

JP2026021267APending Publication Date: 2026-02-10DASSAULT SYSTEMES SA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025117053
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-17
Filing Date
2025-07-11
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

The proliferation of medical imaging modalities leads to ergonomic challenges and lack of standardization for medical professionals, as different modalities provide varying levels of detail in visualizing anatomical and physiological information, necessitating improved solutions for integrating and standardizing medical imaging data.

Method used

A machine learning method that fuses images from multiple medical imaging modalities by aligning and training a function to calculate a fused image, using a dataset of registered images from various modalities, allowing for the reconstruction and transformation of images across different modalities without ad-hoc fusion rules.

Benefits of technology

Enables the integration of complementary medical information from diverse imaging modalities into a single representation, enhancing medical evaluation and reducing the need for additional imaging procedures by transforming images into more practical modalities for comparison and follow-up.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026021267000001_ABST
    Figure 2026021267000001_ABST
Patent Text Reader

Abstract

There is a need for improved solutions for medical imaging.SOLUTION: It particularly relates to a computer-implemented machine learning method of a function configured to receive as input a plurality of registered images of a same patient and each of the different modalities among a predetermined set of medical imaging modalities, and to compute a fused image. The machine learning method includes obtaining a data set including respective images for each patient of the plurality of patients and for each modality of at least a portion of each of the predetermined sets, wherein the respective images for the patient are aligned, and training the function based on the data set.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to the field of computer programs and systems, and more particularly to methods, systems, and data structures related to machine learning and medical imaging. [Background technology]

[0002] Medical imaging involves imaging the inside of a patient's body for clinical analysis and medical intervention, along with visual representations of the patient's internal organs or tissues. Over the past few decades, the proliferation of various imaging devices, sensors, and / or technologies has led to a proliferation of medical imaging modalities. Medical imaging modalities can visualize and / or process different aspects of the same patient's tissues, capturing different medical information depending on the modality being visualized and / or processed. For example, a PET scan can visualize the physiological activity of a tumor relatively clearly, but this medical imaging modality cannot accurately capture the tumor's contours. On the other hand, a CT scan can visualize the physiological activity less clearly than a PET scan, but can visualize the tumor's contours in great detail. Combining these two modalities allows medical professionals to obtain more advanced medical information suitable for use in treatment. However, the large number of medical imaging modalities and the numerous potential combinations depending on the patient's medical history or medical history / infrastructure often plague medical professionals and / or medical processes with ergonomics and / or a lack of standardization. Thus, addressing this medical imaging environment is complex. Summary of the Invention

[0003] In this context, there is a need for improved solutions for medical imaging.

[0004] Thus, there is provided a computer-implemented method (referred to as a "machine learning method") of a function configured to receive as input a plurality of registered images of the same patient and each of different modalities in a predetermined set of medical imaging modalities and to calculate a fused image. The machine learning method includes obtaining a dataset including a respective image for each patient in a plurality of patients and for each modality of at least some of each in the predetermined set, wherein the respective images for a patient are registered. The machine learning method also includes training the function based on the dataset.

[0005] The machine learning method may further include one or more of the following elements:

[0006] The function is configured to iteratively apply a fusion network to a pair of images to compute the fused image, the pair of images comprising two images of a plurality of aligned images in a first iteration, and the pair of images comprising one image of the plurality of aligned images and a result of applying the fusion network in a previous iteration, in each subsequent iteration.

[0007] The fusion network is identical in each iteration.

[0008] The function is further configured to calculate a reconstructed image for each modality of the predetermined set from the fused images.

[0009] The training includes minimizing a loss that includes a sum of reconstruction costs over the images of the dataset.

[0010] the function is configured to receive a variable number of images, including two, as input, and the loss further comprises a sum over the images of the dataset of stability losses, the stability loss being expressed for each corresponding image of each corresponding patient by a cost between (i) a first fusion image calculated by applying the function to all images of the corresponding patient included in the dataset as input, and (ii) a second fusion image calculated by applying the function to the corresponding images and the first fusion image as input.

[0011] The losses further include hostile losses.

[0012] The function is order-dependent for a plurality of input images, and training includes applying the function one or more times to corresponding inputs having a randomized order.

[0013] the function is further configured to receive as input, for a corresponding input image, a corresponding label representative of a modality of the corresponding input image, and optionally, said training comprises applying said function one or more times to corresponding inputs including corresponding fused images and corresponding labels representative of fusion properties of the corresponding fused images.

[0014] The predetermined set of medical imaging modalities may be selected from the group consisting of autorefraction, angioscopy, bone densitometry (US), biomagnetic imaging, bone densitometry (X-ray), color flow Doppler, cinefluoroscopy, colposcopy, computed radiography, cystoscopy, computed tomography, duplex Doppler, digital fluoroscopy, fluoroscopy, digital microscopy, digital subtraction angiography, digital radiography, echocardiography, electrocardiography, cardiac electrophysiology, endoscopy, fluorescein angiography, fiducial, fundus examination, general microscopy, hard copy, hemodynamic waveform, intraoral radiography, intraocular lens data, intravascular optical coherence tomography, intravascular ultrasound, keratometry, lensmetry, laparoscopy, laser surface scanning, magnetic resonance angiography, mammography, magnetic resonance imaging, MR T1 weighted imaging, MR T2-weighted imaging, MR proton density-weighted imaging, MR steady-state free precession, MR effective T2, MR susceptibility-weighted imaging, MR short-tau inversion recovery imaging, MR fluid-attenuated inversion recovery imaging, MR double inversion recovery imaging, MR conventional diffusion-weighted imaging, MR apparent diffusion coefficient imaging, MR diffusion tensor imaging, MR dynamic susceptibility contrast imaging, MR arterial spin contrast imaging, MR dynamic contrast-enhanced imaging, MR blood oxygen level-dependent imaging, MR time-of-flight imaging, MR phase contrast imaging, magnetic resonance spectroscopy, nuclear medicine, axial measurement, optical coherence tomography (non-ophthalmology), ophthalmic photography, ophthalmic mapping, ophthalmic refraction, ophthalmic tomography Includes one or more of the following: layer imaging, ophthalmic visual field, optical surface scan, other, positron emission tomography (PET), panoramic x-ray, respiratory waveform, fluoroscopy, radiology images (traditional film / screen), radiotherapy dose, radiotherapy images, radiotherapy planning, radiotherapy records, radiotherapy structure sets, segmentation, slide microscopy, stereometric relationships, single photon emission computed tomography (SPECT), automated slide stainer, thermography, ultrasound, A-mode US, B-mode US, M-mode US, visual acuity, video fluoroscopy, x-ray angiography, and external camera imaging.

[0015] There is also provided a method of using a function machine-learned by the above-described machine learning method (referred to as the "Method of Use"), which includes inputting a plurality of registered images of the same patient, each of which is a different modality from a predetermined set of medical image modalities, into the function, and also includes calculating a fused image using the inputs with the function.

[0016] The method of use may further include one or more of the following:

[0017] The method of use further comprises outputting and / or displaying the fused image.

[0018] The method of use further comprises reconstructing a corresponding reconstructed image for each of one or more modalities in the predetermined set of medical imaging modalities.

[0019] The method of use further comprises reconstructing a corresponding reconstructed image for each of the one or more modalities in the predetermined set of medical image modalities, the one or more modalities comprising an input modality.

[0020] The method of use further comprises outputting one or more reconstructed images.

[0021] The method of use further comprises displaying one or more reconstructed images.

[0022] Further provided is a computer program comprising instructions for carrying out said machine learning method and / or said method of use.

[0023] A function machine-learned by the machine learning method is also provided.

[0024] There is further provided an apparatus comprising a data storage medium having said computer program and / or said functions recorded thereon.

[0025] The device forms or functions as a non-transitory computer-readable storage medium, such as on a Software as a Service (SaaS) or other server- or cloud-based platform. The device may alternatively include a processor coupled to a data storage medium. The device may thus form all or part of a computer system (e.g., the device is a subsystem of an overall system). The device may further include a graphical user interface coupled to the processor. [Brief explanation of the drawings]

[0026] Non-limiting examples will now be described with reference to the accompanying drawings. [Figure 1] FIG. 1 illustrates a flowchart of an example machine learning method. [Figure 2] FIG. 10 is a flowchart of an example method of use. [Figure 3] FIG. 1 illustrates an example of a system. [Figure 4] FIG. [Figure 5] FIG. [Figure 6] FIG. [Figure 7] FIG. [Figure 8] FIG. [Figure 9] FIG. [Figure 10] FIG. [Figure 11] FIG. [Figure 12] FIG. DETAILED DESCRIPTION OF THE INVENTION

[0027] With reference to the flowchart of FIG. 1 , a computer-implemented method for machine learning a function is proposed. The function is configured to receive as input a plurality of aligned images of the same patient, each image of the plurality of images being of a different modality within a predetermined set of medical imaging modalities. Once machine-learned, the function is configured to calculate a fused image from the input. The machine learning method includes obtaining a dataset S10 and training a function based on the dataset S20. The dataset includes training data for a plurality of patients. In particular, the dataset includes corresponding training examples (i.e., training patterns) for each of the plurality of patients. The training examples for the corresponding patients include corresponding images for at least a corresponding portion of the predetermined set of medical imaging modalities. "At least a portion of a predetermined set of medical imaging modalities" means a portion of the predetermined set of medical imaging modalities (i.e., one or more medical imaging modalities forming a subset of the predetermined set) or all of the predetermined set of medical imaging modalities (i.e., all of the medical imaging modalities in the predetermined set). In this dataset, if any (e.g., if the dataset comprises at least two corresponding images for one or more given patients, each for a given patient), all corresponding images of the same patient (i.e., in the same training example) are aligned (i.e., registered). Images are said to be "aligned" when they have the same dimensions (not necessarily the same resolution) and orientation, and they represent the same anatomy of the patient, which is the same shape and size. In other words, this alignment ensures that any anatomical point is in the same position in all images. Thus, there is no distortion of a patient's anatomical structures across the aligned images.For example, in abdominal images, organs may move depending on the patient's position during image acquisition, but registering several images of the abdomen can correct for such movement. In other words, the same pixel (or the same pixel coordinate) in each of several "registered" images represents the same point on the patient's body.

[0028] Machine learning methods offer improved solutions for medical imaging.

[0029] With particular reference to the flowchart of Figure 2, the present machine learning method can be trained (i.e., machine-learned) once to obtain a machine learning function that can be included in a method of use. The method of use includes inputting multiple images of the same patient into the function (S30) and using the function to calculate (and output) a fused image using the input images (S40). The input multiple images are aligned in S30. Additionally, each of the multiple images input in S30 is an image of a different modality from a predetermined set of medical imaging modalities.

[0030] Machine learning capabilities thus provide a method for merging medical images from different modalities into a shared representation in the form of a fused image that integrates information from multiple images into a single view. Due to the machine learning nature of the proposed approach, the solution provided here does not require ad-hoc fusion rule design and can be applied to a wide range of modalities, which has several applications in medicine.

[0031] The method of use may include, for example, displaying the fused image (graphical representation), for example, on a computer system display. The method of use may further include a medical professional viewing the displayed fused image and, optionally, performing a medical evaluation. Here, a medical evaluation may include making a diagnosis, determining a prognosis, or determining or detecting any medical condition or parameter value related to a medical condition, such as a segmentation and / or measurement of a given body part and / or determining a medical treatment or adjustment. Additionally or alternatively, the method of use may include outputting the fused image to, for example, a computer system or processor for automated processing, such as automatically performing a medical evaluation. The fused image allows for improved medical evaluation, thanks to the fused image representing multiple registered images of the same patient, each of which is a different medical imaging modality, as a single image.

[0032] For example, the predetermined set of medical imaging modalities may include high-resolution image modalities, such as computed tomography (CT) modality (also referred to as "CT scan" modality) or magnetic resonance imaging (MRI) modality, and low-resolution image modalities (i.e., having a resolution lower than that of the high-resolution image modality), such as positron emission tomography (PET) modality, diffusion tensor imaging (DTI) modality, or ultrasound modality. The dataset obtained in 10 may include training examples, each of which includes corresponding images of a high-resolution modality (e.g., CT scan) and a corresponding image of a low-resolution modality (e.g., PET), that are registered. Optionally, for at least some of the training examples, the corresponding high-resolution (e.g., CT scan) image and the corresponding low-resolution (e.g., PET) image represent body tissue of the (same) patient, including tissue of the (same) tumor. The method also includes inputting, in S30, multiple images, including one high-resolution modality (e.g., CT scan) and one low-resolution modality (e.g., PET), that are registered and represent the same patient's anatomy, including the same tumor tissue. In this case, the function calculates, in S40, a fused image, which is highly useful in oncology because it combines, for example, a CT scan, which provides highly detailed images of the patient's anatomical features, with a PET scan, which often has lower resolution but can determine the physiological activity of the tumor. Thus, by combining both types of information in a single representation, the fused information can be used to benefit from the physiological information present in the detailed image, for example, to segment contours or measure the size of a lesion area. The proposed approach thus allows for the merging of complementary medical images from different modalities into a single image.

[0033] "Medical imaging modality" means a type of imaging technique that detects signals within a patient's body using a particular physical method to observe either anatomical structures or physiological events. Images of a particular medical imaging modality are thus generated by transferring biological, structural, and physiological properties of the patient's tissues in intensity space (typically represented by a transfer function) to reflect desired characteristics. TIFF2026021267000002.tif5170). Medical imaging modalities may vary by the physical mechanism used by the medical practitioner, the physical sensor used to capture the image, the parameters of the sensor during image acquisition, the use of contrast agents, a delay between contrast agent injection and image acquisition, or processing of the signal after image acquisition.

[0034] The predetermined set of medical imaging modalities of the present method may include modalities that involve image acquisition by different physical mechanisms, modalities that involve image acquisition by different physical sensors, modalities that involve image acquisition by the use of different contrast agents, modalities that involve different delays between injection of contrast agent and image acquisition, and / or modalities that involve different processing of signals after image acquisition.

[0035] The predetermined set of medical imaging modalities may (eg, further include) one or more (eg, any, any combination, or all) of the following modalities: i.e., autorefraction, angioscopy, bone densitometry (ultrasound), biomagnetic imaging, bone densitometry (x-ray), color perfusion Doppler, cinefluoroscopy, colposcopy, computed radiography, cystoscopy, computed tomography, duplex Doppler, digital fluoroscopy, near-infrared fluoroscopy, digital microscopy, digital subtraction angiography, digital radiography, echocardiography, electrocardiography, cardiac electrophysiology, endoscopy, fluorescein angiography, fiducials, ophthalmoscopy, general microscopy, hard copy, hemodynamic waveforms, intraoral radiography, intraocular lens data, intravascular optical coherence tomography, intravascular ultrasound, corneal curvature measurement, lens measurement, laparoscopy, laser surface scanning, magnetic resonance angiography, mammography, magnetic resonance, MR T1-weighted, MR T2-weighted, MR proton density weighted, MR steady-state free precession, MR effective T2, MR susceptibility weighted, MR Short-tau inversion recovery, MR fluid-attenuated inversion recovery, MR double inversion recovery, MR conventional diffusion weighting, MR apparent diffusion coefficient, MR diffusion tensor, MR dynamic susceptibility contrast, MR arterial blood spin contrast, MR dynamic contrast enhancement, MR blood oxygen level dependent imaging, MR time-of-flight, MR phase contrast, magnetic resonance spectroscopy, nuclear medicine, ocular axis measurement, optical coherence tomography (non-ophthalmic), ophthalmic photography, ophthalmic mapping, ocular refraction, ophthalmic tomography, ocular visual field, optical surface scanning, other, positron emission tomography (PET), panoramic x-ray, respiratory waveform, fluoroscopy, radiographic images (conventional film / screen), radiotherapy dose, radiotherapy images, radiotherapy planning, radiotherapy recording, radiotherapy structure set, segmentation, slide microscopy, volumetric relationships, single photon emission computed tomography (SPECT), automated slide stainer, thermography, ultrasound, A-mode ultrasound (US), B-mode Ultrasound, M-mode ultrasound, visual acuity, swallowing contrast test, X-ray angiography, and external camera photography.

[0036] The table below shows the standardized coding for these modalities. [Table 1] TIFF2026021267000004.tif254170TIFF2026021267000005.tif230170

[0037] The dataset may represent all modalities in a predetermined set of medical imaging modalities. Thus, for each corresponding medical imaging modality in the predetermined set, one or more training examples may include corresponding images of the corresponding modality. In addition, the dataset obtained in S10 may be such that all modalities in the predetermined set are interconnected. In other words, the graph defined below is a connected graph, where each node in the graph corresponds to a corresponding modality in the predetermined set, each modality in the predetermined set has a corresponding node, and an edge is defined to be between two nodes if and only if the nodes correspond to a pair of modalities represented in the same training example (i.e., images of the two modalities for at least one of the same patients are present in the dataset).

[0038] The dataset may include more than 100, 200, or 500 training cases (patients for which data is provided). Additionally or alternatively, the predetermined set of medical imaging modalities may include more than two modalities, such as more than 3, 5, or 10 modalities. Additionally or alternatively, the dataset may include, for each corresponding modality in the predetermined set of medical imaging modalities, more than 100 or 200 training cases including images of the corresponding modality. Additionally or alternatively, the dataset may include, for each corresponding pair of modalities in the predetermined set of medical imaging modalities, more than 100 or 200 training cases including images of each corresponding pair of modalities (i.e., connecting the two modalities of the pair in a training case).

[0039] The training in S20 is performed according to any machine learning method. The training S20 may include minimizing the loss on the dataset by varying the parameters and / or weights of the function. Such variable parameters and / or weights of the function are thus trainable parameters and / or weights of the function. This minimization may be performed by any method, for example, by stochastic gradient descent.

[0040] The function may be further configured to compute a reconstructed image for each modality of the predetermined set from the fused image, in other words, the function is configured and guided by a machine learning method such that it can compute from the input images of S30 not only the fused image in S40 but also a composite image having the format and aspects of any corresponding one of the predetermined set of modalities.

[0041] In such a case, the method of use may further include reconstructing a corresponding reconstructed image for each of one or more modalities in the predetermined set of medical imaging modalities. The one or more modalities for which the method provides a corresponding reconstructed image may be defined in any manner, and may, for example, be user-defined and / or predetermined (e.g., can be bypassed by the user, e.g., by default behavior), and / or may include, for example, one or more (e.g., all) modalities of the input provided in S30, and / or other modalities (e.g., not provided in S30).

[0042] The function calculates each such reconstructed image from the fused image. In other words, the function includes a first component configured to receive the plurality of registered images of S30 as input and to calculate and output the fused image of S40. The function also includes a second component (separate from the first component) configured to receive (at least) the fused image (and optionally no other input, or alternatively other inputs, such as one or more of the plurality of registered images provided in S30) as input and to output a reconstructed image of any modality of a predetermined set of modalities. Optionally, the second component may include corresponding subcomponents for corresponding reconstructed modalities, each subcomponent being separate from each other subcomponent (i.e., each configured to receive at least the fused image as input and output a reconstructed image of the corresponding modality). The first and second components may each include different sets of parameters and / or weights. All of the trainable parameters and / or weights of the first and second components may be varied and set within the same training S20. The subcomponents of the second component may each include different sets of trainable parameters and / or weights. All of the trainable parameters and / or weights of the subcomponents may be varied and set within the same training S20.

[0043] The first component, the second component, and / or each subcomponent of the second component may include or consist of any type of neural network, for example a corresponding convolutional neural network.

[0044] This function thus allows for the reconstruction and enrichment of images of any modality included in S30 based on information contained in the other modalities included in S30. Indeed, for reconstructed images of modalities among those present in S30, the reconstruction is based on fused images such that the reconstructed image incorporates anatomical information captured by the other modalities present in S30. This function can thus be used to enhance each individual image provided in S30.

[0045] Additionally, this function allows transforming images of modalities included in S30 into images of other modalities (not included in S30), thereby making it possible to obtain missing modalities for a given patient. The transformation is performed based on intermediate data, i.e., the fused images calculated in S40. This improves the accuracy of the transformation in the sense that the reconstructed images better represent the patient's true anatomy.

[0046] In fact, each such reconstructed image not only contains information from one modality that was present in S30 transformed into another modality (not present in S30), but also takes into account the other images that were present in S30 (by fusing all images provided in S30 and using the fused image as input for the transformation), so that the reconstructed image also contains information from the other images.

[0047] Furthermore, the proposed approach addresses the scarcity of medical imaging data, which can be a problem for performing effective machine learning. Indeed, even if a dataset does not contain examples of how to transform modality A to modality C, it does so as long as it contains examples of how to transform modality A to modality B and how to transform modality B to modality C. Using the fused images as input for the transformation is sufficient to train a function to accurately transform modality A to modality C. In addition, even if the dataset does contain examples of how to transform modality A to modality C, the function can be trained to perform such transformations in a better way, since training S20 can also benefit from the presence of examples of how to transform modality A to modality B and how to transform modality B to modality C. In other words, in situations where training data on patient data is scarce and expensive imaging techniques are not systematically applied to all patients, using fused images as input to perform modality transformation improves machine learning for modality transformation. Instead of splitting the available data into several smaller datasets and training a specific function for each, the approach proposed here allows for a single training run based on one large dataset (i.e., richer, even if some of the training examples are incomplete) relying on incomplete data.

[0048] In particular, the dataset obtained at S10 may, in some instances, include incomplete data for all of the predetermined set of medical imaging modalities for some of the plurality of patients. In other words, while all medical imaging modalities may be generally represented in the dataset when all patients are considered, not all medical imaging modalities may be represented in the dataset for at least some patients. For example, the plurality of patients may include one or more first patient groups, where the training examples for each first patient include corresponding images for each modality of a corresponding first subset of the predetermined set of medical imaging modalities, and one or more second patient groups, where the training examples for each second patient include corresponding images for each modality of a corresponding second subset of the predetermined set of medical imaging modalities. The first and second subsets may have a non-empty intersection but may be different, i.e., both include one or more medical imaging modalities in common, but at least one of the first and second subsets includes one or more medical imaging modalities that are not included in the other of the first and second subsets.

[0049] This method of use may optionally, additionally, or alternatively include outputting and / or displaying the fused image, outputting one or more reconstructed images, and / or displaying one or more reconstructed images. This method of use may, for example, include displaying (a graphical representation of) one or more reconstructed images, for example, on a display of a computer system. Optionally, this method of use may include displaying several reconstructed images (each of different modalities) on a single screen or on several screens. This method of use may include updating the displayed reconstructed images each time based on the same input provided in S30 by a user selecting several combinations of modalities at different times. Thus, a medical professional can optionally perform an evaluation based on different modalities. This method of use may further include a medical professional viewing the displayed one or more reconstructed images and, optionally, performing a medical evaluation. Additionally or alternatively, this method may include outputting one or more reconstructed images to, for example, a computer system or processor, for example, for automatic processing and execution of a medical evaluation.

[0050] Thus, the machine learning function not only provides a way to merge medical images of different modalities into a shared representation, but also allows for the transformation of images of one or more modalities into images of other modalities. In particular, the machine learning function allows for the reconstruction of an image of the original modality from a complementary image, which is useful, for example, for the transformation of missing modalities.

[0051] The proposed approach allows for the retrieval of the original image using the fused image, which uses available information from one or more images to reconstruct the missing modality, thus enabling modality transformation, which allows access to a representation of the patient's observed state in a modality that may be more practical for comparison or follow-up, without the need to perform additional imaging.

[0052] Looking at the PET / CT example mentioned above, this conversion can be used to calculate a CT from a PET scan, for example, for the purpose of attenuation correction. Attenuation is a phenomenon that reduces the detection capability of a PET scan, and a CT-based density map is used to compensate for the lost detection. In this context, PET-to-CT conversion allows the reconstruction of an approximate density map that is not used for diagnosis but is only used to correct and improve the performance of the PET scan, without the need to perform two data acquisitions.

[0053] Training S20 may include minimizing a loss that is a function of the reconstruction cost for images (e.g., some or all) present in the dataset. In other words, training S20 globally minimizes the reconstruction error generated by the function when fusing a group of aligned images of different modalities representing the same patient tissue and reconstructing each individual image of the corresponding modality from the fused images. The reconstruction cost measures the dissimilarity between the images of a given modality present in the dataset and the reconstructed image of the same given modality output by the function when multiple images from the initial training example are input. This enables unsupervised and therefore simple training. In addition, designing such a loss is simpler than designing ad hoc fusion rules and achieves high generalization ability to various types of images. The designed loss function is not task-specific and can be applied to various image fusion tasks, such as visible / infrared light, overexposed / underexposed, far / near focus, or PET / MRI. Furthermore, relying on fused images for transformation to involve available information from other modalities can reduce or prevent the occurrence of artifacts during image transformation (unlike approaches based on the presence of semantics in the dataset, which is missing here).

[0054] The registered image(s) of at least one (e.g., each) of the training examples obtained in S10 and / or the multiple registered images of the same patient input in S30 may represent the same part of the patient's body, optionally at substantially the same time, and / or the patient's body has not substantially changed anatomically or physiologically during that time period. "Substantially the same time" means that the images represent the patient's body (e.g., were acquired for that patient) at times close enough together that the patient's anatomy and physiology have not substantially changed during that time period, so that the images can be compared. For example, the images of the same patient in S10 and / or the images input in S30 may all have been acquired within a week, two days, or even one day for that patient. Absence of "substantial changes in the patient's anatomy and physiology" means that the patient has not been affected by a significant medical condition that would prevent the images from being registered. Thus, the patient's images can still be registered. For example, a tumor has not developed during the time period between the two images. This concept is known in the field of medical imaging, where techniques for registering images are known.

[0055] Depending on a given set of medical imaging modalities, each corresponding image of the dataset obtained at S10 may be 2D or each may be 3D. Similarly, each corresponding image input at S30 may be 2D or each may be 3D. If the dataset obtained at S10 includes only 2D images, each corresponding image input at S30 may be 2D. In such a case, the fused image may also be 2D. If the dataset obtained at S10 includes only 3D images, each corresponding image input at S30 may be 3D. In such a case, the fused image may also be 3D.

[0056] The fused image (which the function is trained to calculate in S20 and / or calculated in S40) can be 2D or 3D (eg, systematically), for example a 2D pixel image, or a 3D voxel image.

[0057] Additionally or alternatively, the fused image (e.g., systematically) may be a one-channel image, i.e., with unique intensity values ​​that are numbers (e.g., real or integer) that can take on values ​​between a minimum value (e.g., 0) and a maximum value (e.g., 255). The fused image may contain only one such intensity value, e.g., one for each pixel (if the fused image is a 2D pixel image) or one for each voxel (if the fused image is a 3D voxel image). Thus, the function calculates (new) intensity values ​​rather than simply concatenating the intensity values ​​of the images input at S30. The fused image may be a one-channel intensity map, and displaying the fused image may involve performing an affine mapping (e.g., identity mapping) of the intensity domain to a domain of grayscale values, and then calculating and rendering a graphical representation of the affine mapping results. In some examples, the function directly outputs a 2D or 3D map of grayscale values ​​(e.g., values ​​from 0 to 255).

[0058] Alternatively, the fused image can be displayed in color using a transformation of intensity to RGB space (a technique often used in medical imaging), which may contain several channels if the input images contain several channels, or even just one channel.

[0059] The method is computer-implemented, i.e., the steps of the method (i.e., substantially all steps) are performed by at least one computer or any similar system. Thus, the steps of the method are computer-implemented, possibly fully automated or semi-automated. In some examples, triggering of at least some of the steps of the method may be performed through user-computer interaction. The level of user-computer interaction required may depend on the expected level of automation and may be balanced against the need to achieve the user's wishes. By way of example, this level may be user-defined and / or predefined.

[0060] A typical example of a computer implementation of the method is performing the method by a system adapted for this purpose. The system may include a processor and a graphical user interface coupled to a memory on which a computer program including instructions for carrying out the method is recorded. The memory may also store a database. The memory is any hardware adapted for such storage, possibly comprising several physically distinct parts (e.g., one for the program and possibly one for the database).

[0061] FIG. 2 shows an example of a system, where the system is a client computer system, such as a user's workstation.

[0062] The exemplary client computer includes a central processing unit (CPU) 1010 connected to an internal communications bus 1000 and a random access memory (RAM) 1070 also connected to the bus. The client computer further includes a graphical processing unit (GPU) 1110 associated with a video random access memory 1100 connected to the bus. The video RAM 1100 is also known in the art as a frame buffer. A mass storage controller 1020 manages access to mass storage devices such as a hard drive 1030. Mass storage devices suitable for tangibly embodying computer program instructions and data include all forms of non-volatile memory, including, by way of example, semiconductor memory devices such as EPROM, EEPROM, and flash memory devices, magnetic disks such as internal hard disks and removable disks, and magneto-optical disks. Any of the foregoing may be supplemented by, or incorporated in, specially designed ASICs (application-specific integrated circuits). A network adapter 1050 manages access to a network 1060. The client computer may also include a haptic device 1090, such as a cursor control device, a keyboard, etc. A cursor control device is used in the client computer to allow a user to selectively position a cursor at any desired location on the display 1080. In addition, the cursor control device allows a user to select various commands and input control signals. The cursor control device includes several signal generating devices for input control signals to the system. Typically, the cursor control device may be a mouse, and the mouse buttons may be utilized to generate the signals. Alternatively or additionally, the client computer system may include a sensitive pad and / or a sensitive screen.

[0063] A computer program may include computer-executable instructions, including instructions for causing the system to perform the method. The program may be recorded on any data storage medium, including the system's memory. The program may be implemented, for example, in digital electronic circuitry, or in computer hardware, firmware, software, or a combination thereof. The program may be implemented as an apparatus, such as an article of manufacture, tangibly embodied in a machine-readable storage device for execution, for example, by a programmable processor. The method steps may be executed by a programmable processor executing a program of instructions to perform the functions of the method by manipulating input data and generating output. Thus, the processor may be programmable and configurable to receive data and instructions from a data storage system, at least one input device, and at least one output device, and to send data and instructions to the data storage system, at least one input device, and at least one output device. The application program may be implemented in a high-level procedural or object-oriented programming language, or in assembly or machine language, if desired. In any case, the language may be a compiled or interpreted language. The program may be a fully installed program or an update program. The application of the program on the system is in any case reduced to instructions for carrying out the method. The computer program may alternatively be stored and executed on a server in a cloud computing environment, the server being in communication with one or more clients over a network. In such a case, a processing unit executes the instructions contained in the program, thereby causing the method to be carried out in the cloud computing environment.

[0064] This function may be configured to receive a variable number of images as input, so that for a given patient, the function may be applied to and calculate the fused image in S40 regardless of the number of images available in S30. The function may be configured to receive a number of images equal to two as input, and may (at the same time) receive a number of images greater than two as input.

[0065] For example, the function can be configured to compute a fused image by iteratively applying a fusion network to a pair of images. The pair of images includes, in a first iteration, two images from the plurality of aligned images. In each subsequent iteration, the 173 images include one image from the plurality of aligned images and the result of applying the fusion network in the previous iteration (i.e., the output of the fusion network in the previous iteration). This allows for the use of a regular fusion network configured to receive a fixed number (2) of images as input. The function is thus independent of the size of the input images because it operates the same regardless of size, i.e., iterating pair-by-pair and fusing each pair. This facilitates training and allows for arriving at an accurate function.

[0066] The fusion network applied in each iteration can be identical (i.e., the same fusion network is applied in each iteration). In other words, the fusion network forms a single element of the function, which has its trainable parameters and / or weights, which are reused in each iteration of the iterative process. This also facilitates training. As a result of learning a shared representation from several images of different modalities using a single fusion network, there are no restrictions on the number of images or the order of fusion.

[0067] However, the function may be order-dependent for the input images, meaning that the function receives as input vectors with coordinates that each represent a corresponding image, and the function is architected so that the result is not the same depending on the order of the vector coordinates. In such a case, training S20 may include one or more applications of the function, each with a corresponding input having a randomized order. In other words, the order of the images in each training example is randomized for training. This trains the initially (architecturally) order-dependent function to become order-independent for the input images. Thus, once trained, the function does not substantially change depending on the order of the input images.

[0068] The function may be further configured to receive as input for a corresponding input image a corresponding label representing the modality of the corresponding input image, which facilitates training by helping the function to learn by focusing on information other than the modality of a given image.

[0069] For example, in subsequent iterations of application of the fusion network described above, if training S20 includes one or more applications of the function (during loss minimization), each with a corresponding input including a corresponding fused image, the label values ​​for the fused images may be specific and distinct from the label values ​​representing the corresponding modalities of a predetermined set of medical imaging modalities. This helps the function recognize that the input image is a fused image. Optionally, a unique label may be utilized during training S20 to identify all fused images. The unique label thus represents the fusion properties of the input image such that the resulting input image is not distinct from the modality or properties of the original images from which the fusion resulted. This improves training.

[0070] For each patient of the plurality of patients (in the dataset obtained in S120), and for at least one modality of a given set, the corresponding image in the dataset (for said at least one modality) is a captured (i.e., acquired) image of the patient. In other words, the corresponding image is a physical / real image (as opposed to a synthetic image) obtained from a real acquisition of the inside of the patient's body.

[0071] Optionally, all images in the dataset are captured images.

[0072] Alternatively, the dataset may include composite images generated from such captured images, such as in the case of patients for whom captured images are initially missing. Optionally, in such an alternative, a composite image may be generated for each missing modality, or alternatively for only some modalities, thus leaving out the missing modalities.

[0073] During training, the method may include applying a random mask to one or more images of the dataset, e.g., to each image of the dataset, before the corresponding image (to which the random mask is applied) is input to the function (e.g., to a fusion network). Applying a random mask to a given image sets values ​​in the image (e.g., intensity channels in pixels or voxels of the image) to zero, or null, for randomly selected portions of the image (e.g., of a predetermined fixed shape and size, but at random locations). Such a random mask helps the function to be robust to missing data during this method of use. Also, if the dataset includes synthetic data, this helps prevent training S20 from reproducing the mapping used to generate the synthetic data.

[0074] An example method will now be described with reference to Figures 4-12.

[0075] This function may constitute an unsupervised deep learning model for merging images of different modalities into a shared image representation, from which the original image can be reconstructed. This function may be modality agnostic in its design, allowing for variation in the modalities considered and the order of fusion. Various examples of this method described below may be developed on top of a base architecture and may differ in the options implemented for losses, training means, and some architectural choices.

[0076] The base architecture can consist of a unique fusion network that iteratively merges images to build a shared representation, and a reconstruction network that captures the subset of fused images from the desired modalities. This architecture allows for fusing any number of images into a shared representation without losing essential information from each modality.

[0077] This model (i.e., machine-learned function) can utilize registered multimodal data without any additional annotations or ground truth. If the data is not registered, this can be done during preprocessing using existing methods, such as mutual information registration or DRMIME optimization, as described below.

[0078] In specific examples, the proposed method provides a way to merge medical images of different modalities in a shared latent representation and to transform images from one or more modalities to other modalities. The advantages of the solution in such specific examples are as follows:

[0079] The shared representation is learned from several images of different modalities using a single fusion network, so there are no constraints on the number of images or the order of fusion.

[0080] Unlike classical image fusion techniques, the proposed solution does not require ad-hoc fusion rule design and is easily generalizable to a wide range of modalities.

[0081] Unlike image fusion methods that only focus on the quality of the fused image, the proposed solution allows for the reconstruction of the image in the original modality, which is convenient for transferring the complementary image to the missing modality.

[0082] Unlike image transformation methods that focus solely on style transfer or direct transformation of images, the proposed solution provides a latent representation that aggregates information from multiple images into a unique view.

[0083] Unlike the CycleGAN model described in the paper "Unpaired Image-to-Image Translation using Cycle-Consistent Adversarial Networks" (arXiv:1703.10593) by Zhu, J.-Y., Park, T., Isola, P., & Efros, A.A. et al. (2020), which is incorporated herein by reference, where artifacts can occur in image translation due to a lack of semantics in the data, the proposed solution relies on fused images for translation in order to rely on available information from other modalities.

[0084] Data Preparation To learn a shared latent representation of multimodal medical images, the machine learning method may acquire a dataset of 2D or 3D medical images of different modalities at S10. The images may be grouped by patient and registered so that anatomical regions seen in different images match when the images are overlaid.

[0085] If paired data is not available, other techniques such as CycleGAN may optionally be used to generate synthetic data as a means of training and testing the method.

[0086] If paired images corresponding to the same patient are available but are not spatially aligned (i.e., registered), several techniques can be used to reliably register the images, such as mutual information registration (as described in the paper by Xu, R., Chen, Y.-W., Tang, S.-Y., Morikawa, S., & Kurumi, Y., "Parzen-window based normalized mutual information for medical image registration. IEICE Transactions on Information and Systems," 91(1), 132-144. (2008), incorporated herein by reference) or DRMIME optimization (as described in the paper by Nan, A., Tennant, M., Rubin, U., & Ray, N., "DRMIME: Differentiable Mutual Information and Matrix Exponential for Multi-Resolution Image Registration," arXiv:2001.09865 (2020), incorporated herein by reference).

[0087] Neural Network Architecture The proposed approach relies on a core neural network model that can be extended with several options to address the specific problems faced when trying to learn shared representations and modality transformations.

[0088] This core architecture is weighted A unique neural network parameterized by TIFF2026021267000006.tif5170 TIFF2026021267000007.tif5170. This network can be used to Two images in TIFF2026021267000008.tif5170 It takes the stack of TIFF2026021267000009.tif5170 as input and creates a latent shared representation The output of this fusion network is TIFF2026021267000010.tif5170. Share the image of TIFF2026021267000011.tif5170 It can be used repeatedly to fuse to TIFF2026021267000012.tif5170.

[0089] To improve the robustness of the model, this machine learning method optimizes the fusion network so that it is invariant to the fusion order. This may involve training the input image using, for example, the following formula: This can be done by randomly sorting the images according to TIFF2026021267000014.tif6170, where TIFF2026021267000015.tif5170 is This is a random sort of TIFF2026021267000016.tif5170.

[0090] This notation will be omitted in the remainder of the discussion, but the machine learning method can be considered to apply random permutations during training (unless otherwise specified).

[0091] The next proposed approach takes a shared latent representation as input. Receive TIFF2026021267000017.tif5170, weight Multiple reconstructed networks with TIFF2026021267000018.tif5170 Add TIFF2026021267000019.tif5170 and the desired TIFF2026021267000020.tif5170th modality reconstructed image Exports TIFF2026021267000021.tif6170.

[0092] A mathematically equivalent, but potentially advantageous, option from an implementation point of view is to add the modality channel as an input to the same network. There is a way to get all modalities using TIFF2026021267000022.tif5170. This improves the feature extraction part of the network, which can be beneficial in some cases. The previous equation is exactly equivalent to TIFF2026021267000023.tif5170, and this notation will be omitted in the following discussion.

[0093] The machine learning method may include training a model with a cycle consistency loss to produce a reconstructed image that matches the original image. Let TIFF2026021267000024.tif5170 be the cost function to be minimized (e.g., L1 or L2 cost function or perceptual loss), TIFF2026021267000025.tif5170 can be written as follows, where TIFF2026021267000025.tif5170 is the number of patients in the dataset. TIFF2026021267000026.tif14170

[0094] If the dataset does not include all modalities for all patients, TIFF2026021267000027.tif5170th modality TIFF2026021267000028.tif5If exists for the 170th patient TIFF2026021267000029.tif5170, otherwise Indicator function such that TIFF2026021267000030.tif5170 TIFF2026021267000031.tif5170 can be introduced. The loss function is as follows: TIFF2026021267000032.tif14170

[0095] Figure 4 shows an example of the architecture of this model when three images of different modalities are input.

[0096] Then, at inference time, this model Shared representation using TIFF2026021267000033.tif6170 TIFF2026021267000034.tif5170 is trained and acquired using a network with missing modalities. TIFF2026021267000035.tif6170 can be inferred. TIFF2026021267000036.tif6170TIFF2026021267000037.tif6170

[0097] This reconstructed image benefits from the aggregate information of all existing imaging modalities and can benefit greatly from the presence of multiple observations. For example, this model can be used to reconstruct a CT scan using MRI observations for structural information and PET scan observations for tumor physiology. The reconstructed CT scan can then be used for further visualization, segmentation, and comparison by medical professionals.

[0098] Model Options Improving on this core architecture, the following proposes several options for loss design, model architecture, training strategy, and data management.

[0099] Data Management Data acquisition is an issue when processing medical images, and in this framework, two major problems are faced: having unregistered paired data and having no paired data at all. As mentioned earlier, machine learning methods can involve initially registering unregistered data by other means and then processing that data using the same pipeline.

[0100] For unpaired data, base model choices are described here.

[0101] To deal with unpaired image data, i.e., when the dataset does not contain images corresponding to the same patient, or when this information is unavailable, the machine learning method may involve using standard image transformation methods such as CycleGAN to generate synthetic data for the purposes of training the network.

[0102] An efficient training scheme to prevent the model from learning identical mappings to CycleGAN and force it to extract more meaningful information from the shared representation can be to utilize random masking. In this way, information missing from each of the real or synthetic modalities can be inferred from the shared representation and retrieved by the reconstruction network.

[0103] Figure 5 illustrates this option, i.e., the model architecture with CycleGAN for synthetic data generation and random masking. The random masking option can be maintained for training even in the case of paired data to make the model more robust for inferring missing information and missing modalities.

[0104] Loss Design and Training Strategies The reconstructed image can be oriented to match the original image during training by a cycle consistency loss, which may be imperfect at regularizing the neural network but may produce improved generalization results.

[0105] This loss may further include an adversarial loss to explain it. Adversarial training can be used to improve the realism of the reconstructed image, as its input probabilities are based on the modality. Weights to predict whether TIFF2026021267000038.tif5170 is the true image Discrimination network parameterized by TIFF2026021267000039.tif5170 This is done with the introduction of TIFF2026021267000040.tif6170.

[0106] Figure 6 illustrates this learning strategy, i.e., a learning strategy with adversarial training.

[0107] The cost function for this model is The file name will be TIFF2026021267000041.tif15170.

[0108] Referring to Figure 7, another way of thinking about loss design is that the shared representation preferably does not change when fusing in images that have already been seen. To this end, a variation of the following model shown in the figure can be implemented: In TIFF2026021267000042.tif5170 (Fig. 4 Each image has already been fused (TIFF2026021267000043.tif5170) Regarding TIFF2026021267000044.tif5170, Merge with TIFF2026021267000045.tif5170 and the result TIFF2026021267000046.tif6170 Compare with TIFF2026021267000047.tif5170.

[0109] The loss function is The file name will be TIFF2026021267000048.tif14170.

[0110] In other words, the loss is a stability loss. Summation across images in the dataset, TIFF2026021267000049.tif6170 TIFF2026021267000050.tif8170. The stability loss is calculated for each corresponding image of each corresponding patient by applying the function as input to all images of the corresponding patient included in the dataset: (i) a first fusion image TIFF2026021267000051.tif5170, and (ii) a second fused image calculated by applying the function as input to the corresponding image and the first fused image. It is expressed by the cost between TIFF2026021267000052.tif5170.

[0111] A further option is to help the fusion network identify important information by using labels that represent the type of modality used. Function to take TIFF2026021267000053.tif6170 TIFF2026021267000054.tif5170. From the network's perspective, this label is added as a new channel that uniformly contains the given label. Images that have already been fused are labeled -1 to distinguish them from the original image.

[0112] Figure 8 shows the architecture of such a labeled model for the fusion network.

[0113] implementation Fusion and reconstruction networks have been tested with the U-Net architecture (described in the paper "U-Net: Convolutional Networks for Biomedical Image Segmentation" by Ronneberger, O., Fischer, P., & Brox, T., arXiv:1505.04597 (2015), incorporated herein by reference) and the DenseNet architecture ("U2Fusion: A Unified Unsupervised Image Fusion Network" by Xu, H., Ma, J., Jiang, J., Guo, X., & Ling, H., IEEE TRANSACTIONS ON PATTERN ANALYSIS AND MACHINE INTELLIGENCE, n.d., incorporated herein by reference). The U-Net architecture is a type of convolutional network developed to process medical image data, capturing high-level structure in the encoding part and finer details at each level with skip connections.

[0114] The U-Net architecture was found to outperform DenseNet for this model.

[0115] Applicable The model was trained and tested on two datasets representing different situations, which are briefly described below and the results presented.

[0116] Synthetic dataset The first tested dataset is a non-medical synthetic dataset containing 10,000 2D images of spheres. The images were classified into three modalities according to the Phong lighting model: ambient lighting, diffuse lighting, and specular lighting. 8,000 images were used for training, and 1,000 for validation and training.

[0117] Figure 9 shows the results obtained from four examples, where the three images on the left 92 are the original images, the image in the middle 94 is the fused image, and the three images on the right 96 are the acquired / reconstructed images.

[0118] Figure 10 shows the highlight of the multimodal transformation for the same example, where the specular information was masked in the original image but could be obtained using the ambient and diffuse components. The model efficiently performs a fusion of the available data and accurately transforms it into the specular component.

[0119] Medical Datasets The second tested dataset is a medical dataset containing 1250 slices of brain MRI in both T1-weighted and T2-weighted modalities, which were divided into 900 images for training, 100 images for validation, and 250 images for testing.

[0120] Figure 11 shows the results obtained for four examples, where the two left images 112 are the original images, the middle image 114 is the fused image, and the two right images 116 are the acquired / reconstructed images, illustrating the entire process of fusion and reconstruction of the original modalities performed during training.

[0121] Figure 12 shows how fusion can aggregate information from both original images 121; the tumor is clearly seen in the T2 image and detail is maintained in the fused image 123 (circle 122), while cortical detail (circle 124) is clearly derived from the T1 image.

Claims

1. 1. A computer-implemented machine learning method of a function configured to receive as input a plurality of registered images of a same patient and each of different modalities within a predetermined set of medical imaging modalities and to calculate a fused image, the machine learning method comprising: obtaining a dataset including a respective image for each patient of a plurality of patients and for each modality of at least some of each of the predetermined sets, the respective images of the patient are registered; Obtaining a dataset and training the function based on the dataset; machine learning methods, including

2. The function is The fusion network calculates configured to iteratively apply to a pair of images; The pair of images is, in a first iteration, two images of a plurality of aligned images. Equipped with the pair of images comprises, in each subsequent iteration, one image of the plurality of aligned images and a result of applying the fusion network in a previous iteration; The machine learning method of claim 1 .

3. The fusion network is identical in each iteration. The machine learning method of claim 2 .

4. The function calculates from the fused images, the respective modalities of the predetermined set. Regarding the reconstructed image further configured to calculate The machine learning method according to any one of claims 1 to 3.

5. Training costs for the images of the data set minimizing losses, including The machine learning method according to claim 4 .

6. the function is configured to receive a variable number of images as input, including two; The loss is a stability loss. for the images of the data set, Furthermore, The stability loss is calculated by applying the function to all images of the corresponding patient contained in the dataset as input to a first fusion image. and (ii) a second fusion image calculated by applying the function to the corresponding image and the first fusion image as input. for each corresponding image of each corresponding patient, expressed by the cost between The machine learning method according to claim 5 .

7. The loss is a hostile loss further comprising: The machine learning method according to claim 5 or 6.

8. the function is order dependent for a plurality of input images; training includes applying the function one or more times to corresponding inputs having a randomized order; The machine learning method according to any one of claims 1 to 7.

9. the function is further configured to receive as input, for a corresponding input image, a corresponding label indicative of the modality of the corresponding input image; Optionally, said training comprises: applying the function one or more times to corresponding inputs including corresponding fused images and corresponding labels representing fusion characteristics of the corresponding fused images. The machine learning method according to any one of claims 1 to 8.

10. The predetermined set of medical imaging modalities may include autorefraction, angioscopy, bone densitometry (US), biomagnetic imaging, bone densitometry (X-ray), color flow Doppler, cinefluoroscopy, colposcopy, computed radiography, cystoscopy, computed tomography, duplex Doppler, digital fluoroscopy, fluoroscopy, digital microscopy, digital subtraction angiography, digital radiography, echocardiography, electrocardiography, cardiac electrophysiology, endoscopy, fluorescein angiography, fiducial, ophthalmoscopy, general microscopy, hard copy, hemodynamic waveform, intraoral radiography, intraocular lens data, intravascular optical coherence tomography, intravascular ultrasound, keratometry, lensmetry, laparoscopy, laser surface scanning, magnetic resonance angiography, mammography, magnetic resonance imaging, MR T1 weighted imaging, MR imaging, T2-weighted imaging, MR proton density-weighted imaging, MR steady-state free precession, MR effective T2, MR susceptibility-weighted imaging, MR short-tau inversion recovery imaging, MR fluid-attenuated inversion recovery imaging, MR double inversion recovery imaging, MR conventional diffusion-weighted imaging, MR apparent diffusion coefficient imaging, MR diffusion tensor imaging, MR dynamic susceptibility contrast imaging, MR arterial spin contrast imaging, MR dynamic contrast-enhanced imaging, MR blood oxygen level-dependent imaging, MR time-of-flight imaging, MR phase contrast imaging, magnetic resonance spectroscopy, nuclear medicine, axial measurement, optical coherence tomography (non-ophthalmology), ophthalmic photography, ophthalmic mapping, ophthalmic refraction, ophthalmic tomography layer imaging, ophthalmologic visual field, optical surface scan, other, positron emission tomography (PET), panoramic x-ray, respiratory waveform, fluoroscopy, radiographic images (traditional film / screen), radiotherapy dose, radiotherapy images, radiotherapy planning, radiotherapy records, radiotherapy structure sets, segmentation, slide microscopy, stereometric relationships, single photon emission computed tomography (SPECT), automated slide stainers, thermography, ultrasound, A-mode US, B-mode US, M-mode US, visual acuity, video fluoroscopy, x-ray angiography, external camera imaging, The machine learning method according to any one of claims 1 to 9.

11. A method for using a function machine-learned by the machine learning method according to any one of claims 1 to 10, the method comprising: inputting a plurality of registered images of the same patient, each of which is a different modality within a predetermined set of medical imaging modalities, into said function; using said input to calculate a fused image according to said function; Including, How to use.

12. outputting and / or displaying the fused image; and / or for each of one or more modalities in the predetermined set of medical imaging modalities, optionally including the modality of the input, reconstructing a corresponding reconstructed image, and optionally outputting the one or more reconstructed images and / or displaying the one or more reconstructed images; further comprising:

12. The use of claim 11.

13. A computer program comprising code instructions adapted to cause a processor to carry out the machine learning method according to any one of claims 1 to 10 and / or the method of use according to claim 11 or 12, and / or A function machine-learned by the machine learning method according to any one of claims 1 to 10. data structures, including

14. A data storage medium storing the data structure of claim 13.

15. 14. A computer system comprising a processor coupled to a memory storing the data structure of claim 13.