Modality matching of medical images
The method and device use a convolutional neural network to transfer features between medical images of different modalities, addressing inefficiencies in diagnosis by maintaining imaging style consistency and enhancing diagnostic accuracy.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-10-03
- Publication Date
- 2026-04-09
AI Technical Summary
Medical images acquired from different imaging modalities, such as CT and X-ray, often differ in style due to variations in image processing techniques and hardware, requiring surgeons to adapt and leading to inefficiencies in diagnosis.
A method and device using a convolutional neural network to process and transfer features between medical images of different modalities, iteratively computing a loss metric to augment the reference image, preserving anatomical structures.
Enables real-time modality matching of medical images, maintaining consistency in imaging styles and enhancing diagnostic efficiency and accuracy by preserving essential features.
Smart Images

Figure IN2025051608_09042026_PF_FP_ABST
Abstract
Description
MODALITY MATCHING OF MEDICAL IMAGES TECHNICAL FIELD
[0001] The present disclosure generally relates to the field of image modality matching. More particularly, but not exclusively, the present disclosure relates to method and device for modality matching of medical images.BACKGROUND
[0002] The following description includes information that may be useful in understanding the present invention. It is not an admission that any of the information provided herein is prior art or relevant to the presently claimed invention, or that any publication specifically or implicitly referenced is prior art.
[0003] Typically, image-guided surgery (IGS) enables a surgeon to perform a surgical procedure using medical images of anatomical structures of a user or subject (e.g., a patient). The medical images may be acquired from devices of different makes (i.e., same model from different manufacturers or different models from same manufacturer). The devices may produce the images of anatomical structures using different imaging modalities. For example, each medical image may correspond to at least one of imaging modality such as ultrasound, magnetic resonance imaging (MRI), computed tomography (CT), X-ray, fluoroscopy, positron emission tomography (PET), etc.
[0004] When the medical images of same anatomical structure are acquired from two different devices (e.g., CT and X-ray devices) each belonging to a different model or different manufacturer, the images (e.g., CT or X-ray images) may appear distinct due to variations in image processing techniques or hardware used. As a result, the images may differ in imaging style (i.e., imaging modality), and the surgeon must adapt to a range of imaging styles when interpreting the medical images from different imaging modalities.
[0005] Particularly, if the surgeon is more accustomed to a specific imaging style based on their training and experience and if he is presented with the images that differ from their usual imaging style, then it may be challenging for them to analyze the images from unfamiliar imaging modalities and may require additional time and effort potentially impacting diagnostic efficiency and accuracy. The extra effort needed to analyze the images of different modalities may hinder the surgeon’s ability to provide a timely and accurate diagnosis.
[0006] Therefore, it is necessary to transform the medical images from one modality to another to maintain consistency in the imaging styles. Further, there is a need for techniques thatperform image modality matching while preserving essential features of the anatomical structures in the medical images.
[0007] Thus, there is a desire for methods and devices for image modality matching that overcome the above limitations.SUMMARY
[0008] This summary is provided to introduce a selection of concepts, in a simplified format, which are further described in detailed description of the present disclosure. This summary is neither intended to identify key or essential inventive concepts of the disclosure nor is it intended to determine the scope of the disclosure.
[0009] The present disclosure relates to modality matching of images in the field of image- guided surgery. More particularly, the present disclosure relates to method and device for modality matching of medical images. A method for performing modality matching of medical images may comprise: (1) obtaining a first medical image having a first image modality from a first imaging module; and (2) obtaining a second medical image having a second image modality from a second imaging module. The method further comprises: (3) processing the first and second medical images to extract first and second sets of features respectively using a convolutional neural network; (4) transferring the first and second sets of features onto a reference medical image; (5) computing a loss metric based on the extracted first and second sets of features; and (6) augmenting the reference medical image by iteratively performing steps (3) to (5) until the loss metric is above a threshold loss metric. The method includes (7) outputting the augmented reference medical image having features from both the first and second medical images that belong to different image modalities.
[0010] The present disclosure relates to the device for performing modality matching of medical images. The device may comprise: a processor; and a memory communicatively coupled with the processor. The processor is configured to: (1) obtain a first medical image having a first image modality from a first imaging module; and (2) obtain a second medical image having a second image modality from a second imaging module. The processor is further configured to: (3) process the first and second medical images to extract first and second sets of features respectively using a convolutional neural network; (4) transfer the first and second sets of features onto a reference medical image; (5) compute a loss metric based on the extracted first and second sets of features; and (6) augment the reference medical image by iteratively performing steps (3) to (5) until the loss metric is above a threshold loss metric. The processor is further configured to: (7) output the augmented reference medical image havingfeatures from both the first and second medical images that belong to different image modalities. The present disclosure may perform the modality matching of medical images with preserving inherent anatomical structures in real time.
[0011] The foregoing summary is illustrative only and is not intended to be in any way limiting. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features will become apparent by reference to the drawings and the following detailed description.BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The accompanying drawings, which are incorporated in and constitute a part of this disclosure, illustrate exemplary embodiments and together with the description, serve to explain the disclosed principles. Some embodiments of device and / or methods in accordance with embodiments of the present subject matter are now described below, by way of example only, and with reference to the accompanying drawings.
[0013] FIGs. 1A-1B illustrate exemplary representations of imaging devices, in accordance with an example embodiment.
[0014] FIG. 2 illustrates a block diagram representation including a modality matching device, in accordance with an embodiment of the present disclosure.
[0015] FIG. 3A illustrates a schematic workflow performed by the modality matching device, in accordance with an embodiment of the present disclosure.
[0016] FIG. 3B illustrates an exemplary example performed by the modality matching device, in accordance with the present disclosure.
[0017] FIG. 4 illustrates a schematic workflow of a method performed by the modality matching device, in accordance with an embodiment of the present disclosure.
[0018] It should be appreciated by those skilled in the art that any block diagrams herein represent conceptual views of illustrative systems embodying the principles of the present subject matter. Similarly, it will be appreciated that any flowcharts, flow diagrams, state transition diagrams, pseudo code, and the like represent various processes which may be substantially represented in computer readable medium and executed by a computer or processor, whether or not such computer or processor is explicitly shown.DETAILED DESCRIPTION
[0019] Exemplary embodiments are described with reference to the accompanying drawings. Wherever convenient, the same reference numbers are used throughout the drawings to referto the same or like parts. While examples and features of disclosed principles are described herein, modifications, adaptations, and other implementations are possible without departing from the spirit and scope of the disclosed embodiments. It is intended that the following detailed description be considered as exemplary only, with the true scope and spirit being indicated by the following claims. Additional illustrative embodiments are listed below.
[0020] In the present document, the word “exemplary” is used herein to mean “serving as an example, instance, or illustration.” Any embodiment or implementation of the present subject matter described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.
[0021] As used herein, the term “comprising” is not intended to be limiting, but may be a transitional term synonymous with “including,” “containing,” or “characterized by.” The term “comprising” may thereby be inclusive or open-ended and does not exclude additional, unrecited elements or method steps when used in a claim. For instance, in describing a method, “comprising” indicates that the claim is open-ended and allows for additional steps. In describing a device, “comprising” may mean that a named element(s) may be essential for an embodiment or aspect, but other elements may be added and still form a construct within the scope of a claim. In contrast, the transitional phrase “consisting of’ excludes any element, step, or ingredient not specified in a claim. This is consistent with the use of the term throughout the specification.
[0022] Typically, image-guided surgery (IGS) enables a surgeon to perform a surgical procedure using images of anatomical structures of a user or subject (e.g., a patient). The medical images may be acquired from devices of different makes (i.e., same model from different manufacturers or different models from same manufacturer). The devices may produce the images of anatomical structures using different imaging modalities. For example, each image may be a medical image and may correspond to at least one of imaging modalities such as ultrasound, magnetic resonance imaging (MRI), computed tomography (CT), X-ray, fluoroscopy, positron emission tomography (PET), etc. When the medical images of same anatomical structure are acquired from two different devices (e.g., CT and X-ray devices) each belonging to a different model or different manufacturer, the images (e.g., CT or X-ray images) may appear distinct due to variations in image processing techniques or hardware used. As a result, the images may differ in imaging style (i.e., imaging modality), and the surgeon must adapt to a range of imaging styles when interpreting the medical images from different imaging modalities.
[0023] The images may be obtained from one or more imaging devices: (i) prior to the surgery (termed as “pre-operative images”); (ii) during the surgery (termed as “intra-operative images”); or (iii) after the surgery (termed as “post-operative images”).
[0024] FIGs. 1A-1B illustrate exemplary representations of the imaging devices 100A-100B, in accordance with an example embodiment. The intra-operative images may be captured using an intra-operative imaging device 100A (e.g., X-ray fluoroscopy, a C-arm imaging device) as shown in Fig. 1A. The pre-operative images may be captured using a pre-operative imaging device 100B (e.g., a CT scanner) as shown in Fig. IB. The image associated with at least one imaging modality may be captured in a two-dimensional (2D) view, a three-dimensional (3D) view, or other formats.
[0025] When one or more images having different modalities are to be analyzed simultaneously (i.e., during the surgery), it may be challenging for the surgeon to analyze the images from various imaging modalities that are unfamiliar to them. For instance, if the surgeon is more accustomed to a specific imaging style based on their training and experience and if he is presented with the images that differ from his usual imaging style, then it may require additional time and effort potentially impacting diagnostic efficiency and accuracy. The extra effort needed to analyze the images from different modalities may hinder the surgeon's ability to provide a timely and accurate diagnosis. Therefore, it is necessary to transform the medical images from one modality to another to maintain consistency in the imaging styles. Further, there is a need for techniques that perform image modality matching while preserving the essential features of the anatomical structures in the medical images.
[0026] Existing techniques for transforming the images from one modality to another modality often require processing hundreds of data samples, exhibit a quite complicated and time consuming process, and provide inefficient results. Therefore, there is a need for efficient techniques that perform image modality transformation with preserving the essential features of the anatomical structures in the images and require a less time consuming process.
[0027] The present disclosure relates to modality matching of images in the field of image- guided surgery. More particularly, but not exclusively, the present disclosure relates to method and device for modality matching of medical images. The present disclosure may perform the modality matching of medical images while preserving inherent anatomical structures of the subject in real time.
[0028] FIG. 2 illustrates a block diagram representation 200 that includes a modality matching device 208, in accordance with an embodiment of the present disclosure. In an embodiment,the block diagram representation 200 may include a first imaging device 202, a second imaging device 204, a user device 206, the modality matching device 208, etc.
[0029] The first imaging device 202 may capture a first medical image 203 that belong to a first image modality. The first imaging device 202 may correspond to the intra-operative imaging device (e.g., X-ray device). The first medical image 203 may be referred to as a first medical image 203 or an intra-operative image 203. The first imaging device 202 (e.g., C-Arm fluoroscopy device) may capture the first medical image 203 (e.g., X-ray image) in the first image modality (2-D view) during the surgery.
[0030] The second imaging device 204 may capture a second medical image 205 that belong to a second image modality. The second imaging device 204 may correspond to the preoperative imaging device (e.g., CT scan). The second medical image 205 may be referred to as a second image 205 or a pre-operative image 205. The second imaging device 204 (e.g., CT scan device) may capture the second medical image 205 (i.e., CT image) in the second image modality (3-D view) prior to the surgery.
[0031] The user device 206 may be an input device or a display device implemented by a computing device. The computing device may correspond to at least one of a desktop, laptop, tablet computer, smart phone, or any other data processing device or combinations thereof. The user device 206 may include the input device configured to enable the user to receive the medical images of various types and styles. The input device may comprise a touchscreen, a keyboard, a mouse, a trackpad, a motion sensing camera, or other input device.
[0032] In some embodiments, the user device 206 may include the display device configured to enable the user to view and / or display the medical images of various types and styles. The display device may comprise a computer monitor, touchscreen, projector, or other display device. The display device may be configured to view reconstructed medical images acquired by imaging modules, and / or interact with various data stored in the computing device. The user device 206 may provide an enhanced visualization of the images to the surgeon for performing the surgical procedure.
[0033] The modality matching device 208 may be configured to analyse the one or more images that belong to different imaging modalities. In some embodiments, the modality matching device 208 may comprise at least one of a processor 210, a memory 212, a network interface 214, or various other components that interact with the user device 206 and other devices.
[0034] The processor 210 in conjunction with the memory 212 may perform the imagemodality matching techniques and various other functions. The processor 210 may include one or more general-purpose processors and / or one ormore special-purpose processors (e.g., digital signal processors On Chip (SOC) Field Programmable Gate Array (FPGA) processor, etc.). The processor 210 may include at least one or more processors, a suitable logic, circuitry, and / or interfaces, configured to decode and execute any instructions received from at least one or more other electronic devices or server(s). The processor 210 may be configured to execute one or more computer-readable program instructions such as program instructions to carry out any of the functions described in the description. The memory 212 may be communicatively coupled with the processor 210 via the network interface 214. The user device 206 may be configured as a peripheral input or output device. The user device 206 may be integrated with the processor 210, the memory 212, and / or the other modules and enabled to interact with the modality matching device 208.
[0035] In some embodiments, the modality matching device 208 may comprise a first imaging module 216, a second imaging module 218, a Digitally Reconstructed Radiograph (DRR) module 220, a convolutional neural network (CNN) module 224, a loss metric module 226, a regularization module 228, etc. The modality matching device 208 may be implemented using various other modules / units, entities, and provided as a component of a larger system, such as a medical imaging device, health care imaging device, diagnostic medical imaging equipment, and / or in various other forms.
[0036] The first imaging module 216 may receive the first medical image 203 that belong to the first imaging modality from the first imaging device 202. The first medical image 203 may be referred to as a content image. The content image is the image whose structure (e.g , shape)
[0037] The second imaging module 218 may receive the second medical image 205 that belong to the second imaging modality from the second imaging device 204. The second medical image 205 may be referred to as a style image. The style image is the image whose artistic
[0038] The DRR module 220 may involve creating a digital reconstruction of at least one image (e.g., the first image 203 or the second image 205) in one projection view (e.g., 3D volume) to another (e.g., 2D volume). For example, the DRR module 220 may receive the second medical image 205 in second imaging modality from the second imaging device 204 via the second imaging module 218. The DRR module 220 involves projecting the second image 205 (e.g., 3D image) along a defined path (e.g., using a ray casting method) forcalculating intensity values of the second image. The DRR module 220 creates the digital reconstruction of the second medical image 205 and mimics a real radiograph from a specific viewing angle. Thus, the DRR module 220 may provide the projection view of 3D to the projection view of 2D based on the reconstructed medical image. The digital reconstruction of the second medical image 205 in a 3D volume may be emulated to the second medical image 205 in a 2D volume. Therefore, the DRR module 220 may enable comparing and analysing the one or more images when each image belongs to a different projection view. The DRR module 220 may include a computational tool (e.g., Blast Match) for reconstructing the 2-D volume of the style image 205. The DRR module 220 may include comparing multi-modal images for the projection view of one image is made to look like the projection view of another image with involving less computational processing. Since the first medical image 203 and the second medical image 205 are often too dissimilar to be subjected to any kind of modality transfer, the DRR module 220 may serve as a synthetic intermediate for reconstructing the images that having different view of projection.
[0039] The CNN module 224 may include one or more deep neural networks comprising a plurality of weights and biases, activation functions, loss functions, instructions, but not limited thereto. The CNN module 224 may be referred to as convolutional neural network 224 or CNN 224. The CNN module 224 may implement the one or more deep neural networks to map the medical images of a particular style to a target style. The CNN module 224 may store instructions for implementing one or more neural networks, each may be a trained and / or untrained neural network. In some embodiments, the CNN module 224 may include a CNN architecture that comprise various model architectures such as a Visual Geometry Group (VGG-19) model, VGG-16 model, AlexNet model, Inception v3 model, ResNet50, etc.
[0040] In an example embodiment, the CNN architecture that relates to the VGG-19 model may include 19 weighted layers and among the 19 layers, 1-16 layers may be configured to extract a distinctive set of features of one or more imaging modalities and 17-19 layers may be configured to classify the distinctive set of features. The set of features are distinctive and includes image attributes such as edges, shapes, texture, intensities, colors, etc. that are extracted from the images. The processor 210 may be configured with the CNN module 224 to extract the set of features of one or more images of various imaging modalities.
[0041] In an embodiment, the modality matching device 208 may include a loss metric module 226 for computing loss metric involved during the process of image modality matching. The loss metric module 226 may compute that one or more loss metric such as the content loss, thestyle loss, the total variation loss, and the total loss. The loss metric module 226 may provide a weighted combination of the one or more losses. The total variation loss may be computed based on combining the style loss and the content loss. The loss metric module 226 may compute the one or more losses for assessing to reduce drastic variation (e.g., due to noise) between the pixels of the image. The loss metric may be determined to minimize the variation between neighboring pixels of the images. For example, the loss metric module may compare the neighboring pixels of the images, monitor if there is any drastic variation in the intensities, and smoothen the image only if there is drastic variation in the intensities. Thus, the loss metric module 226 may enable reducing the pixel variation in the images.
[0042] In an embodiment, the modality matching device 208 may include a regularization module 228 for smoothening the images. The regularization module 228 may compute regularization weight (w) that is responsible for how smooth the image must appear. In general, the regularization weight may be chosen lager for the style loss and lesser for the content loss. For example, a value of the regularization weight (w) may be equivalent to 106 for the style loss and 1 or 10 for the content loss.
[0043] An increase in the regularization weight may improve in smoothing the images and a decrease in the regularization weight may enable extracting more detailed features but introduce artifacts. For example, if the value of w is 0. 1, the image becomes very smooth (i.e., smoother to the point where all the features are lost) and if the value ofw is 0.0001, the image may able to preserve all the features with producing a less noise. Thus, the value of the regularization weight may be achieved based on trial-and-error approach and optimally chosen for retaining all the features of the images. The regularization weight may be used in computation of the total loss that enables smoothing the image and decide how much smoothing must happen.
[0044] The modality matching device 208 may enable transferring image attributes from one image (e.g., style image) to another image (e.g., content image) to perform the modality matching of the medical images. Thus, the process of transferring image attributes of one imaging modality to another imaging modality may be referred to as modality matching or a style transfer process. The modality matching device 208 may perform the modality matching during the surgery for computationally linking the pre-operative images with the intraoperative images. The modality matching device 208 may perform the style transfer for bridging the gap between the modalities in the one or more images. The modality matching device 208 may involve processing of two images (e.g., style image and content image) andhence reduce computational complexities during integration of the modalities of the one or more images.
[0045] In some embodiments, the modality matching device 208 may communicate with the imaging devices (202, 204) via the user device 206. The block diagram representation 200 may include one or more above-mentioned devices configured to perform the modality matching of the images for assessing anatomical content retention in real-time, in accordance with the present disclosure.
[0046] FIG. 3A illustrates a schematic workflow 300A performed by the modality matching device 208, in accordance with an embodiment of the present disclosure. In an embodiment, the modality matching device 208 may include the processor 210 that is configured to perform the schematic workflow 300A for achieving the modality matching of the images.
[0047] The processor 210 may be configured to obtain the first medical image 203 having the first image modality from the first imaging module 216. The first medical image 203 may correspond to a content image 302. The processor 210 may be configured to obtain the second medical image 205 having the second image modality from the second imaging module 218. The second medical image 205 may correspond to a style image 304.
[0048] The processor 210 may be configured to process the one or more images (i.e., the content image 302 and / or the style image 304) of various imaging modalities. Processing the one or more images may include simulating the 3D image data of 3D volume into a 3D image data of 2D volume using the DRR module 220. The processor 210 may be configured to extract the set of features (e.g., shape, intensity, texture, etc.) of one or more images using the convolutional neural network (i.e., CNN module) 224224.
[0049] The CNN module 224 may include a plurality of blocks 308 arranged in series for extracting the set of features. Each block of the plurality of blocks 308 may comprise a normalization layer and a pooling layer. In an exemplary embodiment, the CNN module 224 configured with the CNN architecture may include the normalization layer. Typically, in deep learning models, a hundreds of images may be inputted into the normalization layer and a lot of data may be utilized. In such cases, the images may be inputted into the normalization layer in batches but not at the same time (i.e., individually one image after another image). In an embodiment of the present disclosure, the CNN module 224 may not be required to normalize hundreds of data and may require to normalize at least two images (i.e., the content image 302 and the style image 304). The process of normalizing the at least two images may be referred to as an instance normalization. The modality matching device 208 may prefer the instancenormalization for normalizing the content image 302 and the style image 304, in accordance with the present disclosure.
[0050] In an exemplary embodiment, the CNN module 224 configured with the CNN architecture may include the pooling layer. The CNN module 224 may perform an average pooling process on the normalized set of features using the pooling layer. The average pooling process may be performed to reduce spatial dimensions of the normalized set of features but retaining dominant features of the set of features. The average pooling may enable averaging of the retained dominant features that are extracted from the plurality of blocks 308. Instead of extracting maximum activation from each pooling layer as max pooling does, the average pooling may compute an arithmetic mean of the normalized set of features in each pooling layer. The averaging of pooled features may enable faster convergence process. The CNN module 224 may be configured with the average pooling CNN architecture. In some embodiments, the average pooling may consider averaging of all the pooled features in the pooling layer and combine many sparse features on average even if many of the pooled features have low magnitude. The average pooling may be performed to reduce an amount of computation carried out for retaining essential information of the normalized set of features.
[0051] The plurality of blocks 308 of the CNN module 224 may comprise a plurality of instance normalization layers (Nl, N2, N3,..Nn-l, Nn) and a plurality of average pooling layers (Pl, P2, P3,..Pn-l, Pn) and each block is composed of NnPn layers.
[0052] In some embodiments, the processor 210 may be configured to extract the one or more distinctive features (set of features) from the plurality of blocks 308 of the CNN module 224. The processor 210 may extract the content features (e.g., shape, edges, comers, etc.) of the content image 302 and the style features (e.g., intensity, texture) of the style image 304. Further, the processor 210 may be configured to extract the style features from each block (e.g., NIP 1, N2P2, N3P3,... Nn-lPn-1, NnPn layers) ofthe plurality ofblocks 308 and the content features from a specific block (e.g., Nn-lPn-1 layer) of the plurality ofblocks 308.
[0053] It may be noted that the style features extracted from all the blocks (e .g ., N 1 P 1 , N2P2, .. , etc.) may include increased spatial resolution of high detailed features whilst the content features extracted from the specific block may include decreased spatial resolution of detailed set of features. For instance, the spatial resolution refers to an amount of detail or information present in the image, measured by a number of pixels used to create the image. Increasing the spatial resolution improves detail and accuracy, whilst decreasing the spatial resolution reduces the file size and processing time. Thus, it may be noted that the content features extracted fromeach block may provide a nonlinear transformations of the content image due to decreased spatial resolution of the set of features. Thus, the processor 210 may be configured to extract the style features from each block and the content features from the specific block among the plurality of blocks 308 for improving the efficiency of the modality matching on the images and assessing the anatomical content retention in real-time, in accordance with the present disclosure.
[0054] The processor 210 may be configured to process the first medical image 203 for extracting the first set of features of the first medical image 203. The first medical image 203 (i.e., the content image 302) may be processed to extract the first set of features (e.g., shape, edge, comer, structure, etc.) using the convolutional neural network 224. The first set of features may correspond to content features. The first set of features may be extracted from the specific block of the plurality of blocks 308 of the convolutional neural network 224. In other words, the processor 210 may be configured to extract the first set of features of the content image 302 from the specific block (e.g., Nn-lPn-1) of the plurality of blocks 308.
[0055] The processor 210 may be configured to process the second medical image 205 for extracting the second set of features of the second medical image 205. The second medical image 205 may be processed to extract the second set of features (e.g., intensity, texture, color, etc.) using the convolutional neural network 224. The intensity represents brightness or lightness of the image and the texture represents visual pattern or surface structure of the image resulting from spatial arrangement of colors or intensities. The second set of features may correspond to style features. The second set of features may be extracted from each block of the plurality of blocks 308 of the convolutional neural network 224. In other words, the processor 210 may be configured to extract the second set of features of the style image 304 from each block (N1P1, N2P2, N3P3,... Nn-lPn-1, NnPn layers) of the plurality of blocks 308.
[0056] The convolutional neural network 224 may perform the instance normalization for generating a normalized first set of features. The instance normalization may be performed based on preprocessing the content features and style features to generate the normalized first set of features. In some embodiments, the processor may be configured to perform the instance normalization over the extracted first set of features. The instance normalization may compute mean and variance for the first set of features across spatial dimensions (height and width) of each feature map. The instance Normalization may normalize the extracted first set of features with each of the features independently. Thus, the instance normalization may be performed based on preprocessing the first set of features to generate a normalized first set of features.
[0057] The processor 210 may be configured to perform the average pooling of the normalized first set of features that provide a first average value. The average pooling may be performed based on computing the first average value of the normalized first set of features. The processor may be configured to apply the first average value onto a reference medical image 306. The reference medical image 306 may be the same as that of the content image 302 at a first iteration. The processor 210 may be configured to process the content image 302 iteratively till the set of features from the required blocks are transferred to the reference medical image (306). Therefore, the content image 302 may be processed to extract and transfer the first set of features onto the reference medical image 306. The processor 210 may be configured to iteratively process the content image 302 until the first set of features from required blocks is transferred to the reference medical image 306.
[0058] In another embodiments, the processor 210 may be configured to perform the instance normalization over the extracted second set of features. The instance normalization may compute mean and variance for the second set of features across the spatial dimensions (height and width) of each feature map. The instance normalization may normalize the extracted second set of features with each of the features independently. Thus, the instance normalization may be performed based on preprocessing the second set of features to generate a normalized second set of features. The processor 210 may be configured to perform the average pooling of the normalized second set of features that provides a second average value. The average pooling may be performed based on computing the second average value of the normalized second set of features. The processor 210 may be configured to apply the second average value onto the reference medical image 306. Thus, the style image 304 may be processed to extract and transfer the second set of features onto the reference medical image 306. The processor 210 may be configured to iteratively process the style image 304 until the second set of features from the required blocks is transferred to the reference medical image 306.
[0059] The reference medical image 306 may include all the content features of the content image 302 and all the style features of the style image 304 pooled into the reference medical image 306 after the end of number of iterations. Therefore, the processor 210 may be configured to process the content image 302 and the style image 304 for extracting and transferring the first and the second set of features onto the reference medical image 306.
[0060] FIG. 3B illustrates an exemplary example 300B performed by the modality matching device 208, in accordance with the present disclosure. In the exemplary example 300B, the modality matching device 208 may be enabled to determine an amount of loss exhibited duringextracting and transferring the first and the second set of features onto the reference medical image 306. The loss metric module 226 of the modality matching device 208 may be enabled to compute a loss metric for determining the amount of loss exhibited by the content image 302 compared with the output image 316. The loss metric may be computed based on the extracted first and second sets of features. The loss metric may include computing a total loss based on a weighted sum of one or more losses. The loss metric module 226 may compute the loss metric for determining contention retention of the anatomical structures in the output image 316 in real time. In an embodiment, the processor 210 may be configured to compute the loss metric based on the extracted first and second sets of features.
[0061] In some embodiments, the processor 210 may be configured to extract a third set of features of the reference medical image 306. The third set of features may include feature maps of the reference medical image 306 at each layer of the plurality of blocks 308.
[0062] The processor 210 may be configured to determine a first loss value 310 based on comparison of the first set of features with the third set of features. The first set of features may include feature maps of the content image 302 at the specific block (i.e., Nn-lPn-1) of the plurality of blocks 308. The first loss value 310 may correspond to a content loss 310. The content loss 310 CLOSS may be determined based on comparing the feature maps of the content image 302 at the specific block of the plurality of blocks 308 and feature maps of the reference medical image 306 from all the layers of the plurality of blocks 308, as indicated in equation 1:-> Equation 1 where:1 - layer of the CNN module 224; i and j - spatial dimensions of the feature maps;F'ij - feature maps of the content image 302 at the specific block of the plurality of blocks 308; andP'ij - feature maps of the reference medical image 306 from all the blocks of the plurality of blocks 308.
[0063] The processor 210 may be configured to determine a second loss value 312 based on comparison of the second set of features with the third set of features. The second loss value 312 may correspond to a style loss 312. The style loss 312 (SLOSS) may be determined based on transposing the same set of feature maps using Gram matrix algorithm as defined in equations 2 and 3:"> Equation 2 -> Equation 3 where:F'ji - Transpose of the same feature maps;G'ij - Gram matrix obtained using the feature maps of the style image at layer 1;A'ij - Gram matrix obtained using the feature maps of the generated image at layer 1;N1 = Number of feature maps at layer 1; andMl = Number of pixels in each feature map at layer 1.
[0064] The processor 210 may be configured to determine atotal variation loss value 314. The total variation loss value 314 may be determined based on computing a variation between the neighbouring pixels of the reference medical image 306. The total variation loss 314 (RLOSS) may be calculated as the sum of the average pooling process of absolute differences between the neighboring pixel values in the output image 316 both horizontally and vertically, as represented in equation 4: pRloss= Zi,j ((N,j +1 - Xjj )2+ (xi+ lij- Xjj)2)2 -> Equation 4 where: x - style transferred output at every epoch having spatial dimensions (i x j); andP - hyperparameter that controls smoothness of the output image 316 (e.g., P=1 or P=2; and P=1 allows to extract more textures and edges).
[0065] The processor 210 may be configured to compute the total loss based on a weighted sum of one or more losses (the first loss value 310, the second loss value 312, and the total variation loss 314).Loss wcClossT wsSlossT wrRloss“^Equation 5 where: wc- regularization weight for computing the content loss 310; ws- regularization weight for computing the style loss 312; and wr- regularization weight for computing the total variation loss 314.
[0066] The regularization weight may be used in computation of the total loss for enabling smoothing the image. Since the present disclosure involves performing a style transfer process which may be detailed below, a lot of style features need to be transferred from the style image 304 to the reference medical image 306. Thus, the regularization module 228 may provide regularization weight (ws) to the style loss and the regularization weight ranges from 0 to 1.
[0067] In some embodiments, the processor 210 may be configured to compute the loss metric based on a weighted sum of the first loss value 310, the second loss value 312, and the total variation loss 314. The loss metric may be computed to provide a measure of the modality matching of the content image 302 and the style image 304. In an exemplary embodiment, the modality matching system may perform style transfer with the objective of minimizing the total loss. The loss metric module 226 may prioritize how much content should be retained from the content image 302. The loss metric module 226 may minimize total variation loss by smoothening the variation in the neighboring pixels to remove noise. Therefore, the loss metric module 226 may ensure (i.e., after every iteration) improving the style features of the style image on retaining the content features of the content image, in accordance with the present disclosure.
[0068] In some embodiments, the processor 210 may be configured to augment the reference medical image 306. The reference medical image 306 may be augmented based on: (i) processing the first and second medical images to extract first and second sets of features; (ii) transferring the first and second sets of features onto the reference medical image 306; and (iii) computing the loss metric based on the extracted first and second sets of features. Thus, the processor 210 may be configured to augment the reference medical image 306 by iteratively performing the steps (i) to (iii) until the loss metric is above a threshold loss metric.
[0069] In the example 300B of the modality matching device 208, the processor 210 may be configured to compare the content image 302 and the augmented reference medical image 316. The augmented reference medical image 316 may correspond to an output image 316. The content image 302 and the output image 316 may be compared based on extracting the set of features of both the images. The set of features are the essential features of the anatomical structures of the medical images and are preserved in accordance with the present disclosure.
[0070] The set of features may be extracted from each block of the plurality of blocks 308 that generate a number of feature maps (e.g., 318a, 318b, 318c, 318d, 318e) and each pair of feature maps may be compared with the one or more images. For example, the CNN module 224 may compare a first feature map of the content image 302 to a first feature map of the output image 316. The convolutional neural network module 224 may provide a difference in comparing each pair in terms of a mean value. The lower the differences in the mean value, higher the content features being preserved. Whenever both the content image 302 and the output image 316 appears to be very similar, the loss metric (i.e., mean value) that obtained after performing the difference may be small.
[0071] In some embodiments, the loss metric may quantify a semantic content loss measure (SCLM) metric to assess the content retention of anatomical structures. The SCLM metric is defined as the sum of absolute differences between the feature maps of the content image 302 and the feature maps of the output image 316 obtained from each blocks of the plurality of blocks 308. For example, the SCLM is plotted for feature maps of the plurality of blocks 308 for defining how much content is being preserved. As could be seen in the Nn-iPn-i block of the plurality of blocks 308, the mean value and all the other values in the feature map (i.e., 318d) are much closer to each other. A lower the SCLM value indicates better content preservation. Thus, the SCLM metric may assess the content retention based on analysing the content features of the content image 302 and style features of the style image 304.
[0072] In an example as shown in Fig. 3B, the SCLM metric is computed for a Coronal anteroposterior (AP) view and a Sagittal lateral-posterior (LP) view images across the plurality of blocks of the VGG-19 model. The mean value is slightly increased in the fifth block and all the other values are not very consistent in the blocks 1 to 4. Particularly, in the fifth block, it is found to appear a high variance in the mean value such as some of the mean values are closer to 0.1, and some other mean values are high as described in Table 1.
[0073] Table 1 :
[0074] Since there is a lot of variation in the mean values in other blocks, it is very reliable to extract the content features from the specific block (e.g., fourth block), in accordance with the present disclosure. Thus, for both the AP images and the LP images, the SCLM value may be found using different blocks and extract content features from the specific block that provide least mean value. Hence, the loss metric module 226 may perform comparison efficiently while preserving the content and performing efficient style transfer. The loss metric module 226 may quantify the loss in terms of metrics for assessing how much the content being lost or content being available.
[0075] In an embodiment, the processor 210 may be configured to output the augmented reference medical image 316 (i.e., the output image 316) having features from both the content image 302 and the style image 304 that belong to different image modalities. Further, theaugmented reference medical image 316 may comprise a style of the style image 304 and retains content of the content image 302. The augmented reference medical image 316 may comprise the set of features (i.e., the content features from the content image 302 and the style features from the style image 304). The set of features are the essential features of the anatomical structures of the medical images. Thus, the processor 210 may be configured to perform modality matching of medical images with preserving the inherent anatomical structures of the one or more medical images. Therefore, it is evident that the modality matching device 208 may enable preserving the output image 316 to quantify anatomical preservation of the content features of the content image 302 and the style features of the style image 304.
[0076] FIG. 4 illustrates a schematic workflow of a method 400 performed by the modality matching device 208, in accordance with an embodiment of the present disclosure. In an embodiment, the method 400 may be performed by the modality matching device 208 for performing modality matching of medical images.
[0077] In an embodiment, the modality matching device 208 may receive one or more images in one or more imaging modalities. Specifically, the modality matching device 208 may obtain at least one of the first medical image 203 from the first imaging module 216. Additionally, the modality matching device 208 may obtain at least one of the second medical image 205 from the second imaging module 218. The modality matching device 208 may receive the second medical image 205 prior to the surgery whilst the first medical image 203 in real time. Further, the modality matching device 208 may receive the first medical image 203 in the 2-D projection and the second medical image 205 in the 3-D projection.
[0078] The method 400 may include (step 402) obtaining the first medical image 203 having the first image modality from the first imaging module 216. The first medical image may correspond to a content image. The method 400 may include (step 404) obtaining the second medical image 205 having a second image modality from a second imaging module 218. The second medical image may correspond to a style image.
[0079] In an embodiment, the modality matching device 208 may transfer both the content image 302 and the style image 304 onto the reference medical image 306. Transferring the content image 302 and the style image 304 onto the reference medical image 306 may represent pooling -in the content features of the content image 302 and the style features of the style image 304 onto the reference medical image 306. The reference medical image 306 may include a synthetic intermediate (e.g., random image, original image, white noise, etc.) same as that ofthe content image 302. The reference medical image 306 may be generated based on similar size of the content image 302 and the style image 304. Further, the reference medical image 306 may be a template or a storage buffer where the content or style features are pooled in.
[0080] In an exemplary example, the content features and the style features may be pooled into the reference medical image 306 in every iteration of the CNN architecture. The pooling of the content features and the style features may be performed for number of iterations to provide the output image 316. For example, at the first iteration, the reference medical image 306 may resemble the content image 302. After the end of number of iterations, the output image 316 may resemble the content image with all the content features and the style features pooled into the reference medical image 306. In other words, the output image 316 may represent the style features being superimposed onto the content image 302. The modality matching device 208 may include transferring the content image 302 and the style image 304 onto the reference medical image 306 for enabling the content image 302 to appear as the style image 304 with preserving all the anatomical details of the content image 302.
[0081] In an exemplary embodiment, the modality matching device 208 may perform the transformation of the content image to the style image and the process of transformation may be termed as style transfer process. The style transfer may be performed for bridging the gap between the modalities for better computational integration of the one or more images.
[0082] In another embodiment, the modality matching device 208 may compare the content image 302 with the style image 304 and store the difference in the one or more distinctive features on to the reference medical image 306. Thus, the modality matching device 208 may compare the content image with the style image for directly bridging disparities between the imaging modalities.
[0083] The method may include (step 406) processing the first medical images to extract first set of features using the convolutional neural network module 224. The first set of features may correspond to content features. The convolutional neural network includes a plurality of blocks 308 arranged in series, each block comprising a convolution layer and an average pooling layer. The method 400 may include extracting the first set of features of the first medical image from each block of the plurality of blocks 308 of the convolutional neural network.
[0084] The method 400 may include performing an instance normalization over the extracted first set of features. The instance normalization may be performed for preprocessing the first set of features to generate a normalized first set of features. The method 400 may include performing an average pooling of the normalized first set of features. The average pooling maybe performed for computing a first average value of the normalized first set of features. The method 400 may include applying the first average value onto the reference medical image 306. The method 400 may include (step 408) transferring the first set of features onto the reference medical image 306.
[0085] The method 400 may include (step 406) processing the second medical images to extract second set of features using a convolutional neural network. The second set of features may correspond to style features. The convolutional neural network includes a plurality of blocks 308 arranged in series, each block comprising a convolution layer and an average pooling layer.
[0086] The method 400 may include extracting the second set of features of the second medical image from a specific block of the plurality of blocks 308 of the convolutional neural network 224.
[0087] The method 400 may include performing an instance normalization over the extracted second set of features. The instance normalization may be performed for preprocessing the second set of features to generate a normalized second set of features. The method 400 may include performing an average pooling of the normalized second set of features. The average pooling may be performed for computing a second average value of the normalized second set of features. The method 400 may include applying the second average value onto the reference medical image 306. The method 400 may include (step 408) transferring the second set of features onto the reference medical image 306.
[0088] The method 400 may include (step 410) computing a loss metric based on the extracted first and second sets of features. The method 400 may include extracting a third set of features of the reference medical image 306. The method 400 may include determining a first loss value 310 based on comparison of the first set of features with the third set of features. The method 400 may include determining a second loss value based on comparison of the second set of features with the third set of features. The method 400 may include determining a total variation loss value based on computing a variation between neighbouring pixels of the reference medical image 306.
[0089] The method 400 may include computing the loss metric based on a weighted sum of the first loss value 310, the second loss value 312, and the total variation loss 314. The loss metric may provide a measure of the modality matching of the first and second images. The method 400 may include (step 412) augmenting the reference medical image 306 by iteratively performing steps (3) to (5) until the loss metric is above a threshold loss metric. The method400 may include (step 414) outputting the augmented reference medical image 316 having features from both the first and second medical images that belong to different image modalities. The augmented reference medical image 316 comprises a style of the style image and retains content of the content image.
[0090] The modality matching device 208 may provide the modality matching on the images with an improved efficiency and assess anatomical content retention in real-time, in accordance with the present disclosure. The images are captured by the imaging module to provide the surgeons an enhanced visualization of anatomical structures of the user for performing the surgical procedure in real time.
[0091] The present disclosure may include performing style transfer on the Coronal anteroposterior (AP) view, Sagittal lateral-posterior (LP) view images, spinal radiographs images, but not limited thereto. The present disclosure may utilize a minimum amount of dataset to learn and perform the modality matching transformation. The present disclosure may include performing with at least one of the content and style images and restricts usage of a large dataset of prior image for performing the modality matching. The present disclosure may disclose a novel modality matching technique that specifically designed for domain adaptation between different modality images, allowing us to compare and work computationally on both the images and preserves anatomical information.
[0092] The present disclosure may include transforming both the content image and the style image into the output image e.g., by superimposing style features of the style image onto the content image. The present disclosure may include transferring the modality of the style image 205 to the modality of the content image 203 and the process may be referred to as style transfer. The present disclosure may include improving efficiency of the modality matching on the images and assessing anatomical content retention in real-time, in accordance with the present disclosure.
[0093] The illustrated steps are set out to explain the exemplary embodiments shown, and it should be anticipated that ongoing technological development will change the manner in which particular functions are performed. These examples are presented herein for purposes of illustration, and not limitation. Further, the boundaries of the functional building blocks have been arbitrarily defined herein for the convenience of the description. Alternative boundaries can be defined so long as the specified functions and relationships thereof are appropriately performed. Alternatives (including equivalents, extensions, variations, deviations, etc., of those described herein) will be apparent to persons skilled in the relevant art(s) based on the teachingscontained herein. Such alternatives fall within the scope and spirit of the disclosed embodiments. Also, the words "comprising," "having," "containing," and "including," and other similar forms are intended to be equivalent in meaning and be open ended in that an item or items following any one of these words is not meant to be an exhaustive listing of such item or items or meant to be limited to only the listed item or items. It must also be noted that as used herein, the singular forms “a”, “an”, and “the” include plural references unless the context clearly dictates otherwise.
[0094] Finally, the language used in the specification may be principally selected for readability and instructional purposes, and it may not have been selected to delineate or circumscribe the inventive subject matter. Accordingly, the disclosure of the embodiments of the disclosure is intended to be illustrative, but not limiting, of the scope of the disclosure.
[0095] With respect to the use of substantially any plural and / or singular terms herein, those having skill in the art can translate from the plural to the singular and / or from the singular to the plural as is appropriate to the context and / or application. The various singular / plural permutations may be expressly set forth herein for sake of clarity.
Claims
WE CLAIM:
1. A method (400) for performing modality matching of medical images, comprising:(1) obtaining (402), from a first imaging module (202), a first medical image (203) having a first image modality;(2) obtaining (404), from a second imaging module (204), a second medical image (205) having a second image modality;(3) processing (406), using a convolutional neural network (224), the first and second medical images to extract first and second sets of features respectively;(4) transferring (408) the first and second sets of features onto a reference medical image (306);(5) computing (410) a loss metric based on the extracted first and second sets of features;(6) augmenting (412) the reference medical image (306) by iteratively performing steps (3) to (5) until the loss metric is above a threshold loss metric; and(7) outputting (414) the augmented reference medical image (316) having features from both the first and second medical images that belong to different image modalities.
2. The method (400) of claim 1, wherein the convolutional neural network (224) includes a plurality of blocks (308) arranged in series, each block comprising a convolution layer and an average pooling layer.
3. The method (400) of claim 2, wherein transferring the first set of features onto the reference medical image (306) comprises: extracting the first set of features of the first medical image from a specific block of the plurality of blocks (308); performing an instance normalization over the extracted first set of features, wherein performing the instance normalization includes preprocessing the first set of features to generate a normalized first set of features; performing an average pooling of the normalized first set of features, wherein performing the average pooling includes computing a first average value of the normalized first set of features; and applying the first average value onto the reference medical image (306).
4. The method (400) of claim 2, wherein transferring the second set of features onto the reference medical image (306) comprises: extracting the second set of features of the second medical image from each block of the plurality of blocks (308); performing an instance normalization over the extracted second set of features, wherein performing the instance normalization includes preprocessing the second set of features to generate a normalized second set of features; performing an average pooling of the normalized second set of features, wherein performing the average pooling includes computing a second average value of the normalized second set of features; and applying the second average value onto the reference medical image (306).
5. The method (400) of claim 1, wherein computing the loss metric comprises: extracting a third set of features of the reference medical image (306); determining a first loss value (310) based on comparison of the first set of features with the third set of features; determining a second loss value (312) based on comparison of the second set of features with the third set of features; determining a total variation loss value (314) based on computing a variation between neighbouring pixels of the reference medical image (306); and computing the loss metric based on a weighted sum of the first loss value (310), the second loss value (312), and the total variation loss (314), wherein the loss metric provides a measure of the modality matching of the first and second images.
6. The method (400) of claim 1, wherein the first medical image comprises a style image and the second medical image comprises a content image, wherein the first set of features comprises style features and the second set of features comprises content features, and wherein the augmented reference medical image (316) comprises a style of the style image and retains content of the content image.
7. A device (208) for performing modality matching of medical images, comprising: a processor (210); anda memory (212) communicatively coupled with the processor (210); wherein the processor (210) is configured to:(1) obtain, from a first imaging module (202), a first medical image (203) having a first image modality;(2) obtain, from a second imaging module (204), a second medical image (205) having a second image modality;(3) process, using a convolutional neural network (224), the first and second medical images to extract first and second sets of features respectively;(4) transfer the first and second sets of features onto a reference medical image (306);(5) compute a loss metric based on the extracted first and second sets of features;(6) augment the reference medical image (306) by iteratively performing steps (3) to(5) until the loss metric is above a threshold loss metric; and(7) output the augmented reference medical image (316) having features from both the first and second medical images that belong to different image modalities.
8. The device (208) of claim 7, wherein the convolutional neural network (224) includes a plurality of blocks (308) arranged in series, each block comprising a convolution layer and an average pooling layer.
9. The device (208) of claim 8, wherein, to transfer the first set of features onto the reference medical image (306), the processor is configured to: extract the first set of features of the first medical image from a specific block of the plurality of blocks (308); perform an instance normalization over the extracted first set of features, wherein performing the instance normalization includes preprocessing the first set of features to generate a normalized first set of features; perform an average pooling of the normalized first set of features, wherein performing the average pooling includes computing a first average value of the normalized first set of features; and apply the first average value onto the reference medical image (306).
10. The device (208) of claim 8, wherein, to transfer the second set of features onto the reference medical image (306), the processor is configured to:extract the second set of features of the second medical image from each block of the plurality of blocks (308); perform an instance normalization over the extracted second set of features, wherein performing the instance normalization includes preprocessing the second set of features to generate a normalized second set of features; perform an average pooling of the normalized second set of features, wherein performing the average pooling includes computing a second average value of the normalized second set of features; and apply the second average value onto the reference medical image (306).
11. The device (208) of claim 7, wherein, to compute the loss metric, the processor is configured to: extract a third set of features of the reference medical image (306); determine a first loss value (310) based on comparison of the first set of features with the third set of features; determine a second loss value (312) based on comparison of the second set of features with the third set of features; determine a total variation loss value (314) based on computing a variation between neighbouring pixels of the reference medical image (306); and compute the loss metric based on a weighted sum of the first loss value (310), the second loss value (312), and the total variation loss (314), wherein the loss metric provides a measure of the modality matching of the first and second images.
12. The device (208) of claim 7, wherein the first medical image comprises a style image and the second medical image comprises a content image, wherein the first set of features comprises style features and the second set of features comprises content features, and wherein the augmented reference medical image (316) comprises a style of the style image and retains content of the content image.