Magnification of medical images using super-resolution
A neural network-based method enhances the resolution of selected medical image portions, improving diagnostic accuracy and efficiency by making subtle structures more visible and reducing system processing load.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- GE PRECISION HEALTHCARE LLC
- Filing Date
- 2024-07-09
- Publication Date
- 2026-07-29
AI Technical Summary
Existing medical imaging techniques, such as CT, struggle with limitations in detectability of subtle findings due to partial volume effects, requiring radiologists to observe multiple 2D images at a higher resolution to identify lesions or nodules, which can be time-consuming and inefficient.
A method involving a neural network that enhances the resolution of selected portions of a 2D medical image by displaying them at a higher resolution in a magnified window, using a bounding box to select the area of interest and applying a convolutional neural network (CNN) for image enhancement.
Improves diagnostic accuracy by making subtle anatomical structures more visible, reducing the processing load on the medical imaging system, and increasing the efficiency of medical review applications by performing higher-resolution image processing on the client device.
Smart Images

Figure 2026525289000001_ABST
Abstract
Description
Technical Field
[0004] ,
[0001] Embodiments of the subject matter disclosed herein relate to medical imaging, and more particularly, to systems and methods for enhancing the resolution of medical images.
Background Art
[0002] Non-invasive imaging techniques enable an image of the internal structure of a patient or object to be obtained without performing an invasive procedure on the patient or object. Specifically, techniques such as computed tomography (CT) use various physical principles such as differences in X-ray transmission through a target volume to acquire image data and reconstruct a tomographic image volume (e.g., a three-dimensional (3D) representation of the interior of a human body or other imaged structure). A two-dimensional (2D) image is extracted from the reconstructed image volume and can be observed by a radiologist on a display device. A radiologist may spend a lot of time observing 2D images of a 3D image volume one by one to search for lesions, nodules, or other medical conditions. In many cases, the findings may be subtle, including limitations in detectability (due to partial volume effects). When searching for lesions, nodules, and other medical conditions, a radiologist may want to observe one or more portions of a 2D image at a higher resolution than that provided by the 2D image.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
[0004] This disclosure includes a method for a medical imaging system, the method comprising the step of displaying a two-dimensional (2D) medical image having a first resolution on the screen of a client device of the medical imaging system, while displaying a selected portion of the medical image at a second resolution, the medical image being reconstructed from image data acquired through the medical imaging system, the second resolution being higher than the first resolution, and the selected portion being selected by a user of the client device. The selected portion may be selected by a user of the medical imaging system, such as a radiologist. In various embodiments, the radiologist defines the selected portion by drawing a bounding box on the 2D medical image using an input device, such as a mouse. In the first embodiment, in order to form a selected portion of the medical image at the second resolution, the 2D medical image is cropped based on the bounding box, and the cropped 2D medical image may be input to a neural network trained to output a 2D medical image with the same field of view but a higher resolution. The neural network can output a higher resolution cropped 2D medical image, which may be displayed on a display panel (e.g., a magnified window) on the screen of the client device. The enlarged window can be superimposed on the 2D medical image at the position of the boundary frame.
[0005] In a second embodiment, to form a selected portion of a medical image at a second resolution, the 2D medical image can be input into a neural network when the 2D medical image is first displayed on the client device to generate a higher-resolution 2D medical image. The higher-resolution 2D medical image can be stored in the client device's memory. When a user draws a bounding box on the 2D medical image displayed on the client device's screen, a second bounding box can be generated on top of the higher-resolution 2D medical image stored in memory. The higher-resolution 2D medical image can be cropped based on the second bounding box to form an image of the selected portion at the second resolution. The image can then be displayed in an enlarged window on the client device's screen. By forming a higher-resolution 2D medical image prior to the selection of a portion, the responsiveness of the medical review application running on the client device can be increased. However, the waiting time for the medical review application running on the client device when a new 2D medical image is displayed on the screen may increase.
[0006] By displaying a selected area of a 2D medical image in a second, higher-resolution magnified window, users can observe anatomical structures within the selected area that are not visible or clear in the first, lower-resolution 2D medical image. For example, an area of increased contrast uptake potentially indicating malignant cells may be too small to be seen in the first-resolution 2D medical image, but may become visible in the second-resolution 2D medical image. Similarly, the boundary of a larger area with high contrast uptake may not be clearly defined in the first-resolution 2D medical image, but may be clearly defined in the second-resolution 2D medical image. Thus, observing areas of increased contrast uptake in a higher-resolution 2D medical image can improve the diagnostic accuracy of 2D medical images.
[0007] Another advantage of the systems and methods disclosed in this book is that higher resolution 2D medical images can be formed on the client device of the medical imaging system, and the processing resources of the medical imaging system and / or image processing system can be avoided during the formation of higher resolution 2D medical images. More precisely, 2D medical images can be displayed on the screen of the client device by a medical examination software application that communicates with the server of the medical imaging system and / or image processing system. 2D medical images are requested from the server by the client device and rendered on the server (e.g., in the backend). 2D medical images are sent to the client device for display, and higher resolution 2D medical images are formed using the computing and memory resources of the client device. In this way, the efficiency of the operation of the medical imaging system and / or image processing system can be increased.
[0008] The advantages and other advantages and features described herein will become immediately apparent from the following “Modes for Carrying Out the Invention,” either alone or in conjunction with the accompanying drawings. It should be understood that the above summary is provided in a simplified form to introduce various concepts that will be further described in the detailed description. Such description is not intended to identify the principal or essential features of the claimed subject matter, and the scope of the claimed subject matter is uniquely defined by the claims following the detailed description. Furthermore, the claimed subject matter is not limited to any embodiment that addresses any of the shortcomings described above or in any part of this disclosure. [Brief explanation of the drawing]
[0009] Various aspects of this disclosure will be better understood by reading the detailed explanation below and referring to the attached drawings.
[0010] [Figure 1] This is a schematic diagram of a computer-aided tomography (CT) imaging system according to one or more embodiments of the present disclosure. [Figure 2]This is a schematic block diagram of an example CT imaging system according to one or more embodiments of the present disclosure. [Figure 3] This is a block diagram of an example of an embodiment of an image processing system configured to increase the resolution of a portion of a 2DCT image according to one or more embodiments of the present disclosure. [Figure 4(A)] This figure shows a first workflow for increasing the resolution of a portion of a 2DCT image according to one or more embodiments of the present disclosure. [Figure 4(B)] This figure shows a second workflow for increasing the resolution of a portion of a 2DCT image according to one or more embodiments of the present disclosure. [Figure 5] This is a block diagram of a client-server architecture used to display 2D medical images on the display screen of a client device according to one or more embodiments of the present disclosure. [Figure 6] This is a block diagram of an example embodiment of a neural network training system for training a neural network to improve the resolution of a portion of a 2DCT image according to one or more embodiments of the present disclosure. [Figure 7] This flowchart illustrates an exemplary method for training a neural network to improve the resolution of a portion of a 2DCT image according to one or more embodiments of the present disclosure. [Figure 8] This flowchart illustrates an exemplary method for increasing the resolution of a portion of a 2DCT image using a neural network according to one or more embodiments of the present disclosure. [Figure 9] This figure shows an exemplary architecture of a neural network used to enhance the resolution of 2DCT images according to one or more embodiments of the present disclosure. [Figure 10] This figure shows an example of a 2DCT image displayed on a display device according to one or more embodiments of the present disclosure, wherein a portion of the 2DCT image is displayed at an enhanced resolution. [Figure 11] This figure shows a simplified version of an exemplary data architecture for an autoencoder CNN used to generate a latent spatial representation of an image volume according to one or more embodiments of the present disclosure. [Figure 12] This figure shows an exemplary data architecture for an encoder that generates a latent spatial representation from an image volume according to one or more embodiments of the present disclosure. [Figure 13] This figure shows an alternative encoder-decoder architecture for image-enhanced CNNs according to one or more embodiments of the present disclosure.
[0011] The drawings illustrate specific aspects of the systems and methods described herein. Together with the following descriptions, the drawings demonstrate and illustrate the structures, methods, and principles described in this book. In the drawings, the dimensions of components may be exaggerated or otherwise altered for clarity. To avoid obscuring aspects of the components, systems, and methods described herein, well-known structures, materials, or operations will not be illustrated or described in detail. [Modes for carrying out the invention]
[0012] This book describes a method and system for improving the quality of selected portions of a 2D medical image by displaying them on a display panel superimposed on the 2D medical image (e.g., a magnified window) at a higher resolution (e.g., super-resolution) than the original 2D medical image. Unlike other medical examination software applications that provide digital magnification of image pixels in a magnified window, the proposed system uses a deep learning (DL) neural network to form a higher-resolution image for display in the magnified window.
[0013] Examples of computer-aided tomography (CT) imaging systems are shown in Figures 1 and 2. Images reconstructed from projection data acquired via the CT imaging system can be post-processed by an image processing system, such as the image processing system shown in Figure 3. In one example, a radiologist can observe the reconstructed image on a client device of the CT imaging system, such as the client device shown in Figure 5. The radiologist can select a portion of the 2D medical image using an input device such as a mouse. For example, the radiologist can draw a bounding box around the region of interest of the 2D medical image. In the first example of post-processing shown in Figure 4(A), the image processing system can crop the 2D medical image based on the bounding box and input the cropped 2D medical image into an enhancement DL model. The enhancement DL model can output a cropped version of the 2D medical image with a higher resolution than the original 2D medical image. The higher-resolution cropped version of the 2D medical image can be displayed at the position of the bounding box in an enlarged window of the display device, as shown in Figure 10. In this way, radiologists can enlarge and reshape a selected portion of the original 2D medical image for higher resolution observation.
[0014] In the second example of post-reconstruction processing shown in Figure 4(B), the original 2D medical image is input into an image enhancement DL model to form an enhanced 2D medical image with a higher resolution than the original 2D medical image. The enhanced 2D image can be stored in the memory of the CT imaging system. If the radiologist draws a bounding box around the region of interest, the image processing system can display the corresponding region of interest in the enhanced 2D medical image at the position of the bounding box in an enlarged window.
[0015] In various embodiments, the image enhancement DL model may be a convolutional neural network (CNN), and this model may have a first encoder-decoder architecture as shown in FIG. 9, or a second alternative encoder-decoder architecture as shown in FIG. 13. When the first encoder-decoder architecture is used, the first 3D latent space representation of the image volume generated as shown in FIG. 12 can be input into the CNN in the decoder part of the CNN. When the second alternative encoder-decoder architecture is used, the second 3D latent space representation of the image volume generated as shown in FIGS. 11 and 13 can be input into the CNN in the input layer of the second encoder part of the CNN. The CNN can be trained according to the method described with respect to FIG. 7 inside the neural network training system shown in FIG. 6. The CNN can be deployed to form higher-resolution images by taking one or more steps of the method described with respect to FIG. 8.
[0016] Figure 1 shows an exemplary CT system 100 configured for CT imaging. Specifically, the CT system 100 is configured to image subjects 112 such as patients, inanimate objects, one or more manufactured parts, and / or foreign bodies present in the body, such as dental implants, stents, and / or contrast agents. In one embodiment, the CT system 100 includes a gantry 102, which may further include at least one X-ray source 104 configured to project a beam of X-ray emission 106 (see Figure 2) used to image the subject 112 lying on a table 114. More precisely, the X-ray source 104 is configured to project the X-ray emission beam 106 toward a detector array 108 located opposite the gantry 102. Although Figure 1 shows a single X-ray source 104, in some embodiments, multiple X-ray sources and detectors can be used to project multiple X-ray emission beams and acquire projection data at different energy levels corresponding to the patient. In some embodiments, the X-ray source 104 can enable dual-energy gemstone spectral imaging (GSI) by fast peak kilovolt (kVp) switching. In some embodiments, the X-ray detector used is a photon counting detector capable of discriminating X-ray photons of different energies. In other embodiments, two sets of X-ray sources and detectors are used to generate a dual-energy projection, with one set at a low kVp and the other at a high kVp. Thus, the methods described herein can be implemented as both single-energy acquisition methods and dual-energy acquisition methods.
[0017] In some embodiments, the CT system 100 further includes an image processor unit 110 configured to reconstruct an image of a target volume of the subject 112 using an iterative or analytical image reconstruction method. For example, the image processor unit 110 can reconstruct an image of a patient's target volume using an analytical image reconstruction approach such as filtered back projection (FBP). As another example, the image processor unit 110 can reconstruct an image of the target volume of the subject 112 using iterative image reconstruction approaches such as advanced statistical iterative reconstruction (ASIR), TrueFidelity (trademark), conjugate gradient (CG), maximum likelihood expectation maximization (MLEM), and model-based iterative reconstruction (MBIR). As described in more detail herein, in some examples, the image processor unit 110 can use both an analytical image reconstruction approach such as FBP in addition to an iterative image reconstruction approach.
[0018] In some CT imaging system configurations, the x-ray source projects a conical x-ray beam that is collimated to be located within a plane of the XYZ Cartesian coordinate system generally referred to as the "imaging plane". The x-ray emission beam passes through an imaging object such as a patient or subject. After being attenuated by the object, the x-ray emission beam impinges on an array of detector elements. The intensity of the attenuated x-ray beam received at the detector array depends on the attenuation of the x-ray emission beam by the object. Each detector element of the array generates a separate electrical signal that serves as a measurement of the x-ray beam attenuation at the detector location. Attenuation measurements from all detector elements are acquired separately to generate a transmission profile.
[0019] In some CT systems, the X-ray source and detector array are rotated by the gantry around the object in the imaging plane so that the angle at which the X-ray beam intersects the object changes constantly. A set of X-ray emission attenuation measurements from the detector array at a single gantry angle, such as projection data, is called a "view." A "scan" of an object includes a set of views formed at different gantry angles, i.e., viewing angles, during one rotation of the X-ray source and detector. Since the advantages of the methods described in this document are also obtainable in medical imaging modalities other than CT, the term "view" as used in this document is not limited to the usage described above in relation to projection data from a single gantry angle. The term "view" is used to mean a single data acquisition whenever there are multiple data acquisitions from different angles, whether from CT, positron emission tomography (PET), or single-photon emission CT (SPECT) acquisition, and / or any other modality including modalities under development, or combinations thereof as a fusion embodiment.
[0020] Projection data is processed to reconstruct images corresponding to two-dimensional slices passing through an object, or, in some examples where the projection data includes multiple views or scans, to reconstruct three-dimensional (3D) images of the object. One method of reconstructing an image from a set of projection data is called filtered back projection in the art. Transmission and emission tomography reconstruction methods also include statistical iterative methods such as maximum likelihood expectation (MLEM) reconstruction and ordered subset expectation reconstruction, alongside iterative reconstruction methods. This process converts attenuation measurements from scans into integers called "CT values" or "Hunsfield units" (HU), which are used to control the brightness of the corresponding pixels in a display device.
[0021] Figure 2 shows an exemplary imaging system 200 similar to the CT system 100 in Figure 1. In view of this disclosure, the imaging system 200 is configured to image a subject 204 (e.g., subject 112 in Figure 1). In one embodiment, the imaging system 200 includes a detector array 108 (see Figure 1). The detector array 108 further includes a plurality of detector elements 202, which together sense an X-ray emission beam 106 (see Figure 2) passing through the subject 204 (patient, etc.) to acquire corresponding projection data.
[0022] In some embodiments, the imaging system 200 is configured to traverse different angular positions around the subject 204 to acquire desired projection data. Thus, the gantry 102 and the components mounted on the gantry 102 may be configured to rotate around a rotation center 206 to acquire projection data at different energy levels, for example. Alternatively, in embodiments where the projection angle relative to the subject 204 changes as a function of time, the mounted components may be configured to move along a general curve rather than along a circular arc.
[0023] As the X-ray source 104 and detector array 108 rotate, the detector array 108 collects data of the attenuated X-ray beam. The data collected by the detector array 108 is preprocessed and calibrated to represent the line integral of the attenuation coefficient of the scanned subject 204. This processed data is generally called a projection.
[0024] In some examples, individual detectors or detector elements 202 of the detector array 108 may include photon counting detectors that register individual photon interactions to one or more energy bins. It should be noted that the methods described herein can also be implemented with energy integration detectors.
[0025] The acquired projection data set may be used for base material decomposition (BMD). During BMD, the measured projections are converted into a set of material density projections. The material density projections can be reconstructed to form a set of material density maps or images for each respective base material, such as a bone map, a soft tissue map, and / or a contrast agent map. These density maps or images can then be associated to form a 3D volumetric image of the base material, such as bone, soft tissue, and / or contrast agent, in the imaged volume.
[0026] Once reconstructed, the base material image formed by the imaging system 200 reveals the internal features of the subject 204 as densities of two base materials. The density image can be displayed to show these features. Conventional approaches to diagnosing medical conditions such as disease states, and more generally, medical events, have involved radiologists or internists considering hard copies or displays of density images to identify prominent features of interest. Such features may include lesions, the dimensions and shapes of specific anatomical structures or organs, and other features that can be identified in the image based on the skill and knowledge of the individual physician.
[0027] In one embodiment, the imaging system 200 includes a control mechanism 208 that controls the movement of components such as the rotation of the gantry 102 and the operation of the X-ray source 104. In some embodiments, the control mechanism 208 further includes an X-ray controller 210 configured to supply power and timing signals to the X-ray source 104. In addition, the control mechanism 208 includes a gantry motor controller 212 configured to control the rotational speed and / or position of the gantry 102 based on imaging requirements.
[0028] In some embodiments, the control mechanism 208 further includes a data acquisition system (DAS) 214 configured to sample analog data received from the detector elements 202 and convert the analog data into digital signals for subsequent processing. The DAS 214 may further be configured to selectively sum analog data from a subset of the detector elements 202 to form a so-called macrodetector, as described in detail herein. The data sampled and digitized by the DAS 214 is transmitted to a computer or computing device 216. In one example, the computing device 216 stores the data in a storage device or mass storage unit 218. The storage device 218 may include, for example, a hard disk drive, a floppy disk drive, a compact disc read / write (CD-R / W) drive, a digital multipurpose disc (DVD) drive, a flash drive, and / or a solid-state storage drive.
[0029] In addition, the computing device 216 controls system operations such as data acquisition and / or processing by providing commands and parameters to one or more of the DAS 214, X-ray controller 210, and gantry motor controller 212. In some embodiments, the computing device 216 controls system operations based on operator input. The computing device 216 receives operator input, including, for example, commands and / or scan parameters, via an operator console 220 coupled to the computing device 216 in terms of operation. The operator console 220 may include a keyboard (not shown) or touchscreen that allows the operator to specify commands and / or scan parameters.
[0030] Figure 2 shows one operator console 220, but more than one operator console may be coupled to the imaging system 200 for purposes such as inputting and outputting system parameters, requesting inspections, plotting data, and / or observing images. Furthermore, in some embodiments, the imaging system 200 can be coupled to a number of displays, printers, workstations, and / or similar devices, which may be located, for example, within a facility or hospital, or in completely different locations, either locally or remotely, and may be coupled via one or more configurable wired and / or wireless networks such as the Internet and / or virtual private network, wireless telephone network, wireless private network, wired private network, wireless wide area network, and wired wide area network.
[0031] In one embodiment, for example, the imaging system 200 may include or be coupled to a medical image storage and communication system (PACS) 224. In one example of an embodiment, the PACS 224 may be further coupled to remote systems such as a radiology information system or a hospital information system, and / or an on-premises network or an external network (not shown), enabling operators at various locations to give commands and parameters and retrieve image data.
[0032] The computing device 216 operates the table motor controller 226 using instructions and parameters given by the operator and / or defined by the system, and the table motor controller 226 can then control the table 114, which may be an electric table. More precisely, the table motor controller 226 can move the table 114 to appropriately position the subject 204 in the gantry 102 in order to acquire projection data corresponding to the target volume of the subject 204.
[0033] As described above, the DAS214 samples and digitizes the projection data acquired by the detector element 202. Subsequently, the image reconstructor 230 performs high-speed reconstruction using the sampled and digitized X-ray data. Figure 2 shows the image reconstructor 230 as a separate entity, but in some embodiments, the image reconstructor 230 may form part of the computing device 216. Alternatively, the image reconstructor 230 may not be present in the imaging system 200, and instead, the computing device 216 may perform one or more functions of the image reconstructor 230. The image reconstructor 230 may be located locally or remotely and may be operationally connected to the imaging system 200 using a wired or wireless network. Specifically, one example of an embodiment may use computing resources of a “cloud” network cluster for the image reconstructor 230.
[0034] In one embodiment, the image reconstructor 230 stores the reconstructed image in the storage device 218. Alternatively, the image reconstructor 230 may transmit the reconstructed image to the computer 216 to generate patient information useful for diagnosis and evaluation. In some embodiments, the computer 216 can transmit the reconstructed image and / or patient information to a client device 232 connected to the computer 216 and / or the image reconstructor 230, and display it on the display of the client device 232. In some embodiments, the reconstructed image may be transmitted from the computer 216 or the image reconstructor 230 to the storage device 218 for short-term or long-term storage.
[0035] The image reconstructor 230 may also include one or more image processing subsystems that can be used to assist in image reconstruction. For example, one or more image processing systems may include an image processing system such as the one shown in Figure 3, which can perform post-reconstruction processing on the 3D image volume and / or the 2D medical image extracted from the 3D image volume. In one embodiment, the image processing system may be configured to increase the resolution of one or more portions of the 2D medical image displayed on a client device such as the client device 232.
[0036] While the methods and systems disclosed herein are described in relation to CT images and CT imaging systems such as CT imaging systems 100 and 200, it should be noted that these methods and systems may also be applied to other forms of medical imaging and imaging systems without departing from the scope of this disclosure. For example, other forms of medical imaging include magnetic resonance imaging (MRI), positron emission tomography (PET) images, single-photon emission tomography (SPECT) images, and / or other types of medical imaging.
[0037] Figure 3 shows an image processing system 302 of a medical imaging system 300 according to one embodiment. In some embodiments, at least a portion of the image processing system 302 is located in a device (e.g., edge device, server, etc.) that is connected to the medical imaging system 300 via wired and / or wireless connections. In some embodiments, at least a portion of the image processing system 302 is located in a separate device (e.g., workstation) that can receive images from the medical imaging system 300 or from a storage device that stores images / data generated by the medical imaging system 300.
[0038] The image processing system 302 includes a processor 304 configured to execute machine-readable instructions stored in a non-transient memory 306. The processor 304 may be single-core or multi-core, and the program executed therein may be configured for parallel or distributed processing. In some embodiments, the processor 304 may optionally include separate components distributed across two or more devices that are remotely located and / or configured for cooperative processing. In some embodiments, one or more aspects of the processor 304 are virtualized and executed by a remotely accessible network computing device configured as a cloud computing configuration.
[0039] The non-transient memory 306 can store the neural network module 308, the network training module 310, the inference module 312, and the medical image data 314. The neural network module 308 may contain a deep learning (DL) model, and the instructions that embody the DL model improve the quality of the medical images (e.g., increase the resolution). This will be explained in more detail later. The neural network module 308 may contain one or more trained and / or untrained neural networks, and may further contain data or metadata related to one or more neural networks, which are stored in these neural networks.
[0040] The training module 310 may include instructions for training one or more neural networks that embody the DL model stored in the neural network module 308. Specifically, the training module 310 may include instructions that, when executed by the processor 304, cause the image processing system 302 to perform one or more steps of method 700 for training one or more neural networks in the training phase. This will be discussed in more detail later with reference to Figures 6 and 7. In some embodiments, the training module 310 includes instructions for embodying one or more gradient descent algorithms, and for applying one or more loss functions and / or training routines to be used to adjust the parameters of one or more neural networks in the neural network module 308. Non-transient memory 306 also stores an inference module 312 that includes instructions for deploying the DL model.
[0041] The non-transient memory 306 further stores medical image data 314. The medical image data 314 may include, for example, medical images acquired via a CT scanner, MRI scanner, spectral imaging scanner, or different imaging modalities.
[0042] In some embodiments, the non-transient memory 306 may include components arranged in two or more devices that are located remotely and / or configured for cooperative processing. In some embodiments, one or more aspects of the non-transient memory 306 may include a remotely accessible network storage device configured as a cloud computing configuration.
[0043] The image processing system 302 may be operationally / communically coupled to a user input device 332. The user input device 332 may include one or more touch screens, keyboards, mice, trackpads, motion-sensing cameras, or other devices configured to enable a user to interact with and manipulate data in the image processing system 302.
[0044] The image processing system 302 may also be operationally / communically coupled to the display device 334. The display device 334 may include one or more display devices using substantially any form of technology. In some embodiments, the display device 334 may include a computer monitor that can display medical images. The display device 334 may be combined with the processor 304, non-transient memory 306, and / or user input device 332 in a common housing, or it may be a peripheral display device that includes a monitor, touch screen, projector, or other display devices known in the art that can enable a user to observe medical images formed by the medical imaging system or to interact with various data stored in the non-transient memory 306.
[0045] The image processing system 302 can be operationally coupled / communically coupled to the CT scanner 336. The CT scanner 336 may be any CT imaging device configured to image subjects such as patients, inanimate objects, one or more manufactured parts, and / or foreign bodies present in the body, such as dental implants, stents, and / or contrast agents. The image processing system 302 can receive CT images from the CT scanner 336, process the received CT images via the processor 304 based on instructions stored in one or more modules of the non-transient memory 306, display the processed CT images on the display device 334, and / or store the received CT images as medical image data 314.
[0046] It should be understood that the image processing system 302 shown in Figure 3 is for illustrative purposes only and not as a limitation. Another suitable image processing system may include more components, fewer components, or different components.
[0047] Figures 4(A) and 4(B) illustrate the workflows in which a radiologist uses a medical examination software application to observe 2D medical images formed by imaging systems such as CT imaging systems 100 and 200 on a client device of the imaging system (e.g., client device 232 in Figure 2). The client device may be a computing device such as a PC. An exemplary client-server architecture used by the medical examination software application will be described again later with respect to Figure 5. In both workflows, the medical examination software application is used to create magnified images of portions of the 2D medical image, which include the anatomical structures of the subject in the 2D medical image at a higher resolution than the resolution of the 2D medical image.
[0048] Referring to Figure 4(A), a first workflow diagram 400 is shown illustrating the radiologist's first workflow, in which the radiologist is observing the first CT image 402. The first CT image 402 may be displayed on the screen of the client device by the radiologist. The radiologist may focus on a specific part 406 of the anatomical structure of the subject in the first CT image 402. More precisely, the radiologist may focus on one or more high-contrast areas, such as a high-contrast area 408, which may indicate a pathological condition. For example, the high contrast may be a result of increased contrast agent uptake due to high metabolic activity in the tissue of the high-contrast area 408.
[0049] A radiologist can select a portion 406 of an anatomical structure of a subject using an input device of a client device such as a mouse, and the selected portion 406 is indicated by a bounding box 404. For example, the bounding box 404 can be generated by the radiologist selecting a mouse button at a first position 440 in the first CT image 402 corresponding to the upper left corner of the bounding box 404, and then dragging the mouse to a second position 442 in the first CT image 402 corresponding to the lower right corner of the bounding box 404. As the radiologist drags the mouse, the lines of the bounding box 404 may appear superimposed on the first CT image 402 on the display device to indicate the position of the bounding box 404. In this way, the radiologist can adjust the first position and / or first dimension of the bounding box 404 based on the second position and / or second dimension of the anatomical feature of interest of the selected portion 406.
[0050] When the radiologist releases the mouse button, a second CT image 410 can be generated in the client device, which is a cropped version of the first CT image 402 that includes anatomical features within the bounding frame 404 but does not include anatomical features outside the bounding frame 404. The second CT image 410 includes a high-contrast region 408. In Figure 4(A), the first dimension of the second CT image 410 is shown to be larger than the second dimension of the bounding frame 404, and individual pixels in the second CT image 410 appear clearer than in the first CT image 402. However, the resolution of the second CT image 410 is equal to the resolution of the first CT image 402. Thus, the second CT image 410 may represent a digital enlargement of the anatomical features within the bounding frame 404. As a result of the digital enlargement, the high-contrast region 408 may be more easily visualized in the second CT image 410 than in the first CT image 402. However, due to pixelation of the high-contrast region 408 in the second CT image 410, the boundary and / or other specific features of the high-contrast region 408 may be less visible in the second CT image 410 than in the first CT image 402.
[0051] To make the boundaries of the high-contrast region 408 and / or other specific features more visible, a second CT image 410 may be input to a trained image enhancement DL model 430. The DL model 430 may be included in a medical examination software application or may be stored in a client device and accessed by the medical examination software application. The image enhancement DL model 430 may be trained within an image processing system, such as the image processing system 302 in Figure 3, to enhance the resolution of the second CT image 410, as will be described later with respect to Figures 6 and 7.
[0052] The image enhancement DL model 430 can output a third CT image 412, which may be a higher-resolution enhanced version of the second CT image 410. The third CT image 412 includes a high-contrast region 408. However, in contrast to the second CT image 410, the third CT image 412 shows a selected portion 406 of the subject's anatomical structure, which includes the high-contrast region 408, at a higher resolution than in the second CT image 410 and the first CT image 402. As a result of showing the high-contrast region 408 at a higher resolution, the boundaries and / or other specific features of the high-contrast region 408 may be more visible to the radiologist on the client device screen.
[0053] Figure 4(B) shows the second workflow of the second alternative workflow, Figure 450, in which the radiologist observes the first CT image 402 on the client device as described above with respect to Figure 4(A). Similar to the first workflow in Figure 4(A), the radiologist focuses particularly on high-contrast areas, such as high-contrast regions 408 of the anatomical structure 406 of the subject.
[0054] In the second workflow, when the radiologist first displays the first CT image 402 on the client device, the image processing system automatically inputs the first CT image 402 into the image enhancement DL model 430. The image enhancement DL model 430 outputs a second CT image 452, which is a higher-resolution version of the first CT image 402. The second CT image 452 can be stored in the client device's memory. In the second workflow, the radiologist can generate a bounding box 404 on the first CT image 402 using a mouse, similar to the first workflow. Once the radiologist has generated a bounding box 404 on the first CT image 402, the medical examination software application can generate a second corresponding bounding box 454 on the second CT image 452. The second CT image 452 can be cropped based on the second corresponding bounding box 454 to form a third CT image 456, which includes a high-contrast area 408. As a result of cropping the third CT image 456 from the higher-resolution second CT image 452, the third CT image 456 shows portions of anatomical structures within the bounding box 404 with higher resolution than in the first CT image 402. Specifically, the high-contrast area 408 is displayed with higher resolution in the third CT image 456 than in the first CT image 402.
[0055] The advantage of the second workflow over the first workflow is that the calculations performed during the formation of higher-resolution images only need to be done once per slice of the corresponding 3D image volume, allowing more frames to be formed per second when the mouse is moved. However, the disadvantage of the second workflow is that the latency of the medical examination software application on the client device can be greater than in the first workflow when changing the camera position (e.g., when selecting a new slice of the corresponding 3D image volume). In addition, the amount of memory consumed by the client device when performing calculations to form and store the (larger) higher-resolution images in the second workflow may be greater than the amount of memory consumed when performing calculations to form and store the (smaller) higher-resolution images in the first workflow that correspond to the bounding box rather than the entire 2D image. Thus, there may be a trade-off between the latency of the medical examination software application between camera positions and the responsiveness of the medical examination software application when the mouse is being manipulated.
[0056] Figure 5 shows an exemplary client-server architecture 500 used by a medical examination application 506 running on a client device 502, the client device 502 in contact with a server 505 hosted on a computing device 504. The client device 502 may be a computer such as a PC or tablet. In various embodiments, the computing device 504 may be an unspecified example of the computing device 216 of the imaging system 200 in Figure 2, and the client device 502 may be an unspecified example of the client device 232 of the imaging system 200.
[0057] A user 540 may interact with the client device 502 via a medical examination application 506. For example, the user 540 may be a radiologist observing medical images (e.g., CT images, MRI images, and PET images, etc.) on the display screen 520 of the client device 502. The medical images may be reconstructed by an image reconstructor of the imaging system (e.g., the image reconstructor 230 of the imaging system 200). The client device 502 further includes a processor 522 and non-transient memory 524. The processor 522 may be configured to execute machine-readable instructions stored in the non-transient memory 524. The processor 522 may be single-core or multi-core, and the program executed therein may be configured for parallel processing or distributed processing.
[0058] To examine a medical image on the display screen 520, the user 540 can select the relevant image volume 536 via the menu of the medical examination application 506. Once an image volume 536 is selected, 2D medical images representing slices of the image volume 536 may be formed in the viewport 508 of the medical examination application 506 based on the selected or default viewpoint / camera angle. The viewport 508 may be displayed on the display screen 520. The user 540 can scan the image volume 536 using the control units of the viewport 508, and in response to the user 540 adjusting one or more of the control units, different 2D medical images corresponding to different slices of the image volume 536 may be displayed in the viewport 508. The different 2D medical images may include different views of the image volume 536, such as sagittal, coronal, and / or axial views, and these views may be displayed simultaneously in the viewport 508.
[0059] The image volume 536 can be stored in the memory 534 of the computing device 504. As a user views the image volume 536, they can request a 2D medical image corresponding to the desired camera angle of the image volume 536 from the server 505. In response to receiving a request for a 2D medical image from the medical examination application 506, the server 505 can use the renderer 530 to form the corresponding 2D medical image. Once the 2D medical image has been rendered by the renderer 530, it can be transmitted to the medical examination application 506. The 2D medical image can then be displayed in the viewport 508. In this way, the processing of the image volume 536 performed by the renderer 530 is carried out using the computing resources of the computing device 504 and memory 534, and the processor 522 and memory 524 of the client device 502 are not used for processing the image volume 536. As a result, the use of computing resources on the client device 502 by the medical examination application 506 is minimized, which can shorten the waiting time for displaying medical images and / or increase the responsiveness of the medical examination application 506.
[0060] When a user examines a 2D medical image displayed in viewport 508, there may be times when the user wants to see a part of the 2D medical image in more detail. For example, the user may want to examine a specific area of the 2D medical image or look for signs of diseased tissue that may be small and / or difficult to see. To enable the user to observe a part of the 2D medical image in more detail, the medical examination application 506 includes a zoom tool 510. For example, the zoom tool 510 may be accessible to the user via a control element in viewport 508, such as a magnifying glass icon.
[0061] To use the zoom tool 510, the user can select a portion of the 2D medical image using a user input device such as a mouse 512 connected to the client device 502. For example, the user can select a button on the mouse 512 and drag the mouse 512 to draw a bounding box around the portion of the 2D medical image they are interested in. In other embodiments, different user input devices may be used.
[0062] Once a bounding box is drawn around the area of interest in the 2D medical image (for example, when the user deselects a button on the mouse 512), the zoom tool 510 can form a second cropped 2D medical image within the bounding box that corresponds to the portion of the 2D medical image. The second cropped 2D medical image can be input to a trained DL model 514, which is trained to form a higher-resolution version of the 2D medical image. In various embodiments, the DL model 514 may be a neural network such as a CNN, a spread CNN, or a generative adversarial network (GAN). An exemplary image-enhancing CNN is described later with respect to Figure 6.
[0063] The DL model 514 can take a second cropped 2D medical image as input and output a higher-resolution version of the second cropped 2D medical image. The higher-resolution version of the second cropped 2D medical image can be displayed to the user 540 as an overlay display 516, which is superimposed on the 2D medical image at the position of the bounding frame. For example, the overlay display 516 may include a display panel or window (e.g., an enlarged window) showing the higher-resolution version of the second cropped 2D medical image. An example of an enlarged window is shown in Figure 10.
[0064] To improve the quality of the high-resolution version of the second cropped 2D medical image, a 3D latent space representation 532 of the image volume 536 may be additionally input to the DL model 514 along with the second cropped 2D medical image as a second input. Before the medical examination application 506 is opened, a convolutional neural network (CNN) embodied by an autoencoder architecture is trained to form a compressed version of the image volume 536, called the latent space of the image volume 536. To generate the latent space, the image data of the image volume 536 may be input to the input layer of the first encoder portion of the autoencoder CNN. The image data may be propagated through a first sequence of convolutional (and pooling) layers of the first encoder portion to form a compressed representation of the image volume 536. This compressed representation may then be propagated through a second sequence of convolutional (and upscaling) layers of the second decoder portion of the autoencoder to reconstruct the image volume 536. The loss, representing the difference between the reconstructed image volume output by the autoencoder CNN and the image volume 536, is backpropagated through the decoder and encoder portions of the autoencoder CNN to update the weights and biases in each layer of the autoencoder CNN. When the loss decreases below a threshold loss, training stops and the compressed representation of the image volume 536 forms a latent space. The latent space representation 532 can be generated from the latent space. For example, the latent space representation 532 may be a portion of the latent space. The latent space representation 532 may contain 3D feature information extracted from the image volume 536 by the first sequence of convolutional layers in the encoder portion.
[0065] Referring now to Figure 11, a simplified version of the exemplary data architecture of the autoencoder CNN 1100 is shown. The autoencoder CNN 1100 has an encoder portion 1101 and a decoder portion 1103. The autoencoder CNN 1100 can be trained to take an image volume 1102 as input and output a reconstructed image volume 1104, the reconstructed image volume 1104 containing image data substantially equivalent to that of image volume 1102. In addition, the reconstructed image volume 1104 may be less noisy and of higher quality than image volume 1102.
[0066] During each iteration of training, image data from image volume 1102 is input to one or more convolutional layers of the encoder portion 1101, such as convolutional layers 1106 and 1108. In each convolutional layer, a compressed (lower dimensionality) version of the image data is generated, and these compressed versions are generated for multiple channels, each representing an input feature map. After each pooling iteration, the number of feature maps increases, allowing for the capture of additional information. After the image data has propagated through the convolutional layers, a compressed representation 1110 of image volume 1102 is generated for multiple channels. The compressed representation 1110 is also called the latent space of image volume 1102.
[0067] Next, the compressed representation 1110 can be propagated through multiple convolutional layers of the decoder section 1103, such as convolutional layers 1112 and 1114. In each of the convolutional layers of the decoder section 1103, the image data of the compressed representation 1110 can be expanded (higher dimensionality) until the reconstructed image volume 1104 is output. The image data of the reconstructed image volume 1104 may have a one-to-one correspondence with the image data of image volume 1102. The difference between the reconstructed image volume 1104 and image volume 1102 can then be backpropagated through the convolutional layers of the decoder section 1103, the compressed representation 1110, and the encoder section 1101 to update the weights and biases of the convolutional layers.
[0068] Training of the autoencoder CNN 1100 can be stopped when the training procedure converges. Once the autoencoder CNN 1100 is trained, the compressed representation 1110 can represent the exact latent space of the image volume 1102, which contains sufficient 3D feature information of the image volume 1102 to reconstruct it. After the autoencoder CNN 1100 is trained, during the inference phase, the coding part 1101 can be used to encode a new image volume to generate a latent space for the new input image volume. The inference phase will be explained in more detail later with reference to Figure 8.
[0069] Returning to Figure 5, once the medical examination application 506 is opened, the client device 502 may request the latent spatial representation 532 from the server 505. The received latent spatial representation 532 can be stored in memory 524. The 3D feature information contained in the latent spatial representation 532 can be used by the DL model 514 (e.g., an image-enhanced CNN) to form a higher-resolution version of the cropped 2D medical image. By including the 3D feature information stored in the latent spatial representation 532 when forming the higher-resolution version, the quality of the higher-resolution version can be improved, and the first higher-resolution version of the cropped 2D medical image generated using the 3D feature information may have higher quality than the second higher-resolution version of the cropped 2D medical image generated without using the 3D feature information. The use of the latent spatial representation of the image volume as input to the image-enhanced CNN will be explained again later with reference to Figure 6.
[0070] After a high-resolution version of the second cropped 2D medical image is displayed in the overlay 516, the user 540 can examine the anatomical structure of the 2D medical image displayed in the overlay 516 in more detail than in the 2D medical image displayed in the viewport 508. The user 540 can also observe different parts of the anatomical structure of the 2D medical image by closing the overlay 516 and redrawing the border. For example, in one embodiment, the user 540 can close the overlay 516 by selecting an icon displayed in the overlay 516 (e.g., an X mark in the upper corner of the overlay 516). The user 540 can re-select the magnification tool 510 by selecting the magnifying glass icon in the medical examination application 506, for example. Once the magnifying glass icon is selected, the cursor displayed in the viewport 508 changes to a magnifying glass icon. The user 540 can draw a new bounding box with the mouse 512 to form a new cropped 2D medical image, and a new high-resolution 2D medical image can be formed by the DL model 514 and displayed in the overlay display 516. In this way, various parts of the original 2D medical image can be displayed in real time with higher resolution within the enlarged window based on the bounding box drawn by the user 540.
[0071] Alternatively, in some embodiments, the selection portion of the 2D medical image does not have to be selected based on a bounding box drawn by the user 540, and the selection portion may be selected based on the field of view (FOV) setting at the position of the mouse 512. In such embodiments, the user 540 can click on the desired location in the 2D medical image using the mouse 512, and a bounding box may be automatically drawn around the desired location based on the FOV (for example, with the desired location as the center of the FOV). The FOV may be adjusted by the user 540 via one or more control units of the zoom tool 510.
[0072] For example, user 540 can identify a small area in a high-contrast 2D medical image that may indicate a high contrast agent uptake rate due to high metabolic activity. In this 2D medical image, the boundary and / or other features of the small area may be blurred or too small to be seen, making the small area not clearly visible. User 540 may want to observe this small area at a higher resolution. User 540 can select this location in the 2D medical image using the mouse 512 (or a different user input device). Once user 540 has selected the location, a bounding box may be automatically drawn around this location based on the current predetermined FOV of the zoom tool 510. User 540 can adjust the FOV using one or more control units of the zoom tool 510, for example, to enlarge or reduce the bounding box. Once the bounding box is generated, a high-resolution 2D medical image corresponding to the cropped 2D medical image is displayed on top of the image 516, as described above. User 540 can observe this small area in a higher resolution 2D medical image, and can observe the boundaries and / or other features of the small area that were not visible in the 2D medical image. As a result of observing the boundaries and / or other features of the small area, User 540 can make a more accurate diagnosis of the subject in the 2D medical image than would have been possible if a higher resolution 2D medical image had not been formed.
[0073] As described above with respect to Figure 4(B), in some embodiments, the 2D medical image can be input to the DL model 514 (e.g., an image-enhanced CNN) before the user 540 forms a bounding frame to form a higher-resolution 2D medical image corresponding to the original 2D medical image. For example, the 2D medical image may be input to the DL model 514 when it is received from the server 505 and displayed in the viewport 508. In such embodiments, once the user 540 forms a bounding frame, the corresponding area of the higher-resolution 2D medical image can be cropped and displayed in the overlay display 516.
[0074] Figure 10 shows an enlarged window 1001 of the display panel of a medical examination application 1000, which may be a non-limiting example of the medical examination application 506 of Figure 5. The medical examination application 1000 may be opened in a client device (e.g., client device 502) of a medical imaging system used by a radiologist to observe medical images of a patient. The client device may communicate with a computing device (e.g., computing device 504) of the medical imaging system. To be clear, a 2D medical image can be requested from the computing device by the client device, and the computing device can form a 2D medical image from the image volume stored in the computing device and send this 2D medical image to the client device for display to the radiologist.
[0075] In Figure 10, the medical examination application 1000 displays a 2D medical image 1002 of a patient in window 1001. Window 1001 includes an upper boundary area 1003 which may contain information such as a patient identifier 1004 and / or other textual information. The upper boundary area 1003 may also contain various control elements (e.g., icons) 1006. A user of the medical examination application 1000 can select one or more of the various control elements 1006 to perform one or more actions in window 1001.
[0076] Specifically, the control element 1006 includes a magnifying glass icon 1008, and when the magnifying glass icon 1008 is selected by the user, an enlargement window 1012 can be generated. The enlargement window 1012 can be superimposed on the 2D medical image 1002. As described above, when the user selects the magnifying glass icon 1008, the appearance of the cursor 1010 displayed in window 1001 changes to a magnifying glass icon. Once the cursor is displayed as a magnifying glass icon, the user can position the magnifying glass icon at a desired location on the 2D medical image 1002 and select a desired part of the patient's anatomical structure displayed on the 2D medical image 1002 so that it is shown in more detail in the enlargement window 1012.
[0077] In the illustrated embodiment, the user positions a magnifying glass cursor to select a button or other control element of a user input device such as a mouse. For example, the user may be interested in examining a high-contrast area 1020 of a 2D medical image 1002 in more detail. In the 2D medical image 1002, the dimensions of the high-contrast area 1020 are small, and the high-contrast area 1020 exhibits pixelation, making the details of the high-contrast area 1020 not clearly visible. As a result of selecting a button at the desired position, the magnification window 1012 displays an area centered on the desired position at a resolution higher than the resolution of the 2D medical image 1002. For example, the desired position indicated by the magnifying glass cursor 1010 in the 2D medical image 1002 may correspond to the center point 1015 of the magnification window 1012. The area centered on the desired position may be defined by the distance 1016 between the center point 1015 and one of the sides of the magnification window 1012 (e.g., the left side, right side, top side, and / or bottom side). In the illustrated embodiment, the enlargement window 1012 has a square shape. In other embodiments, the enlargement window 1012 may have a circular shape or a different shape.
[0078] In the enlarged window 1012, the high-contrast region 1020 is shown with higher resolution, and its extent, shape, and boundaries are more clearly defined than in the 2D medical image 1002. As a result of the higher resolution of the high-contrast region 1020, radiologists can diagnose the patient's condition more accurately and efficiently.
[0079] Referring here to Figure 6, an example of a neural network training system 600 that can be used to train a neural network such as image-enhancing CNN 602 is shown. Image-enhancing CNN 602 can be trained to improve the quality of 2D medical images by forming higher-resolution versions of 2D medical images according to one or more operations, which will be described in more detail later with respect to method 700 in Figure 7. The neural network training system 600 can be embodied by an image processing system such as the image processing system 302 in Figure 3.
[0080] The image enhancement CNN 602 may be stored inside the neural network module 601 of the image processing system. The neural network module 601 may be a non-limiting example of the neural network module 308 of the image processing system 302 in Figure 3. The neural network training system 600 also includes a training module 604, which includes a training dataset containing multiple training data pairs, such as image pairs divided into training image pairs 606 and test image pairs 608. The training module 604 may be a non-limiting example of the training module 310 of the image processing system 302.
[0081] To ensure that sufficient training data is available to prevent overfitting, a certain number of training image pairs (606) and test image pairs (608) can be selected, allowing the image-enhanced CNN 602 to learn to map specific features to training set samples that are not present in the test set.
[0082] Each image pair in training image pair 606 and test image pair 608 includes a 2D input image and a 2D target image. The target image and input image may be images of the same anatomical structure of the subject but with different resolutions, and the target image may have a higher resolution than the input image. In various embodiments, the input image can be formed from the target image. For example, the input image and target image can be formed using a set of reference medical images 612. This set of reference medical images 612 may be formed from a plurality of reference image volumes 611 of various forms, such as CT images, MR images, PET images, SPECT images, and / or different forms of image volumes. In various embodiments, this set of reference image volumes 611 and / or reference medical images 612 may be extracted from the PACS of an imaging system (e.g., PACS224 in Figure 2). Reference images can be selected using various criteria. Specifically, the reference images may be selected from various modalities (e.g., CT, MRI, and PET, etc.) and may include various anatomical structures of the subject. The reference image may have a resolution exceeding a threshold (e.g., high) resolution.
[0083] In various embodiments, the input image may be formed from a target image via a resolution reduction process 614 of the neural network training system 600. The resolution reduction process 614 can reduce the resolution of a reference medical image 612 to form a set of lower-resolution medical images 616. The resolution reduction process 614 can incorporate one or more of various methods and / or techniques for reducing the resolution of the reference medical image 612, which will be further explained later with reference to Figure 7.
[0084] The neural network training system 600 may include a training data generator 610 that can be used to generate training image pairs 606 and test image pairs 608 for the training module 604. Images from this set of reference medical images 612 can be paired by the training data generator 610 with corresponding images from lower-resolution medical images 616 to form an image pair. Once each image pair is generated, it can be assigned to either the training image pair 606 or the test image pair 608.
[0085] In various embodiments, the image-enhanced CNN 602 can take latent spatial representations 605 of reference image volumes 611 corresponding to the input and target images of each training image pair 606 as additional inputs. As described above, the latent spatial representation 605 can be generated from the reference image volume 611, including features extracted from the reference image volume 611. The generation of the latent spatial representation 605 will be described in more detail later with reference to Figure 7.
[0086] In one embodiment, image pairs may be randomly assigned to either the training image pair 606 or the test image pair 608 at a predetermined ratio. For example, image pairs may be randomly assigned to either the training image pair 606 or the test image pair 608 such that 90% of the generated image pairs are assigned to the training image pair 606 and 10% are assigned to the test image pair 608. Alternatively, image pairs may be randomly assigned to either the training image pair 606 or the test image pair 608 such that 85% of the generated image pairs are assigned to the training image pair 606 and 15% are assigned to the test image pair 608. The examples provided herein are for illustrative purposes only, and it should be acknowledged that image pairs may be assigned to the training image pair 606 dataset or the test image pair 608 dataset via different procedures and / or different ratios without departing from the scope of this disclosure.
[0087] The neural network training system 600 may include a validater 620 that validates the performance of the image-enhanced CNN 602 on a portion of the test image pair 608 (e.g., a validation set). The validater 620 can take a partially trained image-enhanced CNN 602 and a validation set of the test image pair 608 as input and output an assessment of the performance of the partially trained image-enhanced CNN 602 on the validation set of the test image pair 608.
[0088] Once the image-enhancing CNN 602 is validated, the trained image-enhancing CNN 622 (e.g., validated image-enhancing CNN 602) can be used to form a set of higher-resolution 2D medical images 634 from a set of acquired 2D medical images 632. For example, the acquired medical images 632 may be acquired by an imaging device 630, which could be a non-limiting example of the CT scanner 336 in Figure 3. The trained image-enhancing CNN 622 may be stored within the inference module 621 of the image processing system (e.g., the inference module 312 in Figure 3). The trained image-enhancing CNN 622 may also be deployed to a client device running a medical examination software application, such as the client device 502 in Figure 5.
[0089] Referring here to Figure 7, a flowchart of method 700 for training an image-enhanced CNN (e.g., a DL model such as DL model 514 in Figure 5) is shown. The image-enhanced CNN may be a non-limiting example of the image-enhanced CNN 602 of the image-enhanced network training system 600 in Figure 6, according to one embodiment. In one embodiment, the operation of method 700 may be stored in the non-transient memory of the image processing system (e.g., a training module such as the training module 310 of the image processing system 302 in Figure 3) and executed by the processor of the image processing system (e.g., the processor 304 of the image processing system 302 in Figure 3). The image-enhanced CNN may be trained on training data which includes one or more sets of image pairs. Each of the one or more sets of image pairs may include medical images of different resolutions, as described later. In some embodiments, one or more sets of image pairs may be stored in a medical image dataset of the image processing system, such as the medical image data 314 of the image processing system 302.
[0090] Method 700 begins in block 702, where it includes the step of obtaining a set of reference image volumes that can be used to generate a training dataset of 2D medical images. These reference image volumes may include various forms of images, such as CT images, PET images, MR images, SPECT images, and / or other forms of images. The reference image volumes may cover a diverse range of different patient types and dimensions, including men, women, and children, presenting with various forms of disease in various different organs and / or parts of the body. The disease can be in various stages of progression. For example, a reference image fiber may include a first image volume containing precancerous tissue reconstructed from the patient at a first time point, a second image volume showing early-stage malignant tumor tissue reconstructed from this patient at a second time point, and a third image volume showing late-stage malignant tumor tissue reconstructed from this patient at a third time point, and so on. The purpose of generating the reference image volumes may be to assemble a collection of both healthy and diseased tissues across a wide range of patients and anatomical sites.
[0091] In block 704, method 700 includes the step of forming 2D reference images from a reference image volume for use in training an image-enhanced CNN. The 2D reference images may be selected from various slices of the reference image volume, including sagittal, coronal, and axial images. Multiple 2D reference images may be selected from a single reference image volume. The 2D reference images may include both diseased and healthy tissue.
[0092] In block 706, method 700 includes the step of forming a lower-resolution image for each of the 2D reference images that are formed. Lower-resolution images can be formed using various techniques, methods, and / or techniques. For example, in one embodiment, the resolution of the 2D reference image can be reduced using one or more diffusion models. In other embodiments, the resolution can be reduced using a vision transformer, or different techniques, techniques, or processes can be used. For example, deep features can be extracted using a CNN backbone, and long-term dependencies between similar local regions of the image can be modeled using a transformer backbone.
[0093] In other words, a resolution reduction process (for example, the resolution reduction process 614 of the neural network training system 600 in Figure 6) can be applied to each 2D reference image to form a corresponding lower-resolution medical image (for example, a lower-resolution medical image 616). In some embodiments, the resolution reduction process can be applied to different 2D reference images to form corresponding lower-resolution medical images. For example, a first resolution reduction process can be applied to a first 2D reference image to form a first corresponding lower-resolution medical image, a second resolution reduction process different from the first can be applied to a second 2D reference image to form a second corresponding lower-resolution medical image, and so on. Alternatively, a first set of multiple resolution reduction processes can be applied to a first 2D reference image to form a first corresponding lower-resolution medical image, a second set of multiple resolution reduction processes different from the first can be applied to a second 2D reference image to form a second corresponding lower-resolution medical image, and so on. In this way, by using and / or combining various processes, it is possible to form a wide range of lower-resolution medical images while maintaining a one-to-one correspondence with 2D reference images.
[0094] Furthermore, in some examples, a certain amount of random noise can be added to some or all of the lower-resolution medical images. For example, the random noise can be added using a random noise generator. The noise may include haze, blur, or other artifacts generated at different scales. The amount of random noise may vary across different lower-resolution medical images. For example, one lower-resolution medical image may have a first amount of noise, while other lower-resolution medical images may have a different amount of noise, which may be a larger or smaller amount.
[0095] In block 708, method 700 includes the step of generating a dataset of training pairs of 2D medical images, each training pair including a higher-resolution 2D reference image as the target image and a corresponding lower-resolution medical image derived from the 2D reference image through a resolution reduction process as the input image. In one embodiment, the input image and target image may be paired by a training data generator, such as the training data generator 610 of the neural network training system 600. Once the image pairs have been created, these image pairs can be divided into training image pairs and test image pairs, as described above with respect to Figure 6.
[0096] In block 710, method 700 includes the step of associating each training pair in the training pair dataset with a latent space representation of a reference image volume corresponding to the target image and input image of the training pair. The latent space representation may include features extracted from the associated reference image volume. The image enhancement CNN may take the latent space representation as additional input. In some embodiments, the latent space representation corresponding to a 2D reference image may be based on a latent space of image volumes stored in the memory of the image processing system (e.g., the training module 310 of the image processing system 302). The latent space may be generated as described above with respect to Figure 11. For example, each reference image volume may be input to an autoencoder CNN to generate a latent space for that reference image volume. Each latent space for each reference image volume may be stored in memory.
[0097] In one embodiment, the latent space representation for a related pair of images includes a first 3D portion of the latent space, which includes the target image of the pair. In another embodiment, the latent space representation can be generated by encoding a second 3D portion of the reference image volume, which includes the target image. For example, a 3D portion of the reference image volume can be extracted and fed into an autoencoder CNN to generate a latent space of the 3D portion of the reference image, and a set of 2D images extracted from the latent space that share the same viewpoint as the target image can be used as the latent space representation.
[0098] Now, moving to Figure 12, Figure 1200 of the encoder shows how the latent space representation 1210 can be generated from the image volume 1202 according to one embodiment. The 3D portion 1204 of the image volume 1202 is first extracted from the image volume 1202. In one embodiment, the boundary of the (3D portion) 1204 may be randomly selected during training. The 3D portion 1204 is then input to an encoder, which may be the encoder portion of a CNN, such as the autoencoder CNN described above. The 3D portion 1204 may be input to multiple convolutional layers of the encoder, including convolutional layers 1206 and 1208. The latent space representation 1210 may be the output of the convolutional layer 1208. In this way, the features of the 3D portion 1204 can be represented as a compressed form in the latent space representation 1210. After training, during the inference phase (described later with respect to Figure 8), the trained encoder can be used to generate the latent space of a new image.
[0099] Returning to Method 700, latent space representations corresponding to both the 2D reference (target) image and the lower-resolution (input) image can be input to the image-enhancing CNN during training, as described later. Including the corresponding latent space representation with each training image pair during training can improve the performance of the image-enhancing CNN when forming higher-resolution images.
[0100] In block 712, method 700 includes the step of training an image-enhanced CNN on training pairs containing latent spatial representations. More specifically, the step of training an image-enhanced CNN on training pairs includes training the image-enhanced CNN to learn to map lower-resolution input 2D medical images to higher-resolution target 2D medical images. In some embodiments, the image-enhanced CNN may include one or more convolutional layers, which in turn include one or more convolutional filters. The convolutional filters may include multiple weights, the values of which are learned during the training procedure. The convolutional filters may correspond to one or more visual features / patterns, thereby enabling the image-enhanced CNN to identify and extract features from medical images. In other embodiments, the image-enhanced CNN may be a different form of neural network, rather than a convolutional neural network.
[0101] The step of training an image-enhancing CNN on an image pair may include repeatedly inputting the input image of each training image pair into the input layer of the image-enhancing CNN. In some embodiments, the intensity value of each pixel of the input image may be input to a separate neuron in the input layer of the image-enhancing CNN. The image-enhancing CNN can map the input image to a corresponding target image by propagating the input image from the input layer through one or more hidden layers until it reaches the output layer of the image-enhancing CNN. In some embodiments, the output of the image-enhancing CNN includes a 2D matrix of values, where each value corresponds to a separate intensity of a pixel in the input image, and the separate intensity of each pixel in the output image forms an output image which is a higher-resolution version of the input image than the input image.
[0102] An image-enhancing CNN may have an encoder-decoder architecture that includes a first encoder portion and a second decoder portion. Latent space representations associated with training pairs can be input to the image-enhancing CNN in the input layer or in the initial feature maps of the second decoder portion. The input of latent space representations is illustrated in Figure 13 below with respect to the exemplary image-enhancing CNN in Figure 9.
[0103] An image-enhancing CNN may be configured to iteratively adjust one or more of its weights to minimize a loss function based on an assessment of the difference between an input image and a target image composed of each pair of training images. In one embodiment, the loss function is a mean absolute error (MAE) loss function, where the differences between the input image and the target image are compared and added pixel by pixel. In another embodiment, the loss function may be a structural similarity index (SSIM) loss function. In yet another embodiment, the loss function may be a minimax loss function, a Wasserstein loss function, or a different loss function. The examples provided herein are for illustrative purposes only, and it should be acknowledged that other forms of loss functions may be used without departing from the scope of this disclosure.
[0104] The weights and biases of an image-enhanced CNN can be adjusted based on the difference between the output image and the target (e.g., ground truth) image of a given pair of images. The difference (or loss) determined by the loss function can be backpropagated through the image-enhanced CNN to update the weights (and biases) of the convolutional layers. In some embodiments, the backpropagation of the loss can occur according to a gradient descent algorithm, where the gradient (first derivative, or an approximation of the first derivative) of the loss function is determined for each weight and bias of the image-enhanced CNN. Each weight (and bias) of the image-enhanced CNN is then updated by adding the negative number of the product of the gradients determined (or approximated) for the weight (or bias) at a predetermined step size. The weight and bias updates may be repeated until the weights and biases of the image-enhanced CNN converge, or until the rate of change of the weights and / or biases of the deep neural network in each iteration of weight adjustment falls below a threshold.
[0105] To avoid overfitting, the training of the image-enhanced CNN can be periodically interrupted to verify the performance of the image-enhanced CNN on parts of the training pair that are not used to train the image-enhanced CNN (e.g., a validation set). In one embodiment, the training of the image-enhanced CNN can be terminated when the performance of the image-enhanced CNN on a test pair converges (e.g., the error rate on the test set converges to or within a threshold value). Convergence can be determined by evaluating the trained image-enhanced CNN on the test pair. In this way, the image-enhanced CNN can be trained to generate copies of the input image with a higher resolution than the input image.
[0106] In some embodiments, the performance assessment of an image-enhanced CNN may include a combination of minimum error rate and quality assessment, i.e., the minimum error rate achieved for each pair of test image pairs, and / or one or more quality assessments, or different functions of other factors for assessing the performance of the image-enhanced CNN. The examples provided herein are for illustrative purposes only, and it should be acknowledged that other loss functions, error rates, quality assessments, or performance assessments may be included without departing from the scope of this disclosure.
[0107] Referring here to Figure 8, a flowchart of Method 800 is shown, which uses an image enhancement CNN such as the image enhancement CNN 602 in Figure 6 and / or the image enhancement CNN 900 in Figure 9 to form a higher-resolution version of a 2D medical image and display it in an enlarged window of a medical image examination application. Method 800 can be performed by a client device such as the client device 502 in Figure 5 of an imaging system such as the imaging systems 100 and 200 in Figures 1 and 2, respectively. Some operations of Method 800 can be stored in the memory of the client device (e.g., memory 524) and executed by the processor of the client device (e.g., processor 522). In various embodiments, the image enhancement CNN can be trained as described above with respect to Method 700 in Figure 7.
[0108] Method 800 begins in block 802, where Method 800 includes the step of receiving a bounding box (e.g., bounding box 404 in Figures 4(A) and 4(B)) from a medical image review application running on a user's computer, such as a radiologist. The user can use the medical image review application to review various 2D medical images extracted from image volumes acquired from a subject using an imaging system. Based on the review, the user can diagnose the condition of the subject. The bounding box contains a selection of the 2D medical image that is displayed in the medical image review application.
[0109] The bounding box may be drawn by the user or generated based on the cursor position in the medical image examination application in response to one or more selected control units of the medical image examination application. For example, the bounding box may be generated when the user selects the magnifying glass icon in the medical image examination application using an input device such as a mouse, and then selects the center point of the selection. Once the center point of the selection is selected, the medical image examination application can generate a bounding box centered on the selection.
[0110] In block 804, method 800 includes the step of cropping a 2D medical image based on a bounding frame. In this way, the selected portion of the 2D medical image can be included in the cropped 2D medical image, and the other portion of the 2D medical image outside the selection can be discarded. The (original) 2D medical image may remain on the screen inside the medical image examination application, and an enlarged version of the selection can be formed using the cropped 2D medical image.
[0111] In block 806, method 800 includes the step of retrieving a latent space representation (e.g., latent space representation 532) of the image volume corresponding to a 2D medical image. In the first embodiment, the latent space representation may be retrieved from the memory of the client device (e.g., memory 524). More precisely, the latent space representation may be generated from the latent space of the image volume previously transmitted from the computing device of the imaging system (e.g., computing device 504 in Figure 5). For example, the latent space of the image volume may be transmitted to the client device when the image volume is selected in a medical image examination application, or it may be transmitted to the client device together with the 2D medical image. The latent space of the image volume may be stored in the memory of the client device. The latent space representation may include 3D portions of different channels of the latent space of the image volume, these 3D portions containing the 2D medical image. This will be explained later with reference to Figure 13.
[0112] In a second embodiment, the latent space representation may be requested from a server (e.g., server 505) running in the computing device when a 2D medical image is requested. The server may extract the image volume from the computing device's memory (e.g., memory 534). The server may extract the 3D portion of the image volume that contains the 2D medical image, encode the 2D medical image using an encoder, and generate the latent space representation described with respect to Figure 12. The server may then send the latent space representation back to the client device.
[0113] In block 808, method 800 includes the step of inputting a cropped 2D medical image and a latent space representation into a trained image-enhancing CNN. The trained image-enhancing CNN may include an encoder portion and a decoder portion. In various embodiments, the step of inputting an acquired 2D medical image into a trained image-enhancing CNN includes inputting the image data of each pixel of the acquired 2D medical image into the corresponding node in the input layer of the encoder portion of the image-enhancing CNN. A first latent space representation generated according to the first embodiment described above may be input into the input layer of a second encoder portion of the image enhancement, where the image-enhancing CNN has an architecture similar to that described with respect to Figure 13. Alternatively, a second latent space representation generated according to the second embodiment described above may be input into the decoder portion described with respect to Figure 9.
[0114] The values of the image data can be multiplied by weights in the corresponding nodes and propagated through various hidden layers (e.g., convolutional layers) to the output layer of the image-enhancing CNN. The output layer may contain nodes corresponding to each pixel of the output 2D medical image, where the output 2D medical image is based on the image data output by each node. The output image may have a higher resolution than the input image.
[0115] In block 810, method 800 includes the step of displaying a higher resolution image output by a trained image-enhanced CNN in an enlarged window of a medical image scrutiny application. The higher resolution image may be displayed in real time while the user is scrutinizing the 2D medical image. For example, in response to the user selecting a selection area, the higher resolution image may be displayed on the screen without any intentional delay. The user may then select different parts of the 2D medical image, which may be displayed within the enlarged window of the screen without any intentional delay.
[0116] In some embodiments, the user can also choose to display conventional bilinear / cubic / spline interpolation in an enlarged window to ensure that artifacts are not formed in higher resolution images.
[0117] Referring to Figure 9, an exemplary image-enhanced CNN 900 is shown, which may be a non-restrictive example of the image-enhanced CNN 602 in Figure 6. The image-enhanced CNN 900 can take a 2D medical image 902a as input and form a corresponding higher-resolution 2D medical image 956b as output. The image-enhanced CNN 900 is divided into an encoding portion (descending portion, elements 902b to 930) and a decoding portion (ascending portion, feature maps 932 to 956a). The image-enhanced CNN 900 includes a series of mappings from a pixel representation of the input image 902b that can be received by the input layer 903, through multiple feature maps, to a pixel representation of the output image 956b that can ultimately be generated by the output layer 956a.
[0118] The various elements that make up the image-enhanced CNN900 are classified in Legend 958. As shown in Legend 958, the image-enhanced CNN900 includes multiple feature maps (and / or duplicated feature maps), each of which can receive input from a preceding feature map, transform / map the received input, and output a result that generates the next feature map. Each feature map may contain multiple neurons, and in some embodiments, each neuron can receive input from a subset of neurons in a preceding layer / feature map, calculate a single output based on the received input, and the output can be propagated to a subset of neurons in the next layer / feature map. Feature maps can be described using spatial dimensions such as length and width (which may correspond to the features of each pixel in the input image), where these dimensions refer to the number of neurons that make up the feature map (e.g., the number of neurons along the length and the number of neurons along the width of a given feature map).
[0119] In some embodiments, the neurons in the feature map can calculate an output by computing the dot product of the received inputs using a set of learned weights (each set of learned weights is also called a filter), where each received input has a unique corresponding learned weight, which is learned during the training of the CNN.
[0120] The transformations / mappings performed by each feature map are indicated by arrows, with each type of arrow corresponding to a separate transformation, as shown in Legend 958. A solid black arrow pointing to the right indicates a convolution, where the output from the grid of feature channels in the previous feature map is mapped to a single feature channel in the current feature map. Each convolution may be followed by an activation function, in one embodiment, which includes a normalized linear unit (ReLU).
[0121] The downward-pointing white arrow indicates pooling, where the maximum value from a 2x2 grid of feature channels is propagated from the previous feature map to a single feature channel in the current feature map, resulting in a reduction to one-eighth of the spatial resolution of the previous feature map. In some examples, this pooling occurs independently for each feature. The pooling can be either maximum pooling or average pooling.
[0122] In the decoder portion of the image-enhanced CNN900, upward-pointing white arrows can indicate an upscaling operation, which involves mapping the output from a single feature channel of the previous feature map to a 2x2 grid of feature channels in the current feature map, thereby increasing the spatial resolution of the previous feature map by an 8-fold. Although not shown in Figure 9, additional upscaling operations or upscaling layers may be included in CNN900.
[0123] The dashed arrow pointing to the right indicates the duplication and trimming of a feature map for concatenation with another feature map that will occur later. Trimming allows the dimensions of the duplicated feature map to match the dimensions of the feature map to which the duplicated feature map is concatenated. If the dimensions of the first duplicated feature map are equal to the dimensions of the second feature map to which the first feature map is concatenated, then trimming is not necessary.
[0124] A long, white, right-pointing triangular arrowhead indicates a 1x1 convolution, where each feature channel of the previous feature map is mapped to a single feature channel of the current feature map; in other words, a one-to-one mapping of feature channels occurs between the previous and current feature maps. In addition, a batch normalization operation can be performed, as indicated by a right-pointing white, arc-shaped arrowhead, where the activation distribution of the input feature map is normalized, and / or a dropout operation can be performed, as indicated by a short, white, right-pointing triangular arrowhead, where random or pseudo-random dropout of input neurons (and their inputs and outputs) may occur during training.
[0125] The feature maps of the image-enhanced CNN900 are illustrated as filled rectangles with height (the vertical length as shown in Figure 9, corresponding to the y-spatial dimension of the xy-plane), width (assumed to be equal in size to the height, corresponding to the x-spatial dimension of the xy-plane; not shown in Figure 9), and depth (the horizontal length as shown in Figure 9, corresponding to the number of features inside each feature channel). Similarly, the image-enhanced CNN900 includes open (unfilled) rectangles corresponding to duplicated and cropped feature maps, where the duplicated feature maps include height (the vertical length as shown in Figure 9, corresponding to the y-spatial dimension of the xy-plane), width (assumed to be equal in size to the height, corresponding to the x-spatial dimension of the xy-plane; not shown in Figure 9), and depth (the length from left to right as shown in Figure 9, corresponding to the number of features inside each feature channel).
[0126] In this way, the image-enhanced CNN900 can enable the mapping of a first image with a first resolution to a second image with a higher resolution. The image-enhanced CNN900 shows the feature map transformation that occurs as the first image propagates through the neuron layers of the convolutional neural network to form the second image. The weights (and biases) of the convolutional layers of the image-enhanced CNN900 are learned during training, as described with respect to Figure 7.
[0127] As explained with respect to Figure 7, a latent spatial representation of the image volume of the 2D medical image 902a may be additionally input to the image enhancement CNN 900 at the start of the decoder portion of the image enhancement CNN 900. For example, the latent spatial representation 960 may be input to feature map 932 along with the output of feature map 930. The latent spatial representation 960 may be the latent space of a portion of the image volume, such as the latent spatial representation 1210 in Figure 12. In one embodiment, the latent spatial representation 1210 is concatenated to the output of feature map 930, and the resulting concatenation is input to feature map 932.
[0128] In other embodiments, the latent space representation 960 may be input to the image-enhanced CNN 900 at the input layer 903. In such embodiments, the latent space representation 960 may be generated from the latent space of the image volume. For example, the latent space of the image volume may be generated as described with respect to Figure 11 and stored in the memory of the image processing system together with the image volume. The latent space representation 960 may include a 3D portion of the latent space containing the 2D medical image 902a. Furthermore, in such embodiments, the encoder-decoder architecture of Figure 9 may be modified to include an additional encoder portion, and the latent space representation 960 may be input to the input layer of the additional encoder portion as shown in Figure 13.
[0129] This disclosure deals with a neural network architecture that includes one or more regularization layers, including batch normalization layers, dropout layers, and other regularization layers known in the art of machine learning that can be used during training to mitigate overfitting and improve training efficiency while reducing training time. The regularization layers are used during CNN training and are deactivated or removed in the post-trained embodiment of the CNN. These layers may be distributed among the layer / feature maps shown in Figure 9, or they may replace one or more of the illustrated layer / feature maps.
[0130] The architecture and configuration of the image-enhancing CNN900 shown in Figure 9 are for illustrative purposes only, not limitations, and it should be understood that other suitable neural networks may be used here to improve the resolution of medical images. Other suitable neural networks may include more or fewer layers and feature maps, etc.
[0131] Referring now to Figure 13, a simplified alternative encoder-decoder architecture 1300 for an image-enhanced CNN, such as image-enhanced CNN900, is shown. The alternative encoder-decoder architecture 1300 includes a first encoder portion 1306, a second encoder portion 1308, and a decoder portion 1312. The first encoder portion 1306 and the second encoder portion 1308 may each be structurally the same as or similar to the encoder portion in Figure 9, and the decoder portion 1312 may also be structurally the same as or similar to the decoder portion in Figure 9.
[0132] An input image 1302 from a training pair may be input to the input layer of the first encoder section 1306. A latent space representation 1304 corresponding to the source image volume from which the input image 1302 is extracted may be input to the input layer of the second encoder section 1308. The latent space representation 1304 may contain multiple extracted 2D images from the latent space 1301 of the image volume (e.g., the compressed representation 1110 in Figure 11), where one 2D image may be extracted from each channel of the latent space 1301. The extracted 2D images may share the same viewpoint (e.g., camera position / angle) as the input image 1302.
[0133] The first output of the first encoder section 1306 and the second output of the second encoder section 1308 can be concatenated and input to the decoder section 1312. The decoder section 1312 can output image 1314, which may be a higher-resolution version of the input image 1302. The loss between image 1314 and the target image of the training pair can then be calculated and backpropagated through the image-enhanced CNN to adjust the weights and biases of the convolutional layers in the decoder section 1312, the first encoder section 1306, and the second encoder section 1308. In this way, the latent space representation 1304 can be encoded by different encoder sections than the input image 1302, and the outputs of both encoder sections can be combined in the decoder section 1312. The quality of image 1314 can be improved by including the 3D feature information stored in the latent space representation 1304 in the formation of the output image 1314 from the input image 1302.
[0134] Thus, users of medical examination software applications can observe desired anatomical structures in 2D medical images displayed in the application with greater detail than that provided in the original 2D medical image, by forming a super-resolution magnified version. This super-resolution image can then be used to diagnose the condition of the patient. By displaying a selected portion of the 2D medical image in a higher resolution within the magnified window, the anatomical features of the patient become clearer to the user, enabling a more accurate diagnosis and leading to better treatment outcomes for the patient. Furthermore, the ability to magnify and enhance medical images, along with the higher resolution provided by the magnified window, allows for a better understanding of complex medical conditions and leads to more informed treatment decisions. The ease of use of the magnified window, combined with increased accuracy in medical diagnosis, allows for faster diagnoses and more rapid treatment of medical conditions. Additionally, it can improve the overall efficiency of the imaging system, reduce the computational load on the system, and free up computing and memory resources for other tasks.
[0135] The technical effect of generating and displaying a portion of a 2D medical image in a higher resolution within an enlarged window is that the anatomical structures contained in that portion can be viewed in more detail, leading to a more accurate and timely diagnosis.
[0136] This disclosure also provides support for a method for a medical imaging system, the method comprising the step of displaying a selected portion of a first 2D medical image at a second resolution higher than the first resolution, while displaying a first two-dimensional (2D) medical image having a first resolution and formed from an image volume acquired through a medical imaging system in a medical examination software application running on a client device of the medical imaging system. In a first example of the method, the selected portion of the medical image includes an area defined by a bounding box generated on the first 2D medical image by a user of the client device. In a second example of the method, which optionally includes the first example, the step of displaying a selected portion of the medical image at a second resolution further comprises displaying a second 2D medical image having a second resolution, the second 2D medical image formed by a neural network trained to form a higher-resolution version of the 2D medical image, in an enlarged window superimposed on the 2D medical image. In a third example of a method that optionally includes one or both of the first and second examples, the step of displaying a selected portion of a medical image at a second resolution further includes inputting the cropped first 2D medical image into a trained neural network so as to crop the first 2D medical image based on a bounding frame to form a second 2D medical image. In a fourth example of a method that optionally includes one or more of the first to third examples or each of them, the step of displaying a selected portion of a medical image at a second resolution further includes inputting the first 2D medical image into a trained neural network so as to form a second 2D medical image, cropping the second 2D medical image based on a bounding frame, and displaying the cropped second 2D medical image in an enlarged window. In a fifth example of a method that optionally includes one or more of the first to fourth examples, or each of them, the method further includes the steps of: requesting a first 2D medical image from a server of a medical imaging system's computing device in a client device and receiving it from the server; and forming a second 2D medical image in the client device based on the first 2D medical image using a trained neural network.In the sixth example of the method, which optionally includes one or more of the first to fifth examples, the trained neural network is a convolutional neural network (CNN) having an encoder-decoder architecture including an encoder portion and a decoder portion. In the seventh example of the method, which optionally includes one or more of the first to sixth examples, the method further includes the step of inputting a latent spatial representation of an image volume as an additional input to the trained neural network so as to form a second 2D medical image. In the eighth example of the method, which optionally includes one or more of the first to seventh examples, the latent spatial representation includes a three-dimensional (3D) portion of the latent space of the image volume corresponding to the first 2D medical image, and this 3D portion is input to the decoder portion of the trained neural network. In the ninth example of the method, which optionally includes one or more of the first to eighth examples, the latent space representation includes multiple 2D images extracted from the latent space of the image volume, each of the multiple 2D images corresponding to a different channel in the latent space, the first 2D medical image being input to the first encoder portion of a trained neural network, and the multiple 2D images being input to a second encoder portion of the trained neural network, which is different from the first encoder portion. In the tenth example of the method, which optionally includes one or more of the first to ninth examples, the 2D images extracted from the latent space of the image volume have the same viewpoint as the first 2D medical image. In the eleventh example of the method, which optionally includes one or more of the first to tenth examples, the first output of the first encoder portion and the second output of the second encoder portion are concatenated and input to the decoder portion.
[0137] The disclosure also provides support for a medical imaging system comprising a computing device that stores the image volume of a subject in the medical imaging system, and a client device running a medical examination software application, the client device including a processor connected to the non-transient memory of the client device, the non-transient memory containing instructions, which, when executed, request a two-dimensional (2D) medical image of the image volume and a latent spatial representation of the image volume from the computing device, display the requested 2D medical image on the screen of the client device, and in response to the user of the client device selecting a magnification tool in the medical examination software application, receive a selected portion of the requested 2D medical image selected by the user, input the selected portion and the latent spatial representation of the image volume into a trained convolutional neural network (CNN) stored in the client device, receive a higher resolution 2D medical image corresponding to the selected portion as the output of the trained CNN, and cause the processor to display the higher resolution 2D medical image in the magnification window of the medical examination software application. In the first example of the system, the selected portion is input to the input layer of the encoder portion of the trained CNN, and the latent space representation is input to the first layer of the decoder portion of the trained CNN. In the second example of the system, which optionally includes the first example, the selected portion is input to the input layer of the first encoder portion of the trained CNN, the latent space representation is input to the second encoder portion of the trained CNN, and the first output of the first encoder portion and the second output of the second encoder portion are concatenated and input to the decoder portion of the trained CNN. In the third example of the system, which optionally includes one or both of the first and second examples, the selected portion is defined by a bounding box created on the screen by the user, and further instructions are stored in non-transient memory. When these instructions are executed, they cause the processor to crop the requested 2D medical image based on the bounding box, input the cropped 2D medical image into the trained CNN to form a higher-resolution 2D medical image, and display the cropped 2D medical image in an enlarged window.In a fourth example of a system that optionally includes one or more of the first to third examples, the selected portion is defined by a bounding box created on the screen by the user, and further instructions are stored in non-transient memory, which, when executed, cause the processor to input the requested 2D medical image into a trained CNN to form a higher resolution 2D medical image, crop the higher resolution 2D medical image based on the bounding box, and display the cropped higher resolution 2D medical image in an enlarged window.
[0138] This disclosure also provides support for a method for a medical imaging system, the method comprising the steps of: forming a plurality of 2D medical images from one or more reference image volumes stored in the memory of the medical imaging system; forming a lower-resolution version of each of the plurality of 2D medical images, wherein each plurality of image pairs comprises a 2D medical image of the plurality of 2D medical images as a target ground truth image and a corresponding lower-resolution version of the 2D medical image as an input image; and associating each image pair with a latent spatial representation of the reference image volume corresponding to the image pair. The method includes the steps of training a convolutional neural network (CNN) on image pairs and latent spatial representations to enhance the resolution of 2D medical images using image pairs, and deploying the trained CNN on a client device of a medical imaging system, wherein the trained CNN is configured to receive 2D medical images of image volumes and latent spatial representations of image volumes acquired from a subject in the medical imaging system as input, and to output a higher resolution version of the 2D medical image, which is displayed in an enlarged window of a medical examination software application running on the client device. In the first example of the method, the trained CNN is further configured to input the 2D medical images to a first encoder portion of the CNN and the latent spatial representations of image volumes to a second encoder portion of the trained CNN. In the second example of the method, which optionally includes the first example, the trained CNN is further configured to input the latent spatial representations of image volumes to a decoder portion of the trained CNN.
[0139] In describing the elements of the various embodiments of this disclosure, the terms singular indefinite article, definite article, "the" and "the foregoing" mean that there is one or more of those elements. The terms "first" and "second" do not indicate any order, quantity, or importance, but are used to distinguish one element from another. The terms "comprising," "including," and "having" are inclusive and mean that additional elements may exist in addition to those described. When terms such as "connected" and "combined" are used herein, one object (e.g., material, element, structure, and component) may be connected to or combined with another object, whether one object is directly connected to or combined with the other, or whether there is one or more intervening objects between the two objects. In addition, it should be understood that any reference in this disclosure to "one embodiment" or "one embodiment" is not to be construed as excluding the existence of additional embodiments that similarly incorporate the described features.
[0140] In addition to the modifications described above, many other variations and alternative configurations can be devised by those skilled in the art without departing from the spirit and scope of this description, and the following claims shall cover such modifications and configurations. Thus, while the information above has been specific and detailed in relation to the most practical and preferred viewpoints and those currently considered, it will be clear to those skilled in the art that many modifications, including but not limited to form, function, mode of operation, and usage, can be made without departing from the principles and concepts described herein. Furthermore, as used herein, the examples and embodiments are for illustrative purposes only in all respects and should not be construed as limiting in any way. [Explanation of Symbols]
[0141] 100 Computational Tomography (CT) Systems 102 Gantry 104 X-ray source 106 X-ray emission beam 108 detector array 112 subjects 114 Tables 200 Imaging Systems 202 detector element 204 subjects 206 Center of rotation 208 Control mechanism 213 214 Data Acquisition System (DAS) 232 Client device (display) 300 Medical Imaging Systems 400 First Workflow Diagram 402 First CT image 404, 454 Boundary frame 406 Specific part 408 High Contrast Range 410, 452 Second CT image 412, 456 Third CT image 430 trained image-enhanced DL models 440 First position 442 Second position 450 Second Workflow Diagram 500 Client-Server Architectures 600 Neural Network Training Systems How to train a 700-image enhanced CNN A method for creating and displaying higher-resolution versions of 800 2D medical images. 900 Image-Enhanced CNN 902b Input image 903 Input Layer 904, 906, 908, 910, 912, 913, 914, 916, 918, 919, 920, 922, 924, 926, 928, 930 Encoding part 932, 934, 936, 938, 940, 942, 944, 946, 948, 950, 952, 954 Decoded portion 956a Output layer 960 Latent space representation 1000 Medical Examination Applications 1001 Window 1002 2D Medical Images 1003 Upper boundary area 1004 Patient Identifier 1006 Control elements 1008 Magnifying glass icon 1010 Cursor 1012 Enlarged Window 1015 center point 1016 distance 1020 High Contrast Range 1100 Autoencoder CNN 1101 Encoder part 1102 Image volume 1103 Decoder section 1104 Reconstructed image volume 1106, 1108 Convolutional Layers 1110 Compressed representation 1112, 1114 Convolutional Layers 1200 encoder 1202 Image volume 1204 3D part 1206, 1208 Convolutional Layers 1210 Latent space representation 1300 Encoder-Decoder Architecture 1301 Latent space 1302 Input Image 1304 Latent space representation 1306 First Encoder Section 1308 Second Encoder Section 1310 1312 Decoder section 1314 Output image
Claims
1. A method for a medical imaging system (700, 800), A method comprising the step (810) of displaying a first two-dimensional (2D) medical image having a first resolution and formed from an image volume acquired through the medical imaging system in a medical examination software application running on a client device of the medical imaging system, while displaying a selected portion of the first 2D medical image, selected by the user of the client device, at a second resolution higher than the first resolution.
2. The method according to claim 1 (700, 800), wherein the selected portion of the first 2D medical image includes an area defined by a bounding box generated in the first 2D medical image by the user via the input device of the client device.
3. The method according to claim 2 (700, 800), wherein the step of displaying the selected portion of the first 2D medical image at the second resolution further includes displaying a second 2D medical image having the second resolution, which is formed by a neural network trained to form a higher-resolution version of the 2D medical image, in an enlarged window superimposed on the 2D medical image (810).
4. The step of displaying the selected portion of the first 2D medical image at the second resolution is: Based on the aforementioned boundary frame, the first 2D medical image is cropped (804), The cropped first 2D medical image is input into the trained neural network to form the second 2D medical image (808) The method according to claim 3 (700, 800), further comprising the above.
5. The step of displaying the selected portion of the first 2D medical image at the second resolution is: The first 2D medical image is input to the trained neural network to form the second 2D medical image. Based on the aforementioned boundary frame, the first 2D medical image is cropped (804), The cropped second 2D medical image is displayed in the enlarged window (810). The method according to claim 3 (700, 800), further comprising the above.
6. The client device requests the first 2D medical image from the server of the medical imaging system's computing device and receives it from the server. The steps of forming the second 2D medical image in the client device based on the first 2D medical image using the trained neural network, and The method according to claim 3, further comprising (700, 800).
7. The method according to claim 3 (700, 800), wherein the trained neural network is a convolutional neural network (CNN) having an encoder-decoder architecture including an encoder portion and a decoder portion.
8. The method according to claim 7 (700, 800), further comprising the step (808) of inputting a latent spatial representation of the image volume as an additional input to the trained neural network to form the second 2D medical image.
9. The method according to claim 8 (700, 800), wherein the latent space representation includes a three-dimensional (3D) portion of the latent space of the image volume corresponding to the first 2D medical image, and the 3D portion is input to the decoder portion of the trained neural network.
10. The method according to claim 8 (700, 800). The latent space representation includes a plurality of 2D images extracted from the latent space of the image volume, each of the plurality of 2D images corresponds to a different channel of the latent space, the first 2D medical image is input to a first encoder portion of the trained neural network, and the plurality of 2D images are input to a second encoder portion of the trained neural network, which is different from the first encoder portion.
11. The method according to claim 10 (700, 800), wherein the 2D image extracted from the latent space of the image volume has the same viewpoint as the first 2D medical image.
12. The method according to claim 10 (700, 800), wherein the first output of the first encoder portion and the second output of the second encoder portion are coupled and input to the decoder portion.
13. A computing device (216, 504) that stores the image volume of a subject in a medical imaging system (100, 200), A client device (502) running a medical examination software application (506), the client device (502) including a processor (522) that is in contact with the non-transient memory (524) of the client device (502) and A medical imaging system (100, 200) comprising, wherein the non-transient memory (524) contains an instruction, and when the instruction is executed, The computer (504) is requested to provide a two-dimensional (2D) medical image (402, 1002) of the image volume (536) and a latent spatial representation (532, 1210) of the image volume (536). The requested 2D medical images (402, 1002) are displayed on the screen (520) of the client device (502). In response to a user of the client device (502) selecting the expansion tool (510) of the medical examination software application (506), The selected portion of the requested 2D medical image (402, 1002) selected by the user is received. The selected portion and the latent spatial representation (532, 1210) of the image volume are input to the trained convolutional neural network (CNN) (514, 1100) stored in the client device (502). A higher resolution 2D medical image (1012) corresponding to the selected portion is received as the output of the trained CNN (514, 1100), A higher resolution 2D medical image (956b) is displayed in the enlarged window (1012) of the medical examination software application (506). A medical imaging system (100, 200) that causes the processor to perform the above-mentioned task.
14. The medical imaging system (100, 200) according to claim 13, wherein the selected portion is input to the input layer of the encoder portion (1101) of the trained CNN (514, 1100), and the latent space representation (532, 1210) is input to the first layer of the decoder portion (1103) of the trained CNN (514, 1100).
15. The selected portion is input to the input layer of the first encoder portion (1306) of the trained CNN (514, 1100), the latent space representation is input to the second encoder portion (1308) of the trained CNN (514, 1100), the first output of the first encoder portion (1306) and the second output of the second encoder portion (1308) are concatenated and input to the decoder portion (1312) of the trained CNN (514, 1100), the medical imaging system (100, 200) according to claim 13.