Method and apparatus for medical image reconstruction using machine learning-based process
By processing common mode images based on machine learning, a synthetic image and attenuation map of the entire field of view is generated, which solves the correction inaccuracy problem caused by the finite field of view in the nuclear imaging system, and achieves more accurate attenuation correction and organ volume calculation.
Patent Information
- Application Number
- CN202510115916.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-01-26
- Filing Date
- 2025-01-24
- Publication Date
- 2025-07-29
AI Technical Summary
Existing nuclear imaging systems have shortcomings in the attenuation correction process, especially due to inaccuracy caused by CT images with limited field of view, which affects the accuracy of organ volume calculations and medical measurements.
The common mode images are processed using a machine learning-based method to generate a final synthetic image with a wider field of view, and a corresponding attenuation map is generated to correct non-attenuation corrected images, and a segmentation mask and final synthetic image are generated in combination with multiple machine learning processes to realize attenuation correction and medical metric calculation of the full view.
Improve the accuracy of attenuation and scattering correction, ensure the accuracy of organ volume calculations, and enhance the integrity and accuracy of medical image reconstruction.
Smart Images

Figure CN120387971A_ABST
Abstract
Description
Technical Field
[0001] Aspects of the present disclosure generally relate to medical diagnostic systems, and more particularly to reconstructing images from nuclear imaging systems for diagnostic and reporting purposes. Background Art
[0002] Nuclear imaging systems can employ various techniques to capture images. For example, some nuclear imaging systems use positron emission tomography (PET) to capture images. PET is a nuclear medicine imaging technique that produces a nuclear image representing the distribution of positron-emitting isotopes within the body. As another example, some nuclear imaging systems use single photon emission computed tomography (SPECT) to capture images. SPECT produces a three-dimensional image of the distribution of a radioactive tracer that is injected into a person's bloodstream and subsequently absorbed by certain tissues. For instance, some nuclear imaging systems employ computed tomography (CT) as a co-modality. CT is an imaging technique that uses x-rays to produce anatomical images. Magnetic resonance imaging (MRI) is another imaging technique that uses magnetic fields and radio waves to generate anatomical and functional images. Some nuclear imaging systems combine images from PET or SPECT and CT scanners during an image fusion process to produce an image that shows information from both the PET or SPECT scan and the CT scan (e.g., PET / CT systems, SPECT CT). Similarly, some nuclear imaging systems combine images from PET or SPECT and MRI scanners to produce an image that shows information from both the PET or SPECT scan and the MRI scan.
[0003] Typically, these nuclear imaging systems capture measurement data and use mathematical algorithms to process the captured measurement data to reconstruct medical images. The measurement data generally requires correction for some photons that have been lost to a sinogram bin (i.e., attenuation correction) or have been misassigned to another sinogram bin (i.e., scatter correction). In some cases, the system generates an attenuation map for attenuation and scatter correction. For example, the system can use a mapping algorithm to generate an attenuation map based on a CT image. The system can then generate a final reconstructed medical image based on the captured measurement data and the attenuation map. However, these attenuation correction processes can suffer from drawbacks.
[0004] For example, to accurately correct non-attenuation corrected (NAC) PET images for attenuation, an attenuation map covering all possible lines of response (LORs) is required. However, the attenuation map generated for a CT image may cover only a subset of the LORs, at least because the CT image is scanned with a limited field of view (FOV). For example, as opposed to generating a CT image that covers the entire patient (e.g., from head to toe), the patient may have been scanned over a specific region to generate the CT image. Thus, the attenuation map may inaccurately correct the NAC PET that may cover the entire patient.
[0005] In addition, these limited FOV CT images may not allow or may reduce the accuracy of various medical metrics, such as those based on organ volume. For example, a CT image that omits a portion of an organ may not allow for an accurate calculation of the organ contribution percentage and may limit the comparison of the calculated contribution percentage with those calculated for other scans (e.g., follow-up or population scans). Thus, there is an opportunity to address these and other deficiencies in the attenuation correction process in nuclear imaging systems. SUMMARY OF THE INVENTION
[0006] Systems and methods are disclosed for correcting non-attenuation corrected (NAC) images (e.g., NAC PET images) based on co-modal images (e.g., computed tomography (CT) images). The systems and methods employ a machine learning-based process that processes the co-modal image with a limited field of view to generate a final synthetic image with a wider field of view (FOV) (e.g., a full view of the subject), and may also generate a corresponding attenuation map. The generated attenuation map can be used to correct the NAC image. For example, one or more machine learning processes can be applied to a partial CT image (e.g., a CT image that captures a portion of a patient) to expand the partial CT image to a "full" CT image that covers the entire patient (e.g., covers the patient from head to toe). In addition, the full CT image can be used to calculate various medical metrics of various organs (e.g., tissues, bones, etc.). For example, the percentage contribution of an organ can be defined as the metabolic tumor volume divided by the volume of the corresponding organ. Since the full CT image includes the entire organ (e.g., the full CT image does not omit any portion of the organ), the percentage contribution of the organ can be accurately calculated based on the full CT image. For example, the bone tumor burden can be calculated relative to the bone region of the patient, which can be fully captured within the full CT image.
[0007] In some embodiments, a computer-implemented method includes: receiving a co-modal image, a non-attenuation-corrected kernel image, and an x-ray image. The method further includes applying a first machine learning process to the co-modal image and generating location data identifying a feature location based on the application of the first machine learning process. Additionally, the method includes applying a second machine learning process to the non-attenuation-corrected kernel image, the x-ray image, and the location data and generating a segmentation mask based on the application of the second machine learning process. The method further includes applying at least a third machine learning process to the segmentation mask, a portion of the co-modal image, the non-attenuation-corrected kernel image, and the x-ray image and generating a plurality of synthetic images based on the application of the third machine learning process. The method further includes applying at least a fourth machine learning process to the plurality of synthetic images and generating a final synthetic image based on the application of the at least fourth machine learning process.
[0008] In some embodiments, a non-transitory computer-readable medium stores instructions that, when executed by at least one processor, cause the at least one processor to perform operations including receiving a co-modal image, a non-attenuation-corrected kernel image, and an x-ray image. The operations further include applying a first machine learning process to the co-modal image and generating location data identifying a feature location based on the application of the first machine learning process. Additionally, the operations include applying a second machine learning process to the non-attenuation-corrected kernel image, the x-ray image, and the location data and generating a segmentation mask based on the application of the second machine learning process. The operations further include applying at least a third machine learning process to the segmentation mask, a portion of the co-modal image, the non-attenuation-corrected kernel image, and the x-ray image and generating a plurality of synthetic images based on the application of the third machine learning process. The operations further include applying at least a fourth machine learning process to the plurality of synthetic images and generating a final synthetic image based on the application of the at least fourth machine learning process.
[0009] In some embodiments, a system includes a memory device storing instructions and at least one processor communicatively coupled to the memory device. The at least one processor is configured to execute the instructions to perform operations including receiving a co-modal image, a non-attenuation-corrected kernel image, and an x-ray image. The operations further include applying a first machine learning process to the co-modal image and generating location data identifying a feature location based on the application of the first machine learning process. Additionally, the operations include applying a second machine learning process to the non-attenuation-corrected kernel image, the x-ray image, and the location data and generating a segmentation mask based on the application of the second machine learning process. The operations further include applying at least a third machine learning process to the segmentation mask, a portion of the co-modal image, the non-attenuation-corrected kernel image, and the x-ray image and generating a plurality of synthetic images based on the application of the third machine learning process. The operations further include applying at least a fourth machine learning process to the plurality of synthetic images and generating a final synthetic image based on the application of the at least fourth machine learning process.
[0010] In some embodiments, a device includes means for receiving a co-modal image, a non-attenuation corrected nuclear image, and an x-ray image. The device further includes means for applying a first machine learning process to the co-modal image and generating position data identifying feature locations based on the application of the first machine learning process. Additionally, the device includes means for applying a second machine learning process to the non-attenuation corrected nuclear image, the x-ray image, and the position data and generating a segmentation mask based on the application of the second machine learning process. The device further includes means for applying at least a third machine learning process to the segmentation mask, a portion of the co-modal image, the non-attenuation corrected nuclear image, and the x-ray image and generating a plurality of synthetic images based on the application of the third machine learning process. The device also includes means for applying at least a fourth machine learning process to the plurality of synthetic images and generating a final synthetic image based on the application of the at least fourth machine learning process. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The following will be apparent from the elements of the various figures, which are provided for illustrative purposes and are not necessarily drawn to scale.
[0012] Figure 1 Illustrates an example of a nuclear image reconstruction system according to some embodiments.
[0013] Figure 2A 、 2B 、3A, 3B, 4A, and 4B illustrate exemplary portions of a Figure 1 nuclear image reconstruction system according to some embodiments.
[0014] Figure 5 Illustrates a block diagram of an example computing device that can perform one or more of the functions described herein.
[0015] Figure 6 is a flowchart of an example method of reconstructing an image according to some embodiments.
[0016] Figure 7 is a flowchart of an example method of generating and fusing synthetic images according to some embodiments. DETAILED DESCRIPTION
[0017] This description of the example embodiments is intended to be read in conjunction with the accompanying drawings, which are considered to be a part of the entire written description. Individuals identified with masculine, feminine, or other gender identities are included within the term regardless of the use of grammatical terms.
[0018] Exemplary embodiments are described with respect to the claimed system and with respect to the claimed method. Additionally, exemplary embodiments are described with respect to methods and systems for image reconstruction and with respect to methods and systems for training functions for image reconstruction. Features, advantages, or alternative embodiments herein may be assigned to other claimed subject matter and vice versa. For example, claims providing a system may be improved with features described or claimed in the context of a method and vice versa. Additionally, the functional features of the described or claimed method are embodied by the target units providing the system. Similarly, claims for methods and systems for training an image reconstruction function may be improved with features described or claimed in the context of methods and systems for image reconstruction and vice versa.
[0019] Various embodiments of the present disclosure may employ machine learning methods or processes to provide clinical information from a nuclear imaging system. For example, embodiments may employ machine learning methods or processes to reconstruct an image based on captured measurement data and provide the reconstructed image for clinical diagnosis. In some embodiments, the machine learning methods or processes are trained to improve the reconstruction of the image.
[0020] Various embodiments of the present disclosure may employ machine learning processes to reconstruct an image based on image data captured by a nuclear imaging system. For example, embodiments may employ machine learning processes applied to different types of image data (e.g., CT image data, NAC PET measurement data, topogram data, optimal image data, etc.) to generate a final composite image (e.g., a final composite CT image), and in some examples, generate a corresponding organ mask. Although each of the different types of image data may capture a limited view of a subject (e.g., a patient), the final composite image captures a more comprehensive view of the subject. For example, while a NAC PET image may capture the entire view of a patient (e.g., head to toe), the corresponding co-modal image may capture only a portion of the patient (e.g., lower body, mid-section, upper body, etc.). As is known in the art, a co-modal image is another image received in addition to the original image (e.g., NAC PET image), such as a CT or MR image. Thus, an attenuation map generated based only on the co-modal image will not cover all parts of the subject like the NAC PET image. One reason for the limited view of the co-modal image is to reduce the amount of dose given to the patient.
[0021] Embodiments described herein can “expand” a co-modal image by generating a final composite image that includes a wider view of the patient, which in some examples is the same view as the NAC PET image. For example, in addition to the NAC PET (or NAC SPECT) image, one or more additional images can be received. The additional images can include, for example, co-modal images (such as CT images), x-ray images (such as scout views), and / or optical images. When considered together, the views of the one or more additional images can correspond to the view of the NAC PET image. For example, while the NAC PET image can cover the entire subject (e.g., from head to toe), the co-modal image can cover only a portion of the subject. Similarly, each of the x-ray image and / or the optimal image can cover a portion of the subject. The portions of the subject covered by each of the co-modal, x-ray, and optical images can together cover the entire subject. As further described herein, one or more machine learning processes can be applied to the NAC PET image and the additional images to generate a final composite image that also covers the entire subject.
[0022] In some examples, an attenuation map (e.g., a “full view” attenuation map) can be generated based on the final composite image. Additionally, the NAC PET image can be corrected for attenuation and scatter based on the generated attenuation map. The generated attenuation map can cover all lines of response, and thus provides a more complete and accurate attenuation and scatter correction of the NAC PET image compared to conventional methods.
[0023] Furthermore and as described herein, one or more machine learning processes can be applied to the NAC PET image and the additional images to generate a segmentation mask (e.g., an organ mask). As recognized in the art, a segmentation mask can identify and / or isolate portions of an image. For example, a segmentation mask can identify and / or isolate portions of an image that include an organ. The segmentation mask can allow for a more accurate calculation of medical metrics, such as organ contribution percentages, because the segmentation mask will identify the corresponding organ as a whole.
[0024] Now referring to the accompanying drawings, Figure 1FIG. 0 illustrates an embodiment of a nuclear imaging system 100. As illustrated, the nuclear imaging system 100 includes an image scanning system 102, an image reconstruction system 104, and one or more data repositories 116. The image scanning system 102 may be a nuclear image scanner, such as a PET / CT scanner or a SPECT / CT scanner. The image scanning system 102 may scan a subject (e.g., a person), and based on the scan, generate NAC measurement data 105 representative of a nuclear scan. The NAC measurement data 105 may be, for example, non-attenuation corrected PET or SPECT raw data, such as sinogram data. The NAC measurement data 105 may represent anything imaged within the FOV of the scanner, which for a PET scanner includes a positron emitting isotope, and for a SPECT scanner includes gamma rays. For example, the NAC measurement data 105 may represent a whole body image scan, such as an image scan from the patient's head to the patient's thigh or toes. The image scanning system 102 may transmit the NAC measurement data 105 to the image reconstruction system 104.
[0025] The image scanning system 102 may also generate additional images. For example, the image scanning system 102 may generate CT image data 103 representative of a CT scan of the patient. For example, the CT scan may have a limited FOV compared to the nuclear scan represented by the NAC measurement data 105. The image scanning system 102 may also capture x-rays of the patient, and may generate scout image data 107 representative of each x-ray. While the CT image data 103 may represent a three-dimensional scan of at least a portion of the patient, the scout image data 107 represents a two-dimensional scan of at least a portion of the patient. In some cases, the image scanning system 102 may capture an optical image (e.g., a 2D or 3D optical image) of at least a portion of the patient, and may generate optical image data 109 representative of each optical image. For example, the optical image may be captured by one or more cameras (e.g., 2D or 3D cameras) of the image scanning system 102. The image scanning system 102 may transmit each of the additional images, including the CT image data 103, the scout image data 107, and the optical image data 109, to the image reconstruction system 104.
[0026] As illustrated, the image reconstruction system 104 includes a synthesis engine 110, an attenuation map generation engine 120, and an image volume reconstruction engine 130. Discussed further below Figure 2A 、 2B 、3A, 3B, 4A, and 4B illustrate Figure 1Exemplary portion of the synthesis engine 110. All or portions of each of the synthesis engine 110, the attenuation map generation engine 120, and the image volume reconstruction engine 130 may be implemented in hardware, such as in one or more field programmable gate arrays (FPGAs), one or more application specific integrated circuits (ASICs), one or more state machines, one or more computing devices, digital circuits, or any other suitable circuitry. In some examples, portions or all of the synthesis engine 110, the attenuation map generation engine 120, and the image volume reconstruction engine 130 may be implemented in software as executable instructions such that when executed by one or more processors, cause the one or more processors to perform the corresponding functions as described herein. The instructions may be stored in, for example, a non-transitory computer readable storage medium and read by one or more processors to perform the functions.
[0027] Reference Figure 1 , the synthesis engine 110 may receive NAC measurement data 105 and any additional images, including one or more of CT image data 103, scout image data 107, and optical image data 109. The synthesis engine 110 may perform any of the operations described herein to generate predicted CT image data 111 representative of a final synthesized CT image, and a corresponding final segmentation mask 113 that identifies organs of the subject. For example, the synthesis engine 110 may receive NAC measurement data 105 along with one or more of CT image data 103, scout image data 107, and optical image data 109, and may apply one or more trained machine learning processes to the NAC measurement data 105 along with the CT image data 103, scout image data 107, and optical image data 109 to generate predicted CT image data 111 and the final segmentation mask 113.
[0028] While the CT image data 103, the scout image data 107, and the optical image data 109 may capture a limited view of the subject relative to the nuclear scan view characterized by the NAC measurement data 105, the final synthesized CT image represents a view of the subject that is similar (e.g., equivalent) to the nuclear scan view. For example, the NAC measurement data 105 may characterize a full view of the subject (e.g., head to thigh, head to toe). In some examples, each of the CT image data 103 and the scout image data 107 may cover a view of a corresponding portion of the subject, which may or may not overlap with each other. The portions of the subject's view covered by the CT image data 103 and the scout image data 107, at least combined, may cover the full view of the subject. For example, the CT image data 103 may cover from the subject's head to the subject's chest region, while the scout image data 107 may cover from the subject's chest region to the subject's toes.
[0029] Figure 2A , 2B , 3A, 3B, 4A, and 4B illustrate exemplary portions of the synthesis engine 110. For example, Figure 2A illustrates a landmark prediction engine 256 that receives CT image data 103 and applies a landmark algorithm to the CT image data 103 to generate position data 257 that characterizes the position of anatomical landmarks. The position data 257 can define points within the CT image data 103 that define each organ, organ feature, etc. The landmark algorithm can be a machine learning model trained on various medical images across various modalities, such as the Automatic Landmark and Parsing of Human Anatomy (ALPHA) model, Siemens proprietary technology. The segmentation mask generation engine 254 receives the position data 257 from the landmark prediction engine 256 and also receives either the NAC measurement data 105 and any of the received additional images, such as the images characterized by the scout data 107 and the optical image data 109. As described herein, either the NAC measurement data 105 and any of the received additional images (such as the images characterized by the scout data 107 and the optical image data 109) can together capture a full view of the subject.
[0030] In addition, the segmentation mask generation engine 254 applies a segmentation process to the NAC measurement data 105, the received scout data 107 and optical image data 109, and the position data 257 to generate an initial segmentation mask 203. The segmentation process can be based on a trained machine learning model, such as a trained encoder-decoder model, a trained deep learning network (e.g., a trained convolutional neural network (CNN)), or any other suitably trained machine learning model. The segmentation mask generation engine 254 can generate features based on the NAC measurement data 105, the received scout data 107 and optical image data 109, and the position data 257 and can input the features into the trained machine learning model. Based on the input features, the trained machine learning model can output the initial segmentation mask 203. In some examples, supervised learning is used to train the machine learning model. For example, the segmentation process can be trained with, e.g., labeled PET measurement data, labeled scout data, and corresponding position data.
[0031] In some examples, the optical image data 109 can characterize an image of the full view of the subject. In these instances, as Figure 2BAs illustrated, the segmentation mask generation engine 202 may apply a segmentation process to the optical image data 109 to generate an initial segmentation mask 203. For example, the segmentation process may be based on a trained deep learning network, such as a trained CNN. The segmentation mask generation engine 202 may generate features based on the optical image data 109 and may input the features into the trained deep learning network. Based on the input features, the trained deep learning network may output the initial segmentation mask 203. In some examples, the deep learning network is trained based on labeled optical images. For example, the optical image may include labels identifying the contours of various organs captured within the optical image.
[0032] In some examples, when the optical image data 109 is received, the synthesis engine 110 determines to use the process described with respect to Figure 2B to generate the initial segmentation mask 203. Otherwise, if the optical image data 109 is not received, the synthesis engine 110 determines to use the process described with respect to Figure 2A to generate the initial segmentation mask 203.
[0033] Return reference Figure 1 , one or more trained machine learning processes may be applied to the NAC measurement data 105 and the initial segmentation mask 203 to generate a corresponding synthetic image (e.g., a predicted synthetic CT image). One or more trained machine learning processes may also be applied to each combination of the CT image data 103 and the initial segmentation mask 203, the localization image data 107 and the initial segmentation mask 203, and optionally the optical image data 109 and the initial segmentation mask 203 to generate a corresponding synthetic image. Each of the generated synthetic images may capture a predetermined view of the subject, such as a full view of the subject.
[0034] For example, Figure 3A illustrates a CT-based extrapolation engine 302 that applies a trained machine learning process to the initial segmentation mask 203 and the CT image data 103 to generate a first synthetic image 303. For example, the machine learning process is trained to extrapolate a CT image from a partial view to a predetermined view (e.g., head to chest, chest to toes, full view, etc.). The trained machine learning process may be based on a deep learning model, such as a diffusion model, a generative adversarial network (GAN), or a CNN. For example, the CT-based extrapolation engine 302 may generate features based on the initial segmentation mask 203 and the CT image data 103 and may input the features into the trained deep learning model. Based on the input features, the trained deep learning model outputs the first synthetic image 303.
[0035] In addition, the optical image-based synthesis generation engine 304 applies a trained machine learning process to the initial segmentation mask 203 and the optical image data 109 to generate a second synthesized image 305. For example, the trained machine learning process can be based on an encoder-decoder model or a deep learning model, such as a diffusion model, a GAN, or a CNN. For example, the optical image-based synthesis generation engine 304 can generate features based on the initial segmentation mask 203 and the optical image data 109, and can input the features into the trained deep learning model. Based on the input features, the trained deep learning model outputs the second synthesized image 305.
[0036] In addition, the PET-based extrapolation engine 306 applies a trained machine learning process to the initial segmentation mask 203 and the NAC measurement data 105 to generate a PET extrapolation image 307 that covers a predetermined view of the subject (e.g., head to chest, chest to toes, full view, etc.). For example, the machine learning process is trained to extrapolate a nuclear image from a partial view to a predetermined view. The trained machine learning process can be based on a deep learning model, such as a diffusion model, a generative adversarial network (GAN), or a CNN.
[0037] The PET-based synthesis generation engine 308 can receive the PET extrapolation image 307, and can apply an additional trained machine learning process to the PET extrapolation image 307 to generate a third synthesized image 309. The additional trained machine learning process can be based on a deep learning model, such as a diffusion model, a generative adversarial network (GAN), or a CNN. For example, the PET-based synthesis generation engine 308 can generate features based on the PET extrapolation image 307, and can input the features into the additional trained deep learning model. Based on the input features, the additional trained deep learning model outputs the third synthesized image 309.
[0038] In addition, the 2D to 3D conversion engine 310 applies a 2D to 3D conversion process to the scout image data 107 to convert the scout image data 107 into a 3D optical image 311. The 2D to 3D conversion process is configured (e.g., trained) to convert a 2D image into a 3D image. For example, the 2D to 3D conversion process can be based on a 2D to 3D conversion algorithm, a computer vision model, or a deep learning model, such as a diffusion model, a generative adversarial network (GAN), or a CNN.
[0039] The localization-image-based extrapolation engine 312 can receive the 3D optical image 311 and the initial segmentation mask 203, and can apply a trained machine learning process to the 3D optical image 311 and the initial segmentation mask 203 to generate a fourth synthetic image 313. The trained machine learning process can be based on a deep learning model, such as a diffusion model, a generative adversarial network (GAN), or a CNN, and is trained to extrapolate the 3D optical image 311 to a predetermined view of the subject based on the initial segmentation mask 203. For example, the localization-image-based extrapolation engine 312 can generate features based on the 3D optical image 311 and the initial segmentation mask 203, and can input the features into the trained deep learning model. Based on the input features, the trained deep learning model outputs the fourth synthetic image 313.
[0040] As illustrated, in some examples, the synthesis engine 110 stores one or more of the generated first synthetic image 303, second synthetic image 305, third synthetic image 309, and fourth synthetic image 313 in the data repository 116.
[0041] Return reference Figure 1 , one or more trained machine learning processes can be applied to the synthetic images to fuse the synthetic images and generate predicted CT image data 111 representative of a final synthetic image. In some instances, the location data is generated based on the CT image data 103 or can be received from the image scanning system 102. The location data can identify anatomical landmarks within the CT image data 103, such as organ feature locations. One or more trained machine learning processes can be applied to the location data and the synthetic images to fuse the synthetic images and generate the predicted CT image data 111. In some examples, the synthesis engine 110 can perform operations to register the predicted CT image data 111 to the NAC measurement data 105 and / or the CT image data 103, thereby generating a corresponding final registered synthetic image (i.e., the predicted CT image data 111 registered to the NAC measurement data 105 and / or the predicted CT image data 111 registered to the CT image data 103).
[0042] For example and with reference Figure 3B, the CT mask generation engine 322 receives the CT image data 103 and generates CT mask data 323 representing a segmentation mask (e.g., an organ mask) based on the CT image data 103. The CT mask generation engine 322 can apply any suitable segmentation algorithm known in the art (e.g., a trained machine learning model) to the CT image data 103 to generate the segmentation mask. As described herein, the CT image data 103 can represent a CT image with a partial view of the subject (e.g., the image scanning device 102 captures a CT scan using a FOV with a partial view of the subject). Similarly, the CT mask data 323 can represent the segmentation mask with the corresponding partial view of the subject.
[0043] In addition, the synthetic image fusion engine 326 receives the CT mask data 323 from the CT mask generation engine 322, and each of the first synthetic image 303, the second synthetic image 305, the third synthetic image 309, and the fourth synthetic image 313 (e.g., from the localization image-based extrapolation engine 312). The synthetic image fusion engine 326 performs operations to fuse each of the first synthetic image 303, the second synthetic image 305, the third synthetic image 309, and the fourth synthetic image 313 (e.g., the predicted synthetic images) to generate predicted CT image data 111 representing a final predicted synthetic image (e.g., a single predicted synthetic CT image), and performs further operations to generate a final segmentation mask 113.
[0044] For example, the synthetic image fusion engine 326 can apply a trained machine learning process to the CT mask data 323 and the first synthetic image 303, the second synthetic image 305, the third synthetic image 309, and the fourth synthetic image 313 to generate the predicted CT image data 111 representing the final predicted synthetic image. For example, the trained machine learning process can be based on a deep learning model (e.g., a diffusion model, a GAN, a CNN).
[0045] The machine learning process can be trained to automatically identify and extract relevant features from each modality, such as the implant detected from the fourth synthetic image 313 (generated based on the scout data 107), and generate a final predicted synthetic image based on the relevant features. For example, the machine learning process can be trained based on labeled synthetic images of various modalities (e.g., PET, optical, scout, CT, etc.). During training, the machine learning process can learn the values of various weights applied to the image data of various modalities, as well as the weights applied to various relevant features detected within the image data. Once trained, the synthetic image fusion engine 326 can input the features generated from the CT mask data 323 and the first synthetic image 303, the second synthetic image 305, the third synthetic image 309, and the fourth synthetic image 313 into the trained machine learning process, and the trained machine learning process can apply the learned weights to the corresponding input features to generate the predicted CT image data 111.
[0046] In addition, the synthetic image fusion engine 326 can apply one or more segmentation processes to the predicted CT image data 111 to generate the final segmentation mask 113. The segmentation process can be based on a trained machine learning model, such as a trained encoder-decoder model, a trained deep learning network (e.g., a trained convolutional neural network (CNN)), or any other appropriately trained machine learning model. For example, the machine learning model can be trained on labeled synthetic images (e.g., synthetic CT images labeled to identify organs and / or organ features (e.g., organ contours, dimensions, etc.)).
[0047] Figure 4A An example of the synthetic image fusion engine 326 is illustrated, which employs a trained encoding-decoding process (e.g., an autoencoding process) to generate the predicted CT image data 111 and the final segmentation mask 113. As illustrated, the encoding engine 402 receives the first synthetic image 303, the second synthetic image 305, the third synthetic image 309, the fourth synthetic image 313, and the CT mask data 323, and applies an encoding process thereto. Based on the application of the encoding process to the synthetic images and the CT mask data 323, the encoding engine 402 generates encoded features 403 and provides the encoded features 403 to each of the CT decoding engine 404 and the segmentation mask decoding engine 406.
[0048] The CT decoding engine 404 applies the decoding process to the encoded feature 403 to generate the predicted CT image data 111. Similarly, the segmentation mask decoding engine 406 applies the decoding process to the encoded feature 403 to generate the final segmentation mask 113. In some examples, for feature propagation purposes among other reasons, one or more skip connections are provided from the encoding engine 402 to one or both of the CT decoding engine 404 and the segmentation mask decoding engine 406. The skip connection allows the encoded features from one layer of the encoder of the encoding engine 402 to be provided to another layer (e.g., the corresponding layer) of the decoder of each decoding engine 404, 406.
[0049] Figure 4B An alternative example of the synthetic image fusion engine 326 is illustrated. In this example, the feature extraction engine 412 receives the first synthetic image 303, the second synthetic image 305, the fourth synthetic image 313, and the CT mask data 323, and applies a feature extraction process to extract features from one or more of the first synthetic image 303, the second synthetic image 305, the fourth synthetic image 313, and optionally the CT mask data 323. For example, the feature extraction engine 412 may apply the feature extraction process to the fourth synthetic image 313 (e.g., the localization-like predicted synthetic image), and based on the application of the feature extraction process, identify the implant within the fourth synthetic image 313. Similarly, for example, the feature extraction engine 412 may apply the feature extraction process to the second synthetic image 305 (e.g., the optical image predicted synthetic image), and based on the application of the feature extraction process, identify the organ position (e.g., the major organ position) within the second synthetic image 305, which may be outside the field of view of the first synthetic image 303 and the third synthetic image 309. The feature extraction engine 412 provides the feature data 413 representing the extracted features to the feature modification engine 414.
[0050] The feature modification engine 414 receives the feature data 413 and the third synthetic image 309 (generated based on the NAC measurement data 105), and modifies the feature data 413 based on the third synthetic image 309. For example, the feature modification engine 414 may apply a trained machine learning process (e.g., a process based on CNN or GAN) to the feature data 413 and the third synthetic image 309, and generate the modified features based on the application of the trained machine learning process. The feature modification engine 414 then generates the predicted CT image data 111 based on the modified features.
[0051] In addition, the segmentation mask generation engine 416 receives the predicted CT image data 111 from the feature modification engine 414 and applies a segmentation process to the predicted CT image data 111 to generate a final segmentation mask 113. The segmentation process can be based on a trained machine learning model, such as a trained encoder-decoder model, a trained deep learning network (e.g., a trained convolutional neural network (CNN)), or any other appropriately trained machine learning model. For example, a machine learning model can be trained on labeled synthetic images (e.g., synthetic images labeled to identify organs and / or organ features (e.g., organ contours, dimensions, etc.)).
[0052] Return reference Figure 3B , in some examples, the synthetic image fusion engine 326 performs operations to register the predicted CT image data 111 to the nuclear image characterized by the NAC measurement data 105. For example, the synthetic image fusion engine 326 can apply a registration process (e.g., an algorithm) to the NAC measurement data 105 and the predicted CT image data 111 to align the NAC measurement data 105 and the predicted CT image data 111 according to anatomical features (e.g., alignment of corresponding organs), and generate PET registration CT image data 327 characterizing the aligned NAC measurement data 105 and predicted CT image data 111.
[0053] Return reference Figure 1 , the attenuation map generation engine 120 receives the extended CT image data 111 from the synthesis engine 110 and generates an attenuation map 121 based on the extended CT image data 111. For example, the attenuation map generation engine 120 can apply a trained neural network process (e.g., a deep convolutional neural network process) to the extended CT image data 111 to generate the attenuation map 121.
[0054] In addition, the image volume reconstruction engine 130 can perform a process to correct the nuclear image characterized by the NAC measurement data 105 based on the attenuation map 121. For example, the image volume reconstruction engine 104 can perform an ordered subset expectation maximization (optionally utilizing time-of-flight and / or point spread function) process or a filtered backprojection process on the NAC measurement data 105 and the attenuation map 121 to generate a final image volume 191. For example, the final image volume 191 can include image data that can be provided for display and analysis. In some examples, the image volume reconstruction engine 104 stores the final image volume 191 in the data repository 116.
[0055] In some examples, the image volume reconstruction engine 104 (e.g., via one or more processors executing instructions, as described herein) trains any of the machine learning processes described herein. For example, the image volume reconstruction engine 104 can input features generated from corresponding training data into any of the machine learning models described herein and can execute the machine learning model to generate output data. For example, the output data can characterize a predicted synthetic CT image and a corresponding segmentation mask. The image reconstruction engine 104 can determine whether the machine learning model is made based on the output data. For example, the image volume reconstruction engine 104 can determine at least one metric based on the predicted synthetic CT image and / or the corresponding segmentation mask. For example, the image volume reconstruction engine 104 can calculate the mean squared error, mean absolute error, and / or adverse loss based on the output predicted synthetic CT image. Additionally, the image volume reconstruction engine 104 can calculate the precision score, recall score, loss function, dice score, and / or F1 score based on the output segmentation mask. If the at least one metric meets the corresponding threshold, the image volume reconstruction engine 104 determines that the training of the machine learning process is complete. Otherwise, if the at least one metric does not meet the corresponding threshold, the image volume reconstruction engine 104 continues to train the machine learning process based on features generated from additional corresponding data.
[0056] Figure 5 Illustrated is a computing device 500 that can be employed by the image reconstruction system 104. The computing device 500 can implement one or more functions of, for example, the image reconstruction system 104 described herein, such as any of the functions described with respect to the synthesis engine 110.
[0057] The computing device 500 can include one or more processors 501, a working memory 502, one or more input / output devices 503, an instruction memory 507, a transceiver 504, one or more communication ports 507, and a display 506, all operatively coupled to one or more data buses 508. The data bus 508 allows communication between the various devices. The data bus 508 can include a wired or wireless communication channel.
[0058] The processor 501 can include one or more different processors, each having one or more cores. Each of the different processors can have the same or different architectures. The processor 301 can include one or more central processing units (CPUs), one or more graphics processing units (GPUs), one or more application specific integrated circuits (ASICs), one or more digital signal processors (DSPs), one or more processing cores, and so on.
[0059] The processor 501 may be configured to perform specific functions or operations by executing code stored in the instruction memory 507 that embodies the functions or operations. For example, the processor 501 may be configured to perform one or more of any of the functions, methods, or operations disclosed herein.
[0060] The instruction memory 507 may store instructions that are accessible (e.g., readable) and executable by the processor 501. For example, the instruction memory 507 may be a non-transitory computer-readable storage medium such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory, removable disk, CD-ROM, any non-volatile memory, or any other suitable memory. For example, the instruction memory 507 may store instructions that, when executed by one or more processors 501, cause the one or more processors 501 to perform one or more of the functions of the image reconstruction system 104, such as one or more of the functions of the synthesis engine 100, one or more of the functions of the attenuation map generation engine 120, and one or more of the functions of the image volume reconstruction engine 130 described herein.
[0061] The processor 501 may store data to and read data from the working memory 502. For example, the processor 501 may store a working instruction set to the working memory 502, such as instructions loaded from the instruction memory 507. The processor 501 may also use the working memory 502 to store dynamic data created during the operation of the computing device 500. For example, the processor 501 may store parameters associated with any of the algorithms and models (e.g., machine learning models) described herein, such as hyperparameters and weights. The working memory 502 may be random access memory (RAM), such as static random access memory (SRAM) or dynamic random access memory (DRAM), or any other suitable memory.
[0062] The input / output device 503 may include any suitable device that allows data input or output. For example, the input / output device 503 may include one or more of a keyboard, touchpad, mouse, stylus, touch screen, physical buttons, speakers, microphones, or any other suitable input or output device.
[0063] The (multiple) communication ports 507 may include, for example, serial ports such as universal asynchronous receiver / transmitter (UART) connections, universal serial bus (USB) connections, or any other suitable communication port or connection. In some examples, the (multiple) communication ports 507 allow programming of executable instructions into the instruction memory 507. In some examples, the (multiple) communication ports 507 allow transmission (e.g., upload or download) of data, such as data uploaded for machine learning model parameters.
[0064] The display 506 can display the user interface 505. The user interface 505 can enable user interaction with the computing device 500. For example, the user interface 505 can be a user interface of an application that permits viewing of the final image volume 191. In some examples, the user can interact with the user interface 505 by using the input / output device 503. In some examples, the display 506 can be a touch screen, on which the user interface 505 is displayed.
[0065] The transceiver 504 permits communication with a network, which is, for example, a Wi-Fi network, Ethernet, a cellular network, or any other suitable communication network. For example, if operating in a cellular network, the transceiver 504 is configured to permit communication with the cellular network. One or more processors 501 are operable to receive data from the network or send data to the network via the transceiver 504.
[0066] Figure 6 is a flowchart of an example method 600 for reconstructing an image. The method can be performed by one or more computing devices (such as the computing device 500) that execute instructions (such as instructions stored in the instruction memory 507).
[0067] Starting at block 602, a partial CT image, a NAC PET image, and an x-ray image are received. For example and as described herein, the image reconstruction system 104 can receive CT image data 103, NAC measurement data 105, and scout image data 107. For example, each of the CT image data 103, NAC measurement data 105, and scout image data 107 can be received from the image scanning system 102 or from the data repository 116. At block 604, location data is generated based on applying a first machine learning process to the partial CT image. The location data identifies a feature location within the partial CT image. For example, as described herein, the image reconstruction system 104 can apply a landmark algorithm to the CT image data 103 to generate location data 257 that characterizes anatomical landmarks.
[0068] Proceeding to block 606, a segmentation mask is generated based on applying a second machine learning process to the NAC PET image, the x-ray image, and the location data. For example, as described herein, the image reconstruction system 104 can apply a segmentation process to the NAC measurement data 105, the scout image data 107, and the location data 257 to generate an initial segmentation mask 203.
[0069] In addition, at block 608, a plurality of synthetic images are generated based on applying at least a third machine learning process to the segmentation mask, the partial CT image, the NAC PET image, and the x-ray image. For example, and as described herein, the image reconstruction system 104 may apply one or more trained machine learning processes to each combination of the CT image data 103 and the initial segmentation mask 203, the NAC measurement data 105 and the initial segmentation mask 203, and the scout image data 107 and the initial segmentation mask 203 to generate corresponding synthetic images, such as the first synthetic image 303, the third synthetic image 305, and the fourth synthetic image 313.
[0070] At block 610, a final synthetic CT image and a corresponding organ mask are generated based on applying at least a fourth machine learning process to the plurality of synthetic images and the CT mask of the partial CT image. For example, as described herein, the image reconstruction system 104 may generate CT mask data 323 representative of a segmentation mask (e.g., an organ mask) based on applying a segmentation algorithm to the CT mask data 323. The CT mask generation engine 322 may apply any suitable segmentation algorithm known in the art (e.g., a trained machine learning model) to generate the segmentation mask. In addition, the image reconstruction system 104 may apply a trained machine learning process to the CT mask data 323, the first synthetic image 303, and the third synthetic image 309 to generate predicted CT image data 111 representative of the final predicted synthetic image. For example, the trained machine learning process may be based on a deep learning model (e.g., a diffusion model, a GAN, a CNN). The final synthetic CT image and the corresponding organ mask may be stored in a data repository, such as the data repository 116.
[0071] Figure 7 is a flowchart of an example method 700 for generating and fusing synthetic images. The method may be performed by one or more computing devices (e.g., the computing device 500) executing instructions (e.g., instructions stored in the instruction memory 507).
[0072] Starting at block 702, a plurality of synthetic images are received. The plurality of synthetic images may include, for example, one or more of the first synthetic image 303, the second synthetic image 305, the third synthetic image 309, and the fourth synthetic image 313. At block 704, features are generated based on applying a feature extraction process to the plurality of synthetic images. For example, as described herein, the image reconstruction system 104 may apply a feature extraction process to extract features from one or more of the first synthetic image 303, the second synthetic image 305, and the fourth synthetic image 313 and generate feature data 413 representative of the extracted features.
[0073] Additionally, at block 706, the features are modified based on applying a trained machine learning process to the features and at least one of the plurality of synthetic images. For example and as described herein, the image reconstruction system 104 may modify the feature data 413 based on the third synthetic image 309. For example, the image reconstruction system 104 may apply a trained machine learning process (e.g., a CNN or GAN-based process) to the feature data 413 and the third synthetic image 309, and generate modified features based on the application of the trained machine learning process. At block 708, a final synthetic CT image (e.g., the predicted CT image data 111) is generated based on the modified features.
[0074] The following is a list of non-limiting illustrative embodiments disclosed herein:
[0075] Illustrative Embodiment 1: A computer-implemented method, comprising:
[0076] Receiving a co-modal image, a non-attenuation corrected kernel image, and an x-ray image;
[0077] Applying a first machine learning process to the co-modal image and generating position data identifying feature locations based on the application of the first machine learning process;
[0078] Applying a second machine learning process to the non-attenuated corrected kernel image, the x-ray image, and the position data, and generating a segmentation mask based on the application of the second machine learning process;
[0079] Applying at least a third machine learning process to the segmentation mask, the co-modal image, the non-attenuation corrected kernel image, and the x-ray image, and generating a plurality of synthetic images based on the application of at least the third machine learning process; and
[0080] Applying at least a fourth machine learning process to the plurality of synthetic images and generating a final synthetic image based on the application of at least the fourth machine learning process.
[0081] Illustrative Embodiment 2: The computer-implemented method according to Illustrative Embodiment 1 further comprises generating a final segmentation mask based on the final synthetic image.
[0082] Illustrative Embodiment 3: The computer-implemented method according to any one of Illustrative Embodiments 1-2 further comprises generating an attenuation map based on the final synthetic image.
[0083] Illustrative Embodiment 4: The computer-implemented method according to any one of Illustrative Embodiments 1-3, comprising:
[0084] Receiving an optical image;
[0085] Applying the second machine learning process to the optical image; and
[0086] Apply the at least third machine learning process to the optical image.
[0087] Exemplary embodiment 5: A computer-implemented method according to any one of exemplary embodiments 1-4, wherein each of the co-modal image, the non-attenuation corrected kernel image, and the x-ray image includes a varying view of a subject.
[0088] Exemplary embodiment 6: A computer-implemented method according to exemplary embodiment 5, wherein the final composite image includes a full view of the subject.
[0089] Exemplary embodiment 7: A computer-implemented method according to any one of exemplary embodiments 1-6, wherein the at least third machine learning process includes three machine learning processes, and the method includes:
[0090] Apply a first one of the three machine learning processes to the co-modal image and the segmentation mask, and generate a first composite image among the plurality of composite images based on the application of the first one of the three machine learning processes;
[0091] Apply a second one of the three machine learning processes to the non-attenuation corrected kernel image and the segmentation mask, and generate a second composite image among the plurality of composite images based on the application of the second one of the three machine learning processes; and
[0092] Apply a third one of the three machine learning processes to the x-ray image and the segmentation mask, and generate a third composite image among the plurality of composite images based on the application of the third one of the three machine learning processes.
[0093] Exemplary embodiment 8: A computer-implemented method according to exemplary embodiment 7, wherein applying the at least fourth machine learning process to the plurality of composite images includes:
[0094] Apply a feature extraction process to the first composite image and the third composite image, and generate features based on the application of the feature extraction process;
[0095] Modify the features based on the second composite image; and
[0096] Generate a final composite image based on the modified features.
[0097] Exemplary embodiment 9: A computer-implemented method according to any one of exemplary embodiments 1-8, wherein applying the at least fourth machine learning process to the plurality of composite images includes:
[0098] Apply an encoding process to the plurality of composite images, and generate encoded features based on the application of the encoding process; and
[0099] Apply a first decoding process to the encoded features and generate a final synthesized image based on the application of the first decoding process.
[0100] Exemplary embodiment 10: The computer-implemented method according to exemplary embodiment 9, including applying a second decoding process to the encoded features and generating a final segmentation mask based on the application of the second decoding process.
[0101] Exemplary embodiment 11: The computer-implemented method according to any one of exemplary embodiments 1-10, including applying a registration process to the final synthesized image and the non-attenuation correction kernel image to generate a final registered synthesized image.
[0102] Exemplary embodiment 12: The computer-implemented method according to any one of exemplary embodiments 1-11, including displaying the final synthesized image.
[0103] Exemplary embodiment 13: The computer-implemented method according to any one of exemplary embodiments 1-12, wherein the non-attenuation correction kernel image is a positron emission tomography (PET) image.
[0104] Exemplary embodiment 14: The computer-implemented method according to any one of exemplary embodiments 1-12, wherein the non-attenuation correction kernel image is a single photon emission computed tomography (SPECT) image.
[0105] Exemplary embodiment 15: A non-transitory computer-readable medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform operations including:
[0106] Receiving a co-modal image, a non-attenuation correction kernel image, and an x-ray image;
[0107] Applying a first machine learning process to the co-modal image and generating position data identifying the feature positions based on the application of the first machine learning process;
[0108] Applying a second machine learning process to the non-attenuated correction kernel image, the x-ray image, and the position data and generating a segmentation mask based on the application of the second machine learning process;
[0109] Applying at least a third machine learning process to the segmentation mask, the co-modal image, the non-attenuation correction kernel image, and the x-ray image and generating a plurality of synthesized images based on the application of at least the third machine learning process; and
[0110] Applying at least a fourth machine learning process to the plurality of synthesized images and generating a final synthesized image based on the application of the at least fourth machine learning process.
[0111] Exemplary Embodiment 16: The non - transitory computer - readable medium storing instructions according to Exemplary Embodiment 15, when executed by the at least one processor, the instructions further cause the at least one processor to perform an operation including generating a final segmentation mask based on the final composite image.
[0112] Exemplary Embodiment 17: The non - transitory computer - readable medium storing instructions according to any one of Exemplary Embodiments 15 - 16, when executed by the at least one processor, the instructions further cause the at least one processor to perform an operation including generating an attenuation map based on the final composite image.
[0113] Exemplary Embodiment 18: The non - transitory computer - readable medium storing instructions according to any one of Exemplary Embodiments 15 - 17, when executed by at least one processor, the instructions further cause at least one processor to perform an operation including the following:
[0114] Receiving an optical image;
[0115] Applying the second machine - learning process to the optical image; and
[0116] Applying the at least third machine - learning process to the optical image.
[0117] Exemplary Embodiment 19: The non - transitory computer - readable medium according to any one of Exemplary Embodiments 15 - 18, wherein each of the co - modal image, the non - attenuation - corrected kernel image, and the x - ray image includes a varying view of the subject.
[0118] Exemplary Embodiment 20: The non - transitory computer - readable medium according to Exemplary Embodiment 19, wherein the final composite image includes a full view of the subject.
[0119] Exemplary Embodiment 21: The non - transitory computer - readable medium according to any one of Exemplary Embodiments 15 - 20, wherein the at least third machine - learning process includes three machine - learning processes, and the computer - readable medium stores instructions that, when executed by the at least one processor, the instructions further cause the at least one processor to perform an operation including the following:
[0120] Applying the first of the three machine - learning processes to the co - modal image and the segmentation mask and generating a first composite image among a plurality of composite images based on the application of the first of the three machine - learning processes;
[0121] Applying the second of the three machine - learning processes to the non - attenuation - corrected kernel image and the segmentation mask and generating a second composite image among a plurality of composite images based on the application of the second of the three machine - learning processes; and
[0122] Apply a third of the three machine learning processes to the x-ray image and the segmentation mask, and generate a third synthetic image among the plurality of synthetic images based on the application of the third of the three machine learning processes.
[0123] Illustrative Example 22: The non-transitory computer-readable medium storing instructions according to Illustrative Example 21, when executed by at least one processor, the instructions further cause the at least one processor to perform operations including the following:
[0124] Apply a feature extraction process to the first synthetic image and the third synthetic image, and generate features based on the application of the feature extraction process;
[0125] Modify the features based on the second synthetic image; and
[0126] Generate a final synthetic image based on the modified features.
[0127] Illustrative Example 23: The non-transitory computer-readable medium storing instructions according to any one of Illustrative Examples 15-22, when executed by at least one processor, the instructions further cause the at least one processor to perform operations including the following:
[0128] Apply an encoding process to the plurality of synthetic images, and generate encoded features based on the application of the encoding process; and
[0129] Apply a first decoding process to the encoded features, and generate a final synthetic image based on the application of the first decoding process.
[0130] Illustrative Example 24: The non-transitory computer-readable medium storing instructions according to Illustrative Example 23, when executed by the at least one processor, the instructions further cause the at least one processor to perform an operation, the operation including applying a second decoding process to the encoded features, and generating a final segmentation mask based on the application of the second decoding process.
[0131] Illustrative Example 25: The non-transitory computer-readable medium storing instructions according to any one of Illustrative Examples 15-24, when executed by the at least one processor, the instructions further cause the at least one processor to perform an operation including applying a registration process to the final synthetic image and the non-attenuation correction kernel image to generate a final registered synthetic image.
[0132] Illustrative Example 26: The non-transitory computer-readable medium storing instructions according to any one of Illustrative Examples 15-25, when executed by at least one processor, the instructions further cause the at least one processor to perform an operation including displaying the final synthetic image.
[0133] Exemplary embodiment 27: A non - transitory computer - readable medium according to any one of exemplary embodiments 15 - 26, wherein the non - attenuation - corrected nuclear image is a positron emission tomography (PET) image.
[0134] Exemplary embodiment 28: A non - transitory computer - readable medium according to any one of exemplary embodiments 15 - 26, wherein the non - attenuation - corrected nuclear image is a single - photon emission computed tomography (SPECT) image.
[0135] Exemplary embodiment 29: A system comprising:
[0136] A database; and
[0137] At least one processor communicatively coupled to the database and configured to:
[0138] Receive a co - modal image, a non - attenuation - corrected nuclear image, and an x - ray image;
[0139] Apply a first machine - learning process to the co - modal image and generate position data identifying a feature location based on the application of the first machine - learning process;
[0140] Apply a second machine - learning process to the non - attenuated corrected nuclear image, the x - ray image, and the position data and generate a segmentation mask based on the application of the second machine - learning process;
[0141] Apply at least a third machine - learning process to the segmentation mask, the co - modal image, the non - attenuation - corrected nuclear image, and the x - ray image and generate a plurality of synthetic images based on the application of at least the third machine - learning process; and
[0142] Apply at least a fourth machine - learning process to the plurality of synthetic images and generate a final synthetic image based on the application of the at least fourth machine - learning process.
[0143] Exemplary embodiment 30: The system according to exemplary embodiment 29, wherein the at least one processor is further configured to generate a final segmentation mask based on the final synthetic image.
[0144] Exemplary embodiment 31: The system according to any one of exemplary embodiments 29 - 30, wherein the at least one processor is further configured to generate an attenuation map based on the final synthetic image.
[0145] Exemplary embodiment 32: The system according to any one of exemplary embodiments 29 - 31, wherein the at least one processor is further configured to:
[0146] Receive an optical image;
[0147] Apply a second machine learning process to the optical image; and
[0148] Apply at least a third machine learning process to the optical image.
[0149] Illustrative Example 33: The system according to any one of Illustrative Examples 29 - 32, wherein each of the co-modal image, the non-attenuation corrected nuclear image, and the x-ray image includes a varying view of the subject.
[0150] Illustrative Example 34: The system according to Illustrative Example 33, wherein the final synthesized image includes a full view of the subject.
[0151] Illustrative Example 35: The system according to any one of Illustrative Examples 29 - 33, wherein the at least third machine learning process includes three machine learning processes, and wherein at least one processor is further configured to:
[0152] Apply a first of the three machine learning processes to the co-modal image and the segmentation mask, and generate a first synthesized image of the plurality of synthesized images based on the application of the first of the three machine learning processes;
[0153] Apply a second of the three machine learning processes to the non-attenuation corrected nuclear image and the segmentation mask, and generate a second synthesized image of the plurality of synthesized images based on the application of the second of the three machine learning processes; and
[0154] Apply a third of the three machine learning processes to the x-ray image and the segmentation mask, and generate a third synthesized image of the plurality of synthesized images based on the application of the third of the three machine learning processes.
[0155] Illustrative Example 36: The system according to Illustrative Example 35, wherein in order to apply at least a fourth machine learning process to the plurality of synthesized images, the at least one processor is further configured to:
[0156] Apply a feature extraction process to the first synthesized image and the third synthesized image, and generate features based on the application of the feature extraction process;
[0157] Modify the features based on the second synthesized image; and
[0158] Generate a final synthesized image based on the modified features.
[0159] Illustrative Example 37: The system according to any one of Illustrative Examples 29 - 36, wherein in order to apply at least a fourth machine learning process to the plurality of synthesized images, the at least one processor is further configured to:
[0160] Apply the encoding process to the plurality of synthetic images and generate encoded features based on the application of the encoding process; and
[0161] Apply a first decoding process to the encoded features and generate a final synthetic image based on the application of the first decoding process.
[0162] Exemplary embodiment 38: The system according to exemplary embodiment 37, wherein the at least one processor is further configured to apply a second decoding process to the encoded features and generate a final segmentation mask based on the application of the second decoding process.
[0163] Exemplary embodiment 39: The system according to any one of exemplary embodiments 29-38, wherein the at least one processor is further configured to apply a registration process to the final synthetic image and the non-attenuation corrected nuclear image to generate a final registered synthetic image.
[0164] Exemplary embodiment 40: The system according to any one of exemplary embodiments 29-39, wherein the at least one processor is further configured to display the final synthetic image.
[0165] Exemplary embodiment 41: The system according to any one of exemplary embodiments 29-40, wherein the non-attenuation corrected nuclear image is a positron emission tomography (PET) image.
[0166] Exemplary embodiment 42: The system according to any one of exemplary embodiments 29-40, wherein the non-attenuation corrected nuclear image is a single photon emission computed tomography (SPECT) image.
[0167] Exemplary embodiment 43: A system, comprising:
[0168] means for receiving co-modal images, non-attenuation corrected nuclear images, and x-ray images;
[0169] means for applying a first machine learning process to the co-modal images and generating position data identifying feature locations based on the application of the first machine learning process;
[0170] means for applying a second machine learning process to the non-attenuation corrected nuclear image, the x-ray image, and the position data and generating a segmentation mask based on the application of the second machine learning process;
[0171] means for applying at least a third machine learning process to the segmentation mask, the co-modal images, the non-attenuation corrected nuclear image, and the x-ray image and generating a plurality of synthetic images based on the application of at least the third machine learning process; and
[0172] Apparatus for applying at least a fourth machine learning process to the plurality of synthetic images and generating a final synthetic image based on the application of the at least fourth machine learning process.
[0173] Illustrative Example 44: The system according to Illustrative Example 43 further includes an apparatus for generating a final segmentation mask based on the final synthetic image.
[0174] Illustrative Example 45: The system according to any one of Illustrative Examples 43-44 further includes an apparatus for generating an attenuation map based on the final synthetic image.
[0175] Illustrative Example 46: The system according to any one of Illustrative Examples 43-45 further includes:
[0176] An apparatus for receiving an optical image;
[0177] An apparatus for applying a second machine learning process to the optical image; and
[0178] An apparatus for applying at least a third machine learning process to the optical image.
[0179] Illustrative Example 47: The system according to any one of Illustrative Examples 43-46, wherein each of the co-modal image, the non-attenuation corrected kernel image, and the x-ray image includes a varying view of the subject.
[0180] Illustrative Example 48: The system according to Illustrative Example 47, wherein the final synthetic image includes a full view of the subject.
[0181] Illustrative Example 49: The system according to any one of Illustrative Examples 43-48, wherein the at least third machine learning process includes three machine learning processes, and the system further includes:
[0182] An apparatus for applying the first of the three machine learning processes to the co-modal image and the segmentation mask and generating a first synthetic image among the plurality of synthetic images based on the application of the first of the three machine learning processes;
[0183] An apparatus for applying the second of the three machine learning processes to the non-attenuation corrected kernel image and the segmentation mask and generating a second synthetic image among the plurality of synthetic images based on the application of the second of the three machine learning processes; and
[0184] An apparatus for applying the third of the three machine learning processes to the x-ray image and the segmentation mask and generating a third synthetic image among the plurality of synthetic images based on the application of the third of the three machine learning processes.
[0185] Exemplary embodiment 50: The system according to exemplary embodiment 49, wherein in order to apply at least a fourth machine learning process to a plurality of synthetic images, the system comprises:
[0186] means for applying a feature extraction process to a first synthetic image and a third synthetic image and generating features based on the application of the feature extraction process;
[0187] means for modifying the features based on a second synthetic image; and
[0188] means for generating a final synthetic image based on the modified features.
[0189] Exemplary embodiment 51: The system according to any one of exemplary embodiments 43 - 50, wherein in order to apply at least a fourth machine learning process to the plurality of synthetic images, the system comprises:
[0190] means for applying an encoding process to the plurality of synthetic images and generating encoded features based on the application of the encoding process; and
[0191] means for applying a first decoding process to the encoded features and generating a final synthetic image based on the application of the first decoding process.
[0192] Exemplary embodiment 52: The system according to exemplary embodiment 51, further comprising means for applying a second decoding process to the encoded features and generating a final segmentation mask based on the application of the second decoding process.
[0193] Exemplary embodiment 53: The system according to any one of exemplary embodiments 43 - 52, further comprising means for applying a registration process to the final synthetic image and a non - attenuation - corrected nuclear image to generate a final registered synthetic image.
[0194] Exemplary embodiment 54: The system according to any one of exemplary embodiments 43 - 53, further comprising means for displaying the final synthetic image.
[0195] Exemplary embodiment 55: The system according to any one of exemplary embodiments 43 - 54, wherein the non - attenuation - corrected nuclear image is a positron emission tomography (PET) image.
[0196] Exemplary embodiment 56: The system according to any one of exemplary embodiments 43 - 54, wherein the non - attenuation - corrected nuclear image is a single - photon emission computed tomography (SPECT) image.
[0197] The previous description of the embodiments is provided to enable any person skilled in the art to practice the present disclosure. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be applied to other embodiments without the use of creative faculty. The present disclosure is not intended to be limited to the embodiments shown herein, but rather is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A computer-implemented method, comprising: Receiving a co-modal image, a non-attenuation correction kernel image, and an x-ray image; Applying a first machine learning process to the co-modal image and generating position data identifying a feature position based on the application of the first machine learning process; Applying a second machine learning process to the non-attenuated correction kernel image, the x-ray image, and the position data and generating a segmentation mask based on the application of the second machine learning process; Applying at least a third machine learning process to the segmentation mask, the co-modal image, the non-attenuation correction kernel image, and the x-ray image and generating a plurality of synthetic images based on the application of at least the third machine learning process; And Applying at least a fourth machine learning process to the plurality of synthetic images and generating a final synthetic image based on the application of at least the fourth machine learning process.
2. The computer-implemented method according to claim 1, further comprising generating a final segmentation mask based on the final synthetic image.
3. The computer-implemented method according to claim 1, further comprising generating an attenuation map based on the final synthetic image.
4. The computer-implemented method according to claim 1, comprising: Receiving an optical image; Applying the second machine learning process to the optical image; And Applying the at least third machine learning process to the optical image.
5. The computer-implemented method according to claim 1, wherein, Each of the co-modal image, the non-attenuation correction kernel image, and the x-ray image includes a varying view of a subject.
6. The computer-implemented method according to claim 5, wherein, The final synthetic image includes a full view of the subject.
7. The computer-implemented method according to claim 1, wherein the at least third machine learning process includes three machine learning processes, and the method comprises: Applying a first of the three machine learning processes to the co-modal image and the segmentation mask and generating a first synthetic image of the plurality of synthetic images based on the application of the first of the three machine learning processes; Applying a second of the three machine learning processes to the non-attenuation corrected kernel image and the segmentation mask and generating a second synthetic image of the plurality of synthetic images based on the application of the second of the three machine learning processes; And Applying a third of the three machine learning processes to the x-ray image and the segmentation mask and generating a third synthetic image of the plurality of synthetic images based on the application of the third of the three machine learning processes.
8. The computer-implemented method according to claim 7, wherein applying the at least fourth machine learning process to the plurality of synthetic images includes: Applying a feature extraction process to the first synthetic image and the third synthetic image and generating features based on the application of the feature extraction process; Modifying the features based on the second synthetic image; And Generating a final synthetic image based on the modified features.
9. The computer-implemented method according to claim 1, wherein applying at least a fourth machine learning process to the plurality of synthetic images includes: Applying an encoding process to the plurality of synthetic images and generating encoded features based on the application of the encoding process; And Applying a first decoding process to the encoded features and generating a final synthetic image based on the application of the first decoding process.
10. The computer-implemented method according to claim 9, comprising applying a second decoding process to the encoded features and generating a final segmentation mask based on the application of the second decoding process.
11. The computer-implemented method according to claim 1, comprising applying a registration process to the final synthetic image and the non-attenuation corrected kernel image to generate a final registered synthetic image.
12. The computer-implemented method according to claim 1, comprising displaying the final synthetic image.
13. The computer-implemented method according to claim 1, wherein, The non-attenuation corrected kernel image is a positron emission tomography (PET) image.
14. The computer-implemented method according to claim 1, wherein, The non-attenuation corrected kernel image is a single photon emission computed tomography (SPECT) image.
15. A non-transitory computer-readable medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform operations comprising: receiving co-modal images, non-attenuation corrected kernel images, and x-ray images; applying a first machine learning process to the co-modal images and generating position data identifying feature positions based on the application of the first machine learning process; applying a second machine learning process to the non-attenuated corrected kernel images, x-ray images, and position data and generating a segmentation mask based on the application of the second machine learning process; applying at least a third machine learning process to the segmentation mask, co-modal images, non-attenuation corrected kernel images, and x-ray images and generating a plurality of synthetic images based on the application of the at least third machine learning process; and applying at least a fourth machine learning process to the plurality of synthetic images and generating a final synthetic image based on the application of the at least fourth machine learning process.
16. The non-transitory computer-readable medium storing instructions according to claim 15, which when executed by the at least one processor, the instructions further cause the at least one processor to perform an operation comprising generating a final segmentation mask based on the final synthetic image.
17. The non-transitory computer-readable medium storing instructions according to claim 15, which when executed by the at least one processor, the instructions further cause the at least one processor to perform an operation comprising generating an attenuation map based on the final synthetic image.
18. The non-transitory computer-readable medium storing instructions according to claim 15, which when executed by the at least one processor, the instructions further cause the at least one processor to perform operations comprising: receiving an optical image; applying the second machine learning process to the optical image; and applying the at least third machine learning process to the optical image.
19. The non-transitory computer-readable medium according to claim 15, wherein, The at least third machine learning process includes three machine learning processes, and the computer-readable medium stores instructions that, when executed by the at least one processor, the instructions further cause the at least one processor to perform operations comprising: applying the first of the three machine learning processes to the co-modal images and the segmentation mask and generating a first synthetic image of the plurality of synthetic images based on the application of the first of the three machine learning processes; Apply the second of the three machine learning processes to the non-attenuation-corrected nuclear image and the segmentation mask, and generate a second synthetic image among the plurality of synthetic images based on the application of the second of the three machine learning processes; and Apply the third of the three machine learning processes to the x-ray image and the segmentation mask, and generate a third synthetic image among the plurality of synthetic images based on the application of the third of the three machine learning processes.
20. A system, comprising: A database; And At least one processor communicatively coupled to the database and configured to: Receive a co-modal image, a non-attenuation-corrected nuclear image, and an x-ray image; Apply a first machine learning process to the co-modal image and generate position data identifying feature positions based on the application of the first machine learning process; Apply a second machine learning process to the non-attenuation-corrected nuclear image, the x-ray image, and the position data, and generate a segmentation mask based on the application of the second machine learning process; Apply at least a third machine learning process to the segmentation mask, the co-modal image, the non-attenuation-corrected nuclear image, and the x-ray image, and generate a plurality of synthetic images based on the application of at least the third machine learning process; And Apply at least a fourth machine learning process to the plurality of synthetic images and generate a final synthetic image based on the application of the at least fourth machine learning process.