Image processing method and apparatus
By identifying and correcting motion artifact regions in optical coherence tomography (OCT) images, high-quality target images are generated, solving the problem of low image quality in ophthalmic diagnosis and improving diagnostic accuracy.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BRIGHTVIEW MEDICAL TECHNOLOGIES (NANJING) CO LTD
- Filing Date
- 2024-04-16
- Publication Date
- 2026-07-21
AI Technical Summary
Optical coherence tomography (OCT) imaging suffers from low image quality in ophthalmic diagnosis, such as the presence of horizontal or vertical dark or bright stripes, retinal structural misalignment, stretching, or distortion, which affects the accuracy of diagnostic results.
By identifying motion artifact regions in optical coherence tomography (OCT) images, selecting the image with the fewest motion artifact regions, and then correcting it using images without motion artifact regions, a high-quality target image is generated for diagnostic purposes.
It effectively improves the image quality of optical coherence tomography (OCT) scans, provides highly accurate diagnostic evidence, and enhances the reliability of ophthalmic diagnosis.
Smart Images

Figure CN118351029B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of medical imaging, and more particularly to an image processing method and apparatus. Background Technology
[0002] Optical coherence tomography (OCT), a non-invasive, high-resolution imaging technique, has been continuously developed and widely applied in the field of medical imaging. OCT can provide micron-level images of tissue cross-sections, which is of great significance for observing biological tissue structures and lesions. OCT angiography (OCTA) can also observe blood flow in the fundus.
[0003] The basic principle of optical coherence tomography (OCT) is to acquire high-resolution images of tissues using optical interferometry. Its characteristics include real-time imaging, non-destructive scanning, and high resolution, which has led to its widespread application in various aspects of medicine. In ophthalmology, OCT can provide detailed information about the structure of the eye, aiding in the early diagnosis and monitoring of eye diseases. For example, combining OCT with fundus imaging provides real-time navigation and monitoring for ophthalmic surgery, improving surgical success rates.
[0004] However, when using optical coherence tomography (OCT / OCTA) imaging for ophthalmic diagnosis, the images may contain low-quality images, such as horizontal or vertical dark or bright stripes, misalignment, stretching or distortion of retinal structures, and uneven brightness, which may affect the accuracy of the diagnostic results. Summary of the Invention
[0005] This application provides an image processing method and apparatus, the purpose of which is to improve the image quality of optical coherence tomography (OCT) images.
[0006] To achieve the above objectives, this application provides the following technical solution:
[0007] An image processing method, comprising:
[0008] A set of scanned images of a specified target is obtained using optical coherence tomography (OCT). The set of scanned images includes multiple three-dimensional images that are sequential in scanning time. Each three-dimensional image includes multiple cross-sectional images. The cross-sectional images are used to characterize the two-dimensional structure of the specified target at a specified level.
[0009] Identify the type of each cross-sectional image in each of the three-dimensional images;
[0010] Based on the cross-sectional image of type anomalous, the motion artifact region in each of the three-dimensional images is identified; the motion artifact region represents the area where there is anomaly in the scanning imaging.
[0011] From all the 3D images containing motion artifact regions, determine the first image with the fewest motion artifact regions.
[0012] The motion artifact region in the first image is corrected using the motion artifact region in the second image to obtain the target image; the second image is any other three-dimensional image in the scanned image set besides the first image; the target image is used as the diagnostic basis for the specified target.
[0013] Optionally, identifying the type of each cross-sectional image in each of the three-dimensional images includes:
[0014] For each of the three-dimensional images, each cross-sectional image of the three-dimensional image is used as input to a first classification model to obtain a first classification result output by the first classification model; the first classification model is pre-trained based on the sample cross-sectional image as input and combined with the pre-labeled type label of the sample cross-sectional image as the training target; the first classification result includes the type of each cross-sectional image.
[0015] Optionally, identifying the type of each cross-sectional image in each of the three-dimensional images includes:
[0016] For each of the three-dimensional images, determine the centroid of each cross-sectional image in the three-dimensional image;
[0017] Based on the centroid of each of the cross-sectional images, determine the coordinates of the centroid mapped in the A-scan direction;
[0018] Based on the coordinates of the centroid mapping in the A-scan direction, determine the gradient value corresponding to each mapped coordinate point;
[0019] Based on the gradient value, the type of each cross-sectional image is determined.
[0020] Optionally, based on the gradient value, the type of each cross-sectional image is determined, including:
[0021] The target threshold is determined based on the gradient value corresponding to the mapped coordinate point of each centroid.
[0022] The type of each cross-sectional image is determined by comparing the gradient value corresponding to the mapped coordinate point of the centroid with the target threshold.
[0023] Optionally, based on the gradient value, the type of each cross-sectional image is determined, including:
[0024] Based on the scanning time and corresponding gradient value of each of the cross-sectional images, a target time-domain signal is generated;
[0025] The target time-domain signal is used as input to the second classification model to obtain the second classification result output by the second classification model; the second classification model is pre-trained based on the sample time-domain signal as input and combined with the pre-labeled type label of the sample time-domain signal as the training target; the second classification result is used to indicate the type of each of the cross-sectional images.
[0026] Optionally, after identifying the type of each cross-sectional image in each of the three-dimensional images, the method further includes:
[0027] Cross-sectional images of type anomalous (i.e., images with motion artifacts) are labeled as 1, and cross-sectional images of type normal (i.e., images without motion artifacts) are labeled as 0. This yields a continuous set of labeled data.
[0028] Inflate the labeled data to obtain the corresponding inflated labels;
[0029] Determine the type of the cross-sectional image corresponding to the expansion label.
[0030] Optionally, after identifying the type of each cross-sectional image in each of the three-dimensional images, the method further includes:
[0031] The system detects whether multiple consecutive images with motion artifact regions meet specified conditions. If they do, the system determines that the type of the multiple consecutive images with motion artifact regions is normal.
[0032] Optionally, the motion artifact region image in the first image is corrected using the motion artifact region image in the second image to obtain the target image, including:
[0033] Obtain the enface images of the first image and the second image, and use them as the first enface image and the second enface image, respectively;
[0034] The positional offset between the first image and the second image is determined based on the similarity between the first enface image and the second enface image;
[0035] The target coordinates are determined based on the position offset and the first coordinate; the first coordinate represents the coordinates that match the region of motion artifact in the first image.
[0036] From the multiple motion-free artifact region images contained in the second image, a target image matching the target coordinates is determined; the motion-free artifact region images include normal cross-sectional images.
[0037] The target image is obtained by replacing the motion artifact region in the first image with the target image.
[0038] Conversely, using the target image to replace the motion artifact region in the first image to obtain the target image includes:
[0039] The target image is used to replace the motion artifact region in the first image to obtain a third image.
[0040] Determine the scanning order of each cross-sectional image in the third image;
[0041] For each cross-sectional image in the third image, the positional deviation of the cross-sectional image in a specified direction is determined based on the similarity between the cross-sectional image and the cross-sectional image of the previous or next frame.
[0042] The coordinates of each cross-sectional image are adjusted according to the positional deviation of each cross-sectional image in a specified direction to obtain a fourth image;
[0043] The fourth image is filtered to obtain the target image.
[0044] Optionally, the method further includes:
[0045] The target image is segmented using a pre-trained generative adversarial network to determine the category label of each pixel in the target image; the category label is used to characterize the constituent parts of the specified target; the generative adversarial network is also used to denoise the target image to improve the clarity of the target image.
[0046] An image processing apparatus, comprising:
[0047] An image acquisition unit is used to obtain a set of scanned images of a specified target using optical coherence tomography (OCT) technology; the set of scanned images includes multiple three-dimensional images that are sequential in scanning time; the three-dimensional images include multiple cross-sectional images; the cross-sectional images are used to characterize the two-dimensional structure of the specified target at a specified level;
[0048] An image classification unit is used to identify the type of each cross-sectional image in each of the three-dimensional images;
[0049] The image recognition unit is used to identify motion artifact regions in each of the three-dimensional images based on cross-sectional images of type anomalous; the motion artifact regions represent areas where there are anomalous scanning images.
[0050] The image filtering unit is used to determine the first image with the fewest motion artifact regions from each of the three-dimensional images containing motion artifact regions.
[0051] An image correction unit is used to correct a motion artifact region in a first image using a motion artifact-free region in a second image to obtain a target image; the second image is any other three-dimensional image in the scanned image set other than the first image; the target image is used as a diagnostic basis for the specified target.
[0052] Optionally, the image classification unit is specifically used for:
[0053] For each of the three-dimensional images, each cross-sectional image of the three-dimensional image is used as input to a first classification model to obtain a first classification result output by the first classification model; the first classification model is pre-trained based on the sample cross-sectional image as input and combined with the pre-labeled type label of the sample cross-sectional image as the training target; the first classification result includes the type of each cross-sectional image.
[0054] Optionally, the image classification unit is specifically used for:
[0055] For each of the three-dimensional images, determine the centroid of each cross-sectional image in the three-dimensional image;
[0056] Based on the centroid of each of the cross-sectional images, determine the coordinates of the centroid mapped in the A-scan direction;
[0057] Based on the coordinates of the centroid mapping in the A-scan direction, determine the gradient value corresponding to each mapped coordinate point;
[0058] Based on the gradient value, the type of each cross-sectional image is determined.
[0059] Optionally, the image classification unit is specifically used for:
[0060] The target threshold is determined based on the gradient value corresponding to the mapped coordinate point of each centroid.
[0061] The type of each cross-sectional image is determined by comparing the gradient value corresponding to the mapped coordinate point of the centroid with the target threshold.
[0062] Optionally, the image classification unit is specifically used for:
[0063] Based on the scanning time and corresponding gradient value of each of the cross-sectional images, a target time-domain signal is generated;
[0064] The target time-domain signal is used as input to the second classification model to obtain the second classification result output by the second classification model; the second classification model is pre-trained based on the sample time-domain signal as input and combined with the pre-labeled type label of the sample time-domain signal as the training target; the second classification result is used to indicate the type of each of the cross-sectional images.
[0065] Optionally, the image classification unit is further configured to:
[0066] Cross-sectional images of type anomalous (i.e., images with motion artifacts) are labeled as 1, and cross-sectional images of type normal (i.e., images without motion artifacts) are labeled as 0. This yields a continuous set of labeled data.
[0067] Inflate the labeled data to obtain the corresponding inflated labels;
[0068] Determine the type of the cross-sectional image corresponding to the expansion label.
[0069] Optionally, the image classification unit is further configured to:
[0070] The system detects whether multiple consecutive images with motion artifact regions meet specified conditions. If they do, the system determines that the type of the multiple consecutive images with motion artifact regions is normal.
[0071] Optionally, the image correction unit is specifically used for:
[0072] Obtain the enface images of the first image and the second image, and use them as the first enface image and the second enface image, respectively;
[0073] The positional offset between the first image and the second image is determined based on the similarity between the first enface image and the second enface image;
[0074] The target coordinates are determined based on the position offset and the first coordinate; the first coordinate represents the coordinates that match the region of motion artifact in the first image.
[0075] From the multiple motion-free artifact region images contained in the second image, a target image matching the target coordinates is determined; the motion-free artifact region images include normal cross-sectional images.
[0076] The target image is obtained by replacing the motion artifact region in the first image with the target image.
[0077] Optionally, the image correction unit is specifically used for:
[0078] The target image is used to replace the motion artifact region in the first image to obtain a third image.
[0079] Determine the scanning order of each cross-sectional image in the third image;
[0080] For each cross-sectional image in the third image, the positional deviation of the cross-sectional image in a specified direction is determined based on the similarity between the cross-sectional image and the cross-sectional image of the previous or next frame.
[0081] The coordinates of each cross-sectional image are adjusted according to the positional deviation of each cross-sectional image in a specified direction to obtain a fourth image;
[0082] The fourth image is filtered to obtain the target image.
[0083] Optionally, the device further includes:
[0084] The image segmentation unit is used to: segment the target image using a pre-trained generative adversarial network to determine the category label of each pixel in the target image; the category label is used to characterize the constituent parts of the specified target; the generative adversarial network is also used to denoise the target image to improve the clarity of the target image.
[0085] The technical solution provided in this application utilizes optical coherence tomography (OCT) to obtain a set of scanned images of a specified target and identifies the type of each cross-sectional image in each three-dimensional image. Based on the cross-sectional images classified as anomalous, motion artifact regions are identified in each three-dimensional image. From all the three-dimensional images, a first image with the fewest motion artifact regions is determined. Motion artifact-free regions in a second image are used to correct the motion artifact-containing regions in the first image to obtain the target image, which is used as a diagnostic basis for the specified target. This application identifies motion artifact-containing and motion artifact-free regions in the three-dimensional images through each cross-sectional image. For the motion artifact-containing regions in the first image, motion artifact-free regions in the second image are used for correction to obtain a target image free of motion artifacts, effectively improving the image quality of the scanned images of the specified target. Attached Figure Description
[0086] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0087] Figure 1 A schematic flowchart of an image processing method provided in an embodiment of this application;
[0088] Figure 2 A flowchart illustrating another image processing method provided in an embodiment of this application;
[0089] Figure 3 A flowchart illustrating another image processing method provided in an embodiment of this application;
[0090] Figure 4 This is a schematic diagram of the structure of a segmentation and denoising network provided in an embodiment of this application;
[0091] Figure 5 A flowchart illustrating another image processing method provided in an embodiment of this application;
[0092] Figure 6 A flowchart illustrating another image processing method provided in an embodiment of this application;
[0093] Figure 7 This is a schematic diagram of the architecture of an image processing device provided in an embodiment of this application;
[0094] Figure 8 This is a schematic diagram of a scanning method provided in an embodiment of this application. Detailed Implementation
[0095] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0096] In this application, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0097] In related technologies, when scanning equipment using optical coherence tomography (OCT) technology (such as time-domain OCT equipment and frequency-domain OCT equipment) scans a designated target (such as the eyeball or retina), motion artifacts occur in the scanned image due to the movement of the designated target, resulting in quality defects in the final OCT / OCTA image. Furthermore, due to the quality defects in the scanned image and the complex structure of the eyeball, it is difficult for operators to distinguish the distribution area of the designated target's tissue structure in the scanned image.
[0098] In the field of ophthalmology, motion artifacts in OCT / OCTA images of the eye can be understood as defects in the scanning results caused by head movement, eye movement, or tremors during the scanning process. These defects can manifest as horizontal or vertical dark or bright stripes, image misalignment, stretching, or distortion.
[0099] Therefore, this application provides an image processing method for eliminating motion artifacts in OCT / OCTA scan images and improving the image quality of the scan images, thereby helping to improve the accuracy of ophthalmic diagnosis.
[0100] Example 1
[0101] like Figure 1 The diagram shown is a flowchart of an image processing method provided in an embodiment of this application, including the following steps.
[0102] S101: Obtain a set of scanned images of a specified target using optical coherence tomography (OCT).
[0103] The scanned image set includes multiple three-dimensional images that are sequential in scanning time. Each three-dimensional image includes multiple cross-sectional images, which are used to characterize the two-dimensional structure of a specified target at a specified level.
[0104] In some examples, three-dimensional images can be used to characterize the three-dimensional structure of a specified target. Multiple OCT scan images can be obtained by repeatedly scanning the specified target in the same scanning area using a time-domain OCT device, or they can be OCTA scan images.
[0105] In related technologies, OCT scanning methods include A-scan and B-scan. In a preset image coordinate system, such as... Figure 8 As shown, the x-axis represents the horizontal coordinate axis, the y-axis represents the vertical coordinate axis, and the z-axis represents the axial coordinate axis. Scanning along the axial coordinate axis is called A-scan, which can obtain information in the depth z-direction. Scanning along the horizontal or vertical coordinate axis and performing continuous axial scans (A-scan) is called B-scan, which can obtain two-dimensional information composed of multiple A-scan information. Combining this with scanning along the third coordinate axis (vertical or horizontal coordinate axis) can obtain three-dimensional information composed of multiple frames of B-scan information.
[0106] It should be emphasized that the cross-sectional image shown in this embodiment can be understood as a B-scan scan image. Figure 8 In the xz plane, optical coherence tomography (OCT) images can be understood as a three-dimensional structure composed of the superposition of multiple B-scan scans.
[0107] It is understandable that the so-called specified layer in the preset image coordinate system can refer to any layer parallel to the xz plane.
[0108] S102: Identify the type of each cross-sectional image in each 3D image.
[0109] One approach is to use image classification to identify the type of each cross-sectional image in each 3D image.
[0110] Optionally, the process of identifying the type of each cross-sectional image in each 3D image can be as follows: For each 3D image, each cross-sectional image of the 3D image is used as input to a first classification model to obtain a first classification result output by the first classification model; the first classification model is pre-trained based on the sample cross-sectional image as input and combined with the pre-labeled type label of the sample cross-sectional image as the training target; the first classification result is used to indicate the type of each cross-sectional image.
[0111] In some examples, the network model used by the first classification model includes, but is not limited to, deep learning models such as convolutional neural networks and deep residual networks.
[0112] In one possible implementation, the first classification model can be regarded as a binary classification model, used to distinguish between images with no motion artifacts (i.e., cross-sectional images of the normal type) and images with motion artifacts (i.e., cross-sectional images of the abnormal type) in each cross-sectional image.
[0113] Optionally, the process for identifying the type of each cross-sectional image in each 3D image can also be found in [reference needed]. Figure 2 and Figure 3 The steps shown are explained below.
[0114] S103: Based on the cross-sectional image of type anomalous, identify the region with motion artifacts in each 3D image.
[0115] Among them, motion artifact regions characterize areas where anomalies exist in the scanned imaging. Furthermore, based on normal cross-sectional images, motion artifact-free regions are identified in each 3D image.
[0116] S104: From the various three-dimensional images containing motion artifact regions, determine the first image with the fewest number of motion artifact regions.
[0117] In this process, the number of abnormal cross-sectional images in each 3D image is counted, and the 3D image with the fewest abnormal cross-sectional images is determined as the first image.
[0118] It should be noted that identifying the first image with the fewest motion artifact regions can effectively reduce the computational load of image processing, thereby improving image processing efficiency. Generally speaking, the fewer the number of images with motion artifact regions, the less computational resources are required for subsequent correction of these regions.
[0119] In some examples, the three-dimensional image that meets the corresponding conditions can also be selected as the first image based on the actual situation. For example, the optical coherence tomography image with the earliest scanning time can be selected as the first image.
[0120] S105: Using the motion artifact-free region image in the second image, correct the motion artifact region image in the first image to obtain the target image.
[0121] The second image is any three-dimensional image in the scanned image set other than the first image, and the target image is used as the basis for diagnosis of the specified target.
[0122] It is important to emphasize that by using the motion artifact-free region of the second image to correct the motion artifact region of the first image, the resulting target image is free of motion artifact regions and can accurately represent the three-dimensional structure of the specified target, thus providing a high-quality diagnostic basis for the specified target.
[0123] In some examples, target images are used as diagnostic criteria for a specified target. In addition to ensuring high image quality, it is also necessary to identify the type of each pixel in the target image to assist in the diagnosis of the specified target. Furthermore, higher resolution image quality can also improve the diagnostic accuracy of the specified target.
[0124] Optionally, after determining the target image, a neural network can be used to perform image segmentation on the target image to determine the category label of each pixel in the target image. The category label is used to characterize the tissue structure of the specified target. In addition, the neural network is also used to denoise the target image to improve the clarity of the target image.
[0125] In some examples, the neural network employs a segmentation-based denoising generative adversarial network, such as... Figure 4 As shown, it includes an enhancement module and a segmentation module. The enhancement module and the segmentation module share a set of encoder and decoder to achieve output fusion between the enhancement module and the segmentation module, that is, to obtain the generative adversarial network to simultaneously achieve image segmentation and image denoising.
[0126] In one possible implementation, the encoder includes one or more convolutional layers with a first specified stride (e.g., the first specified stride is set to 1 or 2), and the decoder includes multiple convolutional layers with a second specified stride (e.g., the second specified stride is set to 1) and residual connections between them. The functional principles and training methods of the decoder and encoder are conventional techniques in the field of machine learning, and will not be elaborated here.
[0127] In one possible implementation, the segmentation module or enhancement module may include one or more upsampling convolutional layers, the number of which can be set according to the actual amount of computing resources available.
[0128] In one possible implementation, using the retina as the designated target, multiple tissue structure layers of the retina can be determined based on the category labels of each pixel in the target image. These layers include the pigment epithelium, rods and cones, external retinal membrane, outer nuclear layer, outer retinal network, inner nuclear layer, inner retinal network, ganglion cell layer, nerve fiber layer, and internal limiting membrane. Furthermore, lesion identification and nerve fiber layer vessel segmentation can be performed based on these multiple retinal layer structures. Therefore, during the training of the generative adversarial network, sample 3D images with pre-labeled pixel categories corresponding to each layer can be used as the training set for the segmentation module.
[0129] like Figure 4 As shown, for a single-frame OCT B-scan cross-sectional image, data enhancement including deformation, brightness adjustment, random regional brightness changes, global and local Gaussian blur, motion blur, and white noise is randomly added to obtain an enhanced single-frame B-scan. The enhanced data is input into a segmentation and denoising generative adversarial network. The segmentation module outputs the image segmentation result, which may include pixel-level resolution regions containing retinal layer structures, blood vessels, and lesions. The enhancement module outputs the single-frame B-scan reconstruction result, which is a clear B-scan image after denoising, signal enhancement, and deblurring.
[0130] In one possible implementation, a semi-supervised training method can be used to train the enhancement module to reconstruct each cross-sectional image in the target image, thereby improving the sharpness of the target image.
[0131] It is important to note that the objective functions of the segmentation module and the augmentation module are trained together. The weights between the objective functions of the two modules can be adjusted to maximize training efficiency and learning effectiveness.
[0132] Understandably, by using AI models (such as generative adversarial networks) to segment target images, the tissue parts of the specified target can be successfully identified and located, further improving the detail and resolution of the target image and providing more accurate information for medical analysis and diagnosis.
[0133] The process described in S101-S105 above uses cross-sectional images in a 3D image to determine the regions with motion artifacts and the regions without motion artifacts in the 3D image. For the regions with motion artifacts in the first image, the regions without motion artifacts in the second image are used for correction, thereby effectively eliminating the regions with motion artifacts in the first image and obtaining a target image that does not contain motion artifacts, thus effectively improving the image quality of the scanned image of the specified target.
[0134] Example 2
[0135] like Figure 2 The diagram shown is a flowchart of another image processing method provided in this application embodiment, including the following steps.
[0136] S201: For each 3D image, determine the centroid of each cross-sectional image in the 3D image.
[0137] The centroid of a cross-sectional image can also be called the center of gravity of the cross-sectional image. According to the mathematical concept of the center of gravity, the pixel value of each point in the cross-sectional image can be understood as the mass at that point.
[0138] In some examples, the calculation process of the centroid of the cross-sectional image can be found in Equation (1).
[0139]
[0140] In formula (1), C represents the coordinates of the centroid, p represents the coordinates of any point in the cross-sectional image, and g represents the feature function of the cross-sectional image, which can characterize the quality of each point (based on pixel values).
[0141] It is understandable that the various cross-sectional images together form a three-dimensional image, which can represent the three-dimensional structure of a specified target in a specified three-dimensional spatial coordinate system.
[0142] S202: Based on the centroids of each cross-sectional image, determine the coordinates of the centroids mapped onto the A-scan direction.
[0143] In some examples, the centroid coordinates of each cross-sectional image are mapped onto an A-scan (such as the z-axis in this embodiment) to obtain discrete mapped coordinates.
[0144] S203: Based on the coordinates of the centroid mapping along the A-scan direction, determine the gradient values corresponding to each mapped coordinate point.
[0145] After determining the mapped coordinates of the centroids, edge detection operators can be used to calculate the gradient value at each mapped coordinate point of the centroid. Edge detection operators include, but are not limited to, the Sobel operator, the Roberts operator, and the Prewitt operator.
[0146] S204: Determine the type of each cross-sectional image based on the gradient value corresponding to the mapped coordinate point of each centroid.
[0147] In this embodiment, after calculating the gradient value corresponding to the mapped coordinate point of each centroid using an edge detection operator, the standard deviation and mean deviation of the gradient values are calculated. Then, the sum of these standard deviations and mean deviations is used to determine the target threshold. Finally, the type of each cross-sectional image is determined based on the comparison between the gradient value corresponding to the mapped coordinate point of the centroid and the target threshold. This is a preferred method provided in this embodiment, but the method for determining the target threshold is not limited to this and can be selected according to the actual situation.
[0148] Optionally, for each cross-sectional image, if the gradient value corresponding to the mapped coordinate point of the centroid is greater than the target threshold, the cross-sectional image is determined to be an anomaly and marked as 1. If the comparison result indicates that the gradient value corresponding to the mapped coordinate point of the centroid is not greater than the target threshold, the cross-sectional image is determined to be normal and marked as 0. This yields a continuous set of labeled data.
[0149] It is understandable that the gradient value corresponding to the centroid mapping coordinate point of each cross-sectional image can characterize the position fluctuation of each cross-sectional image during the scanning imaging process. Based on the comparison result between the gradient value corresponding to the centroid mapping coordinate point and the target threshold, it can be understood as comparing the position fluctuation of each cross-sectional image in the same 3D image. Since the occurrence of motion artifacts is caused by the movement of a specified target, by comparing the position fluctuations between each cross-sectional image, the image with motion artifacts in the same 3D image can be determined.
[0150] To improve the accuracy of motion artifact identification, after determining the presence of motion artifacts in the cross-sectional image based on its centroid, further identification of these artifacts is necessary to eliminate judgment errors. Optionally, the further identification process for motion artifacts can be found in the steps and explanations of Examples 3 and 4.
[0151] Alternatively, machine learning models can be used to compare the gradient values at the centroid mapping coordinates of each cross-sectional image.
[0152] Optionally, the process of determining the type of each cross-sectional image based on the comparison between the gradient values corresponding to the centroid mapping coordinate points of each cross-sectional image can be as follows: generating a target time-domain signal based on the scanning time and corresponding gradient value of each cross-sectional image; using the target time-domain signal as the input of the second classification model to obtain the second classification result output by the second classification model; the second classification model is pre-trained based on the sample time-domain signal as input and combined with the pre-labeled type label of the sample time-domain signal as the training target; the second classification result is used to indicate the type of each cross-sectional image.
[0153] In some examples, the sample time-domain signal includes multiple sample scan times and corresponding sample gradient values. The type labels pre-labeled on the sample time-domain signal can characterize the type of each sample gradient value.
[0154] The processes described in S201-S204 above determine the motion artifact region in the three-dimensional image by comparing the gradient values corresponding to the centroid mapping coordinate points of each cross-sectional image in the same three-dimensional image, thereby achieving effective identification of the motion artifact region.
[0155] Example 3
[0156] like Figure 3 The diagram shown is a flowchart of another image processing method provided in this application embodiment, including the following steps.
[0157] After identifying the type of cross-sectional image, inaccurate identification may occur. For example, an image with motion artifacts may be identified as an image without motion artifacts. That is, in a set of continuous labeled data obtained in Embodiment 1 or Embodiment 2 of this invention, an image that should actually be labeled as 1 may be labeled as 0. Therefore, it is necessary to further identify whether any missed motion artifact images have not been identified. The specific steps are as follows:
[0158] S301: Dilate a set of labeled data for a 3D image containing motion artifact regions to obtain corresponding dilated labels.
[0159] S302: Determine the type of cross-sectional image corresponding to the dilation label.
[0160] In some examples, functions in the OpenCV library (such as `cv2.dilate`) can be called to dilate a set of labeled data from a 3D image. For instance, if a cross-sectional image in a 3D image, after type recognition, yields the following labeled data: 0, 0, 0, 0, 0, 1, 0, 0, 1, 1, 1, 1, 1, 0, 0, the dilated label would be: 0, 0, 0, 0, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 0. It can be seen that after dilation, the labels 0 adjacent to the original sixth, ninth, and thirteenth labels (1) also become 1. This means that images initially identified as motion-artifact-free regions, adjacent to images with motion artifacts, should actually be classified as images with motion artifacts.
[0161] The dilation step in this embodiment enables further identification of motion artifacts, avoids misjudging motion artifact images in 3D images, and effectively improves the accuracy of cross-sectional image type identification results.
[0162] Example 4
[0163] After identifying the type of each cross-sectional image in the 3D image, there may be a situation where multiple consecutive images are of the abnormal type (images with motion artifact regions). If the motion amplitude of the specified target represented by these images is relatively small and will not affect the image quality, then these images with motion artifact regions do not need to be corrected to reduce the amount of computation.
[0164] Optionally, it can be used to detect whether multiple consecutive images with motion artifact regions meet specified conditions. If they do, the type of multiple consecutive images with motion artifact regions is determined to be normal.
[0165] In some examples, the type of these motion artifact region images is further determined by detecting whether multiple consecutive images with motion artifact regions meet specified conditions. These specified conditions can be frame counts or offsets, etc. In this embodiment, a specified condition of 6 frames is set. When the number of consecutive images with motion artifact regions detected is less than or equal to 6 frames, the type of these images with motion artifact regions is determined to be normal. If there is already labeled data, the value is changed from 1 to 0 accordingly, and no correction is needed. When the number of consecutive images with motion artifact regions detected is greater than 6 frames, the type of these images with motion artifact regions is still determined to be abnormal, and correction is required.
[0166] The above detection process, based on the comparison between multiple consecutive images of regions with motion artifacts and specified conditions, achieves further identification of motion artifacts, effectively reducing subsequent operation steps and improving computational efficiency.
[0167] Example 5
[0168] like Figure 5 The diagram shown is a flowchart of another image processing method provided in this application embodiment, including the following steps.
[0169] It should be noted that the motion artifact region image in the first image is corrected using the motion artifact region image in the second image. The correction method can be to directly replace the corresponding frame's motion artifact region image with the motion artifact region image, or to use other processing methods to achieve the correction. This embodiment provides a preferred example of a correction method, with the specific steps as follows:
[0170] S501: Obtain the enface images of the first image and the second image, and use them as the first enface image and the second enface image, respectively.
[0171] After determining the first image and the second image, the two image data are superimposed along the Z-axis to obtain the corresponding first enface image and second enface image (e.g., ...). Figure 8 (xy plane image).
[0172] S502: Determine the positional offset between the first image and the second image based on the similarity between the first enface image and the second enface image.
[0173] In some examples, the NCC (Normalized Cross Correlation) algorithm can be used to calculate the similarity between the first enface image and the second enface image.
[0174] Generally speaking, the higher the similarity between the first and second enface images, the smaller the positional offset between the first and second images; conversely, the lower the similarity between the first and second enface images, the greater the positional offset between the first and second images.
[0175] In one possible implementation, after determining the similarity between the first enface image and the second enface image, the position offset corresponding to the similarity can be queried from a first relation table, which includes multiple sample similarities and the position offset corresponding to each sample similarity.
[0176] S503: Determine the target coordinates based on the position offset and the first coordinate.
[0177] The first coordinate represents the coordinate that matches the region of motion artifacts in the first image.
[0178] In some examples, the positional offset can be represented by a coordinate offset. The first coordinate and the coordinate offset are added together to obtain the target coordinates in the second image corresponding to the first coordinate.
[0179] In some examples, if there are multiple second images, after determining the positional offset between any second image and the first image, the target coordinates corresponding to any second image are calculated based on the positional offset and the first coordinates.
[0180] In some examples, if there are multiple motion artifact regions in the first image, after determining the positional offset between the first and second images, the target coordinates of each motion artifact region in the second image are determined by combining the first coordinates corresponding to each motion artifact region in the first image.
[0181] S504: From the multiple motion-free artifact region images contained in the second image, determine the target image that matches the target coordinates.
[0182] Among them, the image of the region without motion artifacts includes a normal cross-sectional image.
[0183] It should be noted that, due to the positional deviation between the second image and the first image, the target image that meets the conditions can be located from the multiple motion artifact region images contained in the second image using the target coordinates. The so-called conditions refer to the fact that the second image and the motion artifact region image in the first image both point to the same tissue part in the specified target.
[0184] In some examples, there are multiple second images. If any of the multiple motion-free region images contained in a second image cannot cover all the motion-artifact region images in the first image, then the target image is continued to be filtered from the multiple motion-free region images contained in the remaining second images until all the motion-artifact region images in the first image can be matched with the corresponding target image.
[0185] S505: Use the target image with no motion artifacts to replace the image with motion artifacts in the first image to obtain the target image.
[0186] In this method, the target image is obtained by replacing the motion artifact region in the first image with the target image. This target image not only represents the three-dimensional structure of the specified target, but also does not contain motion artifacts. The image quality of the target image is much higher than any three-dimensional image in the scanned image set, thus providing the most objective and accurate diagnostic basis for the specified target.
[0187] The above-described processes S501-S505 can replace the motion artifact region image in the first image with the motion artifact region image in the second image to obtain a high-quality target image and provide accurate diagnostic basis for the specified target.
[0188] Example 6
[0189] like Figure 6 The diagram shown is a flowchart of another image processing method provided in this application embodiment, including the following steps.
[0190] The method of this invention yields an image that does not contain motion artifact regions. However, the images in this image may exhibit jitter, causing positional deviations between images in a specified direction (e.g., the z-axis direction in the image coordinate system), requiring further alignment. The specific steps are as follows:
[0191] S601: Replace the motion artifact region in the first image with the target image to obtain the third image.
[0192] If there are multiple motion artifact region images in the first image, then after determining the target image that matches each motion artifact region image, the target image is used to replace each motion artifact region image in the first image to obtain the third image.
[0193] S602: Determine the scanning order of each cross-sectional image in the third image.
[0194] The scanning order of each cross-sectional image in the third image can be determined based on the way the OCT device scans the specified target.
[0195] Generally speaking, if the third image is obtained by correcting the first image, then the scanning order of each cross-sectional image in the third image is the same as the scanning order of each cross-sectional image in the first image.
[0196] S603: For each cross-sectional image in the third image, determine the positional deviation of the cross-sectional image in a specified direction based on the similarity between the cross-sectional image and the cross-sectional image of the previous or next frame.
[0197] It is understandable that the so-called previous frame cross-sectional image refers to the cross-sectional image that was scanned earlier than the cross-sectional image, and the next frame cross-sectional image refers to the cross-sectional image that was scanned later than the cross-sectional image.
[0198] In some examples, the NCC algorithm can be used to calculate the similarity between the cross-sectional image and the cross-sectional image of the previous or next frame.
[0199] Generally speaking, the higher the similarity between the cross-sectional image and the cross-sectional image of the previous or next frame, the smaller the positional deviation of the cross-sectional image in the specified direction. Conversely, the lower the similarity between the cross-sectional image and the cross-sectional image of the previous or next frame, the greater the positional deviation of the cross-sectional image in the specified direction.
[0200] S604: Adjust the coordinates of each cross-sectional image according to the positional deviation of each cross-sectional image in a specified direction to obtain a fourth image.
[0201] Specifically, the coordinates of each cross-sectional image are adjusted based on the positional deviation of each cross-sectional image. This can be understood as adding the positional deviation of each cross-sectional image in a specified direction to the coordinates of each cross-sectional image, thereby changing the coordinates of each cross-sectional image and obtaining the fourth image.
[0202] In some examples, the first frame cross-sectional image is selected as the initial preceding frame cross-sectional image. Its similarity to the second frame cross-sectional image is calculated to obtain the positional deviation of the second frame cross-sectional image. Then, the coordinates of the second frame cross-sectional image are adjusted based on the positional deviation. This adjusted cross-sectional image is then used as the preceding frame cross-sectional image, and its similarity and positional deviation to the following frame cross-sectional image are calculated. The coordinates of the following frame cross-sectional image are adjusted based on the positional deviation, and this process is repeated to obtain a fourth image composed of images with no height difference along the z-axis. The selection of the initial preceding frame cross-sectional image is not limited to the first frame; it can also be a cross-sectional image located in the middle frame (middle images contain more information). Then, the height difference adjustment is repeated from the middle image to the left and right ends of the preceding and following frames. This method provides higher calculation accuracy.
[0203] S605: Filter the fourth image to obtain the target image.
[0204] Since each cross-sectional image in the fourth image is obtained by adjusting each cross-sectional image in the third image in a specified direction, there may be noise in each cross-sectional image in the fourth image in the specified direction. Therefore, the fourth image needs to be filtered to eliminate the noise in each cross-sectional image in the fourth image in the specified direction and improve the image quality of the target image.
[0205] The processes described in S601-S605 above can eliminate possible positional offsets and noise in the specified directions of each cross-sectional image in the first image, so as to obtain a high-quality target image.
[0206] Example 7
[0207] Corresponding to the image processing methods provided in the above embodiments, this application also provides an image processing apparatus.
[0208] like Figure 7 The diagram shown is a schematic representation of the architecture of an image processing apparatus provided in an embodiment of this application, including the following units.
[0209] The image acquisition unit 100 is used to obtain a set of scanned images of a specified target using optical coherence tomography (OCT) technology; the set of scanned images includes multiple three-dimensional images that are sequential in scanning time; the three-dimensional images include multiple cross-sectional images; the cross-sectional images are used to characterize the two-dimensional structure of the specified target at a specified level.
[0210] Image classification unit 200 is used to identify the type of each cross-sectional image in each of the three-dimensional images.
[0211] Optionally, the image classification unit 200 is specifically configured to: for each of the three-dimensional images, use each cross-sectional image of the three-dimensional image as input to a first classification model to obtain a first classification result output by the first classification model; the first classification model is pre-trained based on the sample cross-sectional image as input and combined with the pre-labeled type label of the sample cross-sectional image as the training target; the first classification result includes the type of each cross-sectional image.
[0212] Optionally, the image classification unit 200 is specifically used for: for each of the three-dimensional images, determining the centroid of each cross-sectional image in the three-dimensional image; determining the coordinates of the centroid mapped in the A-scan direction based on the centroid of each cross-sectional image; determining the gradient value corresponding to each mapped coordinate point based on the coordinates of the centroid mapped in the A-scan direction; and determining the type of each cross-sectional image based on the gradient value.
[0213] Optionally, the image classification unit 200 is specifically used to: determine a target threshold based on the gradient value corresponding to the mapping coordinate point of each centroid; and determine the type of each cross-sectional image based on the comparison result between the gradient value corresponding to the mapping coordinate point of the centroid and the target threshold.
[0214] Optionally, the image classification unit 200 is specifically used to: generate a target time-domain signal based on the scanning time and corresponding gradient value of each of the cross-sectional images; use the target time-domain signal as input to a second classification model to obtain a second classification result output by the second classification model; the second classification model is pre-trained based on the sample time-domain signal as input and combined with the pre-labeled type label of the sample time-domain signal as the training target; the second classification result is used to indicate the type of each of the cross-sectional images.
[0215] Optionally, the image classification unit 200 is further configured to: label cross-sectional images of type abnormal (i.e., images with motion artifacts) as 1, and label cross-sectional images of type normal (i.e., images without motion artifacts) as 0, thereby obtaining a set of continuous labeled data; dilate the labeled data to obtain corresponding dilated labels; and determine the type of cross-sectional image corresponding to the dilated label.
[0216] Optionally, the image classification unit 200 is further configured to: detect whether multiple consecutive images with motion artifact regions meet specified conditions; if they do, determine that the type of the multiple consecutive images with motion artifact regions is normal.
[0217] The image recognition unit 300 is used to identify motion artifact regions in each of the three-dimensional images based on cross-sectional images of the type of anomaly; the motion artifact regions represent areas where there are anomalies in the scanning imaging.
[0218] The image filtering unit 400 is used to determine the first image with the fewest motion artifact region images from each of the three-dimensional images containing motion artifact region images.
[0219] The image correction unit 500 is used to correct the motion artifact region image in the first image using the motion artifact region image in the second image to obtain the target image; the second image is other three-dimensional images in the scanned image set besides the first image; the target image is used as the diagnostic basis for the specified target.
[0220] Optionally, the image correction unit 500 is specifically configured to: acquire enface images of a first image and a second image, respectively serving as the first enface image and the second enface image; determine the positional offset between the first image and the second image based on the similarity between the first enface image and the second enface image; determine target coordinates based on the positional offset and a first coordinate; the first coordinate represents the coordinates that match the motion artifact region image in the first image; determine the target image that matches the target coordinates from a plurality of motion artifact-free region images contained in the second image; the motion artifact-free region images include normal cross-sectional images; and use the target image to replace the motion artifact region image in the first image to obtain the target image.
[0221] Optionally, the image correction unit 500 is specifically configured to: replace the motion artifact region image in the first image with the target image to obtain a third image; determine the scanning order of each cross-sectional image in the third image; for each cross-sectional image in the third image, determine the positional deviation of the cross-sectional image in a specified direction based on the similarity between the cross-sectional image and the cross-sectional image of the previous or next frame; adjust the coordinates of each cross-sectional image based on the positional deviation of each cross-sectional image in the specified direction to obtain a fourth image; and filter the fourth image to obtain the target image.
[0222] The image segmentation unit 600 is used to: perform image segmentation on the target image using a pre-trained generative adversarial network to determine the category label of each pixel in the target image; the category label is used to characterize the constituent parts of the specified target; the generative adversarial network is also used to denoise the target image to improve the clarity of the target image.
[0223] Each of the units described above determines the motion artifact region and the motion artifact-free region in the three-dimensional image through various cross-sectional images in the three-dimensional image. For the motion artifact region in the first image, the motion artifact-free region in the second image is used for correction to obtain a target image without motion artifacts, effectively improving the image quality of the scanned image of the specified target.
[0224] Furthermore, the functions described above in the embodiments of this application can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SOCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0225] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
[0226] While several specific implementation details are included in the foregoing discussion, these should not be construed as limiting the scope of this application. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0227] The above description is merely a preferred embodiment of this application and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this application is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this application.
Claims
1. An image processing method, characterized in that, include: A set of scanned images of a specified target is obtained using optical coherence tomography (OCT). The set of scanned images includes multiple three-dimensional images that are sequential in scanning time. Each three-dimensional image includes multiple cross-sectional images. The cross-sectional images are used to characterize the two-dimensional structure of the specified target at a specified level. Identify the type of each cross-sectional image in each of the three-dimensional images; Based on the cross-sectional image of type anomalous, the motion artifact region in each of the three-dimensional images is identified; the motion artifact region represents the area where there is anomaly in the scanning imaging. From all the 3D images containing motion artifact regions, determine the first image with the fewest motion artifact regions. The motion artifact region in the first image is corrected using the motion artifact region in the second image to obtain the target image; the second image is any other three-dimensional image in the scanned image set besides the first image; the target image is used as the diagnostic basis for the specified target.
2. The method according to claim 1, characterized in that, Identifying the type of each cross-sectional image in each of the three-dimensional images, including: For each of the three-dimensional images, each cross-sectional image of the three-dimensional image is used as input to a first classification model to obtain a first classification result output by the first classification model; the first classification model is pre-trained based on the sample cross-sectional image as input and combined with the pre-labeled type label of the sample cross-sectional image as the training target; the first classification result includes the type of each cross-sectional image.
3. The method according to claim 1, characterized in that, Identifying the type of each cross-sectional image in each of the three-dimensional images, including: For each of the three-dimensional images, determine the centroid of each cross-sectional image in the three-dimensional image; Based on the centroid of each of the cross-sectional images, determine the coordinates of the centroid mapped in the A-scan direction; Based on the coordinates of the centroid mapping in the A-scan direction, determine the gradient value corresponding to each mapped coordinate point; Based on the gradient value, the type of each cross-sectional image is determined.
4. The method according to claim 3, characterized in that, Based on the gradient values, the type of each cross-sectional image is determined, including: The target threshold is determined based on the gradient value corresponding to the mapped coordinate point of each centroid. The type of each cross-sectional image is determined by comparing the gradient value corresponding to the mapped coordinate point of the centroid with the target threshold.
5. The method according to claim 3, characterized in that, Based on the gradient values, the type of each cross-sectional image is determined, including: Based on the scanning time and corresponding gradient value of each of the cross-sectional images, a target time-domain signal is generated; The target time-domain signal is used as input to the second classification model to obtain the second classification result output by the second classification model; the second classification model is pre-trained based on the sample time-domain signal as input and combined with the pre-labeled type label of the sample time-domain signal as the training target; the second classification result is used to indicate the type of each of the cross-sectional images.
6. The method according to claim 1, characterized in that, After identifying the type of each cross-sectional image in each of the three-dimensional images, the method further includes: Cross-sectional images of type anomalous (i.e., images with motion artifacts) are labeled as 1, and cross-sectional images of type normal (i.e., images without motion artifacts) are labeled as 0. This yields a continuous set of labeled data. Inflate the labeled data to obtain the corresponding inflated labels; Determine the type of the cross-sectional image corresponding to the expansion label.
7. The method according to claim 1, characterized in that, After identifying the type of each cross-sectional image in each of the three-dimensional images, the method further includes: The system detects whether multiple consecutive images with motion artifact regions meet specified conditions. If they do, the system determines that the type of the multiple consecutive images with motion artifact regions is normal.
8. The method according to claim 1, characterized in that, Using the motion artifact-free region image in the second image, the motion artifact region image in the first image is corrected to obtain the target image, including: Obtain the enface images of the first image and the second image, and use them as the first enface image and the second enface image, respectively; The positional offset between the first image and the second image is determined based on the similarity between the first enface image and the second enface image; The target coordinates are determined based on the position offset and the first coordinate; the first coordinate represents the coordinates that match the region of motion artifact in the first image. From the multiple motion-free artifact region images contained in the second image, a target image matching the target coordinates is determined; the motion-free artifact region images include normal cross-sectional images. The target image is obtained by replacing the motion artifact region in the first image with the target image.
9. The method according to claim 8, characterized in that, Using the target image to replace the motion artifact region in the first image to obtain the target image includes: The target image is used to replace the motion artifact region in the first image to obtain a third image. Determine the scanning order of each cross-sectional image in the third image; For each cross-sectional image in the third image, the positional deviation of the cross-sectional image in a specified direction is determined based on the similarity between the cross-sectional image and the cross-sectional image of the previous or next frame. The coordinates of each cross-sectional image are adjusted according to the positional deviation of each cross-sectional image in a specified direction to obtain a fourth image; The fourth image is filtered to obtain the target image.
10. The method according to claim 1, characterized in that, The method further includes: The target image is segmented using a pre-trained generative adversarial network to determine the category label of each pixel in the target image; the category label is used to characterize the constituent parts of the specified target; the generative adversarial network is also used to denoise the target image to improve the clarity of the target image.
11. An image processing apparatus, characterized in that, include: An image acquisition unit is used to obtain a set of scanned images of a specified target using optical coherence tomography (OCT) technology; the set of scanned images includes multiple three-dimensional images that are sequential in scanning time; the three-dimensional images include multiple cross-sectional images; the cross-sectional images are used to characterize the two-dimensional structure of the specified target at a specified level; An image classification unit is used to identify the type of each cross-sectional image in each of the three-dimensional images; The image recognition unit is used to identify motion artifact regions in each of the three-dimensional images based on cross-sectional images of type anomalous; the motion artifact regions represent areas where there are anomalous scanning images. The image filtering unit is used to determine the first image with the fewest motion artifact regions from each of the three-dimensional images containing motion artifact regions. An image correction unit is used to correct a motion artifact region in a first image using a motion artifact-free region in a second image to obtain a target image; the second image is any other three-dimensional image in the scanned image set other than the first image; the target image is used as a diagnostic basis for the specified target.