A reconstruction method based on sparse view CT images
By preprocessing sparse-view CT images and using an improved scale-invariant feature transform algorithm to determine the initial point cloud, and combining it with a three-dimensional Gaussian splash model to generate new-view CT images, the problem of low reconstruction accuracy of sparse-view CT images is solved, and high-precision image reconstruction results are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING HANGXING MACHINERY MFG CO LTD
- Filing Date
- 2026-01-21
- Publication Date
- 2026-05-29
AI Technical Summary
Existing technologies suffer from low image accuracy in sparse-view CT image reconstruction, especially in scenarios with strict limitations on scanning speed and cost, such as high-speed security inspection lines, making it difficult to meet the requirements for high-precision detection.
By acquiring and preprocessing multiple sparse viewpoint CT images, an improved scale-invariant feature transform algorithm is used to determine the initial point cloud. Then, a new viewpoint CT image is generated using a three-dimensional Gaussian splash model. Finally, image reconstruction is performed. Marr wavelet functions with multiple scale spatial parameters are used to determine key feature points and to statistically analyze gradient direction histograms to reduce noise and improve image accuracy.
It improves the accuracy of sparse-view CT image reconstruction, generates high-quality new-view images, meets the requirements of high-precision detection, and reduces computational costs and time consumption.
Smart Images

Figure CN122115642A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of three-dimensional reconstruction technology, and in particular to a reconstruction method based on sparse viewpoint CT images. Background Technology
[0002] In recent years, X-ray-based computed tomography (CT) technology has become a core technology in security inspection, medical diagnosis, and industrial non-destructive testing. It projects images of objects from multiple angles and uses reconstruction algorithms to generate three-dimensional images of their internal structures.
[0003] However, in some applications with strict limitations on scanning speed or cost, such as high-speed security inspection lines, it is difficult to perform comprehensive and dense omnidirectional spectral sampling of target objects; typically, only a limited number of sparse spectral slice images can be obtained. This limitation directly leads to problems such as degraded reconstructed image quality and loss of detail, making it difficult to meet the requirements of high-precision detection.
[0004] To overcome the challenges posed by sparse perspectives, existing technologies offer methods based on neural radiation fields. These methods implicitly represent 3D scenes by training neural networks, enabling the generation of new perspectives from sparse inputs. However, these methods rely on voxel rendering techniques, requiring extensive dense sampling to calculate ray integrals. This results in inherent drawbacks such as high computational costs, slow training and rendering speeds, and massive memory consumption, limiting their application in real-time scenarios with high performance requirements.
[0005] Another approach is to first reconstruct a 3D sparse point cloud from a 2D image sequence using traditional multi-view geometric algorithms such as structure-of-motion reconstruction. Although this method can directly generate a clear geometric representation, the point cloud it generates is usually low in density, noisy, and lacks the ability to continuously model the appearance of the scene, such as color and texture. This makes it difficult to directly use it to generate high-quality new perspective images, thus affecting the accuracy of the final CT reconstruction.
[0006] Therefore, there is an urgent need for a new technical solution for reconstructing sparse viewpoint images. Summary of the Invention
[0007] Based on the above analysis, the present invention aims to provide a reconstruction method based on sparse viewpoint CT images to solve the problem of low image accuracy in existing reconstructions based on sparse viewpoint images.
[0008] This invention provides a reconstruction method based on sparse viewpoint CT images, the reconstruction method comprising:
[0009] Multiple sparse-view CT images are acquired, and each sparse-view CT image is preprocessed to obtain multiple preprocessed CT images; wherein, the view interval between two adjacent sparse-view CT images is the same. The initial point cloud is determined based on multiple preprocessed CT images and an improved scale-invariant feature transform algorithm. The initial point cloud is input into the new perspective image generation model to obtain multiple new perspective CT images; the new perspective image generation model is trained based on a three-dimensional Gaussian splash model. Image reconstruction is performed based on multiple sparse view CT images and multiple new view CT images to obtain the reconstructed CT image.
[0010] Based on a further improvement to the above reconstruction method, the determination of the initial point cloud based on multiple preprocessed CT images and an improved scale-invariant feature transform algorithm includes: Key feature points in multiple preprocessed CT images were determined using Marr wavelet functions with multiple scale-space parameters. Match key feature points in all adjacent preprocessed CT images to obtain matching feature point pairs in all adjacent preprocessed CT images; Based on the matching feature point pairs in each pair of adjacent preprocessed CT images, the camera parameters of each pair of adjacent preprocessed CT images are determined, and the camera parameters of multiple preprocessed CT images are obtained. The initial point cloud is determined based on matching feature point pairs in all two adjacent preprocessed CT images and camera parameters of multiple preprocessed CT images.
[0011] Based on a further improvement to the above reconstruction method, the key feature points in multiple preprocessed CT images are determined using Marr wavelet functions with multiple scale-space parameters, including: Each preprocessed CT image is convolved using the Marr wavelet function with multiple scale space parameters to obtain multiple scale space images corresponding to each preprocessed CT image with multiple scale space parameters. Key feature points in each preprocessed CT image are determined using multiple scale-space images corresponding to each preprocessed CT image.
[0012] Based on the further improvement of the above reconstruction method, the expression of the Marr wavelet function is: ; in, , This represents the coordinates of a pixel in each preprocessed CT image. Represents scale-space parameters. This represents the weight values of the convolution kernel.
[0013] Based on a further improvement to the above reconstruction method, the step of determining key feature points in each preprocessed CT image using multiple scale-space images corresponding to each preprocessed CT image includes: For each preprocessed CT image, all pixels are traversed to determine whether the pixel is an extreme point in different scale spatial images; if the pixel is an extreme point in different scale spatial images, then the pixel is a key feature point in each preprocessed CT image. The neighborhood of each key feature point is divided into multiple grids, and the gradient direction histogram in each grid is statistically analyzed to obtain the descriptor of each key feature point.
[0014] Based on further improvements to the above reconstruction method, the multi-layer mesh includes 4×4 mesh and 2×2 mesh.
[0015] Based on a further improvement to the above reconstruction method, the step of statistically analyzing the gradient direction histogram in each mesh layer to obtain the descriptor of each key feature point includes: The gradient direction histogram in each grid layer is used as the descriptor for each key feature point in each grid layer; PCA dimensionality reduction update is performed on the descriptor of each key feature point in each grid layer to preserve the principal energy components; The updated descriptors of each key feature point in each grid layer are concatenated to obtain the descriptor of each key feature point.
[0016] Based on a further improvement of the above reconstruction method, the scale space parameter is 1 or 2.
[0017] Based on the further improvement of the above reconstruction method, the loss function used when training the 3D Gaussian splash model is: ; ; ; ; in, This represents the loss value for each new viewpoint image. Represents the weight parameters. Indicates the first image in each new perspective image Pixel prediction value of each pixel. In each new perspective image, the first The actual pixel value of each pixel. This indicates the number of pixels in each new viewpoint image. , This represents the predicted and actual brightness values of all pixels in each new viewpoint image. , Represents the local minimum constant. , This represents the predicted and actual values of the brightness variance of all pixels in each new viewpoint image. This represents the covariance between the predicted and actual brightness values of all pixels in each new viewpoint image.
[0018] Based on further improvements to the above reconstruction method, the image reconstruction employs any of the following algorithms: Filtered back projection algorithm; Convolution back projection algorithm; Differential-Hilbert backprojection algorithm; Gradient descent algorithm; Maximum likelihood iterative reconstruction algorithm.
[0019] Compared with the prior art, the present invention can achieve at least one of the following beneficial effects: 1. Based on multiple sparse view CT images with the same view interval, an improved scale-invariant feature transform algorithm is used to obtain an initial point cloud with low noise. A new view image generation model trained based on a 3D Gaussian splash model is used to obtain multiple new view CT images. Image reconstruction is performed based on multiple sparse view CT images and multiple new view CT images, which improves the accuracy of the reconstructed image. 2. Using Marr wavelet functions with multiple scale spatial parameters, key feature points in each preprocessed CT image are determined. The neighborhood of each key feature point is divided into multiple grids. The gradient direction histogram in each grid is statistically analyzed to obtain the descriptor of each key feature point, which reduces the noise of each key feature point and provides a basis for further improving the accuracy of the reconstructed image.
[0020] In this invention, the above-described technical solutions can be combined with each other to achieve more preferred combinations. Other features and advantages of this invention will be set forth in the following description, and some advantages may become apparent from the description or be learned by practicing the invention. The objects and other advantages of this invention can be realized and obtained from what is particularly pointed out in the description and drawings. Attached Figure Description
[0021] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts.
[0022] Figure 1 This is a flowchart illustrating a reconstruction method based on sparse viewpoint CT images provided in an embodiment of the present invention. Detailed Implementation
[0023] Preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, which form part of this application and are used together with the embodiments of the present invention to illustrate the principles of the present invention, but are not intended to limit the scope of the present invention.
[0024] One specific embodiment of the present invention discloses a reconstruction method based on sparse viewpoint CT images, such as... Figure 1 As shown, the reconstruction method includes: Step S1: Acquire multiple sparse view CT images and preprocess each sparse view CT image to obtain multiple preprocessed CT images; wherein, the view interval between two adjacent sparse view CT images is the same. Step S2: Determine the initial point cloud based on multiple preprocessed CT images and an improved scale-invariant feature transform algorithm; Step S3: Input the initial point cloud into the new perspective image generation model to obtain multiple new perspective CT images; wherein, the new perspective image generation model is trained based on a three-dimensional Gaussian splash model; Step S4: Reconstruct the image based on multiple sparse view CT images and multiple new view CT images to obtain the reconstructed CT image.
[0025] Specifically, such as Figure 1 As shown, in step S1, when acquiring sparse-view CT images, it is necessary to ensure that the viewing angle interval between two adjacent sparse-view CT images is the same. For example, four sparse-view CT images are acquired, wherein the viewing angle of the first sparse-view image is 0°, the viewing angle of the second sparse-view image is 90°, the viewing angle of the third sparse-view image is 180°, and the viewing angle of the fourth sparse-view image is 270°.
[0026] It is worth noting that the number of sparse view CT images should be set reasonably according to actual cost, time and other requirements, with a minimum of 3 images.
[0027] Specifically, such as Figure 1 As shown, in step S1, each sparse view CT image is preprocessed to reduce noise interference. The preprocessing methods include one or more of image denoising, contrast enhancement, geometric correction and distortion compensation, and data standardization, which are set reasonably according to actual needs.
[0028] Specifically, noise, such as quantum noise and electronic noise, is inevitably introduced during the acquisition of sparse-view CT images. This noise interferes with the accuracy of feature point extraction and matching, resulting in a large number of outliers or inaccuracies in the generated point cloud. Image denoising is then necessary, and methods that can be used include: non-local mean denoising: utilizing the redundant information of all pixels in the sparse-view CT image for denoising, which can effectively preserve edge and texture details; and anisotropic diffusion filtering: smoothing noise while diffusing it along the edge direction of the sparse-view CT image, thereby protecting edge information.
[0029] Specifically, due to differences in scanning conditions or object density, sparse-view CT images may have insufficient contrast, resulting in unclear features. In this case, contrast enhancement can make the edges and internal texture features of the object more prominent, making it easier for feature detection algorithms to extract key points more stably. The methods that can be used include: histogram equalization: redistributing the pixel intensity values of the image to enhance global contrast; contrast-limited adaptive histogram equalization: performing histogram equalization in local areas and limiting the contrast amplification, avoiding the amplification of noise, which is usually better than global histogram equalization.
[0030] Specifically, image distortion caused by misalignment of the X-ray source, detector, and rotation center may result in geometric distortion in the CT system. In this case, geometric correction and distortion compensation can ensure the accurate spatial correspondence of sparse view CT images under different viewpoints. This is the basis for subsequent accurate triangulation to generate point clouds. The method that can be adopted is to use the calibration parameters of the CT system, such as focal length, principal point, distortion coefficient, etc., to correct the original sparse view CT image and eliminate lens distortion and perspective error.
[0031] Specifically, data standardization can be achieved using Z-Score standardization: converting pixel values in sparse viewpoint CT images into a distribution with a mean of 0 and a standard deviation of 1, which can make the data distribution more stable and accelerate model convergence.
[0032] Specifically, such as Figure 1 As shown, in step S2, the initial point cloud is determined by using multiple preprocessed CT images obtained in step S1, combined with the improved scale-invariant feature transformation algorithm.
[0033] Preferably, the step of determining the initial point cloud based on multiple preprocessed CT images and combined with an improved scale-invariant feature transform algorithm includes: Key feature points in multiple preprocessed CT images were determined using Marr wavelet functions with multiple scale-space parameters. Match key feature points in all adjacent preprocessed CT images to obtain matching feature point pairs in all adjacent preprocessed CT images; Based on the matching feature point pairs in each pair of adjacent preprocessed CT images, the camera parameters of each pair of adjacent preprocessed CT images are determined, and the camera parameters of multiple preprocessed CT images are obtained. The initial point cloud is determined based on matching feature point pairs in all two adjacent preprocessed CT images and camera parameters of multiple preprocessed CT images.
[0034] Specifically, multiple scale space parameters and Marr wavelet functions are pre-set.
[0035] Preferably, the expression for the Marr wavelet function is: ; in, , This represents the coordinates of a pixel in each preprocessed CT image. Represents scale-space parameters. This represents the weight values of the convolution kernel.
[0036] Preferably, the scale space parameter is 1 or 2.
[0037] Specifically, when setting the Marr wavelet function, the corresponding convolution kernel weight values are determined under different scale space parameters by combining the pre-set scale space parameters.
[0038] Specifically, key feature points in multiple preprocessed CT images are determined using Marr wavelet functions with multiple scale-space parameters.
[0039] Preferably, the step of using Marr wavelet functions with multiple scale-space parameters to determine key feature points in multiple preprocessed CT images includes: Each preprocessed CT image is convolved using the Marr wavelet function with multiple scale space parameters to obtain multiple scale space images corresponding to each preprocessed CT image with multiple scale space parameters. Key feature points in each preprocessed CT image are determined using multiple scale-space images corresponding to each preprocessed CT image.
[0040] Specifically, the convolution kernel weights corresponding to all pixels in each preprocessed CT image are determined using pre-set scale space parameters and Marr wavelet functions. The pixel value of the pixel is obtained by multiplying the convolution kernel weights by the pixel value. By traversing all scale space parameters, multiple scale space images corresponding to multiple scale space parameters for each preprocessed CT image can be obtained.
[0041] Specifically, after obtaining multiple scale-space images corresponding to each preprocessed CT image, the key feature points in each preprocessed CT image are determined using these multiple scale-space images.
[0042] Preferably, the step of determining key feature points in each preprocessed CT image using multiple scale-space images corresponding to each preprocessed CT image includes: For each preprocessed CT image, all pixels are traversed to determine whether the pixel is an extreme point in different scale spatial images; if the pixel is an extreme point in different scale spatial images, then the pixel is a key feature point in each preprocessed CT image. The neighborhood of each key feature point is divided into multiple grids, and the gradient direction histogram in each grid is statistically analyzed to obtain the descriptor of each key feature point.
[0043] Specifically, each pixel in the preprocessed CT image is traversed to determine whether it is an extremum point in different scale spatial images. For example, the neighborhood of a pixel is set to 7×7, and the pixel is used as the center pixel of the 7×7 neighborhood. It is determined whether the center pixel is an extremum point in the 7×7 neighborhood. If it is an extremum point, then the center pixel is an extremum point in the scale spatial image.
[0044] Specifically, if a pixel is an extreme point in each of the pre-set scale space images, then that pixel is used as a key feature point in the preprocessed CT image.
[0045] Specifically, the neighborhood of key feature points is pre-defined, for example, the neighborhood of key feature points is set to 17×17 or 33×33. After removing the center row and center column, the neighborhood of each key feature point is divided into multiple sub-neighborhoods.
[0046] Preferably, the multi-layered grid includes a 4×4 grid and a 2×2 grid.
[0047] Specifically, the neighborhood of the key feature points is set to 17×17. After removing the center row and center column, the remaining 16×16 pixel area is divided into a 4×4 grid. Each sub-neighborhood is a 4×4 area containing 16 pixels. The remaining 16×16 pixel area is divided into a 2×2 grid. Each sub-neighborhood is an 8×8 area containing 64 pixels.
[0048] Specifically, the neighborhood of the key feature points is set to 33×33. After removing the center row and center column, the remaining 32×32 pixel area is divided into a 4×4 grid, so each sub-neighborhood is an 8×8 area containing 64 pixels; the remaining 32×32 pixel area is divided into a 2×2 grid, so each sub-neighborhood is a 16×16 area containing 256 pixels.
[0049] Specifically, the neighborhood of each key feature point is divided into multiple grids, and the gradient direction histogram in each grid is statistically analyzed to obtain the descriptor of each key feature point.
[0050] Preferably, the step of obtaining a descriptor for each key feature point by statistically analyzing the gradient direction histogram in each grid layer includes: The gradient direction histogram in each grid layer is used as the descriptor for each key feature point in each grid layer; PCA dimensionality reduction update is performed on the descriptor of each key feature point in each grid layer to preserve the principal energy components; The updated descriptors of each key feature point in each grid layer are concatenated to obtain the descriptor of each key feature point.
[0051] Specifically, when the neighborhood of a key feature point is 17×17, the 4×4 grid includes 16 sub-neighborhoods, each containing 16 pixels. The gradient direction histogram for each sub-neighborhood is calculated. For example, there are eight gradient directions: 0°, 45°, 90°, 135°, 180°, 225°, 270°, and 315°. The sum of pixel values for each gradient direction within each sub-neighborhood is calculated and used as the descriptor for that sub-neighborhood. The descriptors of all sub-neighborhoods are combined to form the descriptor for the 4×4 grid layer.
[0052] It is worth noting that this invention utilizes both 4×4 and 2×2 grids to calculate the descriptor of each key feature point in each grid layer.
[0053] Specifically, each key feature point includes 16 sub-neighborhoods in a 4×4 grid, and each sub-neighborhood includes 8 dimensions of description. Therefore, the descriptor of each key feature point in the 4×4 grid includes a total of 128 dimensions of description.
[0054] Similarly, the descriptor for each key feature point in the 2×2 grid includes a total of 32 dimensions.
[0055] Specifically, PCA dimensionality reduction is performed on the descriptor of each key feature point in each grid layer to preserve the principal components of energy. For example, PCA dimensionality reduction is performed on a 32-dimensional vector generated by a 2×2 grid and a 128-dimensional vector generated by a 4×4 grid, respectively, preserving 95% of the principal components of energy, reducing the 32-dimensional vector to 20-dimensional and the 128-dimensional vector to 60-dimensional.
[0056] Specifically, the descriptors of each key feature point in each layer of the grid after the update are concatenated to obtain the descriptor of each key feature point.
[0057] For example, the descriptor for each key feature point (1) 80) is: [0.1, 0.2, 0.3, 0.15, 0.05, 0.08, 0.02, 0.1, 0.12, 0.25, 0.3, 0.1, 0.03, 0.05, 0.05, 0.05, 0.08, 0.15, 0.2, 0.3, 0.1, 0.05, 0.03, 0.04, 0.05, 0.1, 0.2, 0.25, 0.15, 0.05, 0.03, 0.02, 0.05, 0.08, 0.12, 0.2, 0.3, 0.15, 0.03, 0.05, 0.04, 0.06, 0.1, 0.2, 0.22, 0.18, 0.05, 0.03, 0.02, 0.05, 0.06, 0.12, 0.25, 0.3, 0.1, 0.03, 0.04, 0.05, 0.07, 0.1, 0.2, 0.2, 0.15, 0.05, 0.03, 0.02, 0.05, 0.08, 0.15, 0.2, 0.3, 0.1, 0.05, 0.03, 0.04, 0.05, 0.1, 0.2, 0.25, 0.15, 0.04, 0.06).
[0058] Specifically, Marr wavelet functions with multiple scale-space parameters are used to determine key feature points and descriptors of key feature points in multiple preprocessed CT images.
[0059] Specifically, for example, by acquiring four sparse-view CT images, the key feature points and their descriptors in the four preprocessed CT images are determined.
[0060] Specifically, key feature points in all two adjacent preprocessed CT images are matched to obtain matching feature point pairs in all two adjacent preprocessed CT images.
[0061] For example, a preprocessed CT image with a 0° viewing angle and a 90° viewing angle are two adjacent preprocessed CT images; a preprocessed CT image with a 90° viewing angle and a 180° viewing angle are two adjacent preprocessed CT images; and a preprocessed CT image with a 180° viewing angle and a 270° viewing angle are two adjacent preprocessed CT images.
[0062] Specifically, multiple key feature points are obtained from two adjacent preprocessed CT images. The Euclidean distance between the descriptor of each key feature point in the first preprocessed CT image and the descriptors of all key feature points in the second preprocessed CT image is calculated. The key feature point in the second preprocessed CT image corresponding to the shortest Euclidean distance and the key feature point in the first preprocessed CT image are taken as the matching feature point pair in the two preprocessed CT images. Thus, all matching feature point pairs in two adjacent preprocessed CT images are obtained. Furthermore, all matching feature point pairs in two adjacent preprocessed CT images are obtained.
[0063] Specifically, the camera parameters of each pair of adjacent preprocessed CT images are determined based on the matching feature point pairs in each pair of adjacent preprocessed CT images, resulting in camera parameters for multiple preprocessed CT images.
[0064] For example, a preprocessed CT image with a 0° viewing angle is used as a reference image. The extrinsic parameter matrix of the reference image is fixed, and the extrinsic parameter matrices of other viewing angle images are determined based on the reference image.
[0065] It is understandable that the intrinsic parameter matrix K within a CT system is determined by the X-ray detector hardware conditions and can be derived from the physical structure. The essential matrix is calculated using the intrinsic parameter matrix and matching feature point pairs from each pair of adjacent preprocessed CT images. This essential matrix is then decomposed to obtain the extrinsic parameter matrix of another preprocessed CT image. The camera parameters for each preprocessed CT image include both the intrinsic and extrinsic parameter matrices, which will not be elaborated upon here.
[0066] Specifically, the initial point cloud is determined based on the matching feature point pairs in all two adjacent preprocessed CT images and the camera parameters of multiple preprocessed CT images.
[0067] For example, by utilizing the fact that two paired points between two adjacent viewpoint images are the same point in three-dimensional space, and combining the triangulation algorithm to calculate the coordinates of the point in three-dimensional space, point clouds corresponding to multiple preprocessed CT images are obtained, which serve as initial point clouds.
[0068] Specifically, such as Figure 1As shown, in step S3, the initial point cloud is input into the new viewpoint image generation model. The new viewpoint image generation model generates a complete point cloud based on the initial point cloud. Then, combined with the fast differentiable rasterization method, the complete point cloud is rendered from multiple views to generate a two-dimensional image, thus obtaining the new viewpoint CT image corresponding to any new viewpoint.
[0069] It is worth noting that the 3D Gaussian splashing model has significant advantages over traditional methods, such as Neural Radiation Field (NeRF): ① Extremely fast rendering speed: Due to the use of a highly optimized rasterization process, it can achieve real-time rendering at hundreds of frames per second, far exceeding methods such as NeRF that require volumetric rendering; ② High-quality details: It can capture and reproduce very fine visual details, generating new perspective images with high clarity and sharp edges; ③ Explicit representation: The scene is explicitly represented as a set of Gaussian ellipsoids, which makes the geometry represented by the model intuitive and easy to edit after training.
[0070] Specifically, the 3D Gaussian splash model cleverly combines rasterization technology in computer graphics with optimization methods in modern machine learning, providing a powerful solution for fast and high-quality 3D reconstruction and rendering. In step S3, this invention uses the initialized point cloud and the trained new perspective image generation model to obtain multiple new perspective CT images.
[0071] Preferably, the loss function used when training the 3D Gaussian spam model is: ; ; ; ; in, This represents the loss value for each new viewpoint image. Represents the weight parameters. Indicates the first image in each new perspective image Pixel prediction value of each pixel. In each new perspective image, the first The actual pixel value of each pixel. This indicates the number of pixels in each new viewpoint image. , This represents the predicted and actual brightness values of all pixels in each new viewpoint image. , Represents the local minimum constant. , This represents the predicted and actual values of the brightness variance of all pixels in each new viewpoint image. This represents the covariance between the predicted and actual brightness values of all pixels in each new viewpoint image.
[0072] Specifically, the loss function used in this invention for training the 3D Gaussian scattering model can better preserve the edge and texture information of the image, while having a certain tolerance for errors in the reconstruction process. It can optimize the overall structural similarity of the image while preserving image details.
[0073] Specifically, the new perspective image generation model is trained through the following steps: (1) Build a X-ray image acquisition platform and acquire full-view images by changing the position of the X-ray source and detector; (2) Select several sparse view images from the acquired full-view images as the original sparse view images, pair the full-view images with the corresponding sparse view images as a training sample pair, and construct the training set in this way. (3) Use each training sample in the training set to train the new perspective ray image generation model until the loss function of the new perspective ray image generation model satisfies the iteration termination condition, and the training of the model is completed, and the trained new perspective image generation model is obtained.
[0074] Specifically, by inputting sparse viewpoint slice images into a new viewpoint ray image generation model, the new viewpoint ray image generation model outputs a full-viewpoint slice image.
[0075] Specifically, the trained new perspective image generation model is used in step S3 of this invention to obtain multiple new perspective CT images.
[0076] Specifically, such as Figure 1 As shown, in step S4, the multiple sparse view CT images acquired in step S1 and the multiple new view CT images generated in step S3 are combined to perform image reconstruction, resulting in the reconstructed CT image.
[0077] Specifically, the image reconstruction employs any of the following algorithms: Filtered back projection algorithm; Convolution back projection algorithm; Differential-Hilbert backprojection algorithm; Gradient descent algorithm; Maximum likelihood iterative reconstruction algorithm.
[0078] The reconstructed images are obtained by reconstructing multiple sparse-view CT images acquired in step S1 and multiple new-view CT images generated in step S3 using a pre-set image reconstruction algorithm.
[0079] Compared with existing technologies, the reconstruction method based on sparse-view CT images provided in this invention uses multiple sparse-view CT images with the same viewpoint interval. It uses an improved scale-invariant feature transform algorithm to obtain an initial point cloud with lower noise, and uses a new viewpoint image generation model trained on a three-dimensional Gaussian splash model to obtain multiple new viewpoint CT images. Image reconstruction is performed based on multiple sparse-view CT images and multiple new viewpoint CT images, which improves the accuracy of the reconstructed image. At the same time, Marr wavelet functions with multiple scale spatial parameters are used to determine key feature points in each preprocessed CT image. The neighborhood of each key feature point is divided into multiple grids, and the gradient direction histogram in each grid is statistically analyzed to obtain the descriptor of each key feature point, which reduces the noise of each key feature point and provides a foundation for further improving the accuracy of the reconstructed image.
[0080] Those skilled in the art will understand that all or part of the processes of the methods described in the above embodiments can be implemented by a computer program instructing related hardware, and the program can be stored in a computer-readable storage medium. The computer-readable storage medium may be a disk, optical disk, read-only memory, or random access memory, etc.
[0081] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.
Claims
1. A reconstruction method based on sparse viewpoint CT images, characterized in that, The reconstruction method includes: Multiple sparse-view CT images are acquired, and each sparse-view CT image is preprocessed to obtain multiple preprocessed CT images; wherein, the view interval between two adjacent sparse-view CT images is the same. The initial point cloud is determined based on multiple preprocessed CT images and an improved scale-invariant feature transform algorithm. The initial point cloud is input into the new perspective image generation model to obtain multiple new perspective CT images; the new perspective image generation model is trained based on a three-dimensional Gaussian splash model. Image reconstruction is performed based on multiple sparse view CT images and multiple new view CT images to obtain the reconstructed CT image.
2. The reconstruction method according to claim 1, characterized in that, The process of determining the initial point cloud based on multiple preprocessed CT images, combined with an improved scale-invariant feature transform algorithm, includes: Key feature points in multiple preprocessed CT images were determined using Marr wavelet functions with multiple scale-space parameters. Match key feature points in all adjacent preprocessed CT images to obtain matching feature point pairs in all adjacent preprocessed CT images; Based on the matching feature point pairs in each pair of adjacent preprocessed CT images, the camera parameters of each pair of adjacent preprocessed CT images are determined, and the camera parameters of multiple preprocessed CT images are obtained. The initial point cloud is determined based on matching feature point pairs in all two adjacent preprocessed CT images and camera parameters of multiple preprocessed CT images.
3. The reconstruction method according to claim 2, characterized in that, The method of using Marr wavelet functions with multiple scale-space parameters to determine key feature points in multiple preprocessed CT images includes: Each preprocessed CT image is convolved using the Marr wavelet function with multiple scale space parameters to obtain multiple scale space images corresponding to each preprocessed CT image with multiple scale space parameters. Key feature points in each preprocessed CT image are determined using multiple scale-space images corresponding to each preprocessed CT image.
4. The reconstruction method according to claim 3, characterized in that, The expression for the Marr wavelet function is: ; in, , This represents the coordinates of a pixel in each preprocessed CT image. Represents scale-space parameters. This represents the weight values of the convolution kernel.
5. The reconstruction method according to claim 3, characterized in that, The process of determining key feature points in each preprocessed CT image using multiple scale-space images corresponding to each preprocessed CT image includes: For each preprocessed CT image, all pixels are traversed to determine whether the pixel is an extreme point in different scale spatial images; if the pixel is an extreme point in different scale spatial images, then the pixel is a key feature point in each preprocessed CT image. The neighborhood of each key feature point is divided into multiple grids, and the gradient direction histogram in each grid is statistically analyzed to obtain the descriptor of each key feature point.
6. The reconstruction method according to claim 5, characterized in that, The multi-layered grid includes 4×4 grids and 2×2 grids.
7. The reconstruction method according to claim 5, characterized in that, The gradient direction histogram in each grid layer is used to obtain the descriptor for each key feature point, including: The gradient direction histogram in each grid layer is used as the descriptor for each key feature point in each grid layer; PCA dimensionality reduction update is performed on the descriptor of each key feature point in each grid layer to preserve the principal energy components; The updated descriptors of each key feature point in each grid layer are concatenated to obtain the descriptor of each key feature point.
8. The reconstruction method according to any one of claims 3-7, characterized in that, The scale space parameter is 1 or 2.
9. The reconstruction method according to claim 1, characterized in that, The loss function used when training the 3D Gaussian spam model is: ; ; ; ; in, This represents the loss value for each new viewpoint image. Represents the weight parameters. Indicates the first image in each new perspective image Pixel prediction value of each pixel. In each new perspective image, the first The actual pixel value of each pixel. This indicates the number of pixels in each new viewpoint image. , This represents the predicted and actual brightness values of all pixels in each new viewpoint image. , Represents the local minimum constant. , This represents the predicted and actual values of the brightness variance of all pixels in each new viewpoint image. This represents the covariance between the predicted and actual brightness values of all pixels in each new viewpoint image.
10. The reconstruction method according to claim 1, characterized in that, The image reconstruction employs any of the following algorithms: Filtered back projection algorithm; Convolution back projection algorithm; Differential-Hilbert backprojection algorithm; Gradient descent algorithm; Maximum likelihood iterative reconstruction algorithm.