Three-dimensional CT image reconstruction method based on multi-view projection, program product and equipment
Through the three-dimensional CT image reconstruction method of multi-view projection, the dimension expansion and feature fusion network are used to solve the problem of low reconstruction quality under sparse view, and achieve fast and high-quality three-dimensional CT image reconstruction and computing resource optimization.
Patent Information
- Application Number
- CN202510954319.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-07-10
AI Technical Summary
Existing CT imaging technology has low reconstruction quality under sparse viewing conditions. Deep learning methods have bottlenecks in data set acquisition and generalization capabilities. In addition, the number of three-dimensional convolution parameters is large and the computing resource consumption is high, making it difficult to quickly reconstruct three-dimensional images with high quality.
A three-dimensional CT image reconstruction method based on multi-view projection is adopted. Through the dimension expansion network, feature extraction network, implicit neural learning network and feature reconstruction network, depth information is gradually embedded and features are fused to reconstruct the three-dimensional image.
It achieves fast and high-quality reconstruction of three-dimensional CT images from fewer viewing angles, improves reconstruction speed and quality, reduces computing resource consumption, and enhances the generalization ability of the model.
Smart Images

Figure CN120689524A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer tomography, and in particular to a three-dimensional CT image reconstruction method and computing equipment based on multi-view projection. Background Art
[0002] Computed tomography (CT) technology, due to its non-destructive, efficient, and non-contact characteristics, has broad industrial applications. CT technology can accurately depict the internal information of the object being measured. Currently, CT imaging technologies are primarily divided into two categories: traditional reconstruction and deep learning reconstruction. Traditional reconstruction methods can be broadly categorized into two types: analytical and iterative. Using hundreds of X-ray projections, traditional reconstruction algorithms can approximately reconstruct the CT volume. Due to the limitations of the Nyquist sampling theorem, high-quality reconstruction with traditional reconstruction algorithms relies on obtaining complete projections, which significantly limits reconstruction speed and application scenarios. Subsequently, researchers introduced compressed sensing to CT reconstruction and developed a regularization-based reconstruction method. This method incorporates prior information into the reconstruction model, successfully reducing the number of required projection images. However, in some constrained application scenarios, where only a very limited number of projection images are available, such as in-line inspection of industrial products and dynamic testing of transient damage, neither traditional nor regularized methods can achieve high-quality results. Therefore, reconstructing images from ultra-sparsely sampled projections to accelerate the CT imaging process has become a hot research topic.
[0003] Recently, with the rapid development of deep learning, it has been applied to sparse and ultra-sparse CT reconstruction. By using artificially designed neural networks to directly learn the mapping function from X-ray projections to 3D CT volumes, impressive results have been achieved. Compared to traditional regularization-based methods, deep learning methods extract prior information for CT prediction through a data-driven, end-to-end network, making CT reconstruction more feasible.
[0004] However, reconstruction methods based on deep learning also face common limitations in applications. First, it is difficult to obtain large-scale training data sets, which may constitute a significant bottleneck in certain specific application scenarios. In addition, CT reconstruction technology based on deep learning performs poorly in generalization ability and is difficult to adapt to a variety of imaging objects, limiting its wide applicability. Furthermore, current three-dimensional CT reconstruction algorithms of this type generally use three-dimensional convolution as the core computing unit, and realize feature mapping from projection data to voxel space through the spatial dimension perception characteristics of three-dimensional convolution. However, cascaded three-dimensional convolution will cause a sharp increase in the number of parameters and significant video memory consumption, which seriously restricts the reconstructed voxel dimension and reconstruction rate, and also increases the consumption of computing resources and puts higher requirements on memory. Summary of the Invention
[0005] To address existing technical problems, the present invention provides a three-dimensional CT image reconstruction method and computing device based on multi-view projection, which can quickly and high-quality reconstruct 3D CT images using projection images from fewer viewing angles, thereby improving the speed and quality of three-dimensional reconstruction.
[0006] In a first aspect, a three-dimensional CT image reconstruction method based on multi-view projection is provided, comprising: acquiring multiple projection images of a three-dimensional target to be reconstructed collected at different viewpoints and initial volume data corresponding to the three-dimensional target to be reconstructed; utilizing a dimension expansion network in a pre-trained three-dimensional reconstruction model to expand the channel dimension of the projection image to obtain an extended projection image corresponding to each of the projection images; utilizing each of the extended projection images as input to a feature extraction network in the three-dimensional reconstruction model to extract multi-scale feature data corresponding to each of the projection images; utilizing an implicit neural learning network in the three-dimensional reconstruction model based on the initial volume data to obtain multi-scale implicit neural representation data corresponding to the three-dimensional target to be reconstructed; utilizing a feature reconstruction network in the three-dimensional reconstruction model to reconstruct a three-dimensional image based on the multi-scale feature data corresponding to each of the projection images and the multi-scale implicit neural representation data, and outputting a three-dimensional reconstructed target.
[0007] In a second aspect, a computing device is provided, comprising a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the three-dimensional CT image reconstruction method based on multi-view projection as described in any embodiment of the present application.
[0008] In a third aspect, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the three-dimensional CT image reconstruction method based on multi-view projection described in any embodiment of the present application.
[0009] This application obtains projection images from multiple perspectives, expands the projection images through a dimensionality expansion network, gradually embeds the depth information of the CT volume of the to-be-reconstructed 3D target into different channels of the feature map, and obtains an extended projection image corresponding to each projection image; extracts feature maps at different levels in the extended projection image through a feature extraction network, and obtains scale feature data at multiple scales; and based on the initial volume data, uses an implicit neural learning network to represent the CT volume representation of the to-be-reconstructed 3D target, thereby adding the spatial geometric information of the to-be-reconstructed 3D target to the 3D reconstruction model, so as to facilitate the subsequent accurate reconstruction of the 3D image of the to-be-reconstructed 3D target; then uses the feature data of each scale and the implicit neural representation data of each scale as the input of each layer in the feature reconstruction network, continuously reconstructs features through each layer in the feature reconstruction network, fuses features, gradually restores volume details, and finally outputs the 3D reconstructed target. This application can quickly and high-quality reconstruct 3D CT images through projection images from fewer perspectives, thereby improving the 3D reconstruction speed and quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 A diagram illustrating an application environment of a three-dimensional CT image reconstruction method based on multi-view projection in one embodiment;
[0011] Figure 2 is a flow chart of a three-dimensional CT image reconstruction method based on multi-view projection in one embodiment;
[0012] Figure 3 1 is an overall network diagram of a 3D reconstruction model based on multi-view projection in one embodiment;
[0013] Figure 4 This is an example diagram of the network structure of a three-dimensional reconstruction model in one embodiment;
[0014] Figure 5 This is an example diagram of a network during training of a three-dimensional reconstruction model in one embodiment;
[0015] Figure 6 is a schematic diagram of a 3D CT image reconstruction device based on multi-view projection in one embodiment;
[0016] Figure 7 is a schematic diagram of a computing device in one embodiment. DETAILED DESCRIPTION
[0017] The technical solution of the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0018] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this invention pertains. The terms used herein in the specification of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the scope of protection of the present invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0019] In the following description, reference is made to “some embodiments” which describe a subset of all possible embodiments, but it should be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0020] See Figure 1 , is a diagram illustrating the application environment of a three-dimensional CT image reconstruction method based on multi-view projection in one embodiment. The three-dimensional CT image reconstruction method based on multi-view projection is applied to a computing device 10, which can acquire projection images from multiple viewpoints. The projection images can be acquired by a detection device to obtain the three-dimensional target to be reconstructed. The radiation source in the detection device rotates around the target to be measured, acquiring projection images of the target to be measured from different angles. Circumferential X-ray projection images of the target to be reconstructed are acquired, during which relevant parameters such as the tube voltage, tube current, and magnification ratio are kept constant. Projection data from at least two viewpoints are selected from the circumferential projection to form multiple projection images. It can be understood that a projection image is an image representation of the projection data acquired from one viewpoint for the target to be reconstructed. The computing device 10 can reconstruct the target to be reconstructed based on the projection images from multiple viewpoints using the three-dimensional CT image reconstruction method based on multi-view projection.
[0021] See also Figure 2 , is a flow chart of a 3D CT image reconstruction method based on multi-view projection provided in one embodiment of the present application. The 3D CT image reconstruction method based on multi-view projection is applied to a computing device and includes the following steps:
[0022] S20 , acquiring a plurality of projection images of the three-dimensional object to be reconstructed collected at different viewing angles and initial volume data corresponding to the three-dimensional object to be reconstructed.
[0023] In this embodiment, the multiple projection images may be at least two. For example, when there are two projection images, projection images from two mutually orthogonal perspectives may be selected, namely, a first projection image and a second projection image. The initial volume data represents the initial volume data of the three-dimensional object to be reconstructed. For example, if a volume with a resolution of 128 is to be reconstructed, 128 points are uniformly sampled between -1 and 1 in the X-axis, Y-axis, and Z-axis directions.
[0024] S21. Utilize a dimension expansion network in a pre-trained three-dimensional reconstruction model to expand the channel dimension of the projection image to obtain an expanded projection image corresponding to each of the projection images.
[0025] In this embodiment, the 3D reconstruction model is obtained through the training data set. Figure 3 As shown, the 3D reconstruction model includes a dimension expansion network, a feature extraction network, an implicit neural learning network, and a feature reconstruction network. The dimension expansion network is used to expand the number of channels in the channel dimension of the image. As the number of channels gradually increases in a step-by-step manner, the depth information of the CT volume of the 3D target to be reconstructed is gradually embedded into different channels of the feature map. The feature extraction network is used to extract feature maps of different scales of the 3D target to be reconstructed in the projected image. The implicit neural learning network is used to learn the CT volume representation of the 3D target to be reconstructed through implicit neural representation, thereby adding the spatial geometric information of the 3D target to be reconstructed to the 3D reconstruction model, fully learning the prior information, and thus improving the generalization of the model. The feature reconstruction network is used to reconstruct the results of the implicit neural representation learning into results that match the dimensions of each stage of the decoder, and then fuse them with the results of the decoder and the results of the feature reconstruction module itself. By continuously reconstructing and fusing features, the volume details are gradually restored, and finally the 3D reconstructed target is output.
[0026] The dimension expansion network, encoder, decoder, implicit neural learning network, and feature reconstruction network are F de 、F encoder 、F decoder 、F inr 、F fr The mapping function F from multiple projection images fitted by the 3D reconstruction model to the 3D CT volume can be described as follows:
[0027]
[0028] in Represents the network model cascade, Indicates that the network model is in parallel.
[0029] In this embodiment, because the multiple projection images are captured at different viewing angles, the captured ray directions are different. Therefore, a single viewing angle can be selected from the multiple viewing angles as a base viewing angle. Using this base viewing angle as a reference, the projection images at other viewing angles are then expanded in channel dimensions and aligned. This ensures that the expanded projection images corresponding to the respective projection images are at the same coordinates, achieving dimensional uniformity.
[0030] S22: Using each extended projection image as an input to a feature extraction network in the three-dimensional reconstruction model to extract multi-scale feature data corresponding to each projection image.
[0031] In this embodiment, the projected images at each viewing angle correspond to their own dimension expansion network and feature extraction network, which output multi-scale feature data corresponding to each projected image. The multi-scale feature data corresponding to each projected image is then used as input to different layers of the feature reconstruction network. The multi-scale feature data includes feature data at multiple scales.
[0032] S23. Based on the initial volume data, the implicit neural learning network in the three-dimensional reconstruction model is used to obtain multi-scale implicit neural representation data corresponding to the three-dimensional target to be reconstructed.
[0033] In this embodiment, the multi-scale implicit neural representation data includes implicit neural representation data at multiple scales, wherein each scale of implicit neural representation data represents implicit neural representation data of volume data of a three-dimensional object to be reconstructed at that scale.
[0034] S24. Based on the multi-scale feature data and multi-scale implicit neural representation data corresponding to each projection image, the feature reconstruction network in the three-dimensional reconstruction model is used to reconstruct the three-dimensional image and output the three-dimensional reconstruction target.
[0035] In this embodiment, the feature data of each scale and the implicit neural representation data of each scale are respectively used as the input of each layer in the feature reconstruction network. The features are continuously reconstructed and fused through each layer in the feature reconstruction network, and the volume details are gradually restored, and finally the three-dimensional reconstruction target is output.
[0036] In the above embodiment, by acquiring projection images from multiple perspectives, the projection images are expanded through a dimension expansion network, and the depth information of the CT volume of the three-dimensional target to be reconstructed is gradually embedded into different channels of the feature map to obtain an extended projection image corresponding to each projection image; the feature maps at different levels in the extended projection image are extracted through a feature extraction network to obtain scale feature data at multiple scales; and based on the initial volume data, the CT volume representation of the three-dimensional target to be reconstructed is represented by an implicit neural learning network, thereby adding the spatial geometric information of the three-dimensional target to be reconstructed to the three-dimensional reconstruction model, so as to facilitate the subsequent accurate reconstruction of the three-dimensional image of the three-dimensional target to be reconstructed; then, the feature data of each scale and the implicit neural representation data of each scale are respectively used as inputs of each layer in the feature reconstruction network, and the features are continuously reconstructed and fused through each layer in the feature reconstruction network, and the volume details are gradually restored, and finally the three-dimensional reconstructed target is output. The present application can quickly and high-quality reconstruct 3D CT images through projection images from fewer perspectives, thereby improving the three-dimensional reconstruction speed and quality.
[0037] In some embodiments, the dimension expansion network includes a gradient calculation layer, a gradient splicing layer, and a dimension expansion layer. The dimension expansion network in the pre-trained 3D reconstruction model is used to expand the channel dimension of the projection image to obtain an expanded projection image corresponding to each projection image, including:
[0038] Based on the projected image, using a gradient calculation layer, calculate the gradients in the horizontal and vertical directions to obtain a gradient image corresponding to the projected image;
[0039] Using a gradient stitching layer, summing the projected image and the gradient image corresponding to the projected image to obtain a composite image corresponding to the projected image, and performing channel stitching on the projected image, the gradient image corresponding to the projected image, and the composite image corresponding to the projected image to obtain a three-channel initial image corresponding to the projected image;
[0040] Based on multiple cascaded residual blocks in the dimension expansion layer, feature maps in the three-channel initial image are gradually extracted to obtain an extended projected image corresponding to the projected image.
[0041] In this embodiment, for example, the projection image can be an image formed by unprocessed projection data or an image formed by logarithmically processed projection data. The projection image is represented by {I1, I2}∈R H×W×1 , using the Scharr gradient, we get {I3,I4}∈R H×W×1 The Scharr gradient calculation formula is as follows:
[0042]
[0043] Where A represents the image to be processed, represents convolution, G x (i, j) represents the gradient in the horizontal direction, G y (i, j) represents the vertical gradient, G(i, j) represents the gradient image, θ represents the gradient direction, i represents the horizontal coordinate, and j represents the vertical coordinate. The Scharr gradient can more accurately approximate the true gradient of the image, and considers more neighborhood information when calculating the gradient, and has a certain robustness to noise. Through the above method, the gradient image I3 corresponding to I1 and the gradient image I4 corresponding to I2 can be calculated. Then the projected image and the gradient image are summed to obtain the composite image {I1+I3,I2+I4}∈R H×W×1 Finally, {I1,I2}∈R H×W×1 ,{I3,I4}∈R H×W×1 and {I1+I3,I2+I4}∈R H×W×1 Splicing is performed on the channel dimension to generate the initial image {I5,I6}∈RH×W×3 . Then, the three-channel initial image is expanded in number of channels, the channel dimension is encoded into the depth dimension, and the feature channel is used to infer the CT volume depth. The channel number expansion consists of multiple cascaded ResBlock blocks. As the ResBlock blocks are gradually executed, the number of channels of the feature map increases in a step-by-step manner. In this process, the network continues to explore the depth information and gradually embeds the 3D information into different channels of the feature map to obtain the extended projection image corresponding to the projection image. For example, the channel number expansion consists of 4 cascaded ResBlock blocks. As the ResBlock blocks are gradually executed, the number of channels of the feature map increases in a step-by-step manner, namely 3→16→32→64→128. After the three-channel initial image is expanded, the non-reference projection image (projection image captured under a non-reference perspective) is aligned using the alignment network to obtain the extended projection image corresponding to the projection image.
[0044] In the above embodiment, by acquiring a gradient image of the projection image, obtaining a three-channel initial image based on the gradient image and the projection image, and performing dimension expansion on the three-channel initial image, the depth information of the CT volume is gradually embedded into different channels of the feature map during this process, so as to facilitate more accurate reconstruction of the three-dimensional object and improve the reconstruction accuracy.
[0045] In some embodiments, the feature extraction network includes an encoder and a decoder, the decoder is connected to the feature reconstruction network, and the implicit neural learning network is connected to the feature reconstruction network;
[0046] The encoder includes multiple encoding layers, the decoder includes multiple decoding layers, the encoding layers located in the same scale layer are jump-connected to the corresponding decoding layers, the feature reconstruction network includes multiple reconstruction layers with the same number of layers as the decoder and a reconstruction output layer connected to the reconstruction layer, and the input of the decoding layer located in the same scale layer serves as the input of the corresponding reconstruction layer; the feature reconstruction network includes a multi-scale fusion network, the implicit neural learning network is connected to the multi-scale fusion network, the multi-scale fusion network is respectively connected to each layer of the feature reconstruction network, and the multi-scale fusion network is used to output implicit neural representation data of each scale.
[0047] like Figure 4 As shown, Figure 4This is an example diagram of the network structure of a three-dimensional reconstruction model in an embodiment. The multiple projection images include a first projection image and a second projection image. The first projection image corresponds to a dimension expansion network and a feature extraction network. The second projection image corresponds to a dimension expansion network and a feature extraction network. The first projection image is a projection image captured under a reference perspective. The second projection image also corresponds to an alignment network for converting the expanded image and aligning it with the reference perspective to achieve dimensional unification. The encoder includes multiple downsampling layers, such as Figure 4 As shown, it includes four downsampling layers, such as A1, A2, A3 and A4. The decoder includes four decoding layers, each of which includes an upsampling layer and a connection layer. Figure 4 In the
[15] , there are four decoding layers, B4, B3, B2, and B1. Two adjacent upsampling layers are connected by a connection layer. If the feature maps output by A1 and B1 are of the same scale, then A1 and B1 are at the same scale layer. Similarly, A2 and B2 are at the same scale layer, A3 and B3 are at the same scale layer, and A4 and B4 are at the same scale layer. The encoding layer is jump-connected to the corresponding decoding layer. Specifically, the input of the encoding layer at the same scale layer is jump-connected to the output of the downsampling layer in the corresponding decoding layer. For example, the input of A1 is jump-connected to the output of the downsampling layer in B1. The network structures of the dimension expansion network and feature extraction network corresponding to each projection are the same, except for the different inputs.
[0048] The encoder takes the extended projection image with expanded channel number as input and gradually compresses and encodes the high-dimensional feature map. As downsampling proceeds, the size of the feature map is continuously compressed, but the number of feature channels remains unchanged, always matching the depth of the target CT volume data. In this process, the receptive field will gradually expand, and more low-frequency information can be perceived. During the encoding process, feature information of different levels is directly passed to the decoder through jump connections to share feature information and enhance information utilization. Figure 4 As shown in the figure, the encoder consists of four cascaded downsampling layers. Each downsampling layer consists of "2D residual block ResBlock layer → 2D residual block ResBlock layer → pooling layer". Through the gradual downsampling operation, the size of the feature map is gradually reduced.
[0049] The decoder takes the encoded low-dimensional high-level features as input, gradually reconstructs the high-dimensional features and restores the image details. Figure 4 As shown in Figure 1, the decoder includes four concatenated and fused layers and four upsampling layers. The concatenated layers are used to jump-connect the inputs of the encoding layers at the same scale level to the outputs of the corresponding decoding layers. For example, the inputs of A1 and the outputs of the downsampling layers in B1 are concatenated and fused via the concatenated layers in B1.
[0050] Each upsampling layer consists of a sequence of steps: 2D deconvolution layer → 2D ResBlock layer → 2DResBlock layer. Through this gradual upsampling operation, the feature map size gradually increases, while the number of channels remains constant. As upsampling proceeds, the model gradually fuses the decoder's own information with feature information from the corresponding encoder layer, leveraging multi-scale information to gradually restore the target image's structure and details.
[0051] In this embodiment, the implicit neural learning network is connected to the multi-scale fusion network in the feature reconstruction network. The initial volume data is used as the input of the implicit neural learning network. The spatial coordinates of the volume to be reconstructed are first determined, and then the spatial coordinates are positionally encoded using Fourier feature mapping. The result after position encoding is input into a perceptron neural network (MLP) for implicit neural representation learning, learning the continuous implicit neural representation of the CT volume, and embedding the prior information of the image into the network parameters. Finally, the implicit neural representation learning and the decoder are combined through the feature reconstruction module. The perceptron neural network includes multiple linear interpolation modules connected in sequence. The CT volume representation of the object to be reconstructed is learned through implicit neural representation, so that the spatial geometric information of the object to be reconstructed is added to the three-dimensional reconstruction model, the prior information is fully learned, and the generalization of the model is improved.
[0052] The feature reconstruction network consists of a multi-scale feature fusion network and multiple reconstruction layers. The multi-scale feature fusion block uses bilinear interpolation to reconstruct the results of implicit neural representation learning into feature maps that match the dimensions of each decoder stage. A reconstruction layer is an upsampling block, which consists of a concatenation layer, a 2D convolution layer, a 2D deconvolution layer, a 2D convolution layer, a 2D batch normalization layer, and a ReLU layer. The reconstruction output layer includes a concatenation operation and two 2D convolutions.
[0053] like Figure 4 As shown, the feature reconstruction network includes four reconstruction layers. The feature maps output by B4 and C4 have the same scale and are located on the same scale layer. Similarly, B3 and C3 are located on the same scale layer, B2 and C2 are located on the same scale layer, and B1 and C1 are located on the same scale layer. For example, the feature maps output by the multi-scale fusion network are D4, D3, D2, D1, and D0, respectively. The scale of D4 is the same as the scale of the feature map input by C4. Similarly, the scale of D3 is the same as the scale of the feature map input by C3. Similarly, the scale of D0 is the same as the scale of the feature map output by C1.
[0054] The input of B4 corresponding to the first projection image, the input of B4 corresponding to the second projection image, and the scale implicit neural representation data of the same scale as the input of B4 are used as the input of the C4 reconstruction layer, and the process is performed layer by layer until the reconstruction layer C1 outputs the maximum scale feature data. Then, the output of B1 corresponding to the first projection image, the output of B1 corresponding to the second projection image, the final reconstructed feature data output by the final reconstruction layer C1, and the scale implicit neural representation data D0 corresponding to the reconstruction output layer are used as the input of the reconstruction output layer, and the three-dimensional reconstructed target is finally output.
[0055] Optionally, the performing of three-dimensional image reconstruction using a feature reconstruction network in the three-dimensional reconstruction model based on the multi-scale feature data corresponding to each of the projection images and the multi-scale implicit neural representation data, and outputting a three-dimensional reconstruction target includes:
[0056] For a current reconstruction layer at a current layer number, obtaining current scale feature input data corresponding to each projection image, wherein the current scale feature input data is the input of a current decoding layer connected to the current reconstruction layer in a decoder corresponding to each projection image;
[0057] Obtaining an output of a previous reconstruction layer connected to the current reconstruction layer;
[0058] Based on the initial volume data, obtaining current-scale implicit neural representation data corresponding to the current number of layers output by the multi-scale fusion network;
[0059] Based on the current-scale feature input data corresponding to each projection image, the output of the previous reconstruction layer, and the current-scale implicit neural representation data, obtaining the current reconstructed feature data output by the current reconstruction layer, and performing the steps in sequence until the final reconstructed feature data of the final reconstruction layer is obtained;
[0060] Obtaining the maximum scale feature data output by the last decoding layer corresponding to each projection image, thereby obtaining the maximum scale feature data corresponding to each projection image;
[0061] Based on the maximum scale feature data corresponding to each projection image, the scale implicit neural representation data corresponding to the reconstruction output layer and the final reconstruction feature data, the predicted three-dimensional reconstruction target is output through the reconstruction output layer.
[0062] The results of the two decoders are concatenated along the channel dimension with the results of the implicit neural network from the multi-scale feature fusion block and the results of the feature reconstruction module itself. This is then subjected to a 2D convolution operation to effectively fuse the feature information. Through gradual simple upsampling, the feature map size is gradually increased, while the number of channels remains constant. Finally, a concatenation operation and two 2D convolutions are performed to output the 3D reconstructed target.
[0063] In this embodiment, the current reconstructed feature data output by the current reconstructed layer is used as the input of the next reconstructed layer of the current reconstructed layer, so that the next layer is updated to the current reconstructed layer, and the process is executed in sequence until the final reconstructed feature data of the final reconstructed layer is obtained.
[0064] In the above embodiment, by splicing the results of the two decoders with the results of the implicit neural learning network through the multi-scale feature fusion block and the results of the feature reconstruction module itself in the channel dimension, the fusion of features of different scales in the projected image can be achieved, which facilitates the subsequent more accurate reconstruction of the three-dimensional object and improves the reconstruction accuracy.
[0065] In some embodiments, the method further comprises:
[0066] Acquire a training data set, wherein each training sample in the training data set includes a plurality of sample projection images collected from multiple perspectives of a three-dimensional sample target, initial sample volume data corresponding to the three-dimensional sample target, and a three-dimensional reconstructed volume label corresponding to the three-dimensional sample target;
[0067] Based on the training data set, obtaining input training samples of the current iteration, iteratively training the initial 3D reconstruction model to obtain 3D reconstruction samples of the current iteration, and calculating the current total loss value based on the 3D reconstruction volume labels corresponding to the input training samples and the 3D reconstruction samples of the current iteration;
[0068] Based on the current total loss value, determine whether the current iteration meets the iteration termination condition. If the current iteration does not meet the iteration termination condition, continue to obtain input training samples for iterative training until the iteration termination condition is met. The three-dimensional reconstruction model after meeting the iteration termination condition is used as the pre-trained three-dimensional reconstruction model.
[0069] In this embodiment, the 3D reconstruction volume label is obtained by reconstructing the circumferential projection image using a traditional reconstruction algorithm, such as the FDK algorithm. It can be understood that a 3D sample target corresponds to a set of sample projection images and a 3D reconstruction volume label. Figure 5 As shown, Figure 5 This is an example diagram of the network during the training of the 3D reconstruction model in one embodiment. During the training process, the network of the 3D reconstruction model is the same as the trained network. During the training process, the network parameters in the 3D reconstruction model have not been trained yet and need to be updated iteratively. For example, if the 3D sample target includes two sample projection images, the two sample projection images and the initial sample volume data are respectively input into the 3D reconstruction model under training to obtain the 3D reconstruction sample corresponding to the current iteration. The network processing during training is the same as the process of reconstructing the 3D reconstruction target in the actual scene mentioned above, that is, Figure 2The process of reconstructing the 3D reconstructed target is the same as in [1] and will not be repeated here. The initial sample volume data is the initial representation of the volume of the 3D sample target. The iterative termination conditions include, but are not limited to, the number of iterations exceeding a preset number, the current total loss value being less than a preset error value, and so on.
[0070] In this embodiment, in order to effectively improve the reconstruction quality and accuracy and enhance the robustness of the model, the present application constructs a multi-dimensional constrained loss function. The loss function includes reconstruction loss, structural similarity loss, gradient loss and projection loss with clear physical meaning. The reconstruction loss measures the geometric difference between the predicted 3D structure and the real structure. This loss ensures that the model reconstructs a 3D model as accurately as possible in space. The structural similarity loss focuses on the visual quality and structural consistency of the image. It usually evaluates the similarity between the predicted image and the real image, and can better capture the sensitivity of the human eye to image quality. The gradient loss focuses on the edge and texture information of the image, which can help the model better restore the structural information and edge details of the image and avoid blurring. The projection loss with clear physical meaning takes into account the specific physical process of CT imaging, which can provide a clear guide for network learning, allowing the network to learn the physical process of CT imaging, thereby enhancing the generalization and interpretability of the model.
[0071] Optionally, calculating the current total loss value based on the 3D reconstructed volume label corresponding to the input training sample and the 3D reconstructed sample of the current iteration includes:
[0072] Calculating a current reconstruction loss value using a first loss function based on the 3D reconstructed volume label corresponding to the input training sample and the 3D reconstructed sample of the current iteration, wherein the current reconstruction loss value is used to indicate a structural difference between the predicted 3D reconstructed sample and the 3D reconstructed volume label;
[0073] Calculating a current structural similarity loss value using a second loss function based on the 3D reconstructed volume label corresponding to the input training sample and the 3D reconstructed sample of the current iteration, wherein the current structural similarity loss value is used to indicate a similarity difference between the predicted 3D reconstructed sample and the 3D reconstructed volume label;
[0074] Calculating a current gradient loss using a third loss function based on the 3D reconstructed volume label corresponding to the input training sample and the 3D reconstructed sample of the current iteration, wherein the current gradient loss is used to indicate a difference in image detail information between the predicted 3D reconstructed sample and the 3D reconstructed volume label;
[0075] Calculating a current projection loss using a fourth loss function based on the 3D reconstructed volume label corresponding to the input training sample and the 3D reconstructed sample of the current iteration, wherein the current projection loss is used to indicate a physical process loss of CT imaging;
[0076] The current reconstruction loss value, the current structural similarity loss value, the current gradient loss and the current projection loss are weighted to obtain the current total loss value.
[0077] Optionally, the calculating the current gradient loss using a third loss function based on the 3D reconstructed volume label corresponding to the input training sample and the 3D reconstructed sample of the current iteration includes:
[0078] Using a mean square error function, calculating a horizontal gradient loss in a horizontal direction between the 3D reconstructed volume label corresponding to the input training sample and the 3D reconstructed sample of the current iteration, and calculating a vertical gradient loss in a vertical direction between the 3D reconstructed volume label corresponding to the input training sample and the 3D reconstructed sample of the current iteration;
[0079] The current gradient loss is calculated based on the horizontal gradient loss and the vertical gradient loss.
[0080] Optionally, the calculating the current projection loss using a fourth loss function based on the 3D reconstructed volume label corresponding to the input training sample and the 3D reconstructed sample of the current iteration includes:
[0081] Calculating current projection data of the current iteration based on the three-dimensional reconstruction sample of the current iteration;
[0082] Calculating a projection label based on the three-dimensional reconstructed volume label;
[0083] The current projection loss is calculated based on the current projection data and the projection label.
[0084] In this embodiment, the calculation formula of the reconstruction loss is as follows:
[0085] Reconstruction loss: where Y pred is the predicted 3D reconstruction sample, Y truth Represents the 3D reconstruction volume label, where L RE (Y pred ,Y truth ) represents the reconstruction loss.
[0086] Structural similarity loss L RS (Y pred ,Y truth )=1―SSIM(Y pred ,Y truth ), where SSIM(Y pred ,Y truth ) represents Y pred With Y truth The similarity between them.
[0087] Gradient loss where MSE x Represents the horizontal gradient loss, MSE y represents the vertical gradient loss.
[0088]
[0089] Projection loss
[0090] P(x)=∫σ(x)dx
[0091] Total loss L total =λ1L RE +λ2L SSIM +λ3L GL +λ4L PL Where λ1, λ2, λ3, and λ4 are weight coefficients that control the relative importance of different loss terms, and σ(x) is the density value. pred ) represents the projection data of the 3D reconstructed sample, P(Y truth ) represents the projection label corresponding to the 3D reconstructed volume label, where D, H, and W represent length, height, and width.
[0092] In the above embodiment, multi-view projection enables rapid, high-quality reconstruction of 3D CT images. Simultaneously, coupling implicit neural representations and introducing projection losses consistent with the physical principles of CT scanning improve the generalization of the network. Finally, introducing gradient constraints improves the network's utilization of prior information, accelerates model convergence, and enhances the network's ability to recover from edges. This application has broad development prospects and application value in application scenarios where full-view projection is unavailable or near-real-time reconstruction is required.
[0093] On the other hand, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the three-dimensional CT image reconstruction method based on multi-view projection described in any embodiment of the present application.
[0094] Among them, in the computer program product, an optional implementation form of the program module architecture of the computer program that implements each step of the target recognition method can be a three-dimensional CT image reconstruction device based on multi-view projection.
[0095] See also Figure 6, an embodiment of the present application provides a three-dimensional CT image reconstruction device based on multi-view projection, including: an acquisition module 61, used to acquire multiple projection images of a to-be-reconstructed three-dimensional target collected at different viewpoints and initial volume data corresponding to the to-be-reconstructed three-dimensional target; an expansion module 62, used to use a dimension expansion network in a pre-trained three-dimensional reconstruction model to expand the channel dimension of the projection image to obtain an extended projection image corresponding to each of the projection images; an extraction module 63, used to use each of the extended projection images as input of a feature extraction network in the three-dimensional reconstruction model to extract multi-scale feature data corresponding to each of the projection images; an implicit neural learning module 64, used to use an implicit neural learning network in the three-dimensional reconstruction model based on the initial volume data to obtain multi-scale implicit neural representation data corresponding to the to-be-reconstructed three-dimensional target; a reconstruction module 65, used to use the feature reconstruction network in the three-dimensional reconstruction model to reconstruct a three-dimensional image based on the multi-scale feature data corresponding to each of the projection images and the multi-scale implicit neural representation data, and output a three-dimensional reconstructed target.
[0096] Optionally, the expansion module 62 is used to:
[0097] Based on the projected image, using a gradient calculation layer, calculate the gradients in the horizontal and vertical directions to obtain a gradient image corresponding to the projected image;
[0098] Using a gradient stitching layer, summing the projected image and the gradient image corresponding to the projected image to obtain a composite image corresponding to the projected image, and performing channel stitching on the projected image, the gradient image corresponding to the projected image, and the composite image corresponding to the projected image to obtain a three-channel initial image corresponding to the projected image;
[0099] Based on multiple cascaded residual blocks in the dimension expansion layer, feature maps in the three-channel initial image are gradually extracted to obtain an extended projected image corresponding to the projected image.
[0100] Optionally, the feature extraction network includes an encoder and a decoder, the decoder is connected to the feature reconstruction network, and the implicit neural learning network is connected to the feature reconstruction network;
[0101] The encoder includes multiple encoding layers, the decoder includes multiple decoding layers, the encoding layers located in the same scale layer are jump-connected to the corresponding decoding layers, the feature reconstruction network includes multiple reconstruction layers with the same number of layers as the decoder and a reconstruction output layer connected to the reconstruction layer, and the input of the decoding layer located in the same scale layer serves as the input of the corresponding reconstruction layer; the feature reconstruction network includes a multi-scale fusion network, the implicit neural learning network is connected to the multi-scale fusion network, the multi-scale fusion network is respectively connected to each layer of the feature reconstruction network, and the multi-scale fusion network is used to output implicit neural representation data of each scale.
[0102] Optionally, the reconstruction module 65 is further configured to:
[0103] For a current reconstruction layer at a current layer number, obtaining current scale feature input data corresponding to each projection image, wherein the current scale feature input data is the input of a current decoding layer connected to the current reconstruction layer in a decoder corresponding to each projection image;
[0104] Obtaining an output of a previous reconstruction layer connected to the current reconstruction layer;
[0105] Based on the initial volume data, obtaining current-scale implicit neural representation data corresponding to the current number of layers output by the multi-scale fusion network;
[0106] Based on the current-scale feature input data corresponding to each projection image, the output of the previous reconstruction layer, and the current-scale implicit neural representation data, obtaining the current reconstructed feature data output by the current reconstruction layer, and performing the steps in sequence until the final reconstructed feature data of the final reconstruction layer is obtained;
[0107] Obtaining the maximum scale feature data output by the last decoding layer corresponding to each projection image, thereby obtaining the maximum scale feature data corresponding to each projection image;
[0108] Based on the maximum scale feature data corresponding to each projection image, the scale implicit neural representation data corresponding to the reconstruction output layer and the final reconstruction feature data, the predicted three-dimensional reconstruction target is output through the reconstruction output layer.
[0109] Optionally, a training module 66 is further included, for:
[0110] Acquire a training data set, wherein each training sample in the training data set includes a plurality of sample projection images collected from multiple perspectives of a three-dimensional sample target, initial sample volume data corresponding to the three-dimensional sample target, and a three-dimensional reconstructed volume label corresponding to the three-dimensional sample target;
[0111] Based on the training data set, obtaining input training samples of the current iteration, iteratively training the initial 3D reconstruction model to obtain 3D reconstruction samples of the current iteration, and calculating the current total loss value based on the 3D reconstruction volume labels corresponding to the input training samples and the 3D reconstruction samples of the current iteration;
[0112] Based on the current total loss value, determine whether the current iteration meets the iteration termination condition. If the current iteration does not meet the iteration termination condition, continue to obtain input training samples for iterative training until the iteration termination condition is met. The three-dimensional reconstruction model after meeting the iteration termination condition is used as the pre-trained three-dimensional reconstruction model.
[0113] Optionally, the training module 66 is further configured to:
[0114] Calculating a current reconstruction loss value using a first loss function based on the 3D reconstructed volume label corresponding to the input training sample and the 3D reconstructed sample of the current iteration, wherein the current reconstruction loss value is used to indicate a structural difference between the predicted 3D reconstructed sample and the 3D reconstructed volume label;
[0115] Calculating a current structural similarity loss value using a second loss function based on the 3D reconstructed volume label corresponding to the input training sample and the 3D reconstructed sample of the current iteration, wherein the current structural similarity loss value is used to indicate a similarity difference between the predicted 3D reconstructed sample and the 3D reconstructed volume label;
[0116] Calculating a current gradient loss using a third loss function based on the 3D reconstructed volume label corresponding to the input training sample and the 3D reconstructed sample of the current iteration, wherein the current gradient loss is used to indicate a difference in image detail information between the predicted 3D reconstructed sample and the 3D reconstructed volume label;
[0117] Calculating a current projection loss using a fourth loss function based on the 3D reconstructed volume label corresponding to the input training sample and the 3D reconstructed sample of the current iteration, wherein the current projection loss is used to indicate a physical process loss of CT imaging;
[0118] The current reconstruction loss value, the current structural similarity loss value, the current gradient loss and the current projection loss are weighted to obtain the current total loss value.
[0119] Optionally, the training module 66 is further configured to:
[0120] Using a mean square error function, calculating a horizontal gradient loss in a horizontal direction between the 3D reconstructed volume label corresponding to the input training sample and the 3D reconstructed sample of the current iteration, and calculating a vertical gradient loss in a vertical direction between the 3D reconstructed volume label corresponding to the input training sample and the 3D reconstructed sample of the current iteration;
[0121] The current gradient loss is calculated based on the horizontal gradient loss and the vertical gradient loss.
[0122] Optionally, the training module 66 is further configured to:
[0123] Calculating current projection data of the current iteration based on the three-dimensional reconstruction sample of the current iteration;
[0124] Calculating a projection label based on the three-dimensional reconstructed volume label;
[0125] The current projection loss is calculated based on the current projection data and the projection label.
[0126] See also Figure 7 In another aspect of an embodiment of the present application, a computing device 10 is provided, comprising a memory 3011 and a processor 3012. The memory 3011 stores a computer program. When the computer program is executed by the processor, the processor 3012 performs the steps of the multi-view projection-based three-dimensional CT image reconstruction method provided in any of the above embodiments of the present application. The computing device 10 is, for example, a desktop computer, a laptop computer, a tablet computer, a handheld computer, a smart speaker, a server, etc., a mobile phone (e.g., a smartphone, a wireless phone, etc.), a wearable device (e.g., a pair of smart glasses or a smart watch), or a similar device.
[0127] The processor 3012 is the control center, connecting the various components of the entire computer device using various interfaces and lines. It executes the various functions of the computer device and processes data by running or executing software programs and / or modules stored in the memory 3011 and accessing data stored in the memory 3011. Optionally, the processor 3012 may include one or more processing cores. Preferably, the processor 3012 may integrate an application processor and a modem processor, wherein the application processor primarily processes the operating system, user interfaces, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into the processor 3012.
[0128] The memory 3011 can be used to store software programs and modules. The processor 3012 executes various functional applications and data processing by running the software programs and modules stored in the memory 3011. The memory 3011 may mainly include a program storage area and a data storage area. The program storage area may store an operating system, at least one application required for a function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created based on the use of the computer device, etc. In addition, the memory 3011 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 3011 may also include a memory controller to provide the processor 3012 with access to the memory 3011.
[0129] On the other hand, an embodiment of the present application further provides a storage medium storing a computer program. When the computer program is executed by a processor, the processor executes the steps of the three-dimensional CT image reconstruction method based on multi-view projection provided in any of the above embodiments of the present application.
[0130] Those skilled in the art will appreciate that all or part of the processes in the methods provided in the above embodiments can be implemented by instructing the relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0131] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. The scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A three-dimensional CT image reconstruction method based on multi-view projection, characterized in that: The method comprises: Acquire multiple projection images of a three-dimensional object to be reconstructed captured at different viewing angles and initial volume data corresponding to the three-dimensional object to be reconstructed; Using a dimension expansion network in a pre-trained three-dimensional reconstruction model, the projection images are expanded in channel dimension to obtain expanded projection images corresponding to the projection images; Using each of the extended projection images as an input to a feature extraction network in the three-dimensional reconstruction model, and extracting multi-scale feature data corresponding to each of the projection images; Based on the initial volume data, using an implicit neural learning network in a three-dimensional reconstruction model, obtaining multi-scale implicit neural representation data corresponding to the three-dimensional object to be reconstructed; Based on the multi-scale feature data corresponding to each of the projection images and the multi-scale implicit neural representation data, a feature reconstruction network in the three-dimensional reconstruction model is used to perform three-dimensional image reconstruction and output a three-dimensional reconstruction target.
2. The 3D CT image reconstruction method based on multi-view projection according to claim 1, wherein: The dimension expansion network includes a gradient calculation layer, a gradient splicing layer, and a dimension expansion layer. The dimension expansion network in the pre-trained 3D reconstruction model is used to expand the channel dimension of the projection image to obtain an expanded projection image corresponding to each projection image, including: Based on the projected image, using a gradient calculation layer, calculate the gradients in the horizontal and vertical directions to obtain a gradient image corresponding to the projected image; Using a gradient stitching layer, summing the projected image and the gradient image corresponding to the projected image to obtain a composite image corresponding to the projected image, and performing channel stitching on the projected image, the gradient image corresponding to the projected image, and the composite image corresponding to the projected image to obtain a three-channel initial image corresponding to the projected image; Based on multiple cascaded residual blocks in the dimension expansion layer, feature maps in the three-channel initial image are gradually extracted to obtain an extended projected image corresponding to the projected image.
3. The 3D CT image reconstruction method based on multi-view projection according to claim 1, wherein: The feature extraction network includes an encoder and a decoder, the decoder is connected to the feature reconstruction network, and the implicit neural learning network is connected to the feature reconstruction network; The encoder includes multiple encoding layers, the decoder includes multiple decoding layers, the encoding layers located in the same scale layer are jump-connected to the corresponding decoding layers, the feature reconstruction network includes multiple reconstruction layers with the same number of layers as the decoder and a reconstruction output layer connected to the reconstruction layer, and the input of the decoding layer located in the same scale layer serves as the input of the corresponding reconstruction layer; the feature reconstruction network includes a multi-scale fusion network, the implicit neural learning network is connected to the multi-scale fusion network, the multi-scale fusion network is respectively connected to each layer of the feature reconstruction network, and the multi-scale fusion network is used to output implicit neural representation data of each scale.
4. The 3D CT image reconstruction method based on multi-view projection according to claim 3, wherein: The step of reconstructing a three-dimensional image using a feature reconstruction network in the three-dimensional reconstruction model based on the multi-scale feature data corresponding to each of the projection images and the multi-scale implicit neural representation data, and outputting a three-dimensional reconstruction target includes: For a current reconstruction layer at a current layer number, obtaining current scale feature input data corresponding to each projection image, wherein the current scale feature input data is the input of a current decoding layer connected to the current reconstruction layer in a decoder corresponding to each projection image; Obtain the output of the previous reconstruction layer of the current reconstruction layer; Based on the initial volume data, obtaining current-scale implicit neural representation data corresponding to the current number of layers output by the multi-scale fusion network; Based on the current-scale feature input data corresponding to each projection image, the output of the previous reconstruction layer, and the current-scale implicit neural representation data, obtaining the current reconstructed feature data output by the current reconstruction layer, and performing the steps in sequence until the final reconstructed feature data of the final reconstruction layer is obtained; Obtaining the maximum scale feature data output by the last decoding layer corresponding to each projection image, thereby obtaining the maximum scale feature data corresponding to each projection image; Based on the maximum scale feature data corresponding to each projection image, the scale implicit neural representation data corresponding to the reconstruction output layer and the final reconstruction feature data, the predicted three-dimensional reconstruction target is output through the reconstruction output layer.
5. The three-dimensional CT image reconstruction method based on multi-view projection according to any one of claims 1 to 4, characterized in that: The method further comprises: Acquire a training data set, wherein each training sample in the training data set includes a plurality of sample projection images collected from multiple perspectives of a three-dimensional sample target, initial sample volume data corresponding to the three-dimensional sample target, and a three-dimensional reconstructed volume label corresponding to the three-dimensional sample target; Based on the training data set, obtaining input training samples of the current iteration, iteratively training the initial 3D reconstruction model to obtain 3D reconstruction samples of the current iteration, and calculating the current total loss value based on the 3D reconstruction volume labels corresponding to the input training samples and the 3D reconstruction samples of the current iteration; Based on the current total loss value, determine whether the current iteration meets the iteration termination condition. If the current iteration does not meet the iteration termination condition, continue to obtain input training samples for iterative training until the iteration termination condition is met. The three-dimensional reconstruction model after meeting the iteration termination condition is used as the pre-trained three-dimensional reconstruction model.
6. The three-dimensional CT image reconstruction method based on multi-view projection according to claim 5, characterized in that: Calculating the current total loss value based on the 3D reconstructed volume label corresponding to the input training sample and the 3D reconstructed sample of the current iteration includes: Calculating a current reconstruction loss value using a first loss function based on the 3D reconstructed volume label corresponding to the input training sample and the 3D reconstructed sample of the current iteration, wherein the current reconstruction loss value is used to indicate a structural difference between the predicted 3D reconstructed sample and the 3D reconstructed volume label; Calculating a current structural similarity loss value using a second loss function based on the 3D reconstructed volume label corresponding to the input training sample and the 3D reconstructed sample of the current iteration, wherein the current structural similarity loss value is used to indicate a similarity difference between the predicted 3D reconstructed sample and the 3D reconstructed volume label; Calculating a current gradient loss using a third loss function based on the 3D reconstructed volume label corresponding to the input training sample and the 3D reconstructed sample of the current iteration, wherein the current gradient loss is used to indicate a difference in image detail information between the predicted 3D reconstructed sample and the 3D reconstructed volume label; Calculating a current projection loss using a fourth loss function based on the 3D reconstructed volume label corresponding to the input training sample and the 3D reconstructed sample of the current iteration, wherein the current projection loss is used to indicate a physical process loss of CT imaging; The current reconstruction loss value, the current structural similarity loss value, the current gradient loss and the current projection loss are weighted to obtain the current total loss value.
7. The three-dimensional CT image reconstruction method based on multi-view projection according to claim 6, characterized in that: Calculating the current gradient loss using a third loss function based on the 3D reconstructed volume label corresponding to the input training sample and the 3D reconstructed sample of the current iteration includes: Using a mean square error function, calculating a horizontal gradient loss in a horizontal direction between the 3D reconstructed volume label corresponding to the input training sample and the 3D reconstructed sample of the current iteration, and calculating a vertical gradient loss in a vertical direction between the 3D reconstructed volume label corresponding to the input training sample and the 3D reconstructed sample of the current iteration; The current gradient loss is calculated based on the horizontal gradient loss and the vertical gradient loss.
8. The three-dimensional CT image reconstruction method based on multi-view projection according to claim 6, characterized in that: Calculating the current projection loss using a fourth loss function based on the 3D reconstructed volume label corresponding to the input training sample and the 3D reconstructed sample of the current iteration includes: Calculating current projection data of the current iteration based on the three-dimensional reconstruction sample of the current iteration; Calculating a projection label based on the three-dimensional reconstructed volume label; The current projection loss is calculated based on the current projection data and the projection label.
9. A computing device, characterized in that The system comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the three-dimensional CT image reconstruction method based on multi-view projection according to any one of claims 1 to 8.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the three-dimensional CT image reconstruction method based on multi-view projection according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Three-dimensional reconstruction method and system based on implicit function
CN117095132A
Computerized tomography reconstruction method based on deep learning, program product and equipment
CN119478086A
Three-dimensional CT imaging method and device, electronic equipment and storage medium
CN119784931A