A method and apparatus for fast image reconstruction from a ring-shaped dual-source dual-energy source based on neural networks.
By employing a neural network-based method for rapid reconstruction of dual-energy images from a ring-shaped dual-source light source, and utilizing sparse perspective data acquisition combined with multi-loss function training, the problems of long reconstruction time and high radiation in dual-energy CT are solved, achieving high-quality image reconstruction suitable for clinical diagnosis and detection of subtle lesions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- MANTEIA TECH CO LTD
- Filing Date
- 2026-07-01
- Publication Date
- 2026-07-31
AI Technical Summary
Existing dual-energy CT reconstruction methods rely on full-view data, resulting in long scanning times and high radiation doses. Sparse view sampling leads to a decline in image quality, making it difficult to meet clinical diagnostic requirements.
A fast image reconstruction method based on a ring dual-source dual-energy source is adopted. Projection data is collected under sparse viewpoints and the image is reconstructed using a neural network model, including a projection domain sub-network, an analytical reconstruction operator and an image domain sub-network. The model is trained and reconstructed by combining pixel loss, structural similarity loss, dual-energy spectrum difference loss and cross-domain feature alignment loss function.
While reducing scanning time and radiation dose, it improves image quality, meets clinical diagnostic needs, reduces artifacts and structural blur, and is particularly suitable for the diagnosis of subtle lesions in high-sparseness scenarios.
Smart Images

Figure CN122492474A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of medical technology, and more specifically, to a method and apparatus for rapid reconstruction of ring-shaped dual-source dual-energy images based on neural networks. Background Technology
[0002] In computed tomography (CT) imaging, dual-energy CT (DECT) enables tissue identification by acquiring projection data at two different energy spectra. Current dual-energy CT reconstruction methods typically require acquiring projection data from the entire field of view to obtain satisfactory image quality.
[0003] However, full-view data acquisition requires completing the entire scan trajectory, resulting in longer scan times and higher radiation doses to patients. Some studies have attempted to reduce the dose by decreasing the number of sampling views, but reducing the number of sampling views leads to incomplete projection data, and the image quality reconstructed by existing methods is significantly reduced, making it difficult to meet clinical diagnostic requirements. Summary of the Invention
[0004] This application provides a method and apparatus for rapid reconstruction of ring-shaped dual-source dual-energy images based on neural networks, which at least solves the technical problems of existing dual-energy CT reconstruction methods that rely on full-view data, resulting in long scanning time and high radiation dose.
[0005] According to one aspect of the embodiments of this application, a method for fast reconstruction of a ring-shaped dual-source dual-energy image based on a neural network is provided, comprising: acquiring a first type of projection data and a second type of projection data collected by a ring-shaped dual-source dual-energy imaging system under sparse viewing angles, wherein the sparse viewing angles represent that the number of sampling viewing angles is less than the number of viewing angles used in full-view acquisition; the first type of projection data is used to represent the projection signal collected after a ray of a first energy level penetrates the object under test, and the second type of projection data is used to represent the projection signal collected after a ray of a second energy level penetrates the object under test, wherein the second energy level is higher than the first energy level; and reconstructing the image using the first type of projection data and the second type of projection data through a neural network model to obtain a target reconstructed image; wherein the loss function corresponding to the neural network model consists of a pixel loss sub-function, a structural phase loss sub-function, and a structural phase loss sub-function. The loss function is calculated by weighting the similarity loss function, the dual-energy spectrum difference loss function, and the cross-domain feature alignment loss function. Among them: the pixel loss function is used to constrain the pixel-level difference between the target reconstructed image and the ground truth image data, where the ground truth image data is an image reconstructed from full-view projection data; the structural similarity loss function is used to constrain the structural similarity between the target reconstructed image and the ground truth image data; the dual-energy spectrum difference loss function is used to constrain the image difference between different channels in the target reconstructed image to be consistent with the image difference between corresponding different channels in the ground truth image data; the cross-domain feature alignment loss function is used to constrain the consistency between the projection domain feature map extracted by the projection domain sub-network and the image domain feature map extracted by the image domain sub-network, where both the projection domain sub-network and the image domain sub-network are partial sub-networks of the neural network model.
[0006] According to another aspect of the embodiments of this application, a training method for a neural network model for fast reconstruction of a ring-shaped dual-source dual-energy image is also provided. The neural network model includes at least: a projection domain sub-network, an analytical reconstruction operator, and an image domain sub-network, comprising the following steps: obtaining a training dataset, the training dataset including: training projection data collected under sparse perspectives as training input, and ground truth image data as training labels, wherein the ground truth image data is an image reconstructed from full-view projection data, and sparse perspectives represent that the number of sampling perspectives is less than the number of perspectives used when collecting full-view images; based on the training dataset, jointly training the projection domain sub-network, the analytical reconstruction operator, and the image domain sub-network using a loss function; wherein, the loss function... The loss function is calculated by weighting the pixel loss function, structural similarity loss function, dual-energy spectrum difference loss function, and cross-domain feature alignment loss function. Specifically: the pixel loss function constrains the pixel-level differences between the reconstructed target image and the ground truth image data; the structural similarity loss function constrains the structural similarity between the reconstructed target image and the ground truth image data; the dual-energy spectrum difference loss function constrains the image differences between different channels in the reconstructed target image to be consistent with the image differences between corresponding different channels in the ground truth image data; and the cross-domain feature alignment loss function constrains the consistency between the projection domain feature map extracted by the projection domain sub-network and the image domain feature map extracted by the image domain sub-network.
[0007] According to another aspect of the embodiments of this application, a fast image reconstruction device for a ring-shaped dual-source dual-energy image based on a neural network is also provided, comprising: a projection data acquisition unit, configured to acquire a first type of projection data and a second type of projection data collected by a ring-shaped dual-source dual-energy imaging system under sparse viewing angles, wherein the sparse viewing angles represent that the number of sampling viewing angles is less than the number of viewing angles used when acquiring full-view images; the first type of projection data is used to represent the projection signal acquired after a ray of a first energy level penetrates the object under test, and the second type of projection data is used to represent the projection signal acquired after a ray of a second energy level penetrates the object under test, wherein the second energy level is higher than the first energy level; and an image reconstruction unit, configured to perform image reconstruction on the first type of projection data and the second type of projection data through a neural network model to obtain a target reconstructed image, wherein the loss function corresponding to the neural network model consists of a pixel loss subfunction. The loss function is calculated by weighting the structural similarity loss function, the dual-energy spectrum difference loss function, and the cross-domain feature alignment loss function. Specifically: the pixel loss function constrains the pixel-level difference between the reconstructed target image and the ground truth image data, where the ground truth image data is an image reconstructed from full-view projection data; the structural similarity loss function constrains the structural similarity between the reconstructed target image and the ground truth image data; the dual-energy spectrum difference loss function constrains the image difference between different channels in the reconstructed target image to be consistent with the image difference between corresponding different channels in the ground truth image data; and the cross-domain feature alignment loss function constrains the consistency between the projection domain feature map extracted by the projection domain sub-network and the image domain feature map extracted by the image domain sub-network, where both the projection domain sub-network and the image domain sub-network are partial sub-networks of the neural network model.
[0008] According to another aspect of the embodiments of this application, a training device for a neural network model for fast reconstruction of a ring-shaped dual-source dual-energy image is also provided. The neural network model includes at least: a projection domain sub-network, an analytical reconstruction operator, and an image domain sub-network, comprising the following steps: a training dataset acquisition unit, used to acquire a training dataset, the training dataset including: training projection data collected under sparse perspectives as training input, and ground truth image data as training labels, the ground truth image data being an image reconstructed from full-view projection data, sparse perspectives representing that the number of sampling perspectives is less than the number of perspectives used when collecting full-view images; and a joint training processing unit, used to perform joint training processing on the projection domain sub-network, the analytical reconstruction operator, and the image domain sub-network based on the training dataset using a loss function. The network is jointly trained; the loss function is calculated by weighting the pixel loss function, structural similarity loss function, dual-energy spectrum difference loss function, and cross-domain feature alignment loss function. Specifically: the pixel loss function is used to constrain the pixel-level difference between the target reconstructed image and the ground truth image data; the structural similarity loss function is used to constrain the structural similarity between the target reconstructed image and the ground truth image data; the dual-energy spectrum difference loss function is used to constrain the image difference between different channels in the target reconstructed image to be consistent with the image difference between corresponding different channels in the ground truth image data; and the cross-domain feature alignment loss function is used to constrain the consistency between the projection domain feature map extracted by the projection domain sub-network and the image domain feature map extracted by the image domain sub-network. Attached Figure Description
[0009] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0010] Figure 1 This is a schematic diagram of an optional neural network-based fast image reconstruction method for ring-shaped dual-source dual-energy images according to an embodiment of this application;
[0011] Figure 2 This is a flowchart of an optional neural network-based method for rapid reconstruction of ring-shaped dual-source dual-energy images according to an embodiment of this application;
[0012] Figure 3 This is a schematic diagram of the internal processing flow of an optional integrated domain image reconstruction network according to an embodiment of this application;
[0013] Figure 4 This is a schematic diagram of a two-stream residual network structure of an optional synthetic domain image reconstruction network according to an embodiment of this application;
[0014] Figure 5This is a schematic diagram of an optional dual-energy CT scanning configuration with spectral complementarity according to an embodiment of this application;
[0015] Figure 6 This is a schematic diagram of an optional training method for a neural network model for fast reconstruction of a ring-shaped dual-source dual-energy image according to an embodiment of this application;
[0016] Figure 7 This is a schematic diagram of an optional neural network-based ring-shaped dual-light source dual-energy image rapid reconstruction device according to an embodiment of this application;
[0017] Figure 8 This is a schematic diagram of an optional training apparatus for a neural network model for rapid reconstruction of a ring-shaped dual-source dual-energy image, according to an embodiment of this application. Detailed Implementation
[0018] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present application. It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of the present application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the present application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or devices.
[0019] According to an embodiment of this application, a method embodiment for rapid reconstruction of ring-shaped dual-source dual-energy images based on neural networks is provided. It should be noted that the steps shown in the flowcharts can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowcharts, in some cases, the steps shown or described may be executed in a different order. It should also be noted that the information and data collected in this application are authorized by the user or fully authorized by all parties. The collection, storage, use, processing, transmission, provision, disclosure, and application of related data all comply with relevant laws, regulations, and standards, and necessary confidentiality measures are taken. They do not violate public order and good morals, and corresponding operation entry points are provided for users to choose to authorize or refuse. Interfaces are set up between this system and relevant users or institutions, providing users with corresponding operation entry points for users to choose to agree to or refuse automated decision results; if the user chooses to refuse, the process enters the expert decision-making process. According to the embodiments of this application, an image processing system can be used as the execution subject of the fast reconstruction method of ring dual-light source dual-energy image based on neural network in the embodiments of this application. The system can be a software system or an embedded system combining software and hardware. Of course, the execution subject of the method in the embodiments of this application can also be other forms of execution subject. This application does not make any special limitation on the specific form of the execution subject.
[0020] Figure 1 This is a schematic diagram of a fast image reconstruction method based on a neural network for ring-shaped dual-source dual-energy images according to an embodiment of this application, as shown below. Figure 1 As shown, the fast image reconstruction method for ring-shaped dual-source dual-energy images based on neural networks includes the following steps:
[0021] Step S101: Acquire the first type of projection data and the second type of projection data collected by the ring dual-source dual-energy imaging system under sparse viewing angle. The sparse viewing angle indicates that the number of sampling viewing angles is less than the number of viewing angles used when acquiring full-view data. The first type of projection data is used to characterize the projection signal acquired after the first energy level of the ray penetrates the object under test. The second type of projection data is used to characterize the projection signal acquired after the second energy level of the ray penetrates the object under test. The second energy level is higher than the first energy level.
[0022] For example, a ring-shaped dual-source dual-energy imaging system can refer to a ring-shaped rotating gantry system equipped with two X-ray tubes and two corresponding detectors. The two X-ray tubes are installed at a certain angle and can simultaneously emit rays of different energies, which can be, but are not limited to, X-rays, gamma rays, proton beams, or heavy ion beams. The first type of projection data corresponds to the projection signal acquired by the corresponding detector after the lower-energy ray penetrates the object under test, and the second type of projection data corresponds to the projection signal acquired after the higher-energy ray penetrates the object under test. Sparse viewing angles refer to sampling a smaller number of viewing angles than those used in full-view acquisition; that is, only projection data in some angular directions are acquired, rather than continuously acquiring all angles. The image processing system acquires the first and second types of projection data, which is beneficial for providing input data for subsequent neural network completion and reconstruction using the complementarity of dual-energy information. By adopting sparse viewing angle sampling, it is beneficial to significantly reduce scanning time and the radiation dose received by the object under test, and by utilizing the spatial domain consistency and energy domain redundancy of dual-energy data, it is beneficial to improve the quality of the image reconstructed from sparse projections.
[0023] In some embodiments, the image processing system can control the annular dual-source dual-energy imaging system to alternately acquire rays of a first energy level and a second energy level during the rotation of the annular gantry. For example, sampling angles of the first type of projection data are extracted from full-view projection data (e.g., 720 full views) at preset intervals (e.g., extracting one low-energy angle from every eight full views), and sampling angles of the second type of projection data are extracted at angle positions complementary to the sampling angles of the first type of projection data. The image processing system can synchronously receive the projection signals output by the two detectors through a data interface and store them in angular order as first type of projection data and second type of projection data. This alternating acquisition embodiment utilizes the complementary angle information formed by dual-energy alternating sampling, which allows sufficient target structure information to be retained even under sparse sampling, thus helping to avoid the loss of key information caused by sparse acquisition of a single energy spectrum.
[0024] In other embodiments, the image processing system can acquire dual-energy projection data under sparse viewing angles by simulating the geometry of fan-beam imaging. The image processing system presets geometric parameters such as the distance from the X-ray source to the detector, the distance from the X-ray source to the rotation center, and the unit size of the detector. The image processing system determines the total number of viewing angles to be acquired across the entire viewing angle based on the target scanning range, then generates a sparse sampling sequence according to a preset sparsity factor, and controls the imaging system to perform actual scanning according to this sparse sampling sequence. After acquisition, the image processing system performs standardization and initial interpolation preprocessing on the acquired first-type and second-type projection data to generate initial projection data for subsequent neural network model processing. This simulation acquisition embodiment is suitable for flexibly adjusting sparsity according to clinical needs, which is beneficial for achieving different proportions of sparse sampling, thereby matching different dose reduction targets.
[0025] For example, the image processing system can employ the following simulated fan-beam imaging geometry parameters: the distance from the X-ray source to the detector is set to 1270 mm, and the distance from the X-ray source to the rotation center is set to 870 mm. The detector contains 848 detector elements, each with a size of 0.6 mm. The reconstructed image uses a uniform pixel size of 0.50 mm × 0.50 mm. During full-view acquisition, a total of 720 views are acquired, with an angular interval of 0.5 degrees between adjacent views. For sparse view sampling, the image processing system supports two sampling schemes: a 90-view sampling scheme (i.e., extracting one low-energy view from every 8 full views and one high-energy view from a complementary position, corresponding to k=90, n=8) and a 72-view sampling scheme (i.e., extracting one low-energy view from every 10 full views and one high-energy view from a complementary position, corresponding to k=72, n=10). The image processing system can select one of these two sampling schemes to execute based on clinical requirements for radiation dose and reconstruction quality.
[0026] Step S102: Input the first type of projection data and the second type of projection data into the projection domain sub-network in the neural network model, and complete the sparsely sampled missing projection data through the projection domain sub-network to output the target projection data, wherein the target projection data includes the completed first type of projection data and the completed second type of projection data.
[0027] For example, the projection domain subnetwork can refer to a neural network model component used to process sparse projection data. The projection domain subnetwork can complete the viewpoint gaps caused by sparse sampling. The image processing system inputs the acquired first-type and second-type projection data into the projection domain subnetwork. The projection data missing due to sparse sampling can refer to projection information in the angular direction that was not acquired during the sparse viewpoint acquisition process. The projection domain subnetwork utilizes the spatial domain consistency and energy domain redundancy of dual-energy data. It can process the low-energy and high-energy channels separately through its internal dual-stream residual structure and can employ a detail attention mechanism to strengthen the completion weights of small structural regions, thereby outputting the completed target projection data, which includes the completed first-type and second-type projection data. Through the completion processing of the projection domain subnetwork, the image processing system can restore the incomplete projection data acquired under sparse viewpoints to near-full-view complete projection data, which helps improve the integrity of the analytically reconstructed data and reduce image artifacts and structural blurring caused by viewpoint gaps.
[0028] In some embodiments, the projection domain subnetwork can employ a two-stream residual network structure, including a first branch and a second branch. The first branch processes a first type of projection data, and the second branch processes a second type of projection data. The image processing system inputs the first type of projection data and the second type of projection data into their respective branches. The first and second branches have the same structure but independent parameters, allowing each to learn the mapping rules from sparse projection to complete projection. The projection domain subnetwork introduces a channel attention mechanism in the middle layer of the network. This mechanism can automatically identify pixel regions with high gradient magnitudes in the projection data, which may correspond to the edges or microstructures of the measured object, and increase the completion weights of these pixel regions with high gradient magnitudes. This makes the network pay more attention to pixel positions containing rich structural information during training and inference. Through this two-stream residual network structure, the image processing system can effectively recover the edge details and microstructure information missing in sparse sampling, improving the quality of the completed projection.
[0029] In other embodiments, the projection domain subnetwork receives structural prior feedback from the image domain subnetwork during the completion process. For example, the image domain subnetwork extracts anatomical features, such as feature maps of fine structures like microvessels and trabeculae, from the intermediate reconstructed image and converts these image domain features into structural prior projection reference data through a Radon transform-based mapping. The image processing system inputs this structural prior projection reference data along with the original sparse projection data into the projection domain subnetwork as additional guiding information. Utilizing the anatomical feature maps from the image domain subnetwork and the corresponding structural prior projection reference data, the projection domain subnetwork more accurately reconstructs the projection region corresponding to the anatomical structure during completion. This bidirectional feature interaction embodiment forms a closed-loop collaborative enhancement through bidirectional feature interaction between the projection domain and the image domain, which is beneficial for improving the accuracy of projection completion under sparse perspectives, and is particularly suitable for image reconstruction in high-sparseness scenarios.
[0030] Step S103: The target projection data is reconstructed using the analytical reconstruction operator in the neural network model to obtain an intermediate reconstructed image. The analytical reconstruction operator is used to transform the target projection data from the projection domain to the image domain.
[0031] For example, the analytical reconstruction operator can refer to a mathematical transformation operator used to transform projection domain data to the image domain. The image processing system inputs the target projection data output by the projection domain sub-network into the analytical reconstruction operator. For instance, this analytical reconstruction operator can perform weighted sum filtering on the projection data at each acquisition angle, and then back-project it along the ray propagation path to the image space, thereby converting the projection intensity distribution in the projection domain into the attenuation coefficient distribution in the image domain, obtaining an intermediate reconstructed image. Through the transformation of the analytical reconstruction operator, the image processing system can quickly obtain a preliminary image with a certain spatial resolution from the completed projection data. Compared with filtered back-projection, the analytical reconstruction operator in this application can be trained in conjunction with a neural network model, and the internal parameters of the analytical reconstruction operator can be adaptively adjusted in an end-to-end framework, which is beneficial to improving reconstruction accuracy.
[0032] In some embodiments, the analytical reconstruction operator can employ an explicit mathematical expression based on the inverse Radon transform, which may include a weighted summation and filtering operation of the projected data. The image processing system performs reconstruction for each energy channel separately: the completed first-class projected data is filtered by a ramp filter and then back-projected along the ray path to each voxel in the image space to obtain an intermediate reconstructed image of the first energy level; the same operation is performed on the completed second-class projected data to obtain an intermediate reconstructed image of the second energy level. The intermediate reconstructed images of the two energy levels can retain low-energy and high-energy attenuation information respectively, which is beneficial for the dual-energy collaborative processing of subsequent image domain subnetworks. This embodiment utilizes the analytical properties of the inverse Radon transform, which helps to ensure the mathematical interpretability of the reconstruction process and has high computational efficiency, making it suitable for real-time or near-real-time clinical image reconstruction needs.
[0033] In other embodiments, the analytical reconstruction operator can be embedded into the neural network model as a differentiable module. During the training phase, the image processing system uses parameters of the analytical reconstruction operator, such as the filter cutoff frequency and weighting factors, as learnable variables, and performs end-to-end joint optimization together with the parameters of the projection domain subnetwork and the image domain subnetwork. During the inference phase, the image processing system uses the trained analytical reconstruction operator to perform forward computation on the target projection data to obtain the intermediate reconstructed image. This embodiment allows the parameters of the analytical reconstruction operator to be adaptively adjusted according to the statistical characteristics of dual-energy data, which is beneficial for obtaining better intermediate image quality than fixed-parameter reconstruction operators under conditions such as sparse viewpoints and low doses, thereby improving the optimization effect of subsequent image domain subnetworks.
[0034] Step S104: Input the intermediate reconstructed image into the image domain sub-network of the neural network model, and perform edge sharpening and noise suppression on the intermediate reconstructed image through the image domain sub-network to output the target reconstructed image.
[0035] For example, an image domain subnetwork can refer to a neural network model component used to optimize the quality of intermediate reconstructed images. The image domain subnetwork can post-process the analytically reconstructed images to improve sharpness and signal-to-noise ratio. The image processing system inputs the intermediate reconstructed images output by the analytical reconstruction operator into the image domain subnetwork. The image domain subnetwork employs a multi-scale convolutional structure, which can extract anatomical features at different scales from the intermediate reconstructed images and remove noise and artifacts from the images through a residual learning mechanism. Simultaneously, it sharpens and enhances edges, outputting the target reconstructed image. Through the optimization processing of the image domain subnetwork, the image processing system can improve the detail preservation ability and overall visual quality of the reconstructed images, reducing edge blurring and noise amplification problems that may occur in analytical reconstruction under sparse viewpoint conditions.
[0036] In some embodiments, the image domain subnetwork can employ a two-stream residual network structure, including a third branch and a fourth branch. The third branch processes the portion of image data corresponding to the first type of projection data in the intermediate reconstructed image, and the fourth branch processes the portion of image data corresponding to the second type of projection data in the intermediate reconstructed image. The third and fourth branches have the same structure but independent parameters, each learning the mapping rule from the intermediate reconstructed image to the target reconstructed image. The image domain subnetwork uses multi-scale convolution in the encoder part of the network, such as dilated convolution with different dilation rates, to expand the receptive field and capture contextual information at different scales, which is beneficial for simultaneously recovering large-scale anatomical structures and local edge textures. Through this two-stream residual network structure, the image processing system can perform targeted optimization on low-energy and high-energy channel images separately, while maintaining spatial consistency between the two-energy images.
[0037] In other embodiments, the image domain subnetwork not only outputs the reconstructed target image but also feeds back the anatomical features extracted during the optimization process to the projection domain subnetwork, forming a cross-domain collaborative enhancement loop. For example, the image domain subnetwork extracts feature maps related to fine structures from the high-level feature maps of the decoder and maps these feature maps back to the projection domain using Radon transform, generating structural prior projection reference data. The image processing system feeds this structural prior projection reference data back to the projection domain subnetwork to guide it in more accurately completing sparse and missing projection data in subsequent iterations or the next training round. This closed-loop feedback embodiment, through bidirectional feature interaction between the image and projection domains, enables projection domain completion and image domain optimization to mutually promote each other, which is beneficial for obtaining clearer target reconstruction images with fewer artifacts under high-sparse sampling conditions, and is particularly suitable for the early diagnosis and quantitative analysis of subtle lesions.
[0038] Figure 2This diagram illustrates a flowchart of a neural network-based fast image reconstruction method for a ring-shaped dual-source dual-energy imaging system according to an embodiment of this application. First, a ring-shaped dual-source dual-energy imaging system acquires first and second type projection data from a sparse viewpoint. This acquisition process employs a spectral complementary sparse viewpoint acquisition method, where rays of the first and second energy levels are sampled alternately at complementary angles. The acquired sparse projection data then undergoes edge-aware interpolation preprocessing to generate initial projection data that preserves contour information. Next, the initial projection data is fed into a comprehensive domain image reconstruction network, which includes a projection domain subnetwork, an analytical reconstruction operator, and an image domain subnetwork. The projection domain subnetwork completes the missing projection data from the sparse sampling, outputting the target projection data; the analytical reconstruction operator reconstructs the target projection data into an intermediate reconstructed image; and the image domain subnetwork performs edge sharpening and noise suppression on the intermediate reconstructed image, outputting the target reconstructed image. Finally, the target reconstructed image is output.
[0039] In some optional embodiments, before reconstructing the target projection data using the parsing reconstruction operator to obtain an intermediate reconstructed image, the method further includes: the image processing system can obtain a completed projection data term and a cross-domain feature fusion term, wherein the completed projection data term is equal to the target projection data output by the projection domain sub-network, and the cross-domain feature fusion term is equal to the product of the cross-domain feature weights and the structural prior projection reference data fed back from the image domain sub-network; the structural prior projection reference data represents the projection data obtained after mapping the image optimized by the image domain sub-network back to the projection domain through Radon transform; the completed projection data term and the cross-domain feature fusion term are summed to obtain the target data term; the transform operator, self- The algorithm includes an adaptive projection weighting factor, an adaptive total variation regularization coefficient, and a total variation regularization term. The transformation operator is the inverse Radon transform operator, used as the transformation coefficient to transform the target projection data from the projection domain to the image domain. The adaptive projection weighting factor dynamically adjusts the projection weights of the target projection data based on its sparsity. The adaptive total variation regularization coefficient adjusts the regularization intensity of the intermediate reconstructed images based on the sparsity of the target projection data. The total variation regularization term constrains the edge smoothness of the intermediate reconstructed images. The analytical reconstruction operator is determined based on the target data term, the transformation operator, the adaptive projection weighting factor, the adaptive total variation regularization coefficient, and the total variation regularization term.
[0040] For example, the analytical reconstruction operator can be used to transform the target projection data output by the projection domain subnetwork into an intermediate reconstructed image. The image processing system obtains a completed projection data term and a cross-domain feature fusion term, where the completed projection data term is equal to the target projection data output by the projection domain subnetwork, i.e., the completed first-class projection data and second-class projection data, and the cross-domain feature fusion term is equal to the product of the cross-domain feature weights and the structural prior projection reference data fed back from the image domain subnetwork. The structural prior projection reference data represents the projection data obtained by mapping the image optimized by the image domain subnetwork back to the projection domain through Radon transform, and this projection data may contain the recovered anatomical structure information in the image domain. The image processing system sums the completed projection data term and the cross-domain feature fusion term to obtain the target data term. By introducing the cross-domain feature fusion term, the image processing system can inject the structural prior information in the image domain into the projection domain, so that the analytical reconstruction process not only depends on the data completed in the projection domain, but is also guided by the optimization results in the image domain, which is beneficial to improving reconstruction accuracy, especially in high-sparse scenarios.
[0041] For example, the transformation operator is the inverse Radon transform operator, used to transform the target data term from the projection domain to the image domain. An adaptive projection weighting factor is used to dynamically adjust the weights of the projected data based on the sparsity of the target projection data: when the sparsity is high (e.g., high data density but low noise), the weight of the effective projection is increased; when the sparsity is low (e.g., the data itself has a high noise ratio), the baseline weight is maintained to avoid excessive amplification of noise. An adaptive total variation regularization coefficient is used to adjust the regularization strength of the intermediate reconstructed image based on the sparsity of the target projection data; the higher the sparsity, the greater the regularization strength, to suppress artifacts caused by high sparsity. The total variation regularization term is used to constrain the edge smoothness of the intermediate reconstructed image; it is obtained by calculating the square root of the sum of the squares of the differences between adjacent pixels, which can effectively reduce stripe artifacts and maintain sharp edges.
[0042] For example, the process of determining the analytical reconstruction operator can be as follows: multiply the target data term by the adaptive projection weighting factor, then apply the inverse Radon transform operator, and finally multiply the result by the adaptive total variation regularization coefficient and the total variation regularization term to obtain the intermediate reconstructed image. This analytical reconstruction operator is embedded in the neural network model, and its internal parameters can be adaptively adjusted during end-to-end training, thereby improving the reconstruction stability and detail recovery capability in high-sparse scenes.
[0043] By employing an adaptive projection weighting factor and an adaptive total variation regularization term, the reconstruction parameters can be dynamically adjusted according to the sparsity of the projection data, effectively suppressing stripe artifacts and noise caused by sparse sampling. Embedding the analytical reconstruction operator into the neural network for end-to-end training helps to reduce the shortcomings of traditional operators, such as weak anti-interference ability and poor detail recovery.
[0044] For example, the analytical reconstruction operator adopts the form based on the inverse Radon transform and incorporates cross-domain features for optimization. The image processing system determines the analytical reconstruction operator according to the following formula (1):
[0045] Formula (1);
[0046] in, It can represent the intermediate reconstructed image of the output; It can represent the inverse Radon transform operator, i.e., the transform operator, used to transform projection domain data to the image domain; It can represent the target projection data after the projection domain subnetwork is completed, that is, the completed projection data item; It can represent the structural prior projection reference data fed back from the image domain subnetwork. It can represent cross-domain feature weights. It can represent cross-domain feature fusion terms; This can represent the adaptive projection weighting factor, which is used to adjust the weighting factor based on the sparsity of the target projection data. Dynamically adjust projection weights; It can represent the adaptive total variation regularization coefficient, which is used to adjust the regularization strength according to sparsity; It can represent the total variation regularization term, used to constrain the edge smoothness of the intermediate reconstructed image.
[0047] In some embodiments, when calculating the adaptive projection weighting factor, the image processing system first detects the sparsity of the target projection data. When the detected sparsity is less than or equal to a preset threshold (e.g., in high-sparse scenarios, such as sampling ratios of 1 / 8 or 1 / 10), the adaptive projection weighting factor is determined to be greater than the standard weight value of 1, which can be represented as 1 plus a preset coefficient multiplied by (1 minus the sparsity). When the detected sparsity is greater than the preset threshold (i.e., low sparsity, such as near-full-view acquisition), the adaptive projection weighting factor is determined to be equal to the standard weight value of 1. This embodiment, by dynamically adjusting the projection weights, can adaptively enhance the effective projection signal according to the sampling sparsity, thereby improving the reconstruction quality under high sparsity conditions.
[0048] In other embodiments, when calculating the total variation regularization term, the image processing system can calculate the squared difference between each pixel position in the intermediate reconstructed image and its right-hand neighbor and its bottom neighbor, sum the two squares, and then take the square root. The sum of the square roots for all pixel positions yields the total variation regularization term. This regularization term, as part of the reconstruction objective function, can be multiplied by an adaptive total variation regularization coefficient and applied to the reconstruction result. This effectively suppresses stripe artifacts and noise caused by sparse sampling while preserving image edge and texture details. The adaptive total variation regularization coefficient can be a coefficient dynamically adjusted based on the sparsity of the target projection data, used to control the smoothing strength of the total variation regularization term on the intermediate reconstructed image. Higher sparsity results in a larger coefficient, enhancing artifact suppression. The image processing system adjusts the adaptive total variation regularization coefficient according to the sparsity of the target projection data; higher sparsity results in a larger coefficient and stronger regularization, thus achieving a balance between noise suppression and detail preservation.
[0049] In some optional embodiments, the projection domain subnetwork and the image domain subnetwork constitute a cross-domain collaborative enhancement closed-loop architecture; the anatomical structural features extracted from the intermediate reconstructed image by the image domain subnetwork are fed back to the projection domain subnetwork to guide the projection domain subnetwork to complete the projection data missing by sparse sampling, and the anatomical structural features represent the organ structure information extracted from the image domain corresponding to the intermediate reconstructed image and fed back to the projection domain; the target projection data output by the projection domain subnetwork is reconstructed by the analytical reconstruction operator and then input into the image domain subnetwork.
[0050] For example, the cross-domain collaborative enhancement closed-loop architecture can refer to a closed-loop processing flow formed by bidirectional feature interaction between the projection domain sub-network and the image domain sub-network. In this cross-domain collaborative enhancement closed-loop architecture, the target projection data output by the projection domain sub-network is reconstructed by the parsing reconstruction operator and then input into the image domain sub-network, forming a forward path; the image domain sub-network extracts anatomical structural features from the intermediate reconstructed image and feeds these anatomical structural features back to the projection domain sub-network, forming a feedback path. The anatomical structural features represent organ structural information extracted from the image domain for feedback to the projection domain, such as the edge orientation of microvessels, the texture distribution of trabeculae, or the contour boundaries of organs. The image processing system uses the feedback anatomical structural features to guide the projection domain sub-network to complete the projection data missing due to sparse sampling, allowing the projection domain sub-network to refer to the structural priors already recovered in the image domain during completion, thereby improving completion accuracy. Through this closed-loop architecture, the optimization processes of the projection domain and the image domain promote and enhance each other, which is beneficial for preserving subtle anatomical structures under high sparse sampling conditions, thus avoiding structural loss caused by single-domain optimization.
[0051] By enhancing the closed-loop architecture through cross-domain collaboration, the anatomical features extracted by the image domain sub-network are fed back to the projection domain sub-network to guide projection completion, thereby effectively preserving subtle anatomical details and reducing the loss of details caused by single projection domain optimization.
[0052] In some embodiments, the image processing system can extract multi-scale anatomical feature maps, such as edge feature maps or texture feature maps, at the decoder stage of the image domain subnetwork. The image processing system can then map these anatomical feature maps back to the projection domain using Radon transform to obtain structural prior projection reference data. This structural prior projection reference data is then multiplied by cross-domain feature weights and added to the completed projection data term output by the projection domain subnetwork to form the enhanced target data term. This enhanced target data term is then input into the analytical reconstruction operator for reconstruction. This image domain injection embodiment, by injecting the recovered fine structural information from the image domain into the projection domain, can effectively guide the projection domain subnetwork to focus on recovering projection regions related to anatomical structures during completion, thereby improving the preservation of minute lesions or fine tissues in the final reconstructed image.
[0053] In other embodiments, the image processing system maintains this closed-loop architecture during both the training and inference phases. During training, the system can constrain the consistency between the projection domain feature map output by the projection domain sub-network and the image domain feature map extracted by the image domain sub-network using a cross-domain feature alignment loss function, enabling both sub-networks to learn a unified feature representation. During inference, the system fuses the structural prior projection reference data fed back by the image domain sub-network with the completed projection data output by the sub-network for reconstruction. The intermediate reconstructed image is then input back into the image domain sub-network for optimization, forming a complete closed-loop iteration. For highly sparse scenes, the system can perform multiple closed-loop iterations, such as two or three, with projection domain completion and image domain optimization alternating in each iteration to gradually improve reconstruction quality. This multi-round iterative embodiment, through multiple rounds of closed-loop iteration, can gradually recover higher-quality images from extremely sparse projection data, reducing radiation dose and scanning time, and helping to ensure the integrity of image details.
[0054] In some optional embodiments, obtaining the adaptive projection weighting factor includes: when the image processing system detects that the sparsity of the target projection data is less than or equal to a preset threshold, determining that the adaptive projection weighting factor is greater than the standard weight value, where the standard weight value is 1; and when the image processing system detects that the sparsity of the target projection data is greater than the preset threshold, determining that the adaptive projection weighting factor is equal to the standard weight value.
[0055] For example, sparsity can characterize the ratio of the number of sampled viewpoints to the total number of viewpoints in the target projection data output by the projection domain subnetwork; a smaller ratio indicates sparser sampling. The adaptive projection weighting factor can be a coefficient used in the analytical reconstruction operator to dynamically adjust the weights of the projection data, with a standard weight value of 1. The image processing system detects the sparsity of the target projection data. When the detected sparsity is less than or equal to a preset threshold, the image processing system determines that the adaptive projection weighting factor is greater than the standard weight value of 1 to enhance the dominant role of the effective projection signal in reconstruction. When the detected sparsity is greater than the preset threshold, the image processing system determines that the adaptive projection weighting factor is equal to the standard weight value of 1. By dynamically adjusting the projection weighting factor according to the sparsity, the image processing system can enable the analytical reconstruction operator to adaptively adjust the projection weights under input data of different sparsities, which is beneficial for increasing the contribution of the effective signal in high-sparse scenarios, while maintaining the baseline weights to maintain reconstruction stability in low-sparse scenarios.
[0056] In some embodiments, the image processing system can determine the value of the weighting factor by threshold comparison. When the sparsity of the target projection data is detected to be lower than or equal to a preset sparsity threshold, the image processing system sets the adaptive projection weighting factor to a value greater than 1. The specific value can be calculated according to a preset function based on the difference between the sparsity and the threshold. When the sparsity is detected to be higher than the preset threshold, the image processing system sets the adaptive projection weighting factor to 1. This embodiment is beneficial for realizing the binarization switching of the weighting factor, is simple to calculate, and is suitable for real-time reconstruction scenarios.
[0057] In other embodiments, the image processing system can use a continuous function to calculate an adaptive projection weighting factor based on sparsity. The image processing system uses sparsity as an input variable and directly calculates the specific value of the weighting factor through a preset nonlinear mapping relationship, so that the weighting factor smoothly increases as the sparsity decreases. This embodiment helps avoid abrupt changes in reconstruction results caused by threshold jumps and facilitates a smooth transition between different sparsity levels, resulting in more consistent reconstruction performance.
[0058] In some optional embodiments, determining that the adaptive projection weighting factor is greater than the standard weight value when the sparsity of the target projection data is detected to be less than or equal to a preset threshold includes: the image processing system determining that the adaptive projection weighting factor is equal to the preset threshold when the sparsity of the target projection data is detected to be less than or equal to the preset threshold. ,in, A preset coefficient greater than 0. The sparsity of the target projection data.
[0059] For example, preset coefficient This is a constant greater than 0, used to control the sensitivity of the weighting factor to changes in sparsity. Sparsity This can refer to the ratio of the number of sampled viewing angles to the total number of viewing angles in the target projection data output by the projection domain subnetwork, used to characterize the sparsity of the projection data. When the image processing system detects that the sparsity is less than or equal to a preset threshold, it uses formula (2): Calculate the adaptive projection weighting factor. In formula (2) It can reflect the proportion of missing projection data. The smaller The larger the sparsity, the greater the weighting factor becomes, thus increasing as sparsity decreases. In this way, the image processing system can continuously adjust the weighting factor according to the sparsity, enhancing the importance of the effective projection signal in reconstruction under high sparsity, which helps compensate for information loss caused by sparse sampling.
[0060] In some embodiments, the image processing system will preset coefficients Set to a fixed value. When sparsity... As the weighting factor approaches 0 (i.e., extremely sparse sampling), the weighting factor approaches... When sparsity When the weighting factor equals the preset threshold, the weighting factor equals The image processing system can directly calculate the weighting factors without requiring additional learning parameters. This embodiment uses a simple linear formula to facilitate adaptive adjustment of the weighting factors, resulting in high computational efficiency and suitability for real-time reconstruction scenarios.
[0061] In other embodiments, the image processing system will preset coefficients. As learnable parameters, they are optimized end-to-end along with other network parameters during the training process of the neural network model. The image processing system automatically adjusts preset coefficients based on the loss function during the training phase. The values of these values are chosen so that the final weighting factor can be adapted to the statistical characteristics of the training data. During the inference phase, the image processing system uses the pre-trained preset coefficients. The weighting factor is calculated. This embodiment enables the calculation of the weighting factor to be adaptively optimized according to the specific task and data distribution, which is beneficial for obtaining better reconstruction results in different clinical scenarios.
[0062] For example, adaptive projection weighting factors are calculated in exponential form, such as when the sparsity of the target projection data... When the value is less than or equal to the preset threshold, refer to formula (3):
[0063] Formula (3);
[0064] in, Indicates the adaptive projection weighting factor; This is a magnification factor greater than 0, i.e., a preset factor; It can represent exponential parameters; This indicates the sparsity of the target projection data. When the sparsity exceeds a preset threshold... .
[0065] sparsity It not only represents the ratio of the number of sampled viewpoints to the total number of viewpoints, but also reflects the information density and noise level of the projected data. When the data is highly sparsity, there are fewer effective projection signals but more detailed information. In this case, the projection weights are dynamically increased exponentially. When the data is low sparsity, the noise ratio is higher. The baseline weights are maintained to avoid noise amplification.
[0066] Optionally, the adaptive total variation regularization coefficient is calculated using a linear formula:
[0067] Formula (4);
[0068] in, It can represent a preset scaling factor. It can represent the adaptive total variation regularization coefficient. This can represent the sparsity of the target projection data. The formula... This causes the regularization coefficient to increase linearly with increasing sparsity, thereby enhancing the smoothing intensity under high sparsity and effectively suppressing stripe artifacts.
[0069] In some optional embodiments, the adaptive total variation regularization coefficient increases as the sparsity of the target projection data increases. Obtaining the total variation regularization term includes: the image processing system can calculate the sum of squares of the differences between adjacent pixels in the intermediate reconstructed image; take the square root of the sum of squares to obtain the square root result; and use the square root result as the total variation regularization term.
[0070] For example, the sum of squares of the differences between adjacent pixels can refer to the sum of the squared grayscale differences between each pixel in the intermediate reconstructed image and its right-hand neighbor and its bottom neighbor, respectively. The image processing system calculates the sum of squares of the differences between adjacent pixels at each pixel position in the intermediate reconstructed image, and then performs a square root operation on the sum of squares to obtain the square root result for that pixel position. The image processing system accumulates the square root results for all pixel positions and uses the accumulated result as the total variation regularization term. This total variation regularization term is used to measure the degree of local spatial variation in the intermediate reconstructed image and can reflect the edge and texture information of the image. By applying this total variation regularization term as a component of the analytical reconstruction operator to the reconstruction result, the image processing system can effectively suppress stripe artifacts and noise caused by sparse sampling, while maintaining the edge sharpness of the image, which is beneficial to improving the quality of the intermediate reconstructed image. The calculation method of this embodiment is beneficial to directly reflect the gradient magnitude between each pixel point and its neighborhood in the image, which is beneficial to suppressing stripe artifacts and noise caused by sparse sampling during the analytical reconstruction process, while maintaining the edge and texture details of the image. By multiplying the total variation regularization term by the adaptive total variation regularization coefficient and applying it to the reconstruction result, the image processing system can dynamically adjust the smoothing intensity according to the sparsity of the projection data, thereby enhancing the regularization effect under high sparsity.
[0071] For example, the total variation regularization term The formula used to constrain the edge smoothness of the intermediate reconstructed image is as follows:
[0072] Formula (5);
[0073] in, The subscript represents the intermediate reconstructed image output by the analytical reconstruction operator. Indicates an energy channel. Corresponding to the first energy level, This corresponds to the second energy level. It can represent the pixel coordinates in the intermediate reconstructed image. Represents pixels Its right adjacent pixel The square of the grayscale difference, Represents pixels Its adjacent pixels below The square of the grayscale difference.
[0074] In some embodiments, the image processing system can use a discrete difference method to calculate the total variation regularization term. For each pixel position in the intermediate reconstructed image, the image processing system calculates the horizontal difference between that pixel and its right-hand neighbor and the vertical difference between that pixel and its bottom neighbor. The image processing system adds the squares of the horizontal and vertical differences to obtain a sum of squares, and then takes the square root of the sum. The image processing system sums the square root results for all pixel positions as the total variation regularization term. This embodiment, through pixel-by-pixel discrete difference calculation, is direct, computationally efficient, and suitable for integration into deep learning frameworks.
[0075] In other embodiments, the image processing system may use only the sum of the absolute values of the differences between adjacent pixels without performing square root operations when calculating the total variation regularization term. For example, the image processing system calculates the absolute values of the differences in the horizontal and vertical directions, adds them together as the contribution value for that pixel location, and then sums the contribution values for all pixel locations to obtain the total variation regularization term. This embodiment can better maintain the edge sharpness of the image while suppressing noise, making it suitable for reconstruction tasks with high edge preservation requirements. The image processing system can select an appropriate form of total variation based on the noise level of the input data and the reconstruction target to achieve a balance between noise suppression and edge preservation.
[0076] In some optional embodiments, the projection domain subnetwork is a two-stream residual network, including a first branch and a second branch. The first branch is used to process the first type of projection data, and the second branch is used to process the second type of projection data. The projection domain subnetwork adopts a channel attention mechanism to increase the completion weight of pixel regions in the first type of projection data and the second type of projection data whose gradient magnitude is higher than a preset threshold.
[0077] For example, a two-stream residual network can refer to a residual network structure containing two independent parameter branches. The first branch processes the first type of projection data, and the second branch processes the second type of projection data. Channel attention mechanism can refer to a feature recalibration mechanism that automatically learns the weights of each feature channel, allowing the network to focus on channels with greater information content. Image processing systems employ channel attention mechanism to increase the completion weights of pixel regions in the projection data whose gradient magnitudes exceed a preset threshold. Pixel regions with gradient magnitudes exceeding the preset threshold may correspond to the edges or minute structures of the measured object. By increasing the weights of these pixel regions during the completion process, the projection domain subnetwork can more accurately recover missing edge details and minute structural information, which helps improve the quality of the completed projection data, thereby improving the sharpness and structural fidelity of the final reconstructed image.
[0078] In some embodiments, each branch of the projection domain subnetwork can be composed of multiple stacked residual blocks, with a channel attention module embedded in each residual block. The channel attention module first performs global average pooling on the input feature map to obtain a global description vector for each channel; then, it generates channel weight vectors through two fully connected layers; finally, it multiplies the channel weights with the original feature map to achieve feature recalibration in the channel dimension. Through this residual block plus channel attention structure, the image processing system enables the network to adaptively enhance the response of feature channels containing structural information while suppressing noisy channels, thereby effectively improving the accuracy of restoring small structures in the projection.
[0079] In other embodiments, the projection domain subnetwork further introduces a spatial attention mechanism on top of the channel attention mechanism. The spatial attention mechanism is used to identify pixel locations with high gradient magnitudes in the projection data and generate a spatial weight map. The image processing system fuses the outputs of channel attention and spatial attention, enabling the network to simultaneously focus on which feature channels are important and which spatial locations are important. This embodiment, by jointly using channel attention and spatial attention, can more accurately locate areas requiring key completion, which is beneficial for preserving more anatomical details under high-sparse sampling conditions and improving the accuracy of projection completion.
[0080] In some optional embodiments, the image domain subnetwork is a two-stream residual network, including a third branch and a fourth branch. The third branch is used to process the portion of image data in the intermediate reconstructed image that corresponds to the first type of projection data, and the fourth branch is used to process the portion of image data in the intermediate reconstructed image that corresponds to the second type of projection data. The image domain subnetwork adopts a multi-scale convolutional structure to extract anatomical features of different scales from the intermediate reconstructed image.
[0081] For example, a two-stream residual network can refer to a residual network structure containing two independent parameter branches. The third branch processes the portion of image data corresponding to the first type of projection data in the intermediate reconstructed image, and the fourth branch processes the portion of image data corresponding to the second type of projection data. A multi-scale convolutional structure can refer to using convolutional kernels with different receptive fields (e.g., convolutional kernels of different sizes or dilated convolutions with different dilation rates) in the same network layer to extract features, simultaneously capturing image information from local details to global structure. Image processing systems employ multi-scale convolutional structures to extract anatomical features at different scales from intermediate reconstructed images, such as the fine texture of microvessels and the overall contour of organs. By processing dual-energy images separately and extracting multi-scale features, the image domain sub-network can be optimized for the characteristics of both low-energy and high-energy images. Simultaneously, it utilizes multi-scale information to enhance the recovery capability of anatomical structures of different sizes, which is beneficial for improving the detail fidelity and structural integrity of the target reconstructed image.
[0082] In some embodiments, the multi-scale convolutional structure in the image domain sub-network can employ parallel stacking of dilated convolutions with different dilation rates. For example, three dilated convolutional branches with dilation rates of 1, 2, and 4 can be used in the same layer. The output feature maps of each branch are concatenated along the channel dimension and then fused through a 1×1 convolution. Through this multi-scale dilated convolutional structure, the image processing system can capture multi-scale contextual information while maintaining feature map resolution, effectively expanding the receptive field. This embodiment is beneficial for simultaneously reconstructing large-scale organ boundaries and local fine textures, enhancing the image domain sub-network's ability to model complex anatomical structures.
[0083] In other embodiments, a cross-scale feature fusion mechanism is introduced between the two-stream branches of the image domain subnetwork. The image processing system performs bidirectional interaction between the feature maps of the third and fourth branches at various scale levels of multi-scale convolution. For example, detailed features from the low-energy branch are supplemented to the edge restoration process of the high-energy branch, while the overall structural information of the high-energy branch is passed to the low-energy branch to enhance soft tissue contrast. This embodiment, through cross-branch collaboration of dual-energy information, can effectively utilize the complementary advantages of low-energy and high-energy images, which is beneficial to improving the accuracy of material decomposition and the overall quality of the final reconstructed image.
[0084] In some optional embodiments, spectral-spatial joint alignment can also be performed between the projection domain subnetwork and the image domain subnetwork. For example, the image processing system aligns the projection domain feature map extracted by the projection domain subnetwork with the image domain feature map extracted by the image domain subnetwork in both spectral and spatial dimensions: in the spectral dimension, the feature map can be transformed to the frequency domain using a Fast Fourier Transform, and the similarity loss of the frequency domain features can be calculated; in the spatial dimension, the projection domain feature map can be weighted using anatomical structure attention weights (i.e., a spatial weight map generated based on the location of fine structures such as microvessels and trabeculae in the image domain feature map), making the projection domain subnetwork pay more attention to the fine structure regions already restored in the image domain during completion. The image processing system injects the prior knowledge of fine structures in the image domain into the projection domain through anatomical structure attention weights to guide projection completion. This spectral-spatial joint alignment works synergistically with the cross-domain feature alignment loss function to further improve the effect of cross-domain collaborative enhancement.
[0085] Figure 3A schematic diagram of the internal processing flow of the integrated domain image reconstruction network in this embodiment is shown. First, projection data acquired by the ring-shaped dual-source dual-energy imaging system under sparse viewing angles is input to an edge-aware interpolation preprocessing module. This module performs edge-aware interpolation on the sparse projection data to generate initial projection data. Next, the initial projection data is input to the projection domain sub-network. The projection domain sub-network employs a two-stream residual structure, completing the sparse and missing projections through internal residual connections, and outputting the target projection data. Then, the target projection data is input to an analytical reconstruction operator, which is based on the inverse Radon transform and fuses adaptive projection weighting factors and adaptive total variation regularization terms to reconstruct the projection data into an intermediate reconstructed image. Finally, the intermediate reconstructed image is input to the image domain sub-network. The image domain sub-network also employs a two-stream residual structure, sharpening edges and suppressing noise in the intermediate reconstructed image through internal residual connections, and outputting the final target reconstructed image.
[0086] Figure 4 This diagram illustrates a two-stream residual network structure of the integrated domain image reconstruction network in this embodiment. Different colors and shapes distinguish different functional modules. Orange rectangles represent one type of convolutional layer, green rectangles represent residual blocks, yellow rectangles represent another type of convolutional layer, and purple rectangles represent deconvolutional layers. Icon I represents the multi-scale convolutional structure in the image domain sub-network, icon P represents the channel attention mechanism in the projection domain sub-network, and gray dots represent feature concatenation. The input data first passes through the multi-scale convolutional structure in the image domain sub-network to extract image domain anatomical features from different scales. Next, the extracted features enter the channel attention mechanism in the projection domain sub-network. Under the channel attention mechanism, the data is divided into two branches: the left branch processes the second type of projection data, and the right branch processes the first type of projection data. Both branches undergo feature concatenation through multiple deconvolutional layers, residual blocks, and convolutional layers to extract high-energy and low-energy channels. Finally, the features from the two branches are fused through feature concatenation and output as the target reconstructed image or target projection data through the output layer.
[0087] In some optional embodiments, the first type of projection data and the second type of projection data are collected alternately at different angular dimensions, including: the image processing system can extract the sampling angle of the first type of projection data from the full-view projection data at preset intervals, and extract the sampling angle of the second type of projection data at an angle position complementary to the sampling angle of the first type of projection data, so that the first type of projection data and the second type of projection data form spectral complementarity; wherein, the spectral complementarity characterization utilizes the spatial domain structure consistency and energy domain information redundancy of the first type of projection data and the second type of projection data to complement information; the full-view projection data characterizes the projection data continuously collected within the full scanning angle range, and the first type of projection data and the second type of projection data are subsets obtained by sparse sampling from the full-view projection data.
[0088] For example, full-view projection data can refer to projection data continuously acquired within the complete scanning angle range, such as acquiring a projection at regular angular intervals, covering 360 degrees. The preset interval can refer to a fixed step size used for sparse sampling, such as extracting one viewpoint as a sampling point every few full-view data points. The complementary angular position can refer to an angular position offset from the sampling viewpoint of the first type of projection data by half a sampling interval, allowing low-energy and high-energy projections to alternate in the angular dimension. The image processing system extracts the sampling viewpoints of the first type of projection data from the full-view projection data at preset intervals, and extracts the sampling viewpoints of the second type of projection data at angular positions complementary to the sampling viewpoints of the first type of projection data, thus making the first and second types of projection data spectrally complementary. Spectral complementarity characterization utilizes the spatial domain structural consistency and energy domain information redundancy of low-energy and high-energy projection data for information complementarity; that is, since the anatomical structure of the same measured object has the same spatial distribution in both low-energy and high-energy projections, the sparsely sampled dual-energy data can complement each other angularly. By using the above-mentioned alternating acquisition method, the image processing system can make low-energy projection and high-energy projection each cover different angular subsets without increasing the total number of samples. After merging, more comprehensive angular information can be obtained, which helps to avoid the loss of key information caused by sparse acquisition of a single energy spectrum.
[0089] In some embodiments, the image processing system can set a preset interval of 8 full-view angles. For the first type of projection data, the image processing system extracts a sampling viewpoint every 8 full-view angles, that is, extracts the projections corresponding to the 1st, 9th, 17th, ... Nth full-view angles. For the second type of projection data, the image processing system extracts at positions complementary to the sampling viewpoints of the first type of projection data, that is, starting from the 5th full-view angle, extracting a sampling viewpoint every 8 full-view angles. In this way, low-energy projection and high-energy projection alternate on the angle axis, and the total number of sampling viewpoints is approximately 1 / 4 of the full-view angles. This embodiment utilizes the complementary angle information formed by alternating dual-energy sampling, so that sufficient target structure information can still be retained even with sparse sampling.
[0090] In other embodiments, the image processing system can dynamically adjust the preset interval according to clinical requirements for radiation dose and image quality. When a further reduction in radiation dose is needed, the image processing system increases the preset interval, for example, from 8 to 10, corresponding to extracting one sampling viewpoint for every 10 full-view angles, thus reducing the total number of samples. Simultaneously, the image processing system adjusts the complementary positions of the second type of projection data accordingly, maintaining an alternating distribution of low-energy and high-energy projection angles. This embodiment allows for flexible adjustment of sparsity according to different scanning protocols, which is beneficial for balancing radiation dose and reconstruction quality.
[0091] Figure 5A schematic diagram of a dual-energy CT scanning configuration with spectral complementarity is shown in an embodiment of this application. The blue arc represents the scanning trajectory of the first energy level ray, and the pink arc represents the scanning trajectory of the second energy level ray. Alternating short lines on the blue and pink arcs represent sparse sampling positions for projection measurements, i.e., projection data is acquired only at certain angles. The scanning trajectories of the first and second energy levels are alternately distributed in the angular dimension, allowing low-energy and high-energy projections to be acquired at complementary angular positions. This utilizes the spatial domain structural consistency and energy domain information redundancy of the dual-energy images to achieve spectral complementarity. This scanning configuration helps avoid the loss of critical information caused by sparse acquisition of a single energy spectrum.
[0092] In some optional embodiments, the first type of projection data and the second type of projection data are collected alternately in the angular dimension, satisfying the following relationship: the sampling angle of the i-th first type of projection data corresponds to the (i-1)-th projection data in the full-view projection data. The n+1) viewpoints; the sampling viewpoint of the i-th type of second-class projection data corresponds to the (i-1 / 2)-th viewpoint in the full-view projection data. (n+1) viewpoints; where i is an integer greater than or equal to 1, and n is the sampling interval parameter and n is an integer greater than or equal to 1.
[0093] For example, the sampling interval parameter n is an integer greater than or equal to 1, representing the step size of sparse sampling. For instance, n=8 means extracting one sampling viewpoint every 8 full-view angles. i is the index of the sampling viewpoint, i=1,2,3,…,k, where k is the total number of samples for each energy level. The sampling viewpoint of the i-th first type of projection data corresponds to the ((i-1)×n+1)-th viewpoint in the full-view projection data, that is, starting from the 1st viewpoint, one viewpoint is extracted every n views. The sampling viewpoint of the i-th second type of projection data corresponds to the ((i-1 / 2)×n+1)-th viewpoint in the full-view projection data, that is, starting from the ((1-0.5)×n+1)-th viewpoint, one viewpoint is extracted every n views, thus offsetting the sampling viewpoint of the first type of projection data by n / 2 angular positions. In this way, the image processing system can accurately generate a sparsely sampled viewpoint sequence, so that the first type of projection data and the second type of projection data are alternately distributed in the angular dimension, forming complementarity. The sampling scheme of this application embodiment helps to ensure the uniformity and complementarity of low-energy projection and high-energy projection in angular space, providing a structured input for subsequent projection completion using the spatial consistency and energy redundancy of dual-energy data, and facilitating flexible adjustment of parameters according to different sampling densities.
[0094] For example, the image processing system determines the sampling angle of the first type of projection data and the second type of projection data according to the following formulas (6) and (7):
[0095] Formula (6);
[0096] Formula (7);
[0097] in, Indicates the first Sampled values of the first type of projection data, This represents the full-view projection data of the low-energy channel. Indicates the first Sampled values of the second type of projection data, This represents the full-view projection data of the high-energy channel. The sampling interval parameter, The total number of samples for each energy level. When hour, ;when hour, .
[0098] In some embodiments, the image processing system sets the sampling interval parameter n=8, and the total number of viewing angles is 720. That is, the sampling viewing angles for the first type of projection data are: the 1st viewing angle when i=1, the 9th viewing angle when i=2, the 17th viewing angle when i=3, ..., the 713th viewing angle when i=90. The sampling viewing angles for the second type of projection data are: the 5th viewing angle when i=1 ((1-0.5)×8+1)=5th viewing angle, the 13th viewing angle when i=2, ..., the 717th viewing angle when i=90. The image processing system controls the dual light source to alternately acquire data according to this viewing angle sequence, obtaining 90 low-energy viewing angles and 90 high-energy viewing angles, with a total sampling number of 180, only 1 / 4 of the total number of viewing angles. This embodiment is beneficial for accurately describing the mathematical relationship of alternating acquisition and for generating control commands programmatically in actual scanning, reducing system complexity.
[0099] In other embodiments, the image processing system dynamically adjusts the value of the sampling interval parameter n according to clinical requirements for radiation dose. When a further dose reduction is needed, the image processing system sets n=10. In this case, the sampling angles for the first type of projection data are the 1st, 11th, 21st, ...th, and the sampling angles for the second type of projection data are the 6th, 16th, 26th, ...th, reducing the total number of samples to 144 (72 low-energy and 72 high-energy), thus reducing the radiation dose accordingly. When image quality needs to be improved, n=4 is set, increasing the sampling density and raising the total number of samples to 360 (180 low-energy and 180 high-energy). The image processing system automatically calculates the sampling angle sequence for each energy level based on the n value, eliminating the need for manual specification. This embodiment adapts to different sparsity sampling requirements through a unified mathematical formula, which is beneficial for flexibly balancing scan dose and reconstruction quality.
[0100] Among them, dual-energy sparse view projection data With full-view projection data The relationship between them is defined by the following formulas (8) and (9):
[0101] ; Formula (8);
[0102] ; Formula (9);
[0103] When the sampling angle is 90, n When the sampling angle is 72, n .
[0104] In some optional embodiments, before inputting the first type of projection data and the second type of projection data into the projection domain subnetwork, the method further includes: the image processing system can preprocess the first type of projection data and the second type of projection data to generate initial projection data; the preprocessing includes edge-aware interpolation processing to preserve the contour information of the measured object.
[0105] For example, the initial projection data can refer to the projection dataset obtained after interpolating the original acquired sparse projection data. Edge-aware interpolation can be an interpolation method that considers the gradient information of the projection data. Compared with ordinary linear interpolation, edge-aware interpolation can adjust the interpolation weights according to the gradient direction of the pixel neighborhood, maintaining edge sharpness while smoothly transitioning between sparse projections. Before inputting the first and second types of projection data into the projection domain subnetwork, the image processing system performs edge-aware interpolation preprocessing on the first and second types of projection data to generate initial projection data. Through this preprocessing, the image processing system can provide better initial input to the projection domain subnetwork, preserving the contour information of the measured object, which is beneficial to improving the quality of subsequent completion and reconstruction.
[0106] In some embodiments, the image processing system first performs linear interpolation on the sparse projection data to obtain preliminary interpolation results, and then calculates the gradient map of the projection data to identify edge locations. For pixels located near edges, the image processing system employs edge-preserving interpolation methods, such as anisotropic interpolation, increasing the weights along the edge direction and decreasing the weights across the edge direction. The image processing system then weights and fuses the edge-preserving interpolation results with the linear interpolation results according to the gradient magnitude to obtain the initial projection data. This embodiment, through edge-aware processing, can effectively preserve the contour information of the measured object, avoid edge blurring caused by interpolation, and is beneficial for providing a more accurate initial projection structure for the projection domain subnetwork.
[0107] In other embodiments, the image processing system can use a deep learning-based edge-aware interpolation network for preprocessing. This interpolation network takes sparse projection data as input and outputs initial projection data. The edge-aware interpolation network includes an edge detection branch and an interpolation branch. The edge detection branch extracts gradient information from the projection, and the interpolation branch guides the generation of interpolation weights based on the gradient information. During the training phase, the image processing system jointly optimizes this interpolation network with the subsequent main network, enabling the preprocessing process to adaptively learn the optimal edge-preserving strategy. This embodiment can obtain higher quality initial projections than traditional interpolation methods, and is particularly suitable for highly sparse sampling scenarios. It helps reduce the completion burden of the projection domain sub-network, improving the stability and efficiency of the entire reconstruction process.
[0108] See Figure 6 According to another aspect of the embodiments of this application, a training method for a neural network model for fast reconstruction of a ring-shaped dual-source dual-energy image is also provided. The neural network model includes at least: a projection domain sub-network, an analytical reconstruction operator, and an image domain sub-network, including the following steps: Step S601, obtaining a training dataset, the training dataset including: training projection data collected under sparse perspectives as training input, and ground truth image data as training labels, the ground truth image data being an image reconstructed from full-view projection data, sparse perspectives representing that the number of sampling perspectives is less than the number of perspectives used when collecting full-view images; Step S602, based on the training dataset, using a loss function to jointly train the projection domain sub-network, the analytical reconstruction operator, and the image domain sub-network. The trained projection domain subnetwork is used to sparsely sample and complete the missing projection data of the first and second types of input projection data, and output the target projection data. The first type of projection data is used to represent the projection signal collected after the ray of the first energy level penetrates the object under test, and the second type of projection data is used to represent the projection signal collected after the ray of the second energy level penetrates the object under test, where the second energy level is higher than the first energy level. The analytical reconstruction operator is used to transform the target projection data from the projection domain to the image domain and output the intermediate reconstructed image. The image domain subnetwork is used to sharpen the edges and suppress noise in the intermediate reconstructed image and output the target reconstructed image.
[0109] For example, the training dataset can refer to the data set used for training a neural network model, including training input and training labels. The training input is the training projection data collected under sparse perspectives, i.e., low-energy and high-energy projections obtained by sparse perspective sampling, and the training labels are the ground truth image data reconstructed from full-view projection data. Sparse perspectives represent that the number of sampling perspectives is less than the number of perspectives used in full-view acquisition. Joint training can refer to treating the projection domain subnetwork, the analytical reconstruction operator, and the image domain subnetwork as a whole, using the same loss function to simultaneously optimize the parameters of the three components. The image processing system uses the training dataset and employs the loss function to jointly train the projection domain subnetwork, the analytical reconstruction operator, and the image domain subnetwork. Through joint training, the projection domain subnetwork, the analytical reconstruction operator, and the image domain subnetwork can adapt to each other, enabling the projection domain subnetwork to learn to fill in missing projections from sparse projections, the analytical reconstruction operator to learn to quickly transform the filled projection data into an intermediate image, and the image domain subnetwork to learn to sharpen edges and suppress noise in the intermediate image, ultimately outputting a high-quality target reconstruction image. This end-to-end joint training helps avoid local optima problems that may result from phased training, thereby improving overall reconstruction performance.
[0110] In some embodiments, the image processing system can generate a training dataset through simulation. The system acquires a large amount of dual-energy projection data across all viewpoints and the corresponding reconstructed images as ground truth. Then, the system extracts sparse projections from the projection data according to a preset sparse sampling scheme (e.g., extracting one low-energy viewpoint for every eight viewpoints and one high-energy viewpoint at complementary positions) as training input. The system can use this paired data to jointly train the three components. This embodiment can generate large-scale, diverse training data, which is beneficial for improving the model's generalization ability and reconstruction accuracy.
[0111] In other embodiments, the image processing system can employ a strategy combining staged pre-training and joint fine-tuning. The system first trains the projection domain subnetwork separately using paired data, enabling it to learn to complete missing projections. Then, it trains the image domain subnetwork while fixing the parameters of the projection domain subnetwork. Finally, it performs joint fine-tuning of the projection domain subnetwork, the analytical reconstruction operator, and the image domain subnetwork together. This embodiment uses staged pre-training to obtain better initial parameters for each component, which is beneficial for stable convergence and is particularly suitable for scenarios with limited training data or deep network structures. After training, the image processing system can deploy the model in a clinical environment to rapidly reconstruct sparse bi-energy projections acquired in real time.
[0112] In some optional embodiments, the loss function is calculated by weighting a pixel loss sub-function, a structural similarity loss sub-function, a dual-energy spectrum difference loss sub-function, and a cross-domain feature alignment loss sub-function, wherein: the pixel loss sub-function is used to constrain the pixel-level differences between the target reconstructed image and the ground truth image data; the structural similarity loss sub-function is used to constrain the structural similarity between the target reconstructed image and the ground truth image data; the dual-energy spectrum difference loss sub-function is used to constrain the image differences between different channel parts in the target reconstructed image to be consistent with the image differences between corresponding different channel parts in the ground truth image data; and the cross-domain feature alignment loss sub-function is used to constrain the consistency between the projection domain feature map extracted by the projection domain sub-network and the image domain feature map extracted by the image domain sub-network.
[0113] For example, the pixel loss function is used to calculate the grayscale difference at each pixel position between the reconstructed image and the ground truth image data, reflecting the pixel-level approximation between the two images. The structural similarity loss function measures the similarity between the reconstructed image and the ground truth image data in three dimensions: brightness, contrast, and structure; the closer the value is to 1, the more similar the two images are. The dual-energy spectrum difference loss function constrains the pixel difference between the low-energy and high-energy channels in the reconstructed image to be consistent with the pixel difference between the low-energy and high-energy channels in the ground truth image data, thereby maintaining the physical correlation of the dual-energy spectrum. The cross-domain feature alignment loss function constrains the consistency between the projection domain feature map output by the projection domain sub-network and the image domain feature map output by the image domain sub-network, promoting information alignment between the projection domain and the image domain. The image processing system trains the network using a weighted combination of the above four loss functions, which is beneficial for simultaneously obtaining high pixel accuracy, high structural fidelity, accurate dual-energy spectrum relationships, and cross-domain feature consistency, thereby improving the overall reconstruction quality.
[0114] In some embodiments, the image processing system can set the weights of the four loss sub-functions to equal values, making each loss equally important during training. The image processing system calculates each loss value in each iteration, sums them according to their weights to obtain the total loss, and then backpropagates to update the network parameters. This embodiment simplifies the weight tuning process, is suitable for tasks with similar loss magnitudes, and can balance different optimization objectives.
[0115] In other embodiments, the image processing system dynamically adjusts the weights of each loss function based on the validation set performance. When the structural similarity index on the validation set is low, the image processing system increases the weight of the structural similarity loss function; when the dual-spectral difference loss is high, the image processing system increases the weight of the dual-spectral difference loss function. By adaptively adjusting the weights, the image processing system strengthens constraints on weak points in the network during training. This embodiment is beneficial for optimizing network performance for specific clinical tasks such as soft tissue imaging or bone imaging, and improving the performance of reconstructed images in the target application.
[0116] In some optional embodiments, the calculation of the pixel loss sub-function includes: performing channel pixel loss calculation operations on the first channel and the second channel respectively, including: calculating the absolute difference between the gray value of each pixel position in the current channel part of the target reconstructed image and the gray value of the corresponding pixel position in the corresponding channel part of the ground image data; summing the absolute differences of all pixel positions to obtain the channel pixel loss of the current channel; adding the channel pixel loss of the first channel to the channel pixel loss of the second channel, and then dividing the sum by the total number of pixels to obtain the average absolute error as the pixel loss sub-function; wherein, the first channel corresponds to the first type of projection data, and the second channel corresponds to the second type of projection data.
[0117] For example, the pixel loss sub-function is used to measure the pixel-level difference between the reconstructed image and the ground truth image data. The image processing system performs channel pixel loss calculations for the first channel (corresponding to the first type of projection data, i.e., the low-energy channel) and the second channel (corresponding to the second type of projection data, i.e., the high-energy channel). For each channel, the image processing system calculates the absolute difference between the gray value of each pixel position in the current channel portion of the reconstructed image and the gray value of the corresponding pixel position in the corresponding channel portion of the ground truth image data. Then, it sums the absolute differences of all pixel positions to obtain the channel pixel loss for that channel. The image processing system adds the channel pixel loss of the first channel to the channel pixel loss of the second channel, and then divides the sum by the total number of pixels, i.e., the total number of pixels in the first and second channels, to obtain the mean absolute error as the pixel loss sub-function. This pixel loss sub-function directly supervises the pixel-by-pixel error between the reconstructed image output by the network and the ground truth image, which helps to ensure the basic reconstruction accuracy and enables the network to learn accurate pixel gray-level distributions.
[0118] For example, the pixel loss function uses the mean absolute error form, and the calculation formula is as follows:
[0119] Formula (10);
[0120] in, It can represent a pixel loss function; It can represent an energy channel index. This indicates the first channel, corresponding to the first type of projection data, i.e., the low-energy channel. This indicates the second channel, corresponding to the second type of projection data, namely the high-energy channel; It can represent the total number of pixels in a single image; It can represent pixel coordinates; It can represent the pixel position of the current channel portion in the target reconstructed image. grayscale value; This represents the grayscale value of the corresponding pixel location in the corresponding channel portion of the true image data. The image processing system first calculates the absolute difference between each pixel location in the low-energy channel and the high-energy channel separately, sums the values for all pixel locations in each channel, then adds the sums for the two channels together, and finally divides by the total number of pixels. The mean absolute error is obtained as the pixel loss sub-function.
[0121] In some embodiments, the image processing system can use mean absolute error (MAE) as the pixel loss function, assigning equal weight to the error at all pixel locations. This MAE calculation method is insensitive to outliers and can stably guide network training. During training, the image processing system weights and combines the pixel loss function with other loss functions to jointly optimize the network parameters. This embodiment is simple to implement, applicable to most reconstruction tasks, and can effectively reduce pixel-level errors in reconstructed images.
[0122] In some embodiments, the image processing system can use mean squared error (MSE) instead of mean absolute error (MAE) to calculate the pixel loss function. The system calculates the square of the gray-level difference at each pixel location, sums the squares over all pixel locations, and divides the sum by the total number of pixels to obtain the MSE as the pixel loss function. MSE penalizes large errors more severely, prompting the network to reduce areas of significant deviation more quickly. This embodiment is suitable for tasks sensitive to large local errors in the reconstructed image and helps improve the peak signal-to-noise ratio (PSNR). The image processing system can choose either MSE or MAE as the form of the pixel loss function based on specific application requirements.
[0123] In some optional embodiments, the calculation of the structural similarity loss sub-function includes: performing channel structural similarity loss calculation operations on the first channel and the second channel respectively, including: calculating the product of brightness similarity, contrast similarity and structural similarity between the target reconstructed image of the current channel and the corresponding channel part in the ground truth image data to obtain the structural similarity index of the current channel; subtracting the structural similarity index from 1 to obtain the channel structural similarity loss of the current channel; adding the channel structural similarity loss of the first channel and the channel structural similarity loss of the second channel and dividing by 2 to obtain the structural similarity loss sub-function.
[0124] For example, the structural similarity index is used to measure the similarity between two images in three dimensions: brightness, contrast, and structure. Its value ranges from 0 to 1, with values closer to 1 indicating greater similarity. The image processing system calculates the channel structural similarity loss for the first and second channels respectively. For each channel, the image processing system calculates the product of the brightness similarity, contrast similarity, and structural similarity between the target reconstructed image and the corresponding channel in the ground truth image data for that channel, obtaining the structural similarity index for that channel. Then, the image processing system subtracts this structural similarity index from 1 to obtain the channel structural similarity loss for the current channel. The image processing system adds the channel structural similarity loss of the first channel to the channel structural similarity loss of the second channel and divides by 2 to obtain the structural similarity loss sub-function. By minimizing this structural similarity loss sub-function, the image processing system can guide the network to recover the overall structure, contrast, and brightness information of the image, which is beneficial for improving the visual quality and structural fidelity of the reconstructed image and reducing the insufficient constraint of local structure due to pixel loss.
[0125] For example, the formula for calculating the structural similarity loss function is as follows:
[0126] Formula (11);
[0127] in, Represents the structural similarity loss function; Indicates the first channel. Indicates the second channel; This represents the structural similarity index between the reconstructed image of the current channel and the corresponding channel in the ground truth image data. This structural similarity index is a product of brightness similarity, contrast similarity, and structural similarity, and its value ranges from 0 to 1. The image processing system calculates the structural similarity index for each channel separately. The value is then added together with the results of the two channels and divided by 2 to obtain the structural similarity loss function.
[0128] In some embodiments, when calculating the structural similarity index, the image processing system can use a Gaussian window to perform a weighted average of image patches to estimate the mean, variance, and covariance of local regions. The image processing system calculates the structural similarity index for all image patches in the entire image and then takes the average value as the structural similarity index for that channel. This embodiment can more accurately evaluate local structural similarity, which is beneficial for the network to recover the texture and edge details of the image.
[0129] In other embodiments, the image processing system can use multi-scale structural similarity instead of single-scale structural similarity. The image processing system calculates structural similarity indices at different downsampling scales, and then performs a weighted average of the structural similarity index values at each scale to obtain the multi-scale structural similarity index. Multi-scale structural similarity can simultaneously evaluate the structural similarity of images at multiple resolutions, and has the ability to constrain both global structure and local details. This embodiment is beneficial for the network to simultaneously maintain a large range of anatomical contours and minute tissue textures, and is particularly suitable for dual-energy CT reconstruction tasks that require the simultaneous recovery of macroscopic and microscopic structures.
[0130] In some optional embodiments, the calculation method of the dual-energy spectrum difference loss sub-function includes: calculating the pixel difference between the first channel portion and the second channel portion in the target reconstructed image to obtain a reconstructed dual-energy difference map; calculating the pixel difference between the first channel portion and the second channel portion in the ground truth image data to obtain a ground truth dual-energy difference map; calculating the absolute difference between the reconstructed dual-energy difference map and the ground truth dual-energy difference map for each pixel position, summing the absolute differences of all pixel positions, dividing the summation result by the total number of pixels, and obtaining the average absolute error as the dual-energy spectrum difference loss sub-function.
[0131] For example, the reconstructed dual-energy difference map can refer to the image formed by the gray-level difference between the first channel portion and the second channel portion at the same pixel position in the target reconstructed image, reflecting the difference distribution between the dual-energy images predicted by the network. The ground-value dual-energy difference map can refer to the image formed by the pixel difference between the corresponding first channel portion and the second channel portion in the ground-value image data, reflecting the difference distribution between the dual-energy images in the real physical process. The image processing system calculates the pixel difference between the first channel portion and the second channel portion in the target reconstructed image to obtain the reconstructed dual-energy difference map, and calculates the pixel difference between the first channel portion and the second channel portion in the ground-value image data to obtain the ground-value dual-energy difference map. Then, the image processing system calculates the absolute difference at each pixel position in the reconstructed dual-energy difference map and the ground-value dual-energy difference map, sums the absolute differences at all pixel positions, and divides the sum by the total number of pixels to obtain the mean absolute error as the dual-energy spectrum difference loss sub-function. This dual-energy spectrum difference loss function can directly constrain the predicted dual-energy difference to be consistent with the actual dual-energy difference, which is beneficial for the network to learn the correct dual-energy spectrum physical relationship, improve the accuracy of material decomposition and the synergistic utilization effect of dual-energy information.
[0132] For example, the formula for calculating the difference loss function in dual-energy spectroscopy is as follows:
[0133] Formula (12);
[0134] in, This represents the difference loss function between the two energy spectra; It can represent the pixel difference between the first channel portion and the second channel portion at the same pixel position in the target reconstructed image, forming a reconstruction dual-energy difference map; This can represent the pixel difference between the first and second channels of the ground truth image data at the same pixel location, forming a ground truth dual-energy difference map. The image processing system calculates the absolute difference between the two difference values at each pixel location, sums the absolute differences for all pixel locations, and then divides by the total number of pixels. The mean absolute error is obtained as the dual-energy spectrum difference loss function.
[0135] In some embodiments, the image processing system can calculate the dual-energy spectral difference loss using the mean absolute error (MAE) form, assigning equal weight to the error at all pixel locations. This MAE form can stably guide the network to reduce the overall bias of the dual-energy difference map, making it suitable for most dual-energy CT reconstruction tasks. The image processing system weightedly combines this loss function with other loss functions to jointly optimize the network parameters. This embodiment is simple to implement and can effectively constrain the consistency of difference between low-energy and high-energy images, improving the physical fidelity of the dual-energy spectrum.
[0136] In other embodiments, the image processing system may use mean squared error (MSE) instead of mean absolute error (MAE) to calculate the dual-energy spectral difference loss. The system calculates the square of the reconstructed difference versus the ground truth difference at each pixel location, sums the squared values across all pixel locations, and divides the sum by the total number of pixels to obtain the MSE as a sub-function of the dual-energy spectral difference loss. The MSE penalizes large deviations more severely, prompting the network to fit the dual-energy difference values more accurately. This embodiment is suitable for tasks requiring high accuracy in dual-energy difference calculations, such as quantitative substance decomposition and electron density mapping, and helps improve the accuracy of dose calculation. The image processing system can select a suitable error form based on specific application requirements.
[0137] In some optional embodiments, the calculation of the cross-domain feature alignment loss sub-function includes: performing channel feature alignment loss calculation operations on the first channel and the second channel respectively, including: obtaining the projection domain feature map of the current channel output by the projection domain sub-network; obtaining the image domain feature map of the current channel output by the image domain sub-network; obtaining the cross-domain feature weights of the current channel, where the cross-domain feature weights are learnable parameters used to map the image domain feature map to the feature space of the projection domain feature map; multiplying the cross-domain feature weights by the image domain feature map to obtain a weighted image domain feature map; calculating the absolute difference of each pixel position in the projection domain feature map and the weighted image domain feature map, summing the absolute differences of all pixel positions to obtain the channel feature alignment loss of the current channel; adding the channel feature alignment loss of the first channel to the channel feature alignment loss of the second channel, and then dividing the sum by the total number of pixels to obtain the cross-domain feature alignment loss sub-function.
[0138] For example, the projection domain feature map can refer to the feature map extracted by the projection domain sub-network during the processing of projection data, and may contain signal distribution and structural information in the projection domain. The image domain feature map can refer to the feature map extracted by the image domain sub-network during the processing of intermediate reconstructed images, and may contain anatomical structural information in the image domain. The cross-domain feature weights are learnable parameters, such as a scalar or vector, used to map the numerical range of the image domain feature map to the feature space of the projection domain feature map, making the projection domain feature map and the image domain feature map comparable. The image processing system performs channel feature alignment loss calculation for the first channel and the second channel respectively. For each channel, the image processing system obtains the projection domain feature map output by the projection domain sub-network, obtains the image domain feature map output by the image domain sub-network, obtains the cross-domain feature weights of the current channel, multiplies the cross-domain feature weights by the image domain feature map to obtain the weighted image domain feature map, then calculates the absolute difference between the projection domain feature map and the weighted image domain feature map for each pixel position, and sums the absolute differences of all pixel positions to obtain the channel feature alignment loss for the current channel. The image processing system adds the channel feature alignment loss of the first channel to the channel feature alignment loss of the second channel, and then divides the sum by the total number of pixels to obtain the cross-domain feature alignment loss sub-function. The cross-domain feature alignment loss sub-function constrains the consistency between the projection domain features and the image domain features, which is beneficial to promoting feature alignment between the projection domain sub-network and the image domain sub-network. This allows the anatomical structure information recovered from the image domain to be utilized during projection domain completion, forming an effective cross-domain collaborative enhancement.
[0139] For example, the formula for calculating the cross-domain feature alignment loss function is as follows:
[0140] Formula (13);
[0141] in, This represents the cross-domain feature alignment loss function; Indicates the first channel; Indicates the second channel; This represents the projection domain feature map of the current channel output by the projection domain subnetwork, where... These are the coordinates on the feature map of the projection domain; This can represent the image domain feature map of the current channel output by the image domain subnetwork, where, Represents the coordinates on the image domain feature map; This represents the cross-domain feature weight of the current channel. It is a learnable parameter used to map the image domain feature map to the feature space of the projected domain feature map. This represents the total number of pixels in the feature map. The image processing system multiplies the cross-domain feature weights by the image domain feature map to obtain a weighted image domain feature map. Then, it calculates the absolute difference between the projected domain feature map and the weighted image domain feature map for each pixel position. The absolute differences for all pixel positions are summed to obtain the channel feature alignment loss for the current channel. The channel feature alignment losses for the two channels are added together and then divided by the total number of pixels. This yields the cross-domain feature alignment loss function.
[0142] In some embodiments, the image processing system sets the cross-domain feature weights as a learnable scalar parameter, which is optimized along with other network parameters during training. For each channel, the image processing system multiplies the image domain feature map by the cross-domain feature weights (i.e., the scalar parameter) and calculates the absolute difference with the projected domain feature map. The above-described embodiment of feature alignment via scalars achieves feature alignment through simple scaling, has a small number of parameters, is easy to train, and can effectively adjust image domain features to a numerical range close to that of projected domain features.
[0143] In other embodiments, the image processing system sets the cross-domain feature weights as a learnable convolutional kernel, mapping the channel count, size, or numerical distribution of the image domain feature map to the space of the projected domain feature map through convolution operations. The image processing system inputs the image domain feature map into the convolutional layer corresponding to the cross-domain feature weights to obtain the transformed feature map, and then calculates the feature alignment loss with the projected domain feature map. The above embodiments that achieve feature alignment through convolutional kernels can provide stronger feature mapping capabilities, handle situations where there are significant differences between the projected domain feature map and the image domain feature map, and are beneficial for establishing more accurate alignment relationships between complex cross-domain features, further improving the effect of cross-domain collaborative enhancement.
[0144] See Figure 7According to another aspect of the embodiments of this application, a fast reconstruction device for ring-shaped dual-source dual-energy images based on neural networks is also provided, comprising: a projection data acquisition unit, configured to acquire a first type of projection data and a second type of projection data acquired by a ring-shaped dual-source dual-energy imaging system under sparse viewing angles, wherein sparse viewing angles characterize the number of sampling viewing angles being less than the number of viewing angles used in full-view acquisition; the first type of projection data is used to characterize the projection signal acquired after a ray of a first energy level penetrates the object under test, and the second type of projection data is used to characterize the projection signal acquired after a ray of a second energy level penetrates the object under test, wherein the second energy level is higher than the first energy level; and a projection data processing unit, configured to process the first type of projection data... The second type of projection data is input into the projection domain subnetwork of the neural network model. The projection domain subnetwork completes the sparsely sampled and missing projection data, and outputs the target projection data. The target projection data includes the completed first type of projection data and the completed second type of projection data. The reconstruction processing unit is used to reconstruct the target projection data using the analytical reconstruction operator in the neural network model to obtain the intermediate reconstructed image. The analytical reconstruction operator is used to transform the target projection data from the projection domain to the image domain. The image processing unit is used to input the intermediate reconstructed image into the image domain subnetwork of the neural network model. The image domain subnetwork performs edge sharpening and noise suppression on the intermediate reconstructed image and outputs the target reconstructed image.
[0145] Optionally, the reconstruction processing unit further includes: a first acquisition subunit, used to acquire the completed projection data term and the cross-domain feature fusion term, wherein the completed projection data term is equal to the target projection data output by the projection domain sub-network, and the cross-domain feature fusion term is equal to the product of the cross-domain feature weights and the structural prior projection reference data fed back from the image domain sub-network; the structural prior projection reference data represents the projection data obtained after mapping the image optimized by the image domain sub-network back to the projection domain through Radon transform; a summation subunit, used to sum the completed projection data term and the cross-domain feature fusion term to obtain the target data term; and a second acquisition subunit, used to acquire the transformation operator, the adaptive projection weighting factor, and the adaptive total... The system comprises a variation regularization coefficient and a total variation regularization term. The transformation operator is the inverse Radon transform operator, used as the transformation coefficient to transform the target projection data from the projection domain to the image domain. An adaptive projection weighting factor is used to dynamically adjust the projection weights of the target projection data based on its sparsity. An adaptive total variation regularization coefficient is used to adjust the regularization intensity of the intermediate reconstructed image based on the sparsity of the target projection data. The total variation regularization term constrains the edge smoothness of the intermediate reconstructed image. An operator determination subunit is used to determine the analytical reconstruction operator based on the target data term, the transformation operator, the adaptive projection weighting factor, the adaptive total variation regularization coefficient, and the total variation regularization term.
[0146] Optionally, the projection domain subnetwork and the image domain subnetwork constitute a cross-domain collaborative enhancement closed-loop architecture; the anatomical structure features extracted from the intermediate reconstructed image by the image domain subnetwork are fed back to the projection domain subnetwork to guide the projection domain subnetwork to complete the projection data missing by sparse sampling, and the anatomical structure features represent the organ structure information extracted from the image domain corresponding to the intermediate reconstructed image and fed back to the projection domain; the target projection data output by the projection domain subnetwork is reconstructed by the analytical reconstruction operator and then input into the image domain subnetwork.
[0147] Optionally, the second acquisition subunit includes: a first detection module, configured to determine that the adaptive projection weighting factor is greater than the standard weight value when the sparsity of the target projection data is less than or equal to a preset threshold, wherein the standard weight value is 1; and a second detection module, configured to determine that the adaptive projection weighting factor is equal to the standard weight value when the sparsity of the target projection data is greater than the preset threshold.
[0148] Optionally, the first detection module is specifically used to: determine the adaptive projection weighting factor equal to a preset threshold when the sparsity of the detected target projection data is less than or equal to a preset threshold. ,in, A preset coefficient greater than 0. The sparsity of the target projection data.
[0149] Optionally, the adaptive total variation regularization coefficient increases with the increase of the sparsity of the target projection data; the second acquisition subunit further includes: a sum of squares calculation module, used to calculate the sum of squares of the differences between adjacent pixels in the intermediate reconstructed image; a square root module, used to take the square root of the sum of squares to obtain the square root result; and a regularization term determination module, used to use the square root result as the total variation regularization term.
[0150] Optionally, the projection domain subnetwork is a two-stream residual network, including a first branch and a second branch. The first branch is used to process the first type of projection data, and the second branch is used to process the second type of projection data. The projection domain subnetwork adopts a channel attention mechanism to increase the completion weight of pixel regions in the first type of projection data and the second type of projection data whose gradient magnitude is higher than a preset threshold.
[0151] Optionally, the image domain subnetwork is a two-stream residual network, including a third branch and a fourth branch. The third branch is used to process the portion of image data in the intermediate reconstructed image that corresponds to the first type of projection data, and the fourth branch is used to process the portion of image data in the intermediate reconstructed image that corresponds to the second type of projection data. The image domain subnetwork adopts a multi-scale convolutional structure to extract anatomical structural features at different scales from the intermediate reconstructed image.
[0152] Optionally, the projection data acquisition unit includes: an alternating acquisition subunit, used to extract the sampling angle of the first type of projection data from the full-view projection data at preset intervals, and to extract the sampling angle of the second type of projection data at an angle position complementary to the sampling angle of the first type of projection data, so that the first type of projection data and the second type of projection data form spectral complementarity; wherein, the spectral complementarity characterization utilizes the spatial domain structure consistency and energy domain information redundancy of the first type of projection data and the second type of projection data to achieve information complementarity; the full-view projection data characterizes the projection data continuously acquired within the complete scanning angle range, and the first type of projection data and the second type of projection data are subsets obtained by sparse sampling from the full-view projection data.
[0153] Optionally, the alternating acquisition sub-units satisfy the following relationship: the sampling viewpoint of the i-th type of first-class projection data corresponds to the (i-1)-th type of full-view projection data. The n+1) viewpoints; the sampling viewpoint of the i-th type of second-class projection data corresponds to the (i-1 / 2)-th viewpoint in the full-view projection data. (n+1) viewpoints; where i is an integer greater than or equal to 1, and n is the sampling interval parameter and n is an integer greater than or equal to 1.
[0154] Optionally, the ring-shaped dual-source dual-energy image fast reconstruction device based on neural networks further includes: a preprocessing unit, used to preprocess the first type of projection data and the second type of projection data before inputting the first type of projection data and the second type of projection data into the projection domain subnetwork to generate initial projection data; the preprocessing includes edge-aware interpolation processing to preserve the contour information of the measured object.
[0155] See Figure 8According to another aspect of the embodiments of this application, a training method for a neural network model for fast reconstruction of a ring-shaped dual-source dual-energy image is also provided. The neural network model includes at least: a projection domain sub-network, an analytical reconstruction operator, and an image domain sub-network. The method includes the following steps: a training dataset acquisition unit, used to acquire a training dataset, which includes: training projection data collected under sparse perspectives as training input, and ground truth image data as training labels. The ground truth image data is an image reconstructed from full-view projection data, where sparse perspectives represent a number of sampling perspectives less than the number of perspectives used during full-view acquisition; and a joint training processing unit, used to, based on the training dataset, apply a loss function to the training data to train the projection domain sub-network, the analytical reconstruction operator, and the image domain sub-network. The projective operator and the image domain subnetwork are jointly trained. The trained projection domain subnetwork is used to sparsely sample and fill in missing projection data from the first and second types of input projection data, and output the target projection data. The first type of projection data represents the projection signal collected after a ray of the first energy level penetrates the object under test, and the second type of projection data represents the projection signal collected after a ray of the second energy level penetrates the object under test, where the second energy level is higher than the first energy level. The analytical reconstruction operator is used to transform the target projection data from the projection domain to the image domain and output the intermediate reconstructed image. The image domain subnetwork is used to sharpen the edges and suppress noise in the intermediate reconstructed image and output the target reconstructed image.
[0156] Optionally, the loss function is calculated by weighting the pixel loss function, the structural similarity loss function, the dual-energy spectrum difference loss function, and the cross-domain feature alignment loss function, wherein: the pixel loss function is used to constrain the pixel-level difference between the target reconstructed image and the ground truth image data; the structural similarity loss function is used to constrain the structural similarity between the target reconstructed image and the ground truth image data; the dual-energy spectrum difference loss function is used to constrain the image difference between different channel parts in the target reconstructed image to be consistent with the image difference between corresponding different channel parts in the ground truth image data; and the cross-domain feature alignment loss function is used to constrain the consistency between the projection domain feature map extracted by the projection domain sub-network and the image domain feature map extracted by the image domain sub-network.
[0157] Optionally, the joint training processing unit includes: a pixel loss calculation subunit, used to perform the calculation operation of the pixel loss subfunction. The calculation operation of the pixel loss subfunction includes: performing channel pixel loss calculation operations on the first channel and the second channel respectively. The channel pixel loss calculation operation includes calculating the absolute difference between the gray value of each pixel position in the current channel part of the target reconstructed image and the gray value of the corresponding pixel position in the corresponding channel part of the ground image data; summing the absolute differences of all pixel positions to obtain the channel pixel loss of the current channel; adding the channel pixel loss of the first channel to the channel pixel loss of the second channel, and then dividing the sum by the total number of pixels to obtain the average absolute error as the pixel loss subfunction; wherein, the first channel corresponds to the first type of projection data, and the second channel corresponds to the second type of projection data;
[0158] The structural similarity loss calculation subunit is used to perform the calculation operation of the structural similarity loss subfunction. The calculation operation of the structural similarity loss subfunction includes: performing channel structural similarity loss calculation operations on the first channel and the second channel respectively. The channel structural similarity loss calculation operation includes calculating the product of brightness similarity, contrast similarity and structural similarity between the target reconstructed image of the current channel and the corresponding channel part in the ground truth image data to obtain the structural similarity index of the current channel; subtracting the structural similarity index from 1 to obtain the channel structural similarity loss of the current channel; adding the channel structural similarity loss of the first channel and the channel structural similarity loss of the second channel and dividing by 2 to obtain the structural similarity loss subfunction.
[0159] The dual-energy spectrum difference loss calculation subunit is used to perform the calculation operation of the dual-energy spectrum difference loss sub-function. The calculation operation of the dual-energy spectrum difference loss sub-function includes: calculating the pixel difference between the first channel part and the second channel part in the target reconstructed image to obtain the reconstructed dual-energy difference map; calculating the pixel difference between the first channel part and the second channel part in the ground truth image data to obtain the ground truth dual-energy difference map; calculating the absolute difference of each pixel position between the reconstructed dual-energy difference map and the ground truth dual-energy difference map, summing the absolute differences of all pixel positions, dividing the summation result by the total number of pixels, and obtaining the mean absolute error as the dual-energy spectrum difference loss sub-function.
[0160] The cross-domain feature alignment loss calculation subunit is used to perform the calculation operation of the cross-domain feature alignment loss sub-function. The calculation operation of the cross-domain feature alignment loss sub-function includes: performing channel feature alignment loss calculation operations on the first channel and the second channel respectively. The channel feature alignment loss calculation operation includes: obtaining the projection domain feature map of the current channel output by the projection domain sub-network; obtaining the image domain feature map of the current channel output by the image domain sub-network; obtaining the cross-domain feature weights of the current channel, which are learnable parameters used to map the image domain feature map to the feature space of the projection domain feature map; multiplying the cross-domain feature weights with the image domain feature map to obtain the weighted image domain feature map; calculating the absolute difference of each pixel position in the projection domain feature map and the weighted image domain feature map, summing the absolute differences of all pixel positions to obtain the channel feature alignment loss of the current channel; adding the channel feature alignment loss of the first channel to the channel feature alignment loss of the second channel, and then dividing the sum by the total number of pixels to obtain the cross-domain feature alignment loss sub-function.
[0161] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, which stores a computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located executes the above-described method for rapid reconstruction of ring-shaped dual-light source dual-energy images based on neural networks.
[0162] According to another aspect of the embodiments of this application, an electronic device is also provided, including one or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to execute the above-described method for rapid reconstruction of ring-shaped dual-light source dual-energy images based on neural networks.
[0163] According to another aspect of the embodiments of this application, a computer program product is also provided, including a computer program or instructions, which, when executed by a processor, implement the above-described method for rapid reconstruction of ring-shaped dual-source dual-energy images based on neural networks.
[0164] The sequence numbers of the embodiments described above are merely for descriptive purposes and do not represent the superiority or inferiority of the embodiments. In the above embodiments of this application, the descriptions of each embodiment have their own emphasis; for parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments. It should be understood that the disclosed technical content in the several embodiments provided in this application can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units can be a logical functional division, and in actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Furthermore, the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces; the indirect coupling or communication connection between units or modules can be electrical or other forms. The units described as separate components may or may not be physically separated; the components shown as units may or may not be physical units, that is, they can be located in one place or distributed across multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution in this embodiment.
[0165] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes: USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, optical disks, and other media capable of storing program code. The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A neural network-based fast reconstruction method for annular dual-source dual-energy images, characterized in that, include: The first type of projection data and the second type of projection data are acquired by a ring dual-source dual-energy imaging system under sparse viewing angle, wherein the sparse viewing angle indicates that the number of sampling viewing angles is less than the number of viewing angles used in full-view acquisition; the first type of projection data is used to characterize the projection signal acquired after the ray of the first energy level penetrates the object under test, and the second type of projection data is used to characterize the projection signal acquired after the ray of the second energy level penetrates the object under test, wherein the second energy level is higher than the first energy level. Image reconstruction is performed on the first type of projection data and the second type of projection data using a neural network model to obtain the target reconstructed image; wherein, the loss function corresponding to the neural network model is obtained by weighted calculation of pixel loss sub-function, structural similarity loss sub-function, dual-energy spectrum difference loss sub-function, and cross-domain feature alignment loss sub-function, wherein: The pixel loss function is used to constrain the pixel-level differences between the target reconstructed image and the ground truth image data, where the ground truth image data is an image reconstructed from full-view projection data; the structural similarity loss function is used to constrain the structural similarity between the target reconstructed image and the ground truth image data; the dual-energy spectrum difference loss function is used to constrain the image differences between different channel parts in the target reconstructed image to be consistent with the image differences between corresponding different channel parts in the ground truth image data; the cross-domain feature alignment loss function is used to constrain the consistency between the projection domain feature map extracted by the projection domain sub-network and the image domain feature map extracted by the image domain sub-network, where both the projection domain sub-network and the image domain sub-network are partial sub-networks of the neural network model.
2. The method according to claim 1, characterized in that, Image reconstruction is performed on the first type of projection data and the second type of projection data using a neural network model to obtain the target reconstructed image, including: The first type of projection data and the second type of projection data are input into the projection domain sub-network in the neural network model. The projection domain sub-network is used to complete the sparsely sampled projection data and output the target projection data. The target projection data includes the completed first type of projection data and the completed second type of projection data. The target projection data is reconstructed using the analytical reconstruction operator in the neural network model to obtain an intermediate reconstructed image, wherein the analytical reconstruction operator is used to transform the target projection data from the projection domain to the image domain; The intermediate reconstructed image is input into the image domain subnetwork of the neural network model. The image domain subnetwork performs edge sharpening and noise suppression on the intermediate reconstructed image and outputs the target reconstructed image.
3. The method according to claim 2, characterized in that, The method further includes: Obtain the complete projection data term and the cross-domain feature fusion term, wherein the complete projection data term is equal to the target projection data output by the projection domain sub-network, and the cross-domain feature fusion term is equal to the product of the cross-domain feature weights and the structural prior projection reference data fed back from the image domain sub-network; the structural prior projection reference data represents the projection data obtained by mapping the image optimized by the image domain sub-network back to the projection domain through Radon transform. The target data item is obtained by summing the completed projection data item and the cross-domain feature fusion item. The transformation operator, adaptive projection weighting factor, adaptive total variation regularization coefficient, and total variation regularization term are obtained. The transformation operator is an inverse Radon transform operator, used as transformation coefficients to transform the target projection data from the projection domain to the image domain. The adaptive projection weighting factor is used to dynamically adjust the projection weights of the target projection data according to its sparsity. The adaptive total variation regularization coefficient is used to adjust the regularization intensity of the intermediate reconstructed image according to its sparsity. The total variation regularization term is used to constrain the edge smoothness of the intermediate reconstructed image. The analytical reconstruction operator in the neural network model is determined based on the target data item, the transformation operator, the adaptive projection weighting factor, the adaptive total variation regularization coefficient, and the total variation regularization term.
4. The method according to claim 2, characterized in that, The projection domain subnetwork and the image domain subnetwork constitute a cross-domain collaborative enhancement closed-loop architecture; the anatomical structure features extracted from the intermediate reconstructed image by the image domain subnetwork are fed back to the projection domain subnetwork to guide the projection domain subnetwork to complete the projection data missing by sparse sampling; the anatomical structure features represent the organ structure information extracted from the image domain corresponding to the intermediate reconstructed image and fed back to the projection domain. The target projection data output by the projection domain subnetwork is reconstructed by the analytical reconstruction operator and then input into the image domain subnetwork.
5. The method according to claim 3, characterized in that, Obtain the adaptive projection weighting factor, including: If the sparsity of the target projection data is detected to be less than or equal to a preset threshold, it is determined that the adaptive projection weighting factor is greater than the standard weight value, where the standard weight value is 1. If the sparsity of the target projection data is detected to be greater than the preset threshold, the adaptive projection weighting factor is determined to be equal to the standard weight value.
6. The method according to claim 5, characterized in that, If the sparsity of the target projection data is detected to be less than or equal to a preset threshold, it is determined that the adaptive projection weighting factor is greater than the standard weight value, including: If the sparsity of the target projection data is detected to be less than or equal to a preset threshold, the adaptive projection weighting factor is determined to be equal to... ,in, The preset coefficient is greater than 0. The sparsity of the target projection data.
7. The method according to claim 3, characterized in that, The adaptive total variation regularization coefficient increases with the sparsity of the target projection data. Obtaining the total variation regularization term includes: Calculate the sum of squares of the differences between adjacent pixels in the intermediate reconstructed image; Take the square root of the sum of squares to obtain the square root result; The square root result is used as the total variation regularization term.
8. The method according to claim 1, characterized in that, The projection domain subnetwork is a two-stream residual network, including a first branch and a second branch. The first branch is used to process the first type of projection data, and the second branch is used to process the second type of projection data. The projection domain subnetwork adopts a channel attention mechanism to increase the completion weight of pixel regions in the first type of projection data and the second type of projection data whose gradient magnitude is higher than a preset threshold.
9. The method according to claim 2, characterized in that, The image domain subnetwork is a two-stream residual network, including a third branch and a fourth branch. The third branch is used to process the portion of image data in the intermediate reconstructed image that corresponds to the first type of projection data, and the fourth branch is used to process the portion of image data in the intermediate reconstructed image that corresponds to the second type of projection data. The image domain subnetwork adopts a multi-scale convolutional structure to extract anatomical structural features at different scales from the intermediate reconstructed image.
10. The method according to claim 1, characterized in that, The first type of projection data and the second type of projection data are collected alternately at different angles and dimensions, including: The sampling angles of the first type of projection data are extracted from the full-view projection data at preset intervals, and the sampling angles of the second type of projection data are extracted at angle positions that are complementary to the sampling angles of the first type of projection data, so that the first type of projection data and the second type of projection data form spectral complementarity. The spectral complementarity characterization utilizes the spatial domain structure consistency and energy domain information redundancy of the first type of projection data and the second type of projection data to complement each other; the full-view projection data characterizes the projection data continuously acquired within the full scanning angle range, and the first type of projection data and the second type of projection data are subsets obtained by sparse sampling from the full-view projection data.
11. The method according to claim 10, characterized in that, The first type of projection data and the second type of projection data are collected alternately in the angular dimension, satisfying the following relationship: The sampling viewpoint of the i-th type of first-class projection data corresponds to the (i-1)-th viewpoint in the full-view projection data. (n+1) perspectives; The sampling viewpoint of the i-th type of second-order projection data corresponds to the (i-1 / 2)-th viewpoint in the full-view projection data. (n+1) perspectives; Where i is an integer greater than or equal to 1, and n is the sampling interval parameter and n is an integer greater than or equal to 1.
12. The method according to claim 2, characterized in that, Before inputting the first type of projection data and the second type of projection data into the projection domain subnetwork, the method further includes: The first type of projection data and the second type of projection data are preprocessed to generate initial projection data; the preprocessing includes edge-aware interpolation to preserve the contour information of the object being measured.
13. A training method for a neural network model for fast reconstruction of a ring-shaped dual-source dual-energy image, wherein the neural network model includes at least: The projection domain subnetwork, analytical reconstruction operator, and image domain subnetwork are characterized by comprising the following steps: Obtain a training dataset, which includes: training projection data collected under sparse perspectives as training input, and ground truth image data as training labels. The ground truth image data is an image reconstructed from full-view projection data. The sparse perspectives represent that the number of sampling perspectives is less than the number of perspectives used when collecting full-view data. Based on the training dataset, the projection domain subnetwork, the analytical reconstruction operator, and the image domain subnetwork are jointly trained using a loss function. The loss function is calculated by weighting the pixel loss function, the structural similarity loss function, the dual-energy spectrum difference loss function, and the cross-domain feature alignment loss function, wherein: The pixel loss function is used to constrain the pixel-level differences between the target reconstructed image and the ground truth image data; the structural similarity loss function is used to constrain the structural similarity between the target reconstructed image and the ground truth image data; the dual-energy spectrum difference loss function is used to constrain the image differences between different channel parts in the target reconstructed image to be consistent with the image differences between corresponding different channel parts in the ground truth image data; the cross-domain feature alignment loss function is used to constrain the consistency between the projection domain feature map extracted by the projection domain sub-network and the image domain feature map extracted by the image domain sub-network.
14. The training method according to claim 13, characterized in that, The calculation method of the pixel loss sub-function includes: The channel pixel loss calculation operation is performed on the first channel and the second channel respectively, including: calculating the absolute difference between the gray value of each pixel position in the current channel part of the target reconstructed image and the gray value of the corresponding pixel position in the corresponding channel part of the ground truth image data, summing the absolute differences of all pixel positions to obtain the channel pixel loss of the current channel; The channel pixel loss of the first channel is added to the channel pixel loss of the second channel, and the sum is divided by the total number of pixels to obtain the mean absolute error as the pixel loss sub-function; wherein, the first channel corresponds to the first type of projection data, and the second channel corresponds to the second type of projection data. The first type of projection data is used to characterize the projection signal collected after the ray of the first energy level penetrates the object under test, and the second type of projection data is used to characterize the projection signal collected after the ray of the second energy level penetrates the object under test. The second energy level is higher than the first energy level.
15. The training method according to claim 13, characterized in that, The calculation method of the structural similarity loss sub-function includes: The channel structure similarity loss is calculated for the first channel and the second channel respectively, including: calculating the product of brightness similarity, contrast similarity and structural similarity between the target reconstructed image of the current channel and the corresponding channel in the ground truth image data to obtain the structural similarity index of the current channel; subtracting the structural similarity index from 1 to obtain the channel structure similarity loss of the current channel. The structural similarity loss subfunction is obtained by adding the channel structure similarity loss of the first channel to the channel structure similarity loss of the second channel and then dividing by 2.
16. The training method according to claim 13, characterized in that, The calculation method of the dual-energy spectrum difference loss sub-function includes: Calculate the pixel difference between the first channel portion and the second channel portion in the target reconstructed image to obtain the reconstructed dual-energy difference map; Calculate the pixel difference between the first channel portion and the second channel portion in the ground truth image data to obtain the ground truth dual-energy difference map; Calculate the absolute difference between the reconstructed dual-energy difference map and the true dual-energy difference map for each pixel position, sum the absolute differences for all pixel positions, divide the sum by the total number of pixels, and obtain the mean absolute error as the dual-energy spectrum difference loss sub-function.
17. The training method according to claim 13, characterized in that, The calculation method of the cross-domain feature alignment loss sub-function includes: The calculation of channel feature alignment loss is performed on the first channel and the second channel respectively, including: obtaining the projection domain feature map of the current channel output by the projection domain sub-network; obtaining the image domain feature map of the current channel output by the image domain sub-network; obtaining the cross-domain feature weights of the current channel, wherein the cross-domain feature weights are learnable parameters used to map the image domain feature map to the feature space of the projection domain feature map; multiplying the cross-domain feature weights with the image domain feature map to obtain a weighted image domain feature map; calculating the absolute difference between the projection domain feature map and the weighted image domain feature map for each pixel position, and summing the absolute differences of all pixel positions to obtain the channel feature alignment loss of the current channel; The channel feature alignment loss of the first channel is added to the channel feature alignment loss of the second channel, and the sum is divided by the total number of pixels to obtain the cross-domain feature alignment loss sub-function.
18. A ring-shaped dual-light source dual-energy image rapid reconstruction device based on neural networks, characterized in that, include: The projection data acquisition unit is used to acquire a first type of projection data and a second type of projection data collected by the annular dual-light source dual-energy imaging system under sparse viewing angles, wherein the sparse viewing angles represent that the number of sampling viewing angles is less than the number of viewing angles used when acquiring full-view data; the first type of projection data is used to represent the projection signal acquired after the first energy level of the ray penetrates the object under test, and the second type of projection data is used to represent the projection signal acquired after the second energy level of the ray penetrates the object under test, wherein the second energy level is higher than the first energy level. The image reconstruction unit is used to reconstruct the target image from the first type of projection data and the second type of projection data using a neural network model. The loss function corresponding to the neural network model is calculated by weighting a pixel loss sub-function, a structural similarity loss sub-function, a dual-energy spectrum difference loss sub-function, and a cross-domain feature alignment loss sub-function. The pixel loss function is used to constrain the pixel-level differences between the target reconstructed image and the ground truth image data, where the ground truth image data is an image reconstructed from full-view projection data; the structural similarity loss function is used to constrain the structural similarity between the target reconstructed image and the ground truth image data; the dual-energy spectrum difference loss function is used to constrain the image differences between different channel parts in the target reconstructed image to be consistent with the image differences between corresponding different channel parts in the ground truth image data; the cross-domain feature alignment loss function is used to constrain the consistency between the projection domain feature map extracted by the projection domain sub-network and the image domain feature map extracted by the image domain sub-network, where both the projection domain sub-network and the image domain sub-network are partial sub-networks of the neural network model.
19. A training apparatus for a neural network model for fast reconstruction of a ring-shaped dual-source dual-energy image, the neural network model comprising at least: The projection domain subnetwork, analytical reconstruction operator, and image domain subnetwork are characterized by comprising the following units: The training dataset acquisition unit is used to acquire a training dataset, which includes: training projection data collected under sparse perspectives as training input, and ground truth image data as training labels. The ground truth image data is an image reconstructed from full-view projection data. The sparse perspectives represent that the number of sampling perspectives is less than the number of perspectives used when collecting full-view data. The joint training processing unit is used to jointly train the projection domain subnetwork, the analytical reconstruction operator, and the image domain subnetwork based on the training dataset and using a loss function. The loss function is calculated by weighting the pixel loss function, the structural similarity loss function, the dual-energy spectrum difference loss function, and the cross-domain feature alignment loss function, wherein: The pixel loss function is used to constrain the pixel-level differences between the target reconstructed image and the ground truth image data; the structural similarity loss function is used to constrain the structural similarity between the target reconstructed image and the ground truth image data; the dual-energy spectrum difference loss function is used to constrain the image differences between different channel parts in the target reconstructed image to be consistent with the image differences between corresponding different channel parts in the ground truth image data; the cross-domain feature alignment loss function is used to constrain the consistency between the projection domain feature map extracted by the projection domain sub-network and the image domain feature map extracted by the image domain sub-network.