A method and apparatus for fast image reconstruction from a single light source and a dual-energy plate based on neural networks.

By performing pixel-by-pixel calibration and fusion reconstruction of the projection data of the dual-energy plate imaging system using a neural network-based method, the problem of poor image reconstruction accuracy caused by detector signal crosstalk was solved, and higher quality image reconstruction and material decomposition were achieved.

CN122492475APending Publication Date: 2026-07-31MANTEIA TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
MANTEIA TECH CO LTD
Filing Date
2026-07-01
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

In existing dual-energy plate imaging systems, energy spectrum crosstalk between the signals of the front and rear layers of the detector leads to poor image reconstruction accuracy, a problem that current technologies have not been able to effectively solve.

Method used

A neural network-based approach is adopted, which uses a dual-layer detector imaging system to acquire projection data and performs pixel-by-pixel adaptive calibration through an energy spectrum response calibration subnetwork to separate crosstalk components in the signal. Finally, a dual-energy sensing fusion reconstruction network is used to perform image fusion reconstruction.

Benefits of technology

It effectively removes crosstalk between detector signals, improves the quality of image reconstruction and the accuracy of material decomposition, and enhances image reconstruction precision.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122492475A_ABST
    Figure CN122492475A_ABST
Patent Text Reader

Abstract

This application discloses a method and apparatus for rapid image reconstruction using a single-source dual-energy plate based on a neural network, relating to the field of medical technology. The method includes: acquiring first projection data and second projection data using a dual-layer detector imaging system; inputting the first and second projection data into a trained energy spectrum response calibration sub-network; adaptively calibrating the first and second projection data pixel-by-pixel using the energy spectrum response calibration sub-network to obtain first target projection data and second target projection data; the first target projection data represents the calibrated projection data representing the true attenuation distribution of photons at the first energy level, and the second target projection data represents the calibrated projection data representing the true attenuation distribution of photons at the second energy level; and generating a reconstructed target image based on the first and second target projection data. This application solves the problem of poor image reconstruction accuracy caused by crosstalk between the front and rear layers of the detector in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of medical technology, and more specifically, to a method and apparatus for rapid image reconstruction from a single light source dual-energy plate based on a neural network. Background Technology

[0002] In the field of image-guided radiotherapy, dual-energy imaging technology utilizes radiation information from both high and low energy bands, which is beneficial for distinguishing different substances such as bone and soft tissue.

[0003] However, in existing dual-energy plate imaging systems, the dual-layer detector consists of a front-layer scintillator and a rear-layer scintillator. The front-layer scintillator is sensitive to low-energy photons, while the rear-layer scintillator is sensitive to high-energy photons. Due to incomplete energy spectrum separation, some high-energy components are mixed into the front-layer signal, and some low-energy components remain in the rear-layer signal; that is, there is energy spectrum crosstalk between the front and rear layers. The presence of crosstalk causes the acquired low-energy and high-energy projection data to be confused with each other, directly reducing the accuracy of subsequent dual-energy matter decomposition and image reconstruction.

[0004] There is currently no effective solution to the above problems. Summary of the Invention

[0005] This application provides a method and apparatus for fast image reconstruction using a single-source dual-energy plate based on neural networks, which at least solves the technical problem of poor image reconstruction accuracy caused by crosstalk between the front and rear layers of the detector in the prior art.

[0006] According to one aspect of the embodiments of this application, a method for fast image reconstruction of a single-source dual-energy plate based on a neural network is provided, comprising: acquiring first projection data and second projection data using a dual-layer detector imaging system, wherein the first projection data corresponds to the signal output by the detector layer on the side closer to the X-ray source in the dual-layer detector, and the second projection data corresponds to the signal output by the detector layer on the side farther from the X-ray source in the dual-layer detector; inputting the first projection data and second projection data into a trained energy spectrum response calibration sub-network, and performing pixel-by-pixel adaptive calibration on the first projection data and second projection data through the energy spectrum response calibration sub-network to obtain first target projection data and second target projection data; wherein the first target projection data is used to characterize the calibrated projection data representing the true attenuation distribution of photons at a first energy level, and the second target projection data is used to characterize the calibrated projection data representing the true attenuation distribution of photons at a second energy level, wherein the first energy level is lower than the second energy level; and by adjusting the first projection data and second target projection data, the method for fast image reconstruction of a single-source dual-energy plate based on a neural network is provided, comprising: acquiring first projection data and second projection data using a dual-layer detector imaging system, wherein the first projection data corresponds to the signal output by the detector layer on the side closer to the X-ray source, and the second ... A target projection data is reconstructed to obtain an initial reconstructed image of the first energy level. A second target projection data is reconstructed to obtain an initial reconstructed image of the second energy level. The first and second initial reconstructed images of the first and second energy levels are decomposed using learnable energy spectrum basis functions. This decomposition includes linearly weighted summation of the first and second initial reconstructed images of the first and second energy levels to generate a first material image and a second material image. The first material image represents prominent skeletal components, and the second material image represents prominent soft tissue components. The first, second, first, and second initial reconstructed images of the first and second energy levels are combined into a four-channel input and fed into a dual-energy perception fusion reconstruction network. This network fuses energy and material feature information using a dual-branch feature encoder and an attention fusion module to output a target reconstructed image.

[0007] According to another aspect of the embodiments of this application, a training method for a neural network model for fast reconstruction of images from a single-source dual-energy plate is also provided, comprising: acquiring projection data containing energy spectrum crosstalk and projection data without energy spectrum crosstalk; using the projection data containing energy spectrum crosstalk as training input data and the projection data without energy spectrum crosstalk as training labels to perform supervised training on the neural network to obtain an energy spectrum response calibration sub-network; inputting the projection data containing energy spectrum crosstalk into the energy spectrum response calibration sub-network for pixel-by-pixel adaptive calibration to obtain first target projection data and second target projection data for training; reconstructing the first target projection data for training to obtain an initial reconstructed image of the first energy level for training; and further reconstructing the first target projection data for training. The second target projection data is reconstructed to obtain an initial reconstructed image of the second energy level used for training; the initial reconstructed image of the first energy level and the initial reconstructed image of the second energy level used for training are subjected to learnable energy spectrum basis function decomposition, including: generating a first material image and a second material image used for training by linearly weighted summing the initial reconstructed image of the first energy level and the initial reconstructed image of the second energy level used for training; the initial reconstructed image of the first energy level, the initial reconstructed image of the second energy level, the first material image and the second material image used for training are combined into a four-channel input as simulation data, and a dual-energy sensing fusion reconstruction network is trained based on the simulation data.

[0008] According to another aspect of the embodiments of this application, a fast image reconstruction device for a single-source dual-energy plate based on a neural network is also provided, comprising: an acquisition unit, configured to acquire first projection data and second projection data using a dual-layer detector imaging system, wherein the first projection data corresponds to the signal output by the detector layer on the side closer to the X-ray source in the dual-layer detector, and the second projection data corresponds to the signal output by the detector layer on the side farther from the X-ray source in the dual-layer detector; a data processing unit, configured to input the first projection data and second projection data into a trained energy spectrum response calibration subnetwork, and perform pixel-by-pixel adaptive calibration on the first projection data and second projection data through the energy spectrum response calibration subnetwork to obtain first target projection data and second target projection data; wherein the first target projection data is used to characterize the calibrated projection data representing the true attenuation distribution of photons at the first energy level, and the second target projection data is used to characterize the calibrated projection data representing the true attenuation distribution of photons at the second energy level; and an image reconstruction unit, configured to generate a target reconstruction image based on the first target projection data and the second target projection data. The process includes: reconstructing the first target projection data to obtain a first energy level initial reconstructed image; reconstructing the second target projection data to obtain a second energy level initial reconstructed image; performing learnable energy spectrum basis function decomposition on the first and second energy level initial reconstructed images, wherein the energy spectrum basis function decomposition includes: generating a first material image and a second material image by performing linear weighted summation on the first and second energy level initial reconstructed images; wherein the first material image is used to characterize the image of prominent skeletal components, and the second material image is used to characterize the image of prominent soft tissue components; combining the first energy level initial reconstructed image, the second energy level initial reconstructed image, the first material image, and the second material image into a four-channel input, and inputting it into a dual-energy perception fusion reconstruction network, wherein the dual-energy perception fusion reconstruction network performs fusion processing on energy feature information and material feature information through a dual-branch feature encoder and an attention fusion module, and outputs the target reconstructed image.

[0009] According to another aspect of the embodiments of this application, a training device for a neural network model for fast image reconstruction of a single-source dual-energy plate is also provided, comprising: a data acquisition unit for acquiring projection data containing energy spectrum crosstalk and projection data without energy spectrum crosstalk; a first training unit for supervising training of the neural network using projection data containing energy spectrum crosstalk as training input data and projection data without energy spectrum crosstalk as training labels to obtain an energy spectrum response calibration subnetwork; the training device is further configured to input projection data containing energy spectrum crosstalk into the energy spectrum response calibration subnetwork for pixel-by-pixel adaptive calibration to obtain first target projection data and second target projection data for training; and to reconstruct the first target projection data for training to obtain the training data for... The first energy level initial reconstruction image is used to reconstruct the second target projection data, resulting in a second energy level initial reconstruction image for training. The first and second energy level initial reconstruction images used for training are then subjected to learnable energy spectrum basis function decomposition, including: generating a first and second material image for training by linearly weighted summing the first and second energy level initial reconstruction images; the first, second energy level initial reconstruction images, the first material image, and the second material image used for training are combined into a four-channel input as simulation data, and a dual-energy sensing fusion reconstruction network is trained based on the simulation data.

[0010] In this embodiment, firstly, a dual-layer detector imaging system is used to acquire first projection data corresponding to the detector layer closer to the X-ray source and second projection data corresponding to the detector layer farther from the X-ray source. Both the first and second projection data contain crosstalk components from each other's energy levels. Secondly, the first and second projection data are input into a trained energy spectrum response calibration subnetwork. This subnetwork performs pixel-by-pixel adaptive calibration on the first and second projection data, directly outputting first target projection data representing the true attenuation distribution of photons at the first energy level and second target projection data representing the true attenuation distribution of photons at the second energy level. Finally, a target reconstruction image is generated based on the calibrated first and second target projection data. Because the calibration process can separate crosstalk components pixel-by-pixel in the projection domain, the projection data used for subsequent reconstruction is closer to the ideal pure low-energy and high-energy attenuation distribution, thus helping to avoid crosstalk interference on reconstruction accuracy, improving the quality of image reconstruction and the accuracy of matter decomposition, and thus solving the technical problem of poor image reconstruction accuracy caused by crosstalk between the front and rear layers of the detector in the prior art. Attached Figure Description

[0011] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0012] Figure 1 This is a schematic diagram of an optional neural network-based method for rapid image reconstruction of a single-source dual-energy plate according to an embodiment of this application;

[0013] Figure 2 This is a flowchart of an optional method for fast image reconstruction of a single-source dual-energy plate based on a neural network, according to an embodiment of this application.

[0014] Figure 3 This is another overall flowchart of an optional neural network-based fast image reconstruction of a single-source dual-energy plate according to an embodiment of this application;

[0015] Figure 4 This is a schematic diagram of an optional single-source dual-energy plate imaging system according to an embodiment of this application;

[0016] Figure 5 This is a schematic diagram of an optional training method for a neural network model for fast reconstruction of images from a single-source dual-energy plate, according to an embodiment of this application.

[0017] Figure 6 This is a schematic diagram of an optional neural network-based single-source dual-energy plate image rapid reconstruction device according to an embodiment of this application;

[0018] Figure 7 This is a schematic diagram of an optional training apparatus for a neural network model for rapid reconstruction of images from a single-source dual-energy plate, according to an embodiment of this application. Detailed Implementation

[0019] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0020] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0021] According to an embodiment of this application, a method embodiment of a fast image reconstruction method for a single-source dual-energy plate based on a neural network is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0022] It should be noted that the information collected in this application (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for display, data used for analysis, etc.) are information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of this data all comply with relevant laws, regulations, and standards, necessary confidentiality measures have been taken, and they do not violate public order and good morals. Corresponding access points are provided for users to choose to authorize or refuse. For example, interfaces are set up between this system and relevant users or organizations, providing users with corresponding access points to choose to agree to or refuse automated decision-making results; if the user chooses to refuse, the process proceeds to the expert decision-making stage.

[0023] According to the embodiments of this application, an image processing system can be used as the execution subject of the fast image reconstruction method for a single light source dual-energy plate based on neural networks in the embodiments of this application. The system can be a software system or an embedded system combining software and hardware. Of course, the execution subject of the method in the embodiments of this application can also be other forms of execution subject, such as devices, equipment, etc. It should be known by those skilled in the art that this application does not particularly limit the specific form of the execution subject.

[0024] Figure 1 This is a schematic diagram of a fast image reconstruction method for a single-source dual-energy plate based on a neural network according to an embodiment of this application, as shown below. Figure 1 As shown, the method includes the following steps:

[0025] Step S101: Acquire first projection data and second projection data using a dual-layer detector imaging system. The first projection data corresponds to the signal output by the detector layer on the side closer to the X-ray source in the dual-layer detector, and the second projection data corresponds to the signal output by the detector layer on the side farther from the X-ray source in the dual-layer detector.

[0026] For example, a dual-layer detector imaging system can refer to an imaging device composed of at least two detector layers arranged sequentially along the direction of X-ray beam propagation. The first projection data corresponds to the signal output by the detector layer closer to the X-ray source. This detector layer is highly sensitive to lower-energy photons, primarily absorbing them and outputting a signal reflecting the low-energy attenuation distribution. The second projection data corresponds to the signal output by the detector layer farther from the X-ray source. This detector layer is highly sensitive to higher-energy photons, primarily absorbing high-energy photons that have penetrated the previous layer and outputting a signal reflecting the high-energy attenuation distribution. The image processing system is connected to the dual-layer detector imaging system via a data interface. During the multi-angle scanning process of the X-ray source and the dual-layer detector around the target object, it synchronously receives the projection data output by the previous and next layers at each acquisition angle and stores it according to the acquisition angle and channel index to obtain the first and second projection data.

[0027] In some embodiments, the image processing system can control the X-ray source to continuously emit a beam with a preset tube voltage and tube current, while simultaneously controlling the dual-layer detector imaging system to perform exposure acquisition at a fixed frame rate. After the X-ray source and the dual-layer detector complete scanning around the target object within a predetermined angle range, the image processing system reads the electrical signals generated by the front and rear layers in real time from the detector's data output port. After analog-to-digital conversion and preprocessing (e.g., dark field correction, gain correction, and bad pixel replacement), it generates first projection data and second projection data respectively, and stores the first projection data and second projection data in angular order as a two-dimensional projection sequence.

[0028] In other embodiments, the image processing system may also employ a triggered synchronous acquisition mode. The image processing system sends a synchronization pulse signal to the drive controller of the rotating gantry, triggering the X-ray source to emit a beam instantaneously at each preset rotation angle position, and simultaneously triggering the dual-layer detector imaging system to perform a single exposure acquisition. At each trigger moment, the image processing system reads the projection signals output from the front and rear layers respectively, forming first projection data and second projection data. This acquisition mode is beneficial for obtaining projection data with accurate angular positions, reducing the impact of angular errors caused by continuous rotational motion on subsequent image reconstruction.

[0029] Step S102: Input the first projection data and the second projection data into the trained energy spectrum response calibration sub-network, and perform pixel-by-pixel adaptive calibration on the first projection data and the second projection data through the energy spectrum response calibration sub-network to obtain the first target projection data and the second target projection data; wherein, the first target projection data is used to characterize the projection data representing the true attenuation distribution of photons at the first energy level after calibration, and the second target projection data is used to characterize the projection data representing the true attenuation distribution of photons at the second energy level after calibration, and the first energy level is lower than the second energy level.

[0030] For example, the energy spectrum response calibration subnetwork can refer to a lightweight convolutional neural network used for pixel-by-pixel adaptive correction of the raw projection data acquired by a dual-layer detector imaging system. The energy spectrum response calibration subnetwork can take a two-channel two-dimensional image composed of first and second projection data as input and output a two-channel two-dimensional image composed of first and second target projection data. The first target projection data can characterize the true attenuation distribution of lower-energy photons after calibration as they pass through the target object, and the second target projection data can characterize the true attenuation distribution of higher-energy photons after calibration. The image processing system can input the acquired first and second projection data into the trained energy spectrum response calibration subnetwork in an angular sequence, and the network can independently perform calibration operations on each pixel position, outputting calibrated projection data.

[0031] In some embodiments, the energy spectrum response calibration subnetwork can employ a multi-layer convolutional structure, such as five convolutional layers, each containing a convolutional kernel, batch normalization, and activation function. The image processing system concatenates the first and second projection data into a two-channel tensor along the channel dimension and feeds it into the energy spectrum response calibration subnetwork. The network extracts local spatial features around each pixel location through layer-by-layer convolution operations and performs a nonlinear transformation on the original signal based on the learned parameters, outputting the calibration results for both channels. Because the network performs calibration independently for each pixel location—that is, the calibration parameters for each pixel location can be different—it can adapt to the spatial distribution differences of crosstalk coefficients caused by variations in the thickness and tissue composition of the target object.

[0032] In other embodiments, the energy spectrum response calibration subnetwork has been pre-trained using supervised learning. During training, a paired dataset is generated using Monte Carlo simulation. This paired dataset includes projection data with energy spectrum crosstalk (as training input) and corresponding ideal projection data without energy spectrum crosstalk (as training labels). The image processing system directly calls the trained network parameters, inputting the real-time acquired first and second projection data into the network for forward inference to quickly obtain the first and second target projection data. This calibration process requires no iterative optimization and can be completed in a single forward propagation, which is beneficial for meeting the speed requirements of real-time image processing in clinical settings.

[0033] Step S103: Generate a target reconstruction image based on the first target projection data and the second target projection data.

[0034] For example, the target reconstruction image can refer to three-dimensional volume data or a two-dimensional cross-sectional image obtained by processing calibrated dual-energy projection data. The target reconstruction image can reflect the attenuation characteristics of the internal tissue of the target object at different energy levels. After obtaining the first target projection data and the second target projection data, the image processing system performs a reconstruction operation on the first target projection data and the second target projection data, converting the data in the projection domain into image data in the spatial domain to generate the target reconstruction image. The reconstruction operation can include, but is not limited to, one or more of analytical reconstruction algorithms, iterative reconstruction algorithms, or neural network-based image generation algorithms.

[0035] In some embodiments, the image processing system can independently reconstruct the first target projection data and the second target projection data in three dimensions to obtain a first energy level reconstructed image and a second energy level reconstructed image. Subsequently, the image processing system fuses the first energy level reconstructed image and the second energy level reconstructed image, for example, through weighted summation or material decomposition based on a physical model, to generate a target reconstructed image including tissue composition information. This implementation is suitable for application scenarios that require the separate preservation of low-energy and high-energy image features, which is beneficial for subsequent material identification and analysis.

[0036] In other embodiments, the image processing system can also directly feed the first target projection data and the second target projection data as dual-channel inputs into a trained image reconstruction network. This image reconstruction network can employ an encoder-decoder structure, taking the calibrated projection data as input and learning an end-to-end mapping from projection to image to output the reconstructed target image in one step. The image processing system rapidly completes the reconstruction process from projection to image through the network's forward inference computation. This implementation method helps reduce manual parameter adjustments in traditional reconstruction algorithms and shortens computation time while maintaining image quality.

[0037] Figure 2The present invention illustrates an overall flowchart of a method for rapid image reconstruction using a single-source dual-energy plate based on a neural network according to an embodiment of the present application. The image processing system first acquires first projection data and second projection data, then performs pixel-by-pixel adaptive calibration and performs analytical initial reconstruction through the energy spectrum response calibration sub-network. Next, the reconstructed initial image is decomposed by a learnable basis function and then fed into the dual-energy sensing fusion reconstruction network for fusion reconstruction. Finally, the target reconstructed image is output for image-guided radiotherapy applications.

[0038] In some optional embodiments, pixel-by-pixel adaptive calibration is used to remove spectral crosstalk between the front and back layers of the dual-layer detector, which includes a second energy level signal component mixed into the front layer signal of the dual-layer detector and a first energy level signal component remaining in the back layer signal of the dual-layer detector.

[0039] For example, spectral crosstalk can refer to the signal mixing phenomenon that occurs in a dual-layer detector imaging system due to incomplete energy spectrum separation between the detector layer (front layer) closer to the X-ray source and the detector layer (rear layer) farther from the X-ray source. For instance, the front layer signal may contain a second energy level (higher energy) signal component that should have been received by the rear layer, while the rear layer signal may contain a first energy level (lower energy) signal component that should have been absorbed by the front layer. The image processing system can perform pixel-by-pixel adaptive calibration on the first and second projection data through an energy spectrum response calibration subnetwork. This helps to separate the pure first and second energy level attenuation distributions from the mixed signal, thereby reducing the impact of spectral crosstalk on the accuracy of subsequent image reconstruction.

[0040] In some embodiments, the energy spectrum response calibration subnetwork takes first and second projection data as input and extracts local spatial features around each pixel location through convolutional layers within the network. The network learns a nonlinear mapping relationship from the original signal containing crosstalk to an ideal signal without crosstalk. For each pixel location, the network adaptively estimates the crosstalk coefficient at that pixel location based on the pixel's neighborhood information and the signal strength of that pixel in the first and second projection data, such as the proportion of high-energy components in the previous layer signal and the proportion of low-energy components in the subsequent layer signal. Then, it subtracts the crosstalk component from the original signal through linear or nonlinear transformation, outputting calibrated first and second target projection data. Since the crosstalk coefficient varies with the thickness of the target object and the spatial distribution of tissue components, pixel-by-pixel processing can adapt to different levels of crosstalk at different locations.

[0041] In other embodiments, the image processing system may employ a pre-trained spectral response calibration subnetwork, which has learned the mapping from crosstalk-included projections to crosstalk-free projections on a large amount of simulated data. The training data is generated through Monte Carlo simulations, simulating the real signal responses of target objects of varying thicknesses and tissue compositions in a dual-layer detector imaging system. During training, the spectral response calibration subnetwork automatically adjusts its internal parameters to minimize the difference between the calibration result output and the ideal projection for any input first and second projection data. In actual operation, the image processing system inputs real-time acquired projection data into the spectral response calibration subnetwork, which quickly outputs calibrated first and second target projection data through a single forward propagation, effectively removing spectral crosstalk without introducing additional iterative calculations.

[0042] In some optional embodiments, generating a target reconstruction image based on first target projection data and second target projection data includes: an image processing system can reconstruct the first target projection data to obtain a first energy level initial reconstruction image; reconstruct the second target projection data to obtain a second energy level initial reconstruction image; and generate a target reconstruction image based on the first energy level initial reconstruction image and the second energy level initial reconstruction image.

[0043] For example, the first energy level initial reconstructed image can refer to the three-dimensional volume data obtained by performing a three-dimensional reconstruction algorithm on the calibrated first target projection data. The first energy level initial reconstructed image can reflect the attenuation distribution of lower energy level photons inside the object under test. The second energy level initial reconstructed image can refer to the three-dimensional volume data obtained by performing the same three-dimensional reconstruction algorithm on the calibrated second target projection data. The second energy level initial reconstructed image can reflect the attenuation distribution of higher energy level photons inside the object under test. The image processing system independently performs reconstruction operations on the first target projection data and the second target projection data, respectively, to obtain two initial reconstructed images: the first energy level initial reconstructed image and the second energy level initial reconstructed image. Then, further processing is performed based on the first energy level initial reconstructed image and the second energy level initial reconstructed image to generate the target reconstructed image.

[0044] In some embodiments, the image processing system can employ a filtered back-projection algorithm to reconstruct the first target projection data and the second target projection data, respectively. For each energy level, the image processing system first filters the projection data at each acquisition angle, for example, using a ramp filter or a window function filter. Then, it back-projects the filtered projection data along the ray propagation path to each voxel in the three-dimensional image space, and accumulates the back-projection results from all angles to generate the initial reconstructed image of the first energy level and the initial reconstructed image of the second energy level, respectively. The image processing system then performs a weighted fusion of the initial reconstructed image of the first energy level and the initial reconstructed image of the second energy level, for example, by weighted summation according to a fixed ratio, to obtain a target reconstructed image including dual-energy information. This implementation has high computational efficiency and is suitable for real-time or near-real-time imaging scenarios with high requirements for reconstruction speed.

[0045] In other embodiments, the image processing system can use the first energy level initial reconstructed image and the second energy level initial reconstructed image as input data for subsequent neural network processing. The image processing system quickly obtains two relatively coarse-quality initial reconstructed images containing basic anatomical structural information using an analytical reconstruction algorithm, and then feeds these two initial reconstructed images into a pre-trained dual-energy-sensor fusion reconstruction network. This dual-energy-sensor fusion reconstruction network optimizes the initial reconstructed images through end-to-end mapping, outputting a higher-quality target reconstructed image. This implementation combines the efficiency of analytical reconstruction with the optimization capabilities of neural networks, which helps to improve the clarity and contrast of the final image while maintaining high processing speed.

[0046] In some optional embodiments, an initial reconstructed image of the first energy level is obtained by reconstructing the first target projection data; an initial reconstructed image of the second energy level is obtained by reconstructing the second target projection data, including: the image processing system can perform the following reconstruction operations on the first target projection data and the second target projection data respectively: weighted sum filtering of the target projection data at each acquisition angle to obtain filtered projection data corresponding to the acquisition angle; back-projecting the filtered projection data corresponding to each acquisition angle along the ray propagation path to each voxel in the three-dimensional image space to obtain the back-projection result of the acquisition angle; and accumulating the back-projection results of all acquisition angles to generate the corresponding initial reconstructed image of the first energy level or the initial reconstructed image of the second energy level.

[0047] For example, the reconstruction operations performed by the image processing system on the first target projection data and the second target projection data can refer to a series of computational steps to convert the projection domain data into a spatial domain image. Here, the projection data can refer to the one-dimensional or two-dimensional intensity distribution recorded by the detector after the ray passes through the target object at each acquisition angle; weighting can refer to assigning different weights according to the projection quality or geometric position at different acquisition angles to compensate for non-uniform sampling or correct the redundancy of the projection data; filtering can refer to performing frequency domain or spatial domain transformation operations on the projection data to enhance high-frequency components to compensate for the blurring effect of the imaging system and suppress noise. Through weighting and filtering, the signal-to-noise ratio and edge sharpness of the projection data can be improved. By back-projecting the filtered projection data along the ray propagation path to each voxel in the three-dimensional image space and accumulating the back-projection results from all acquisition angles, a complete voxel attenuation coefficient distribution can be gradually constructed, thereby generating a three-dimensional image reflecting the internal structure of the target object. The image processing system independently performs the above reconstruction operations on the first target projection data and the second target projection data, generating initial reconstructed images of the first and second energy levels respectively, which is beneficial for providing basic input data for subsequent dual-energy fusion processing.

[0048] In some embodiments, the image processing system can employ angle-related weighting coefficients during weighted processing. For example, acquisition angles with high projection data quality, such as those with short ray penetration paths or minimal scattering, can be assigned larger weights; while acquisition angles with low projection data quality can be assigned smaller weights. The filtering process employs a filtering function combining smoothing and sharpening to suppress noise while preserving edge details. During backprojection, the image processing system pre-calculates the geometric path of the ray from the ray source to the detector pixel and determines the voxel index and intersection length of each ray. Then, the filtered projection values ​​are weighted and distributed to the corresponding voxels according to the intersection length. After all angle processing is completed, the image processing system normalizes the accumulated result for each voxel, which helps reduce the magnification differences caused by different numbers of angles, thereby generating an initial reconstructed image with a uniform response. This embodiment helps improve image contrast and geometric accuracy while maintaining reconstruction efficiency.

[0049] In other embodiments, the image processing system can employ a reconstruction workflow under cone-beam geometry. For each acquisition angle, the image processing system first performs cone-beam weighting on the projection data, for example, by adjusting the weights based on the angle between the ray and the rotation center axis, and then performs two-dimensional filtering. The filtered projection data is then distributed to each voxel along the ray path by calculating the intersection length of the ray and the voxel mesh or by interpolation. After accumulating the back-projection contributions from all angles for each voxel, the image processing system also performs normalization to reduce magnification differences caused by inconsistencies in the number of projection angles. This processing workflow is suitable for rapid 3D reconstruction under cone-beam acquisition geometry, achieving acceptable computational efficiency while ensuring reconstruction accuracy.

[0050] In some optional embodiments, generating a target reconstructed image based on a first energy level initial reconstructed image and a second energy level initial reconstructed image includes: an image processing system performing learnable energy spectrum basis function decomposition on the first energy level initial reconstructed image and the second energy level initial reconstructed image, wherein the energy spectrum basis function decomposition includes: generating a first material image and a second material image by performing a linear weighted summation on the first energy level initial reconstructed image and the second energy level initial reconstructed image; wherein the first material image is used to represent an image with prominent skeletal components, and the second material image is used to represent an image with prominent soft tissue components; combining the first energy level initial reconstructed image, the second energy level initial reconstructed image, the first material image, and the second material image into a four-channel input; and feeding the four-channel input into a trained dual-energy perception fusion reconstruction network, wherein the dual-energy perception fusion reconstruction network fuses energy feature information and material feature information through a dual-branch feature encoder and an attention fusion module to output the target reconstructed image.

[0051] For example, learnable energy spectrum basis function decomposition can refer to decomposing the initial reconstructed image of the first energy level and the initial reconstructed image of the second energy level into images with two different material components through linear transformation. This learnable energy spectrum basis function decomposition process can adopt two linear weighted summation methods: the first material image is equal to the product of the first coefficient multiplied by the initial reconstructed image of the first energy level plus the product of the second coefficient multiplied by the initial reconstructed image of the second energy level; the second material image is equal to the product of the third coefficient multiplied by the initial reconstructed image of the first energy level plus the product of the fourth coefficient multiplied by the initial reconstructed image of the second energy level. Here, the first material image is used to represent the image highlighting the skeletal component, and the second material image is used to represent the image highlighting the soft tissue component. The first, second, third, and fourth coefficients are learnable parameters. The initial values ​​of these four coefficients can be determined based on the physical attenuation characteristics of bone and soft tissue, and are jointly optimized end-to-end with the subsequent dual-energy sensing fusion reconstruction network during training. The image processing system combines the initial reconstructed image of the first energy level, the initial reconstructed image of the second energy level, the first material image, and the second material image into a four-channel input tensor along the channel dimension. This four-channel input tensor can be considered a multidimensional data structure formed by sequentially stacking these four images along the channel dimension. This four-channel input can simultaneously provide two perspectives for understanding the data: an energy perspective such as low-energy and high-energy decay distribution, and a material perspective such as the approximate distribution of bones and soft tissues. This facilitates subsequent network fusion and reconstruction by fully utilizing dual-energy information.

[0052] In some embodiments, the image processing system can feed four-channel input into a trained dual-energy sensing fusion reconstruction network. This network includes a dual-branch feature encoder. The energy branch of the dual-branch feature encoder can be used to receive the first and second initial reconstructed images from the four-channel input; the matter branch of the dual-branch feature encoder is used to receive the first and second matter images from the four-channel input. The energy and matter branches have the same structure but independent parameters, extracting multi-scale features from both energy and matter perspectives. An attention fusion module performs bidirectional cross-attention enhancement on the features of the energy and matter branches at each encoding level, achieving pixel-by-pixel information interaction between energy and matter features. The fused features are gradually restored to the original image size by the decoder, outputting the target reconstructed image. This embodiment can fully utilize the complementarity of energy and matter feature information, improving the detail clarity and matter discrimination capability of the reconstructed image.

[0053] In other embodiments, the image processing system can perform end-to-end optimization of the learnable energy spectrum basis function decomposition process. During the training phase, the image processing system incorporates the basis function decomposition coefficients along with the parameters of the dual-energy sensing fusion reconstruction network into gradient backpropagation. The dual-energy sensing fusion reconstruction network automatically adjusts the decomposition coefficients during training to ensure that the resulting first and second material images match the optimal target of the downstream fusion reconstruction task. After training, the image processing system fixes the decomposition coefficients and, during actual inference, quickly performs a linear weighted summation to obtain a four-channel input, which is then forward-propagated through the dual-energy sensing fusion reconstruction network to output the target reconstructed image.

[0054] In some optional embodiments, the learnable energy spectrum basis function decomposition employs an explicit four-parameter linear transformation, where the first material image is equal to the first coefficient multiplied by the initial reconstructed image of the first energy level plus the second coefficient multiplied by the initial reconstructed image of the second energy level; the second material image is equal to the third coefficient multiplied by the initial reconstructed image of the first energy level plus the fourth coefficient multiplied by the initial reconstructed image of the second energy level; wherein the first, second, third, and fourth coefficients are learnable parameters, and the initial values ​​of the first, second, third, and fourth coefficients are determined based on the physical attenuation characteristics of bone and soft tissue.

[0055] For example, an explicit four-parameter linear transformation can refer to using four learnable coefficients—a first coefficient, a second coefficient, a third coefficient, and a fourth coefficient—to perform a weighted summation of the initial reconstructed images of the first and second energy levels, respectively, to generate images of two different material compositions. For instance, the image of the first material is equal to the sum of the products of the first coefficient multiplied by the initial reconstructed image of the first energy level and the second coefficient multiplied by the initial reconstructed image of the second energy level; the image of the second material is equal to the sum of the products of the third coefficient multiplied by the initial reconstructed image of the first energy level and the fourth coefficient multiplied by the initial reconstructed image of the second energy level. The first, second, third, and fourth coefficients are learnable parameters, and their initial values ​​are pre-calculated based on the physical attenuation characteristics of bone and soft tissue. For example, physical attenuation characteristics can refer to the ratio of the attenuation coefficients of bone and soft tissue to low-energy and high-energy rays under typical energy spectra. By employing an explicit linear transformation, it is beneficial to maintain physical interpretability, meaning that the transformed image can roughly correspond to the distribution of bones and soft tissues. At the same time, the four coefficients can be optimized end-to-end during training to adapt to actual deviations caused by factors such as detector response and scattering.

[0056] For example, the learnable energy spectrum basis function decomposition can generate the first and second material images using the following linear transformations: First material image = First coefficient × Initial reconstructed image of the first energy level + Second coefficient × Initial reconstructed image of the second energy level; Second material image = Third coefficient × Initial reconstructed image of the first energy level + Fourth coefficient × Initial reconstructed image of the second energy level. Here, the first, second, third, and fourth coefficients are learnable parameters. The initial values ​​of the first, second, third, and fourth coefficients are calculated based on the physical attenuation characteristics of bone and soft tissue; for example, the first coefficient can be set to 1.47, the second coefficient to -0.52, the third coefficient to -0.38, and the fourth coefficient to 1.32. These learnable parameters are automatically optimized and adjusted during training along with the dual-energy sensing fusion reconstruction network.

[0057] Figure 3 This paper illustrates another overall flowchart of fast image reconstruction of a single-source dual-energy plate based on neural networks in an embodiment of this application, including: acquiring first projection data and second projection data; preprocessing the acquired first projection data and second projection data; performing pixel-by-pixel adaptive calibration through an energy spectrum response calibration sub-network; obtaining initial reconstructed images of the first energy level and the second energy level through analytical reconstruction; performing learnable energy spectrum basis function decomposition; introducing dose sensitivity weighted training during the training phase; and finally performing fusion reconstruction through a dual-energy sensing fusion reconstruction network to output the target reconstructed image.

[0058] In some embodiments, the image processing system can perform the four-parameter linear transformation immediately after the initial reconstructed images of the first and second energy levels are generated. The image processing system reads the current values ​​of the four learnable coefficients at the current moment and calculates the pixel values ​​of the first and second material images for each pixel location. Since the transformation is a linear weighted summation, the calculation speed is fast, which helps maintain the high efficiency of the entire image processing flow and meets the needs of real-time or near-real-time image reconstruction. During the training phase, the image processing system uses the four coefficients as the parameters to be optimized in the neural network and updates them together with other network parameters through the backpropagation algorithm. In the early stage of training, the four coefficients can be set to theoretical initial values ​​calculated from the physical decay characteristics of bones and soft tissues, for example, based on the mass decay coefficients of two substances at low and high energies, obtained by solving a system of two linear equations. As training progresses, the coefficients gradually deviate from the theoretical values, but always maintain a linear relationship, so that the decomposed material image still has a clear physical meaning, that is, reflects a certain equivalent material distribution. This embodiment can combine physical priors with data-driven optimization, which is beneficial to improving the decomposition accuracy.

[0059] In other embodiments, the image processing system can also allow different initialization strategies for the coefficients of the four-parameter linear transformation in different application scenarios, depending on the specific clinical task objectives. For example, for reconstruction tasks that require highlighting bone contrast, the initial values ​​of the first and second coefficients can be biased towards enhancing the difference between bone and soft tissue; for tasks requiring detailed visualization of soft tissue, a different initialization combination can be used. The image processing system takes the task type as an additional input and adaptively adjusts the range of initial coefficient values ​​based on the task label during training. After training, the image processing system selects the corresponding coefficient set for inference based on the actual task. This embodiment is beneficial for further enhancing the flexibility and specificity of learnable basis function decomposition, and for achieving better material decomposition results under different clinical needs.

[0060] In some optional embodiments, since the coefficient matrix always maintains a linear transformation form, even if the optimized coefficients deviate from the strict theoretical physical values, the corresponding equivalent material information can still be obtained by reverse calculation. For example, the low-energy map and high-energy map reflect the absorption intensity of the same target object at different energy levels, while the first material image and the second material image can reflect the approximate composition of that part. After combining the low-energy map, high-energy map, first material image, and second material image into a four-channel input, the dual-energy sensing fusion reconstruction network simultaneously obtains two perspectives of understanding data: energy perspective and material perspective. This allows the network to make more accurate judgments and reconstructions by making fuller use of dual-energy information.

[0061] In some optional embodiments, the dual-energy sensing fusion reconstruction network includes a dual-branch feature encoder, which comprises an energy branch and a matter branch; the energy branch is used to receive the first energy level initial reconstructed image and the second energy level initial reconstructed image from the four-channel input; the matter branch is used to receive the first matter image and the second matter image from the four-channel input; the energy branch and the matter branch have the same structure but independent parameters, and both adopt a multi-level progressively reduced coding structure to extract multi-scale features of the image.

[0062] For example, a dual-branch feature encoder can refer to two independent feature extraction paths in a dual-energy sensing fusion reconstruction network used to process energy feature information and matter feature information respectively. This dual-branch feature encoder includes an energy branch and a matter branch. The energy branch receives the first and second initial reconstructed images of the first and second energy levels from the four-channel input, i.e., a low-energy decay distribution image and a high-energy decay distribution image; the matter branch receives the first and second matter images from the four-channel input, such as a skeletal approximation image and a soft tissue approximation image. The energy and matter branches can adopt the same network structure, for example, both using a four-level progressively shrinking convolutional coding structure, but the convolutional kernel parameters of the energy and matter branches are independent, each learning how to extract features from different inputs. Each branch, through a multi-level progressively shrinking coding structure, can extract multi-scale image features from local details to global semantics layer by layer, providing rich feature representations for subsequent cross-branch information fusion. The independent design of the energy and matter branches helps maintain the independence of energy and matter feature information, avoiding premature mixing of information in the early stages of feature extraction.

[0063] In some embodiments, both the energy branch and the matter branch can employ a four-level coding structure, with each level containing a convolutional layer, a batch normalization layer, and an activation function. The first level maintains the original image resolution and is used to extract high-frequency detail features, such as edges and textures. The second level halves the feature map size through convolution or pooling operations with a stride of 2, and is used to extract mid-scale structural features, such as organ boundaries. The third level further reduces the size to extract a wider range of anatomical structural features, such as the overall morphology of organs. The fourth level compresses the feature map to its minimum size to extract global contextual features, such as the overall decay distribution trend of the target object. The image processing system saves the multi-scale features output from each level of the energy branch and the matter branch separately for use by the subsequent attention fusion module. This multi-level coding structure allows the network to simultaneously capture both detailed and semantic information of the image, improving the quality of the reconstructed image.

[0064] In other embodiments, while maintaining isomorphic structures, the energy branch and the matter branch allow for adjustments to the specific configuration of each layer based on the characteristics of the input data. For example, the number of channels in the convolutional layers of the energy branch can be set to the same as that of the matter branch, or different numbers of channels can be set according to the information content differences between the energy image and the matter image, while maintaining symmetry in the overall number of encoding levels and the compression ratio of each level. During training, the parameters of the energy branch and the matter branch are updated independently. The energy branch can focus on learning feature patterns in the energy domain, such as the decay variation law under different energies, while the matter branch focuses on learning feature patterns in the matter domain, such as the spatial distribution law of bones and soft tissues. After training, the image processing system runs the energy branch and the matter branch in parallel during actual inference, extracting features separately, and then feeding them into the subsequent attention fusion module for interaction. This embodiment allows for independent parameter optimization while maintaining the symmetry of the branch structure, which is beneficial for fully leveraging the respective representational capabilities of the energy branch and the matter branch.

[0065] In some optional embodiments, the dual-energy sensing fusion reconstruction network further includes an attention fusion module, which performs the following operations at each scale level of the encoder: adding learnable energy spectral position codes to the energy branch features extracted from the energy branch and the material branch features extracted from the material branch, respectively. The energy spectral position codes are used to identify the energy channel type or material channel type corresponding to the feature; performing bidirectional cross-attention enhancement on each pixel position in the initial reconstruction image of the first energy level and the initial reconstruction image of the second energy level, including: performing a cross-attention query from the energy branch features to the material branch features to enable the energy branch features to obtain relevant information about the corresponding pixel position in the material branch features; performing a cross-attention query from the material branch features to the energy branch features to enable the material branch features to obtain relevant information about the corresponding pixel position in the energy branch features; and concatenating the energy branch features and material branch features after the bidirectional cross-attention enhancement, and compressing them into a unified fusion feature through convolution.

[0066] For example, the attention fusion module can refer to the main component in a dual-energy sensing fusion reconstruction network used to achieve fine interaction between energy features and material features. This attention fusion module performs the following operations at each scale level of the encoder: First, it adds learnable energy spectral position codes to the energy branch features extracted from the energy branch and the material branch features extracted from the material branch, respectively. The energy spectral position code can refer to a trainable parameter tensor of the same size as the feature map, used to identify the energy type (e.g., low energy or high energy) or material type (e.g., bone or soft tissue) corresponding to each feature channel. Then, it performs bidirectional cross-attention enhancement at each pixel location: on the one hand, the energy branch features query the material branch features, calculate the similarity weights between the energy branch features and the material branch features at the corresponding pixel location, and weight the relevant information in the material branch features to enhance the energy branch features; on the other hand, the material branch features query the energy branch features to enhance the material features in a similar manner. Finally, the bidirectional cross-attention enhanced energy branch features and material branch features are concatenated along the channel dimension and compressed into a unified fusion feature through a convolutional layer. By adding energy spectrum position encoding to the energy branch features and the matter branch features, the dual-energy sensing fusion reconstruction network can clearly distinguish the physical meaning of different input channels; through bidirectional cross attention, the energy branch features and the matter branch features can achieve point-by-point information exchange at each pixel position, thereby enabling the dual-energy sensing fusion reconstruction network to make more accurate judgments by utilizing both energy feature information and matter feature information.

[0067] In some embodiments, the image processing system can set an attention fusion module at each scale level of the encoder. For the energy branch features and matter branch features of the current level, the image processing system first adds energy spectral location encoding to the energy branch features and matter branch features. The energy spectral location encoding uses a trainable matrix with the same size as the feature map; the encoding matrix for the energy branch is used to identify energy features, and the encoding matrix for the matter branch is used to identify matter features. After adding the encoding, the image processing system calculates cross-attention: it calculates the correlation matrix between the energy branch features and the matter branch features in spatial location, and weights and aggregates the information with high correlation in the matter branch features into the energy branch features; at the same time, it symmetrically calculates the attention from the matter branch features to the energy branch features. After the bidirectional attention calculation is completed, the image processing system concatenates the energy branch features and the matter branch features, and compresses the number of channels to the same number of input feature channels through a 1×1 convolutional layer so that it can be input to the next layer encoder. This embodiment realizes fine-grained cross-branch information interaction at each level, which is beneficial for fusing energy and matter information at multiple scales.

[0068] In other embodiments, the image processing system employs a multi-head attention mechanism in its attention fusion module. Specifically, the system projects energy branch features and material branch features to multiple different subspaces (i.e., multiple heads), independently calculates bidirectional cross-attention within each subspace, and then concatenates the attention results from multiple heads before outputting them through a linear transformation. Multi-head attention can simultaneously capture various correlation patterns between energy branch features and material branch features from different representation subspaces, such as large-scale anatomical structure correspondences and local detail correspondences, enhancing the expressive power of feature fusion. After completing the multi-head attention calculation, the image processing system also performs feature concatenation and channel compression to obtain unified fused features. This embodiment can enrich the diversity of attention interactions without increasing the feature map size, which is beneficial for improving the modeling ability of dual-energy sensing fusion reconstruction networks for complex anatomical structures.

[0069] In some optional embodiments, the dual-energy sensing fusion reconstruction network further includes a decoder. The decoder adopts a multi-level progressive amplification structure symmetrical to the dual-branch feature encoder, and splices the feature maps output by each level of the encoder with the feature maps of the corresponding level of the decoder through skip connections. The dual-energy sensing fusion reconstruction network has three outputs: the first output is the target reconstructed image; the second output is a first energy level enhanced image and a second energy level enhanced image, where the first energy level enhanced image is used to represent the image after enhancing the initial reconstructed image at the first energy level, and the second energy level enhanced image is used to represent the image after enhancing the initial reconstructed image at the second energy level; the third output is a first matter decomposition image and a second matter decomposition image, where the first matter decomposition image is used to represent the image obtained after optimizing the first matter image by matter decomposition, and the second matter decomposition image is used to represent the image obtained after optimizing the second matter image by matter decomposition.

[0070] For example, the decoder can employ a multi-level progressive upscaling structure symmetrical to the dual-branch feature encoder to progressively restore the fused feature map to the size of the original input image. The decoder can use skip connections to concatenate the feature maps output from each level of the encoder with the feature maps of the corresponding levels of the decoder along the channel dimension. Skip connections can refer to directly passing high-resolution detail features saved by the encoder during downsampling to the same level of the decoder to compensate for edge and texture information lost during upsampling.

[0071] For example, the dual-energy sensing fusion reconstruction network can have three outputs: the first output is the target reconstructed image, which is a high-quality reconstruction result after fusion optimization; the second output includes a first-energy-level enhanced image and a second-energy-level enhanced image, where the first-energy-level enhanced image is the result of enhancing the initial reconstructed image at the first energy level, and the second-energy-level enhanced image is the result of enhancing the initial reconstructed image at the second energy level; the third output includes a first-material decomposition image and a second-material decomposition image, where the first-material decomposition image is the result of optimizing the material decomposition of the first material image, and the second-material decomposition image is the result of optimizing the material decomposition of the second material image. The three outputs share the fusion features extracted by the decoder, but generate their respective results through different output convolutional layers.

[0072] In some embodiments, the decoder can employ a four-stage progressively enlarged structure symmetrical to the encoder. Each stage of the decoder first doubles the size of the feature map through an upsampling operation, which can be, but is not limited to, bilinear interpolation or transposed convolution. Then, the feature map output from the corresponding level of the encoder is concatenated with the upsampled feature map along the channel dimension, followed by feature fusion and dimensionality reduction through convolutional layers. Through skip connections, the detailed information such as organ edges and microstructures preserved by the encoder in early levels can directly assist the decoder in recovering fine structures. Starting from the final output feature map of the decoder, the image processing system generates the target reconstruction image, the dual-energy enhanced image, and the material decomposition image through three independent output convolutional layers, respectively. This embodiment is beneficial for preserving local details while recovering the global structure, thus improving the quality of the final output image.

[0073] In other embodiments, consistency constraints can be introduced among the three outputs of the dual-energy sensing fusion reconstruction network. During network training, in addition to monitoring the differences between each output and its corresponding label, the image processing system can also additionally calculate the structural similarity loss between the first-energy-level augmented image and the target reconstructed image, and the feature consistency loss between the matter decomposition image and the target reconstructed image. Through the network's multi-task learning mechanism, the image processing system enables the shared decoder to learn a unified feature representation capable of simultaneously supporting the three output tasks. During actual inference, the image processing system can select one or more of the output results for subsequent use. This embodiment leverages the mutual constraints between multiple tasks, allowing the network to optimize the target reconstructed image while simultaneously ensuring the accuracy of the dual-energy augmented image and the matter decomposition image, thereby improving the prediction accuracy of each output result.

[0074] In some optional embodiments, the image processing system can perform post-processing operations on the reconstruction results before outputting the target reconstructed image. These post-processing operations include window width and level adjustment and format conversion. The image processing system supports outputting one or more of the following results: the target reconstructed image, the enhanced first energy level image and second energy level image, the optimized first and second material decomposition images, and the coefficient matrix optimized by learnable basis function decomposition. The coefficient matrix can be used to infer the corresponding equivalent material information. The above output results can be used for positioning verification, dose calculation, or treatment verification in radiotherapy.

[0075] In some optional embodiments, the energy spectrum response calibration subnetwork is a multilayer convolutional neural network. The input of the energy spectrum response calibration subnetwork is a two-channel two-dimensional image composed of first projection data and second projection data, and the output of the energy spectrum response calibration subnetwork is a two-channel two-dimensional image composed of first target projection data and second target projection data.

[0076] For example, the input to the spectral response calibration subnetwork is a two-channel two-dimensional image composed of first and second projection data, where each channel corresponds to the projection data output by a detector layer. The output of the spectral response calibration subnetwork is a two-channel two-dimensional image composed of first and second target projection data, where the two output channels represent the calibrated low-energy and high-energy projection distributions, respectively. This spectral response calibration subnetwork adopts a multi-layer convolutional structure, and the pixel-by-pixel nonlinear transformation facilitates the mapping from the original projection with crosstalk to the ideal projection without crosstalk. The image processing system can input the acquired projection data into the spectral response calibration subnetwork in angular order. The spectral response calibration subnetwork independently performs calibration operations for each angle and each pixel position, and outputs calibrated projection data.

[0077] In some embodiments, the spectral response calibration subnetwork can be composed of five convolutional layers stacked sequentially. The first convolutional layer maps the input two-channel image to an intermediate feature map, for example, 32 channels; the second convolutional layer further extracts higher-level features, for example, 64 channels; the third and fourth convolutional layers continue to increase the number of channels, for example, 128 channels, to enhance the expressive power of the spectral response calibration subnetwork; the fifth convolutional layer maps the feature map back to the two-channel output. Each convolutional layer is followed by a batch normalization layer and an activation function, such as a modified linear unit, to accelerate convergence and increase nonlinearity. The image processing system feeds the first projection data and the second projection data into the five-layer convolutional network one by one according to the acquisition angle, and the five-layer convolutional network outputs the first target projection data and the second target projection data corresponding to the angle. The spectral response calibration subnetwork of this embodiment has a simple structure, a small number of parameters, and a fast computation speed, which is beneficial to meeting the clinical needs for real-time image processing.

[0078] In other embodiments, the spectral response calibration subnetwork employs a residual connection structure. For example, the spectral response calibration subnetwork adds skip connections between multiple convolutional layers, directly adding the input to the output. This allows the spectral response calibration subnetwork to learn the residual mapping between the input and output rather than the complete mapping, helping to alleviate the gradient vanishing problem in deep network training and making it easier for the spectral response calibration subnetwork to learn small changes from crosstalk-containing projections to crosstalk-free projections. During the training of the spectral response calibration subnetwork in the image processing system, the residual connections enable gradients to propagate more smoothly back to shallower layers, thereby accelerating convergence and improving calibration accuracy. After training, the image processing system uses the residual network to calibrate the projection data during actual inference, outputting first target projection data and second target projection data. This embodiment is beneficial for improving the trainability and calibration effect of the spectral response calibration subnetwork, and is particularly suitable for scenarios where the degree of crosstalk varies significantly with spatial location.

[0079] In some alternative embodiments, the dual-layer detector includes a front scintillator on the side closer to the radiation source and a rear scintillator on the side farther from the radiation source. The front scintillator is sensitive to photons of a first energy level and is made of cesium iodide, while the rear scintillator is sensitive to photons of a second energy level and is made of gadolinium oxysulfate.

[0080] For example, the front scintillator can be made of cesium iodide (CsI), which has high sensitivity to low-energy photons and can effectively absorb low-energy photons and convert them into visible light signals. The back scintillator can be made of gadolinium oxysulfide (GOS), which has high sensitivity to high-energy photons and can effectively absorb high-energy photons that have penetrated the front layer and convert them into visible light signals. The image processing system reads the first projection data through the detector layer corresponding to the front scintillator and the second projection data through the detector layer corresponding to the back scintillator, thereby acquiring both low-energy and high-energy projection information in a single exposure.

[0081] For example, Figure 4 A schematic diagram of a single-source dual-energy plate imaging system is shown. The rays emitted from the single source penetrate the object being scanned and are received by a dual-layer detector. The front scintillator, located closer to the ray source, is sensitive to photons at the first energy level and outputs first projection data; the rear scintillator, located further away from the ray source, is sensitive to photons at the second energy level and outputs second projection data. Figure 4 Energy spectral crosstalk in the diagram refers to the phenomenon where a second energy level component is mixed into the preceding signal, while a first energy level component remains in the following signal. The system acquires first and second projection data from different viewpoints through multi-angle projection, providing input for subsequent energy spectral calibration and image reconstruction.

[0082] In some embodiments, the front scintillator can be a cesium iodide columnar crystal structure. The cesium iodide columnar crystal grows along the beam propagation direction, and the columnar structure helps guide visible light along the columnar direction to the photodiode array below, reducing lateral light diffusion and thus improving spatial resolution. The thickness of the front scintillator can be configured to absorb most low-energy photons while allowing high-energy photons to penetrate. The rear scintillator is a gadolinium oxysulfate particle coating. The gadolinium oxysulfate particles are fixed to the substrate with an adhesive, forming a thicker coating to absorb high-energy photons that have penetrated the front layer. The image processing system reads the electrical signals generated by the front and rear layers respectively to obtain first projection data and second projection data. This embodiment utilizes the complementary absorption characteristics of cesium iodide and gadolinium oxysulfate in different energy bands, which is beneficial for achieving simultaneous acquisition of dual-energy information.

[0083] In other embodiments, the front scintillator can be cesium iodide doped with thallium to enhance visible light conversion efficiency. The rear scintillator is gadolinium oxysulfide doped with praseodymium to provide a faster light decay rate, suitable for scenarios requiring rapid imaging. A transparent isolation layer can be placed between the front and rear scintillators to prevent visible light generated by the front scintillator from entering the rear detector layer and causing optical crosstalk. The image processing system separates and corrects the front and rear signals during post-processing, which helps reduce potential optical crosstalk. This embodiment, by optimizing the scintillator doping composition and adding an isolation layer, helps improve the signal-to-noise ratio and imaging speed of the dual-energy detector.

[0084] In some optional embodiments, the energy spectrum response calibration subnetwork is trained by generating a training dataset using a Monte Carlo simulation method, wherein the training dataset includes projected data with energy spectrum crosstalk and corresponding projected data without energy spectrum crosstalk; and using the projected data with energy spectrum crosstalk as training input data and the projected data without energy spectrum crosstalk as training labels to perform supervised training on the energy spectrum response calibration subnetwork.

[0085] For example, the training of the energy spectrum response calibration subnetwork can utilize Monte Carlo simulation to generate a training dataset. Monte Carlo simulation can be a numerical computation method based on random sampling statistics, capable of simulating the physical processes of ray-matter interaction. The training dataset includes projection data with energy spectrum crosstalk and corresponding projection data without energy spectrum crosstalk. The projection data with energy spectrum crosstalk can simulate the front and back layer signals output by an actual dual-layer detector imaging system, which may contain energy mixing components between the front and back layers; the projection data without energy spectrum crosstalk can simulate pure low-energy projections and pure high-energy projections under ideal conditions where no crosstalk exists. The image processing system uses the projection data with energy spectrum crosstalk as training input data and the projection data without energy spectrum crosstalk as training labels to supervise the training of the energy spectrum response calibration subnetwork. By minimizing the difference between the network output and the training labels, the energy spectrum response calibration subnetwork progressively learns the mapping law from crosstalk-containing projections to crosstalk-free projections.

[0086] In some embodiments, the Monte Carlo simulation method can be based on known digital human body models. The image processing system first constructs a set of digital human body models containing different tissues and body types, such as bones, soft tissues, and lung tissue. For each digital human body model, the image processing system simulates the ideal projection of an X-ray energy spectrum (e.g., tube voltage 120 kV) passing through the model, without energy spectrum crosstalk, as a training label; simultaneously, it simulates the output signal of the same X-ray energy spectrum after passing through a dual-layer detector, containing energy spectrum crosstalk, as the training input. The image processing system generates multiple sets (e.g., 3000 sets) of paired data, covering different scan geometries, different patient body types, and different anatomical locations. The image processing system uses these paired data to batch train the energy spectrum response calibration subnetwork, using a mean squared error loss function to measure the difference between the network output and the label. This embodiment facilitates the energy spectrum response calibration subnetwork learning to adapt to changes in crosstalk levels with patient body type and tissue pathways.

[0087] In other embodiments, the Monte Carlo simulation method is combined with actual phantom scan data for hybrid enhancement. The image processing system first generates a large amount of simulated paired data through Monte Carlo simulation to cover a diverse physical parameter space, such as different tube voltages, different energy thresholds, and different detector response characteristics. Subsequently, the image processing system scans phantoms of known materials using a real dual-layer detector imaging system, such as stepped phantoms with different thicknesses and equivalent atomic numbers, obtaining a small number of real paired projections with crosstalk and ideal projections calculated theoretically, for example, based on the known material composition and thickness of the phantom, using theoretical attenuation formulas to calculate the ideal projection. The image processing system merges the simulated and real data to form a training dataset and trains the energy spectrum response calibration subnetwork. This embodiment combines the richness of simulated data with the accuracy of real data, which is beneficial for improving the generalization ability of the energy spectrum response calibration subnetwork in practical clinical applications.

[0088] In some optional embodiments, the energy spectrum crosstalk satisfies the following relationship model: the first projection data equals the first calibration coefficient multiplied by the first target projection data plus the second calibration coefficient multiplied by the second target projection data; the second projection data equals the third calibration coefficient multiplied by the first target projection data plus the fourth calibration coefficient multiplied by the second target projection data; wherein, the first calibration coefficient is used to characterize the proportion coefficient of the first target projection data that is not affected by crosstalk, the fourth calibration coefficient is used to characterize the proportion coefficient of the second target projection data that is not affected by crosstalk; the second calibration coefficient is used to characterize the proportion coefficient of the second energy level component mixed in the first projection data, and the third calibration coefficient is used to characterize the proportion coefficient of the first energy level component remaining in the second projection data; the first calibration coefficient, the second calibration coefficient, the third calibration coefficient, and the fourth calibration coefficient vary with the thickness of the object being measured and the spatial distribution of the tissue composition, and different pixel positions correspond to different proportion coefficients.

[0089] For example, the first and second projection data are not ideal monoenergetic projections, but rather each contains signal components from the other's energy level. Therefore, a system of linear equations composed of four calibration coefficients is used to characterize the mixing relationship between the actual output and the ideal projection. By establishing this relationship model, the originally unknown and spatially variable crosstalk problem can be transformed into a pixel-by-pixel estimation problem of the four calibration coefficients. This allows the energy spectrum response calibration subnetwork to adaptively learn these calibration coefficients for each pixel location (corresponding to different ray penetration paths), thereby separating the pure low-energy and high-energy projection data from the mixed signal. Using this physical model to guide network design improves the accuracy and flexibility of crosstalk correction, avoids correction errors caused by using fixed coefficients, and provides more accurate input data for subsequent dual-energy matter decomposition and image reconstruction. The image processing system learns the spatial distribution of coefficients in this relationship model through the energy spectrum response calibration subnetwork to achieve pixel-by-pixel adaptive calibration.

[0090] For example, in a dual-layer detector, the relationship between the output of the front layer and the output of the back layer and the ideal pure low-energy projection and pure high-energy projection can be referenced as follows:

[0091] ;

[0092] ;

[0093] in, This represents the first projection data, i.e., the output from the previous layer; This represents the second projection data, i.e., the output from the subsequent layer; This represents the projection data of the first target, i.e., the pure low-energy projection; This represents the projection data of the second target, i.e., the pure high-energy projection; Indicates the first calibration factor; This represents the second calibration coefficient, i.e., the crosstalk coefficient, which can ideally be zero. This represents the third calibration coefficient, also known as the crosstalk coefficient, which can ideally be zero. This represents the fourth calibration factor.

[0094] In some embodiments, the image processing system can use Monte Carlo simulation to obtain the spatial distribution of calibration coefficients for different patient body shapes and different scan sites. For example, the image processing system establishes a set of digital human body models containing various anatomical structures such as soft tissues of different thicknesses and bones of different densities. For each digital human body model, the image processing system simulates both ideal crosstalk-free projection and actual crosstalk-included projection. By solving a system of linear equations for each pixel location, the image processing system derives the first, second, third, and fourth calibration coefficients corresponding to that pixel location. The image processing system uses the derived coefficient distribution as prior knowledge to assist in verifying the training effect of the energy spectrum response calibration subnetwork. This embodiment facilitates an intuitive understanding of the variation of crosstalk coefficients with spatial location, providing a theoretical basis for network design.

[0095] In other embodiments, the spectral response calibration subnetwork can implicitly learn the spatial distribution of the four calibration coefficients without explicitly outputting each coefficient. Through end-to-end supervised training, the spectral response calibration subnetwork directly learns the mapping relationship from the first and second projection data to the first and second target projection data. During training, the convolution kernel parameters within the spectral response calibration subnetwork can automatically encode the complex nonlinear relationships between the first, second, third, and fourth calibration coefficients at different pixel positions. During actual inference, the image processing system does not need to explicitly calculate each calibration coefficient; the network directly outputs the calibrated projection data. This embodiment helps avoid the large amount of computation required for explicitly solving the coefficients, improves calibration speed, and still adapts to the characteristic that crosstalk coefficients vary with spatial location.

[0096] In some optional embodiments, the four-channel input is fed into a trained dual-energy sensing fusion reconstruction network. The dual-energy sensing fusion reconstruction network fuses energy feature information and matter feature information through a dual-branch feature encoder and an attention fusion module to output a target reconstructed image. This includes: an image processing system dividing the three-dimensional image to be reconstructed into multiple two-dimensional slices along the cross-sectional direction, wherein the three-dimensional image to be reconstructed is three-dimensional image data composed of the initial reconstructed image of the first energy level and the initial reconstructed image of the second energy level; the four-channel input corresponding to each two-dimensional slice is fed into the dual-energy sensing fusion reconstruction network, which fuses energy feature information and matter feature information in each two-dimensional slice through a dual-branch feature encoder and an attention fusion module to obtain the reconstruction result of each two-dimensional slice; and the reconstruction results of all two-dimensional slices are stacked along the cross-sectional direction to recover the target reconstructed image.

[0097] For example, the cross-sectional direction can refer to the direction perpendicular to the long axis of the target object in the 3D image to be reconstructed, such as the axial vertical plane from head to toe in human imaging. Dual-energy sensing fusion reconstruction networks are computationally more efficient when processing on 2D slices, and the optimization of 2D convolution operations by graphics processors can be superior to that of 3D convolution. By decomposing the 3D reconstruction task into independent 2D slice reconstruction subtasks, the image processing system can feed the four-channel input data required for each slice into the dual-energy sensing fusion reconstruction network slice by slice. This dual-energy sensing fusion reconstruction network outputs the corresponding reconstruction result independently for each slice. After processing all slices, the slice results can be stacked along the cross-sectional direction to recover the target 3D image. This helps avoid the huge memory consumption and computational load caused by directly using 3D convolutional networks, enabling the network to handle large 3D volumes. Simultaneously, the 2D slices are independent of each other, facilitating parallel acceleration, which is beneficial for meeting the needs of rapid reconstruction in radiotherapy scenarios.

[0098] In some embodiments, the image processing system allows overlapping regions between adjacent 2D slices when segmenting 3D image data. For example, the image processing system selects slices at fixed intervals (e.g., every other slice), each slice's thickness comprising multiple original voxel layers, with the overlapping portion between adjacent slices used for smooth transition during subsequent stitching. The image processing system independently runs a dual-energy sensing fusion reconstruction network on each 2D slice to obtain the corresponding 2D reconstructed slice. During the stacking stage, the image processing system uses a weighted average fusion on the overlapping regions, for example, assigning weights according to their distance from the overlap boundary, to eliminate stitching seams. This embodiment helps reduce blocky artifacts that may be introduced by independent slice processing, improving the continuity and visual consistency of the 3D reconstruction results.

[0099] In other embodiments, the image processing system employs a three-dimensional sliding window strategy instead of one-time global segmentation. The system defines a fixed-size three-dimensional window within the three-dimensional volume and extracts voxel data within the window into several two-dimensional slices along the cross-sectional direction. After each two-dimensional slice is processed by a dual-energy sensing fusion reconstruction network, the system returns the output reconstructed slice to its original position within the window. The system traverses the entire three-dimensional volume with a fixed step size using a sliding window, accumulating and averaging the voxel values ​​in overlapping regions. This embodiment is suitable for scenarios with large three-dimensional volumes that cannot be stored in GPU memory all at once. By combining block processing with overlap averaging, it can reconstruct large-volume three-dimensional images with limited hardware resources while maintaining reconstruction quality.

[0100] See Figure 5 According to another aspect of the embodiments of this application, a training method for a neural network model for rapid image reconstruction of a single-source dual-energy plate is also provided, comprising: an image processing system acquiring projection data containing energy spectrum crosstalk and projection data without energy spectrum crosstalk; using the projection data containing energy spectrum crosstalk as training input data and the projection data without energy spectrum crosstalk as training labels, supervising training of the neural network to obtain an energy spectrum response calibration sub-network; the energy spectrum response calibration sub-network being used to perform pixel-by-pixel adaptive calibration of the input first projection data and second projection data, and outputting first target projection data and second target projection data; wherein, the first projection data corresponds to the signal output by the detector layer on the side closer to the X-ray source in the dual-layer detector, and the second projection data corresponds to the signal output by the detector layer on the side farther from the X-ray source in the dual-layer detector; the first target projection data being used to characterize the calibrated projection data representing the true attenuation distribution of photons at the first energy level, and the second target projection data being used to characterize the calibrated projection data representing the true attenuation distribution of photons at the second energy level; the first target projection data and the second target projection data being used to generate a target reconstruction image.

[0101] For example, training the spectral response calibration subnetwork requires acquiring projection data containing spectral crosstalk (i.e., signals containing energy mixtures collected by the actual detector) and projection data without spectral crosstalk (i.e., ideal, pure projections uncontaminated by crosstalk). By using projection data containing spectral crosstalk as training input and projection data without spectral crosstalk as training labels for supervised training, the spectral response calibration subnetwork can learn the nonlinear mapping law for separating pure low-energy and high-energy components from mixed signals from a large amount of paired data. This facilitates the spectral response calibration subnetwork to adaptively remove crosstalk for different pixel locations, and helps overcome the shortcomings of fixed-coefficient correction methods in adapting to changes in patient body size and tissue composition, thereby improving the accuracy of the final target reconstructed image and the accuracy of material decomposition.

[0102] In some embodiments, training data can be generated entirely through Monte Carlo simulation. The image processing system first constructs a set of digital human body models containing various anatomical structures such as different thicknesses and tissue proportions. For each digital human body model, the image processing system simulates the ideal projection of the X-ray source's energy spectrum after passing through the human body model, without energy spectrum crosstalk, as a training label; simultaneously, it simulates the output signal of the same energy spectrum after passing through the actual response of a dual-layer detector, containing energy spectrum crosstalk, as the training input. The image processing system generates thousands of paired data sets and uses the mean squared error loss function to batch train the energy spectrum response calibration sub-network. This embodiment can cover a wide range of physical parameter spaces, enabling the network to learn to adapt to changes in crosstalk levels with patient body shape and tissue pathways, and does not require a large amount of real scan data, thus reducing training costs.

[0103] In other embodiments, training data can be generated by combining Monte Carlo simulation with real phantom scanning. The image processing system first generates most of the simulated paired data through Monte Carlo simulation to cover diverse energy spectra and geometric parameters. Subsequently, a standard phantom with known material composition, such as a stepped phantom, is scanned using a real dual-layer detector imaging system to acquire a small amount of real projection data with crosstalk. The corresponding ideal projection is then calculated based on the phantom's material thickness and theoretical attenuation formula as training labels. The image processing system merges the simulated and real data to form a training dataset for supervised training of the energy spectrum response calibration subnetwork. This embodiment combines the richness of simulated data with the accuracy of real data, which is beneficial for improving the network's generalization ability in practical clinical applications and reducing the deviation between the simulation and real systems.

[0104] In some optional embodiments, the method further includes: acquiring simulation data and a reference image used when generating the simulation data, wherein the simulation data is an image reconstructed from projection data generated during the acquisition process of a simulated dual-energy plate imaging system after calibration by an energy spectrum response calibration subnetwork; training a dual-energy sensing fusion reconstruction network using the simulation data as training input data and the reference image as training labels, the dual-energy sensing fusion reconstruction network being used to reconstruct and fuse the input first target projection data and second target projection data, and outputting a target reconstruction image.

[0105] For example, simulation data can refer to projection data generated during the acquisition process of a dual-energy plate imaging system, calibrated by the energy spectrum response calibration subnetwork, and then reconstructed using initial data. The reference image can be a high-quality original image used when generating the simulation data, such as one selected from a known high-quality CT dataset. By simulating a real imaging process to generate a large amount of paired data—simulation images and reference images—the dual-energy sensing fusion reconstruction network can learn an end-to-end mapping from a coarse initial image to a high-quality reference image under supervised learning. This facilitates training the network using simulation without acquiring a large amount of real clinical paired data, enabling the dual-energy sensing fusion reconstruction network to reconstruct and fuse the input calibrated projection data, outputting a high-quality target reconstructed image.

[0106] In some embodiments, the image processing system can first select a batch of reference images from a publicly available high-quality CT image database. For each reference image, the image processing system simulates the physical processes of the dual-energy plate imaging system, including the X-ray source energy spectrum, detector response, and acquisition geometry, to generate projection data containing energy spectrum crosstalk. Next, the image processing system inputs this projection data into a pre-trained energy spectrum response calibration sub-network for calibration, obtaining calibrated projection data. The image processing system performs analytical reconstruction on the calibrated projection data, such as a cone-beam reconstruction algorithm, to obtain a simulated image. The image processing system uses the simulated image as training input and the reference image as training labels to perform supervised training on the dual-energy sensing fusion reconstruction network. This embodiment can generate large-scale, diverse training data, which is beneficial for improving the network's generalization ability and reconstruction accuracy.

[0107] In other embodiments, the image processing system combines Monte Carlo simulation with real phantom scan data for hybrid enhancement when generating simulation data. The system first generates most of the simulation-paired data through Monte Carlo simulation, covering different patient body types, anatomical locations, and energy spectrum parameters. Then, the system scans a standard phantom of known material using a real dual-layer detector imaging system to obtain real projection data. After calibration and initial reconstruction by the energy spectrum response calibration subnetwork, a simulation image is obtained, using a high-quality theoretical image or a high-precision CT scan image of the phantom as a reference image. The system merges the simulation data with the real data, using both as training data for the dual-energy sensing fusion reconstruction network. This embodiment combines the richness of simulation data with the accuracy of real data, which helps reduce the deviation between the simulation and real systems and improves the network's adaptability in clinical applications.

[0108] In some optional embodiments, simulation data is generated through the following steps: simulating the data acquisition process of a dual-layer detector imaging system based on a reference image to generate projection data containing energy spectrum crosstalk; inputting the projection data containing energy spectrum crosstalk into an energy spectrum response calibration subnetwork for pixel-by-pixel adaptive calibration to obtain first target projection data and second target projection data for training; reconstructing the first target projection data for training to obtain an initial reconstructed image of the first energy level for training, and reconstructing the second target projection data to obtain an initial reconstructed image of the second energy level for training; performing learnable energy spectrum basis function decomposition on the initial reconstructed images of the first and second energy levels for training, including: generating a first material image and a second material image for training by linearly weighted summing the initial reconstructed images of the first and second energy levels for training; combining the initial reconstructed images of the first and second energy levels for training, the first material image, and the second material image for training into a four-channel input as simulation data.

[0109] For example, the simulation data generation process can mimic the complete signal chain of a real dual-energy imaging system. Starting from a reference image, it goes through projection domain crosstalk addition, energy spectrum calibration, analytical reconstruction, and basis function decomposition to finally obtain four-channel input data corresponding to the real clinical process. Through a controllable simulation environment, a large amount of paired training data is provided to the dual-energy sensing fusion reconstruction network, simulating the four-channel input and reference image. This allows the network to learn the mapping relationship from a coarse initial image to a high-quality reference image, while avoiding dependence on a large amount of real clinical paired data. This facilitates the generation of diverse training samples covering different anatomical sites, patient body types, and scanning parameters, thereby improving the generalization ability and reconstruction accuracy of the dual-energy sensing fusion reconstruction network.

[0110] In some embodiments, the image processing system can use Monte Carlo simulation to generate projection data containing spectral crosstalk. For example, the image processing system selects a high-resolution reference image, such as one obtained from an existing high-quality CT database, and sets the X-ray spectral parameters and dual-layer detector response characteristics. The image processing system simulates X-ray penetration through a digital phantom corresponding to the reference image, generating the front and rear layer raw signals for each acquisition angle according to the dual-plate imaging geometry, i.e., projection data containing spectral crosstalk. Subsequently, the image processing system inputs this projection data into a pre-trained spectral response calibration sub-network for calibration, obtaining the first and second target projection data for training. The image processing system performs analytical reconstruction on the calibrated projection data, such as a cone-beam reconstruction algorithm, to obtain the first and second initial reconstructed images of the first and second energy levels for training. Next, the image processing system performs learnable spectral basis function decomposition on these two initial reconstructed images, generating the first and second material images for training, such as a bone approximation and a soft tissue approximation, through linear weighted summation. Finally, the image processing system combines the low-energy image, high-energy image, bone image, and soft tissue image into a four-channel input as simulation data. The simulation data generation process in this embodiment is conducive to fully reproducing the real clinical treatment path and to enabling the network to learn an input distribution consistent with actual applications.

[0111] In other embodiments, the image processing system can introduce data augmentation operations during the simulation data generation process. After generating projection data containing energy spectrum crosstalk based on the reference image, the image processing system can apply random noise such as Gaussian noise or Poisson noise to the projection data to simulate low-dose acquisition conditions, or perform random angle downsampling of the projection data to simulate sparse acquisition scenarios. Furthermore, the image processing system can also apply random elastic deformation to the initial reconstructed image to simulate geometric changes caused by different patient positions. After these augmentation operations, the image processing system continues to perform subsequent basis function decomposition and four-channel combination to generate diverse simulation data. This embodiment helps to make the dual-energy sensing fusion reconstruction network more adaptable to noise, angle loss, and geometric deformation, improving the reconstruction stability of the network under non-ideal acquisition conditions.

[0112] In some optional embodiments, the training process of the energy spectrum response calibration subnetwork and the dual-energy sensing fusion reconstruction network includes the following training stages: First stage: Training the energy spectrum response calibration subnetwork separately using projection data with energy spectrum crosstalk and corresponding projection data without energy spectrum crosstalk; Second stage: Fixing the parameters of the energy spectrum response calibration subnetwork, using simulation data and reference images as training data, and jointly training the coefficients of the dual-energy sensing fusion reconstruction network and the learnable basis function decomposition based on the standard loss function without dose sensitivity weighting; Third stage: After the training in the second stage is completed, jointly adjusting the energy spectrum response calibration subnetwork and the dual-energy sensing fusion reconstruction network based on the dose sensitivity weighted loss function.

[0113] For example, the three-stage training strategy refers to dividing the training process of the spectral response calibration sub-network and the dual-energy sensing fusion reconstruction network into three sequentially executed stages. Adopting a three-stage training strategy is beneficial for gradually decomposing complex optimization problems and avoiding convergence difficulties or mutual interference that may occur when multiple tasks are learned simultaneously. The first stage trains the spectral response calibration sub-network separately, allowing it to focus on separating clean low-energy and high-energy projections from crosstalk projections, unaffected by subsequent reconstruction and decomposition tasks. The second stage fixes the already trained spectral response calibration sub-network and trains only the dual-energy sensing fusion reconstruction network and basis function decomposition coefficients, allowing the main network to focus on learning the end-to-end mapping from a coarse initial image to a high-quality reference image and the optimal basis function decomposition coefficients under stable input data distribution conditions. The third stage, after the second stage training is completed, jointly fine-tunes the two networks. At this point, the networks have good basic performance, and introducing a dose-sensitivity weighted loss function can enhance the network's reconstruction accuracy in clinically critical areas such as the target area and organs at risk. Simultaneously, the spectral response calibration sub-network is fine-tuned to better suit the needs of the main network. This three-stage training strategy helps avoid gradient conflicts, improve training efficiency, reduce the risk of overfitting, and achieve better reconstruction quality and material decomposition accuracy than end-to-end one-time training.

[0114] In some embodiments, the first stage uses paired data (e.g., 3000 sets) generated by Monte Carlo simulation, with mean squared error as the loss function, to train the spectral response calibration sub-network until convergence. The second stage fixes the network parameters of the spectral response calibration sub-network and uses simulation data (e.g., 5000 sets of four-channel inputs and corresponding reference images), with a weighted combination of mean squared error and structural similarity loss as the standard loss function, to train the dual-energy sensing fusion reconstruction network and the four coefficients of the learnable basis function decomposition. The third stage resets the parameters of the spectral response calibration sub-network to trainable, and uses a dose-sensitivity weighted loss function, including a weighted combination of pixel loss, structural similarity loss, dual-energy consistency loss, and matter decomposition loss, to jointly fine-tune the entire model. This three-stage training in this embodiment helps ensure that the accuracy of crosstalk calibration is not affected by subsequent tasks, while also adapting the final model to the clinical needs of radiotherapy.

[0115] In some other embodiments, the second phase uses a standard loss function without dose sensitivity weighting, but introduces an additional validation set monitoring mechanism. The image processing system calculates the dose sensitivity weighted loss on the validation set after each training round, but does not use it for backpropagation. When the dose sensitivity weighted loss on the validation set no longer decreases for several consecutive rounds, the image processing system stops the second phase of training. When the dose sensitivity weighted loss function is enabled in the third phase, a smaller learning rate is used for fine-tuning to avoid destroying the general features learned in the second phase. This embodiment, by monitoring validation set metrics and reducing the fine-tuning learning rate, helps to steadily enhance the accuracy of clinically critical regions while maintaining the network's general reconstruction capabilities.

[0116] In some optional embodiments, the dose sensitivity weighted loss function uses a dose sensitivity map generated based on the target area and organ at risk contours as spatial weights; the dose sensitivity map is a Gaussian attenuation weight map generated based on the target area contours and organ at risk contours drawn by doctors in historical radiotherapy cases; wherein, the weight values ​​inside the target area and organ at risk are higher than the weight values ​​far away from the treatment area.

[0117] For example, in the dose sensitivity map, the weight values ​​of the target area and organs at risk are higher than those of areas far from the treatment area, and the weight values ​​decrease in a Gaussian manner from the boundary of the target area and organs at risk outwards. This is beneficial for training the dual-energy sensing fusion reconstruction network to prioritize areas that have a greater impact on the radiotherapy dose distribution, such as the tumor target area and adjacent important organs, while appropriately reducing the accuracy requirements for other non-critical areas. This helps to focus the network's learning ability on clinically critical areas, improve the reconstruction accuracy and material decomposition accuracy of these areas, and thus provide more reliable image data for subsequent dose calculation and target localization. At the same time, it avoids wasting network capacity in uniform optimization of the entire image, thereby improving training efficiency and adaptive performance.

[0118] In some embodiments, the dose sensitivity map can be generated based on structural contour files in a radiotherapy planning system. The image processing system reads the target region contour and organ-at-risk contour corresponding to each training sample and creates a weight matrix of the same size as the image in a three-dimensional image space. For voxels located within the target region or organ-at-risk area, the weight value is set to the highest value (e.g., 1.0). For voxels located outside the boundary, the image processing system calculates the Euclidean distance from the voxel to the nearest contour boundary and calculates the weight value according to a Gaussian function, such that the weight decreases with increasing distance (e.g., the weight decays to 0.5 at a distance of 5 mm and to 0.1 at a distance of 10 mm). The image processing system sets a non-zero lower limit weight (e.g., 0.01) for voxels far from the treatment area, which helps to maintain a basic optimization objective across the entire map. The dose sensitivity map generated in this embodiment can smoothly transition from high-weight regions to low-weight regions, which is beneficial for stable network convergence.

[0119] In other embodiments, dose sensitivity maps can be customized based on the dose distribution characteristics of different radiotherapy techniques. For example, for intensity-modulated radiotherapy (IMRT) or volumetric rotational IMRT, the image processing system can directly extract the prescription dose value for each voxel based on the dose distribution curve in the actual treatment plan, and normalize the prescription dose as the weight value of the dose sensitivity map. Regions with higher prescription doses have greater weights, while regions with zero prescription doses have a weight set to a preset lower limit. This embodiment can directly use the distribution of dose calculation results without relying on contour drawing, which can more accurately reflect the importance of each voxel in the treatment plan, thereby guiding the network to prioritize the optimization of image quality in high-dose regions.

[0120] In some optional embodiments, the dose sensitivity weighted loss function comprises a weighted combination of the following four loss functions: a pixel loss function, used to calculate the difference between the target reconstructed image and the reference image output by the dual-energy sensing fusion reconstruction network pixel by pixel; a structural similarity loss function, used to measure the similarity between the target reconstructed image and the reference image in terms of structure and contrast; a dual-energy consistency loss function, used to constrain the enhanced first energy level image and the second energy level image output by the dual-energy sensing fusion reconstruction network to maintain consistency with the target reconstructed image; and a matter decomposition loss function, used to constrain the consistency between the matter decomposition image output by the dual-energy sensing fusion reconstruction network and the reference matter image, wherein the reference matter image is the corresponding matter image decomposed from the reference image.

[0121] For example, the training of the dual-energy perception fusion reconstruction network can be constrained from multiple dimensions by weighted combination of loss functions. The pixel loss function can directly supervise the pixel-by-pixel difference between the target reconstructed image and the reference image, ensuring basic reconstruction accuracy; the structural similarity loss function can measure the similarity of images in brightness, contrast, and structure, guiding the network to maintain clear edge and texture information; the dual-energy consistency loss function can constrain the consistency between the enhanced low-energy and high-energy images and the fusion reconstructed image, which is beneficial for achieving synergistic optimization of energy information; the material decomposition loss function supervises the consistency between the material decomposition image and the reference material image, improving the accuracy of bone and soft tissue decomposition. The weighted combination of the four loss functions can comprehensively improve the pixel-level accuracy, structural fidelity, energy consistency, and material decomposition capability of the reconstructed image, enabling the network to simultaneously output high-quality reconstructed images, enhanced dual-energy images, and accurate material decomposition images, meeting the needs of radiotherapy image guidance for multiple types of output results.

[0122] See Figure 6 According to another aspect of the embodiments of this application, a fast image reconstruction device for a single-source dual-energy plate based on a neural network is also provided, comprising: an acquisition unit, configured to acquire first projection data and second projection data using a dual-layer detector imaging system, wherein the first projection data corresponds to the signal output by the detector layer on the side closer to the X-ray source in the dual-layer detector, and the second projection data corresponds to the signal output by the detector layer on the side farther from the X-ray source in the dual-layer detector; a data processing unit, configured to input the first projection data and second projection data into a trained energy spectrum response calibration subnetwork, and perform pixel-by-pixel adaptive calibration on the first projection data and second projection data through the energy spectrum response calibration subnetwork to obtain first target projection data and second target projection data; wherein the first target projection data is used to characterize the calibrated projection data representing the true attenuation distribution of photons at the first energy level, and the second target projection data is used to characterize the calibrated projection data representing the true attenuation distribution of photons at the second energy level; and an image reconstruction unit, configured to generate a target reconstruction image based on the first target projection data and the second target projection data.

[0123] Optionally, the pixel-by-pixel adaptive calibration in the data processing unit is used to remove energy spectrum crosstalk between the front and back layers of the dual-layer detector. The energy spectrum crosstalk includes the second energy level signal component mixed in the front layer signal of the dual-layer detector and the first energy level signal component remaining in the back layer signal of the dual-layer detector.

[0124] Optionally, the image reconstruction unit includes: a first reconstruction subunit, used to obtain an initial reconstructed image of a first energy level by reconstructing the projection data of a first target; a second reconstruction subunit, used to obtain an initial reconstructed image of a second energy level by reconstructing the projection data of a second target; and a fusion subunit, used to generate a target reconstruction image based on the initial reconstructed image of the first energy level and the initial reconstructed image of the second energy level.

[0125] Optionally, the first reconstruction subunit and the second reconstruction subunit each include: a projection processing module, used to perform weighted sum and filtering processing on the target projection data of each acquisition angle to obtain filtered projection data corresponding to the acquisition angle; a back projection module, used to back project the filtered projection data corresponding to each acquisition angle along the ray propagation path to each voxel in the three-dimensional image space to obtain the back projection result of the acquisition angle; and an accumulation module, used to accumulate the back projection results of all acquisition angles to generate the corresponding first energy level initial reconstruction image or second energy level initial reconstruction image.

[0126] Optionally, the fusion subunit includes: a basis function decomposition module, used to perform learnable energy spectrum basis function decomposition on the first energy level initial reconstructed image and the second energy level initial reconstructed image, wherein the energy spectrum basis function decomposition includes: generating a first material image and a second material image by linearly weighted summing the first energy level initial reconstructed image and the second energy level initial reconstructed image; wherein the first material image is used to represent the image of prominent skeletal components, and the second material image is used to represent the image of prominent soft tissue components; a channel combination module, used to combine the first energy level initial reconstructed image, the second energy level initial reconstructed image, the first material image and the second material image into a four-channel input; and a network inference module, used to feed the four-channel input into a trained dual-energy perception fusion reconstruction network, wherein the dual-energy perception fusion reconstruction network performs fusion processing on energy feature information and material feature information through a dual-branch feature encoder and an attention fusion module, and outputs the target reconstructed image.

[0127] Optionally, the basis function decomposition module employs an explicit four-parameter linear transformation. The first material image is equal to the first coefficient multiplied by the initial reconstructed image of the first energy level plus the second coefficient multiplied by the initial reconstructed image of the second energy level; the second material image is equal to the third coefficient multiplied by the initial reconstructed image of the first energy level plus the fourth coefficient multiplied by the initial reconstructed image of the second energy level. Here, the first, second, third, and fourth coefficients are learnable parameters, and the initial values ​​of the first, second, third, and fourth coefficients are determined based on the physical attenuation characteristics of bones and soft tissues.

[0128] Optionally, the dual-energy sensing fusion reconstruction network includes a dual-branch feature encoder, which comprises an energy branch and a matter branch; the energy branch is used to receive the first energy level initial reconstruction image and the second energy level initial reconstruction image from the four-channel input; the matter branch is used to receive the first matter image and the second matter image from the four-channel input; the energy branch and the matter branch have the same structure but independent parameters, and both adopt a multi-level progressively reduced coding structure to extract multi-scale features of the image.

[0129] Optionally, the dual-energy sensing fusion reconstruction network also includes an attention fusion module. The attention fusion module performs the following operations at each scale level of the encoder: adding learnable energy spectral position codes to the energy branch features extracted from the energy branch and the material branch features extracted from the material branch, respectively. These energy spectral position codes are used to identify the energy channel type or material channel type corresponding to the feature. It performs bidirectional cross-attention enhancement on each pixel location in the initial reconstructed image of the first energy level and the initial reconstructed image of the second energy level, including: performing a cross-attention query from the energy branch features to the material branch features to obtain relevant information about the corresponding pixel location in the material branch features; performing a cross-attention query from the material branch features to the energy branch features to obtain relevant information about the corresponding pixel location in the energy branch features; and concatenating the energy branch features and material branch features after bidirectional cross-attention enhancement, compressing them into a unified fusion feature through convolution.

[0130] Optionally, the dual-energy sensing fusion reconstruction network also includes a decoder. The decoder adopts a multi-level progressive amplification structure symmetrical to the dual-branch feature encoder. The feature maps output by each level of the encoder are spliced ​​with the feature maps of the corresponding levels of the decoder through skip connections. The dual-energy sensing fusion reconstruction network has three outputs: the first output is the target reconstructed image; the second output is the first energy level enhanced image and the second energy level enhanced image. The first energy level enhanced image is used to represent the image after enhancing the initial reconstructed image at the first energy level, and the second energy level enhanced image is used to represent the image after enhancing the initial reconstructed image at the second energy level; the third output is the first material decomposition image and the second material decomposition image. The first material decomposition image is used to represent the image obtained after optimizing the first material image by material decomposition, and the second material decomposition image is used to represent the image obtained after optimizing the second material image by material decomposition.

[0131] Optionally, the energy spectrum response calibration subnetwork in the data processing unit is a multi-layer convolutional neural network. The input of the energy spectrum response calibration subnetwork is a two-channel two-dimensional image composed of the first projection data and the second projection data, and the output of the energy spectrum response calibration subnetwork is a two-channel two-dimensional image composed of the first target projection data and the second target projection data.

[0132] Optionally, the dual-layer detector includes a front scintillator on the side closer to the radiation source and a rear scintillator on the side farther from the radiation source. The front scintillator is sensitive to photons of the first energy level and is made of cesium iodide, while the rear scintillator is sensitive to photons of the second energy level and is made of gadolinium oxysulfate.

[0133] Optionally, the energy spectrum response calibration subnetwork is trained in the following way: a training dataset is generated using the Monte Carlo simulation method, wherein the training dataset includes projected data with energy spectrum crosstalk and corresponding projected data without energy spectrum crosstalk; the energy spectrum response calibration subnetwork is trained under supervision using the projected data with energy spectrum crosstalk as the training input data and the projected data without energy spectrum crosstalk as the training labels.

[0134] Optionally, the energy spectrum crosstalk satisfies the following relationship model: the first projection data equals the first calibration coefficient multiplied by the first target projection data plus the second calibration coefficient multiplied by the second target projection data; the second projection data equals the third calibration coefficient multiplied by the first target projection data plus the fourth calibration coefficient multiplied by the second target projection data; wherein, the first calibration coefficient is used to characterize the proportion coefficient of the first target projection data that is not affected by crosstalk, the fourth calibration coefficient is used to characterize the proportion coefficient of the second target projection data that is not affected by crosstalk; the second calibration coefficient is used to characterize the proportion coefficient of the second energy level component mixed in the first projection data, and the third calibration coefficient is used to characterize the proportion coefficient of the first energy level component remaining in the second projection data; the first calibration coefficient, the second calibration coefficient, the third calibration coefficient, and the fourth calibration coefficient vary with the thickness of the measured object and the spatial distribution of the tissue composition, and different pixel positions correspond to different proportion coefficients.

[0135] Optionally, the network inference module includes: a slice segmentation submodule, used to segment the 3D image to be reconstructed into multiple 2D slices along the cross-sectional direction, wherein the 3D image to be reconstructed is 3D image data composed of the initial reconstructed image of the first energy level and the initial reconstructed image of the second energy level; a slice inference submodule, used to send the four-channel input corresponding to each 2D slice into the dual-energy sensing fusion reconstruction network, the dual-energy sensing fusion reconstruction network performs fusion processing on the energy feature information and material feature information in each 2D slice through a dual-branch feature encoder and an attention fusion module to obtain the reconstruction result of each 2D slice; and a stacking submodule, used to stack the reconstruction results of all 2D slices along the cross-sectional direction to recover the target reconstructed image.

[0136] See Figure 7According to another aspect of the embodiments of this application, a training device for a neural network model for rapid reconstruction of images from a single-source dual-energy plate is also provided, comprising: a data acquisition unit for acquiring projection data containing energy spectrum crosstalk and projection data without energy spectrum crosstalk; a first training unit for supervising training of the neural network using projection data containing energy spectrum crosstalk as training input data and projection data without energy spectrum crosstalk as training labels to obtain an energy spectrum response calibration sub-network; the energy spectrum response calibration sub-network for performing pixel-by-pixel adaptive calibration on the input first projection data and second projection data, and outputting first target projection data and second target projection data; wherein, the first projection data corresponds to the signal output by the detector layer on the side closer to the X-ray source in the dual-layer detector, and the second projection data corresponds to the signal output by the detector layer on the side farther from the X-ray source in the dual-layer detector; the first target projection data is used to characterize the calibrated projection data representing the true attenuation distribution of photons at the first energy level, and the second target projection data is used to characterize the calibrated projection data representing the true attenuation distribution of photons at the second energy level; the first target projection data and the second target projection data are used to generate a target reconstruction image.

[0137] Optionally, the training device further includes: a simulation data acquisition unit for acquiring simulation data and a reference image used when generating the simulation data, wherein the simulation data is an image reconstructed from projection data generated during the acquisition process of a simulated dual-energy plate imaging system after calibration by an energy spectrum response calibration subnetwork; and a second training unit for training a dual-energy sensing fusion reconstruction network using the simulation data as training input data and the reference image as training labels, wherein the dual-energy sensing fusion reconstruction network is used to reconstruct and fuse the input first target projection data and second target projection data, and output a target reconstruction image.

[0138] Optionally, the simulation data acquisition unit includes: a projection simulation subunit, used to simulate the data acquisition process of a dual-layer detector imaging system based on a reference image, generating projection data containing energy spectrum crosstalk; a calibration subunit, used to input the projection data containing energy spectrum crosstalk into an energy spectrum response calibration subnetwork for pixel-by-pixel adaptive calibration, obtaining first target projection data and second target projection data for training; a reconstruction subunit, used to reconstruct the first target projection data for training, obtaining an initial reconstructed image of the first energy level for training, and to reconstruct the second target projection data for training, obtaining an initial reconstructed image of the second energy level for training; a decomposition subunit, used to perform learnable energy spectrum basis function decomposition on the initial reconstructed image of the first energy level and the initial reconstructed image of the second energy level for training, including: generating a first material image and a second material image for training by performing a linear weighted summation on the initial reconstructed image of the first energy level and the initial reconstructed image of the second energy level for training; and a combination subunit, used to combine the initial reconstructed image of the first energy level, the initial reconstructed image of the second energy level, the first material image, and the second material image for training into a four-channel input as simulation data.

[0139] Optionally, the first training unit and the second training unit include the following sub-units: a first-stage training sub-unit, used to train the energy spectrum response calibration sub-network separately using projection data with energy spectrum crosstalk and corresponding projection data without energy spectrum crosstalk; a second-stage training sub-unit, used to fix the parameters of the energy spectrum response calibration sub-network, using simulation data and reference images as training data, and jointly training the coefficients of the dual-energy sensing fusion reconstruction network and the learnable basis function decomposition based on a standard loss function without dose sensitivity weighting; and a third-stage training sub-unit, used to jointly adjust the energy spectrum response calibration sub-network and the dual-energy sensing fusion reconstruction network based on a dose sensitivity weighted loss function after the second-stage training is completed.

[0140] Optionally, the dose sensitivity weighted loss function uses a dose sensitivity map generated based on the target area and organ at risk contours as spatial weights; the dose sensitivity map is a Gaussian attenuation weight map generated based on the target area contours and organ at risk contours drawn by doctors in historical radiotherapy cases; wherein, the weight values ​​inside the target area and organ at risk are higher than the weight values ​​far away from the treatment area.

[0141] Optionally, the third-stage training subunit includes a weighted combination of the following four loss modules: a pixel loss module, used to calculate the difference between the target reconstructed image and the reference image output by the dual-energy sensing fusion reconstruction network pixel by pixel; a structural similarity loss module, used to measure the similarity between the target reconstructed image and the reference image in terms of structure and contrast; a dual-energy consistency loss module, used to constrain the enhanced first-energy-level image and the second-energy-level image output by the dual-energy sensing fusion reconstruction network to maintain consistency with the target reconstructed image; and a matter decomposition loss module, used to constrain the consistency between the matter decomposition image output by the dual-energy sensing fusion reconstruction network and the reference matter image, wherein the reference matter image is the corresponding matter image obtained by decomposition from the reference image.

[0142] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, which stores a computer program, wherein when the computer program is executed, the device where the computer-readable storage medium is located executes the above-described method for fast reconstruction of single-source dual-energy plate images based on neural networks.

[0143] According to another aspect of the embodiments of this application, an electronic device is also provided, including one or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to execute the above-described neural network-based single-source dual-energy plate image fast reconstruction method.

[0144] According to another aspect of the embodiments of this application, a computer program product is also provided, including a computer program or instructions, which, when executed by a processor, implement the above-described method for rapid reconstruction of single-source dual-energy plate images based on neural networks.

[0145] The sequence numbers of the embodiments in this application are merely for description and do not represent the superiority or inferiority of the embodiments. In the above embodiments of this application, the descriptions of each embodiment have their own emphasis; parts not described in detail in a certain embodiment can be referred to in the relevant descriptions of other embodiments. It should be understood that the disclosed technical content in the several embodiments provided in this application can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Furthermore, the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces; the indirect coupling or communication connection of units or modules can be electrical or other forms.

[0146] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated units described above can be implemented in hardware or as software functional units.

[0147] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.

[0148] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A neural network-based single-source dual-energy panel image fast reconstruction method, characterized in that, include: A dual-layer detector imaging system is used to acquire first projection data and second projection data, wherein the first projection data corresponds to the signal output by the detector layer on the side closer to the X-ray source in the dual-layer detector, and the second projection data corresponds to the signal output by the detector layer on the side farther from the X-ray source in the dual-layer detector. The first projection data and the second projection data are input into a trained energy spectrum response calibration subnetwork. The first projection data and the second projection data are then subjected to pixel-by-pixel adaptive calibration by the energy spectrum response calibration subnetwork to obtain first target projection data and second target projection data. The first target projection data is used to characterize the calibrated projection data representing the true attenuation distribution of photons at the first energy level, and the second target projection data is used to characterize the calibrated projection data representing the true attenuation distribution of photons at the second energy level. The first energy level is lower than the second energy level. By reconstructing the projection data of the first target, an initial reconstructed image of the first energy level is obtained; by reconstructing the projection data of the second target, an initial reconstructed image of the second energy level is obtained. Learnable energy spectrum basis function decomposition is performed on the first energy level initial reconstructed image and the second energy level initial reconstructed image, wherein the energy spectrum basis function decomposition includes: generating a first material image and a second material image by performing a linear weighted summation on the first energy level initial reconstructed image and the second energy level initial reconstructed image; wherein the first material image is used to characterize the image of prominent skeletal components, and the second material image is used to characterize the image of prominent soft tissue components; The first energy level initial reconstruction image, the second energy level initial reconstruction image, the first matter image, and the second matter image are combined into a four-channel input and input to the dual-energy sensing fusion reconstruction network. The dual-energy sensing fusion reconstruction network fuses energy feature information and matter feature information through a dual-branch feature encoder and an attention fusion module to output the target reconstruction image.

2. The method according to claim 1, characterized in that, The pixel-by-pixel adaptive calibration is used to remove spectral crosstalk between the front and back layers of the dual-layer detector. The spectral crosstalk includes a second energy level signal component mixed into the front layer signal of the dual-layer detector and a first energy level signal component remaining in the back layer signal of the dual-layer detector.

3. The method according to claim 1, characterized in that, By reconstructing the first target projection data, an initial reconstructed image of the first energy level is obtained; by reconstructing the second target projection data, an initial reconstructed image of the second energy level is obtained, including: Perform the following reconstruction operations on the first target projection data and the second target projection data respectively: The target projection data at each acquisition angle is weighted and filtered to obtain the filtered projection data corresponding to that acquisition angle. The filtered projection data corresponding to each acquisition angle is back-projected along the ray propagation path to each voxel in the three-dimensional image space to obtain the back-projection result of that acquisition angle. The back projection results of all acquired angles are summed to generate the corresponding first-energy-level initial reconstructed image or second-energy-level initial reconstructed image.

4. The method according to claim 1, characterized in that, The learnable energy spectrum basis function decomposition employs an explicit four-parameter linear transformation. The first material image is equal to the first coefficient multiplied by the initial reconstructed image of the first energy level plus the second coefficient multiplied by the initial reconstructed image of the second energy level; the second material image is equal to the third coefficient multiplied by the initial reconstructed image of the first energy level plus the fourth coefficient multiplied by the initial reconstructed image of the second energy level; wherein, the first coefficient, the second coefficient, the third coefficient, and the fourth coefficient are learnable parameters, and the initial values ​​of the first coefficient, the second coefficient, the third coefficient, and the fourth coefficient are determined based on the physical attenuation characteristics of bone and soft tissue.

5. The method according to claim 1, characterized in that, The dual-energy sensing fusion reconstruction network includes a dual-branch feature encoder, which comprises an energy branch and a matter branch. The energy branch is used to receive the first energy level initial reconstruction image and the second energy level initial reconstruction image from the four-channel input. The matter branch is used to receive the first matter image and the second matter image from the four-channel input. The energy branch and the matter branch have the same structure but independent parameters, and both adopt a multi-level progressively shrinking coding structure to extract multi-scale features of the image.

6. The method according to claim 5, characterized in that, The dual-energy sensing fusion reconstruction network also includes an attention fusion module, which performs the following operations at each scale level of the encoder: Learnable energy spectrum position codes are added to the energy branch features extracted from the energy branch and the material branch features extracted from the material branch, respectively. The energy spectrum position codes are used to identify the energy channel type or material channel type corresponding to the feature. The bidirectional cross-attention enhancement of features is performed on each pixel location in the first energy level initial reconstructed image and the second energy level initial reconstructed image, including: performing a cross-attention query from the energy branch feature to the matter branch feature, so as to enable the energy branch feature to obtain relevant information of the corresponding pixel location in the matter branch feature; and performing a cross-attention query from the matter branch feature to the energy branch feature, so as to enable the matter branch feature to obtain relevant information of the corresponding pixel location in the energy branch feature. The energy branch features and material branch features, after completing the bidirectional cross-attention enhancement, are concatenated and compressed into a unified fusion feature through convolution.

7. The method according to claim 1, characterized in that, The dual-energy sensing fusion reconstruction network also includes a decoder, which adopts a multi-level progressive amplification structure symmetrical to the dual-branch feature encoder. It concatenates the feature maps output from each level of the encoder with the feature maps from the corresponding levels of the decoder through skip connections. The dual-energy sensing fusion reconstruction network has three outputs: The first output is the reconstructed image of the target; The second output is a first energy level enhanced image and a second energy level enhanced image. The first energy level enhanced image is used to characterize the image after enhancing the initial reconstructed image of the first energy level, and the second energy level enhanced image is used to characterize the image after enhancing the initial reconstructed image of the second energy level. The third output consists of a first material decomposition image and a second material decomposition image. The first material decomposition image is used to characterize the image obtained after material decomposition optimization of the first material image, and the second material decomposition image is used to characterize the image obtained after material decomposition optimization of the second material image.

8. The method according to claim 1, characterized in that, The energy spectrum response calibration subnetwork is a multi-layer convolutional neural network. The input of the energy spectrum response calibration subnetwork is a two-channel two-dimensional image composed of the first projection data and the second projection data. The output of the energy spectrum response calibration subnetwork is a two-channel two-dimensional image composed of the first target projection data and the second target projection data.

9. The method according to claim 1, characterized in that, The dual-layer detector includes a front scintillator on the side closer to the radiation source and a rear scintillator on the side farther from the radiation source. The front scintillator is sensitive to photons of a first energy level and is made of cesium iodide, while the rear scintillator is sensitive to photons of a second energy level and is made of gadolinium oxysulfate.

10. The method according to claim 1, characterized in that, The energy spectrum response calibration subnetwork was trained in the following manner: A training dataset is generated using the Monte Carlo simulation method, wherein the training dataset includes projected data with energy spectrum crosstalk and corresponding projected data without energy spectrum crosstalk; The energy spectrum response calibration subnetwork is trained under supervision using the projection data with energy spectrum crosstalk as training input data and the projection data without energy spectrum crosstalk as training labels.

11. The method according to claim 2, characterized in that, The energy spectrum crosstalk satisfies the following relationship model: The first projection data is equal to the first calibration coefficient multiplied by the first target projection data plus the second calibration coefficient multiplied by the second target projection data; The second projection data is equal to the third calibration coefficient multiplied by the first target projection data plus the fourth calibration coefficient multiplied by the second target projection data; Wherein, the first calibration coefficient is used to characterize the proportion coefficient of the first target projection data that is not affected by crosstalk, and the fourth calibration coefficient is used to characterize the proportion coefficient of the second target projection data that is not affected by crosstalk; the second calibration coefficient is used to characterize the proportion coefficient of the second energy level component mixed in the first projection data, and the third calibration coefficient is used to characterize the proportion coefficient of the first energy level component remaining in the second projection data; the first calibration coefficient, the second calibration coefficient, the third calibration coefficient, and the fourth calibration coefficient vary with the thickness of the object being measured and the spatial distribution of the tissue composition, and different pixel positions correspond to different proportion coefficients.

12. The method according to claim 1, characterized in that, The four-channel input is fed into a dual-energy sensing fusion reconstruction network. This network fuses energy and matter feature information using a dual-branch feature encoder and an attention fusion module, outputting the reconstructed target image, including: The three-dimensional image to be reconstructed is divided into multiple two-dimensional slices along the cross-sectional direction, wherein the three-dimensional image to be reconstructed is three-dimensional image data composed of the first energy level initial reconstructed image and the second energy level initial reconstructed image. The four-channel input corresponding to each two-dimensional slice is fed into the dual-energy sensing fusion reconstruction network. The dual-energy sensing fusion reconstruction network fuses the energy feature information and material feature information in each two-dimensional slice through a dual-branch feature encoder and an attention fusion module to obtain the reconstruction result of each two-dimensional slice. The reconstruction results of all two-dimensional slices are stacked along the cross-sectional direction to recover the target reconstructed image.

13. A training method for a neural network model for fast reconstruction of images from a single-source dual-energy plate, characterized in that, include: Acquire projection data with and without energy spectrum crosstalk; Using the projection data with energy spectrum crosstalk as training input data and the projection data without energy spectrum crosstalk as training labels, the neural network is trained under supervision to obtain the energy spectrum response calibration sub-network. The projection data containing energy spectrum crosstalk is input into the energy spectrum response calibration subnetwork for pixel-by-pixel adaptive calibration to obtain the first target projection data and the second target projection data used for training. The first target projection data used for training is reconstructed to obtain the first energy level initial reconstructed image used for training; the second target projection data is reconstructed to obtain the second energy level initial reconstructed image used for training. The learnable energy spectrum basis function decomposition of the first energy level initial reconstructed image and the second energy level initial reconstructed image used for training includes: generating a first material image and a second material image used for training by performing a linear weighted summation on the first energy level initial reconstructed image and the second energy level initial reconstructed image used for training. The initial reconstructed image of the first energy level, the initial reconstructed image of the second energy level, the first material image, and the second material image used in the training are combined into a four-channel input as simulation data, and a dual-energy sensing fusion reconstruction network is trained based on the simulation data.

14. The training method according to claim 13, characterized in that, The training process of the energy spectrum response calibration subnetwork and the dual-energy sensing fusion reconstruction network includes the following training stages: Phase 1: Using the projection data with energy spectrum crosstalk and the corresponding projection data without energy spectrum crosstalk, train the energy spectrum response calibration sub-network separately. Second stage: Fix the parameters of the energy spectrum response calibration subnetwork, use the simulation data and reference image as training data, and jointly train the coefficients of the dual-energy sensing fusion reconstruction network and the learnable basis function decomposition based on the standard loss function without dose sensitivity weighting. The third stage: After the training in the second stage is completed, the energy spectrum response calibration subnetwork and the dual-energy sensing fusion reconstruction network are jointly adjusted based on the dose sensitivity weighted loss function.

15. The training method according to claim 14, characterized in that, The dose sensitivity weighted loss function uses a dose sensitivity map generated based on the target area and organ at risk contours as spatial weights; the dose sensitivity map is a Gaussian attenuation weight map generated based on the target area contours and organ at risk contours drawn by doctors in historical radiotherapy cases; wherein, the weight values ​​inside the target area and organ at risk are higher than the weight values ​​far away from the treatment area.

16. The training method according to claim 14, characterized in that, The dose sensitivity weighted loss function comprises a weighted combination of the following four loss functions: A pixel loss function is used to calculate the difference between the target reconstructed image and the reference image output by the dual-energy sensing fusion reconstruction network on a pixel-by-pixel basis. A structural similarity loss function is used to measure the structural and contrast similarity between the target reconstructed image and the reference image; A dual-energy consistency loss function is used to constrain the enhanced first-energy-level image and second-energy-level image output by the dual-energy-sensing fusion reconstruction network to maintain consistency with the target reconstructed image. The material decomposition loss function is used to constrain the consistency between the material decomposition image output by the dual-energy sensing fusion reconstruction network and the reference material image, wherein the reference material image is the corresponding material image decomposed from the reference image.

17. A device for rapid image reconstruction from a single light source using a dual-energy plate based on a neural network, characterized in that, include: The acquisition unit is used to acquire first projection data and second projection data using a dual-layer detector imaging system. The first projection data corresponds to the signal output by the detector layer on the side closer to the X-ray source in the dual-layer detector, and the second projection data corresponds to the signal output by the detector layer on the side farther from the X-ray source in the dual-layer detector. The data processing unit is used to input the first projection data and the second projection data into a trained energy spectrum response calibration sub-network, and perform pixel-by-pixel adaptive calibration on the first projection data and the second projection data through the energy spectrum response calibration sub-network to obtain first target projection data and second target projection data; wherein, the first target projection data is used to characterize the projection data representing the true attenuation distribution of photons at the first energy level after calibration, and the second target projection data is used to characterize the projection data representing the true attenuation distribution of photons at the second energy level after calibration. An image reconstruction unit is configured to generate a target reconstruction image based on the first target projection data and the second target projection data. The unit includes: reconstructing the first target projection data to obtain a first energy level initial reconstruction image; reconstructing the second target projection data to obtain a second energy level initial reconstruction image; performing learnable energy spectrum basis function decomposition on the first and second energy level initial reconstruction images, wherein the energy spectrum basis function decomposition includes: generating a first material image and a second material image by performing a linear weighted summation on the first and second energy level initial reconstruction images; wherein the first material image is used to characterize images of prominent skeletal components, and the second material image is used to characterize images of prominent soft tissue components; combining the first energy level initial reconstruction image, the second energy level initial reconstruction image, the first material image, and the second material image into a four-channel input and inputting it into a dual-energy sensing fusion reconstruction network, wherein the dual-energy sensing fusion reconstruction network fuses energy feature information and material feature information through a dual-branch feature encoder and an attention fusion module, and outputs the target reconstruction image.

18. A training device for a neural network model for fast image reconstruction from a single-source dual-energy plate, characterized in that, include: The data acquisition unit is used to acquire projection data with energy spectrum crosstalk and projection data without energy spectrum crosstalk. The first training unit is used to supervise the training of the neural network by using the projection data with energy spectrum crosstalk as training input data and the projection data without energy spectrum crosstalk as training labels, so as to obtain the energy spectrum response calibration sub-network. The training device is also used to input the projection data containing energy spectrum crosstalk into the energy spectrum response calibration subnetwork for pixel-by-pixel adaptive calibration to obtain the first target projection data and the second target projection data used for training. The first target projection data used for training is reconstructed to obtain the first energy level initial reconstructed image used for training; the second target projection data is reconstructed to obtain the second energy level initial reconstructed image used for training. The learnable energy spectrum basis function decomposition of the first energy level initial reconstructed image and the second energy level initial reconstructed image used for training includes: generating a first material image and a second material image used for training by performing a linear weighted summation on the first energy level initial reconstructed image and the second energy level initial reconstructed image used for training. The initial reconstructed image of the first energy level, the initial reconstructed image of the second energy level, the first material image, and the second material image used in the training are combined into a four-channel input as simulation data, and a dual-energy sensing fusion reconstruction network is trained based on the simulation data.