Machine learning based lag correction for front-layer detectors in two-layer systems

The system addresses lag artifacts in X-ray imaging by using a back layer detector with lag reduction circuitry and a machine learning model to enhance spectral imaging accuracy.

JP2025539399APending Publication Date: 2025-12-05KONINKLIJKE PHILIPS NV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025530727
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-12-02
Filing Date
2023-11-30
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Lag artifacts in X-ray-based imaging, particularly in spectral imaging, result from delayed signal enhancement in detector pixels, leading to artifacts in 2D and 3D imaging, which existing hardware solutions fail to fully mitigate.

Method used

A system utilizing a back layer detector with lag reduction circuitry and a machine learning model processes input images from both front and back layer detectors to correct acquisition lag, employing a trained artificial neural network for improved accuracy.

Benefits of technology

The system effectively reduces acquisition lag artifacts, enhancing the quality of spectral imaging by accurately distinguishing between lag effects and spectral distribution, thereby improving material-specific imaging algorithms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025539399000001_ABST
    Figure 2025539399000001_ABST
Patent Text Reader

Abstract

The present invention relates to a system SYS and associated method for acquisition lag correction in X-ray imaging. The system has an input interface IN for receiving an input including an input image acquired by a front layer detector FLD in a multi-energy X-ray imaging device IA. A corrector component CC processes the input image into a corrected version that represents an estimate of image information at an acquisition lag lower than the acquisition lag experienced for the input image. Specifically, the effects of lag and associated artifacts are removed, and the corrected version is an estimate of the input image in an acquisition setup where such acquisition lag did not exist.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a system for acquisition lag correction in X-ray based imaging, an imaging device including such a system, a training system for training machine learning models for use in such a system, a training data generator system for generating training data for use in such a training system, related methods, computer program elements, and computer readable media. [Background technology]

[0002] For both 2D and 3D imaging, lag artifacts are a significant problem in X-ray-based imaging. They result from delayed signal enhancement in detector pixels, especially in detector regions that are in the shadow of attenuating objects and are subsequently illuminated by the direct X-ray beam. Another frequent source of lag artifacts is a sequence in which several images acquired at high X-ray doses are followed by low-dose images. In such situations, there is delayed signal spillover from the high-dose images that interferes with the low-dose images. Both lag effects lead to artifacts in 2D imaging, such as cone-beam CT (CBCT) volumes, and in tomographic 3D imaging.

[0003] Some such artifacts can be mitigated by applying hardware solutions, such as the backside reset light system described in T. Ducourant et al., "Latest advancements in state-of-the-art aSi-based X-ray Flat Panel Detectors," published in Proc. SPIE 10573, Medical Imaging (2018).

[0004] However, artifacts have been found to remain, especially in spectral imaging, which is multi-energy imaging in which multiple sets of projection data are acquired at different radiation energy ranges. These sets can be computationally combined to obtain images in 2D or 3D with material-specific contrast that are useful, among other things, in therapy and diagnosis. Summary of the Invention [Problem to be solved by the invention]

[0005] There may be a need for improved x-ray imaging, and in particular improved spectral imaging. [Means for solving the problem]

[0006] The objects of the present invention are achieved by the subject matter of the independent claims, with further embodiments being incorporated in the dependent claims. It is noted that the below-described aspects of the invention apply equally to an imaging device including the system, a training system for training a machine learning model for use in the system, a training data generator system for generating training data for use in the training system, related methods, computer program elements, and computer-readable media.

[0007] According to a first aspect of the present invention, an input interface for receiving an input including an input image acquired by a front layer detector in the multi-energy x-ray imaging device and an additional image acquired by a back layer detector in the multi-energy x-ray imaging device; A system for acquisition lag correction in X-ray imaging is provided, comprising: the back layer preferably having circuitry (such as a backside reset light or other hardware solution) for reducing time lag effects, which is preferably activated when an additional input image is acquired by the back layer detector, the back layer detector having circuitry for reducing time lag effects, the additional image being acquired by the back layer detector while said circuitry is activated.

[0008] Further, the system has a corrector component configured to process the input image and the additional input image to generate a corrected input image representing an estimate for the image information of the front layer detector at an acquisition lag that is smaller than the acquisition lag experienced for the input image.

[0009] Therefore, two-channel input data is provided to the corrector component, as explained below, which allows for better performance and especially higher accuracy.

[0010] In some embodiments, the processing by the corrector component includes scaling the input image by a scaling factor obtained based on a series of prior input images and, preferably, additional prior input images previously acquired by respective front and back layer detectors of the multi-energy X-ray imaging device.

[0011] In an embodiment, the corrector component is implemented based on a trained machine learning model.

[0012] In some embodiments, a machine learning ML model is previously trained based on a supervised or unsupervised learning setup.

[0013] In an embodiment, the trained ML model comprises an artificial neural network.

[0014] In an embodiment, an artificial neural network is configured for multi-scale processing. For example, an artificial neural network (NN) of U-net type architecture may be used. Multi-scale processing using filtering or other techniques, not necessarily NN type, is also contemplated herein. In an embodiment, the ML model is trained in a supervised setup based on training data including training input data and associated ground truth data, where the training input data includes prior images acquired by a front-layer detector and the ground truth data includes prior images acquired by said or other back-layer detector having circuitry for reducing time lag effects, with such circuitry activated.

[0015] Such circuitry for reducing the temporal lag effect may include, for example, the above-mentioned backside reset optical facility coupled to the distal surface of the back layer detector. As described further herein, in multi-layer detectors, the front layer detector generally cannot be provided with such lag correction circuitry. The training input data may include prior images acquired by a front layer detector (or a single layer detector) in which no lag correction circuitry is present (or, in the case of a single layer detector, present but not activated), and / or prior images acquired by a single layer detector in which a lag correction circuitry is present and activated. Thus, in effect, the training input data may include front layer prior image data having a temporal lag effect.

[0016] Thus, in an embodiment, a two-channel training data input is used, using front-layer and back-layer images. This allows for better accounting for artifacts: generally, observed artifacts are the result of at least two effects: i) lag caused by the way the detector signal is amplified with signal overspill from one detection event to another, and ii) the spectral distribution of the detected signal caused by the various materials of the imaged object and their interactions. The proposed two-channel ML approach represents a regularization scheme that facilitates the model learning to distinguish between the two effects i) and ii), eliminating or reducing i) and better preserving ii) in a purer form. This facilitates better spectral processing by the spectral imaging algorithm, as envisioned downstream. Such spectral imaging algorithms include material decomposition, decomposition into Compton and photoelectric images, and calculation of any one of virtual monoenergetic images, virtual contrast-only images, virtual non-contrast images, and any other where multi-energy information is required. The two-channel approach can also be used during deployment, post-training.

[0017] In an embodiment, the front layer detector is one of the single layer detector x-ray imaging devices, or the front layer detector is one of the above or other multi-energy x-ray imaging devices from which the back layer detector has been removed.

[0018] In particular, the effects of lag and associated artifacts are removed, and the corrected version is therefore an estimate for the input image in an acquisition setup where such acquisition lag was not present (or was negligible).

[0019] The proposed system addresses the problem of acquisition lag-based artifacts in projection images and in their reconstruction in the image domain. Unlike back-layer detectors, in multi-layer spectral detector systems (such as dual-layer detector systems), the front-layer detector modules cannot be equipped with the above-mentioned backside reset light or similar hardware-based lag artifact solutions because the X-ray shadows of their components, such as power lines, feed through into the associated back-layer X-ray images and cause their own artifacts there. The proposed system solves this problem computationally.

[0020] The system can be used in all imaging protocols and is expected to be most relevant in angiographic DSA protocols and CBCT scans with asymmetric or non-isocentric subjects where lag artifacts are most severe.

[0021] The system is not limited to use in dual-energy imaging, but can also be used in a spectral imaging setup of three or more layers, with only the last layer being equipped with a backside reset light or similar hardware-based lag error correction function.

[0022] The back layer of a dual-layer detector system, thanks to its hardware-based lag error correction function, records no or much less lag artifacts compared to the front layer, which gives a strong indication of whether the front layer data suffers from delayed signal enhancement or attenuation during temporal (video) 2D or 3D (e.g., CBCT) imaging.

[0023] In another aspect, an imaging device including a system is provided, wherein the multi-energy X-ray imaging device includes a front layer detector and a back layer detector, preferably as part of a dual or multi-layer detector system.

[0024] In yet another aspect, a training system configured to train an ML model based on training data is provided.

[0025] In yet another aspect, a training data generator system configured to generate training data from which a model is trained by a training system is provided.

[0026] In yet another aspect, a computer-implemented method for acquisition lag correction in x-ray imaging is provided, the method comprising: receiving an input including an input image acquired by a front layer detector in a multi-energy x-ray imaging device and an additional input image acquired by a back layer detector of the multi-energy x-ray imaging device, the back layer detector having circuitry for reducing time lag effects, the additional image being acquired by the back layer detector while the circuitry is activated; processing the input image and the additional input image to generate a corrected input image representing an estimate for image information of the front layer detector at an acquisition lag lower than the acquisition lag experienced for the input image; It has.

[0027] In another aspect, the use of a corrected version of a computer-implemented reconstruction operation and / or a computer-implemented spectral image processing algorithm is provided, preferably where an associated back layer image (preferably acquired with active lag reduction circuitry) is used together with the corrected front layer image.

[0028] In yet another aspect, a method for training an ML model based on training data is provided.

[0029] In yet another aspect, a method for generating training data for use in a method training method is provided.

[0030] A method for generating training data may include: i) imaging a phantom from multiple relative directions, each in the same imaging geometry, with and / or without activation of such lag reduction circuitry (LRC) for the detector of a single layer detector, and / or ii) imaging the phantom using a multi-layer detector with the phantom repositioned.

[0031] The method may be used for supervised or unsupervised learning.

[0032] The method may further include imaging a physical phantom with and without activation of the lag reduction circuitry in the detector back layer.

[0033] To enable the scaling embodiment described above, the "lag-free" back layer signal is compared to the front layer signal to estimate the lag component when imaging the phantom. The front layer lag artifact can be estimated by recording a sequence of 2D images both with and without back-side lag correction for the back layer. By comparing the deviation between the front layer and back layer images with and without lag correction for the 2D image sequence, the correct scaling correction to account for the front layer detector lag artifact can be estimated.

[0034] In yet another aspect, there is provided a computer program element configured, when executed by at least one processing unit, to cause the processing unit to perform any one of the methods described above.

[0035] In yet another aspect, at least one computer-readable medium having stored thereon program elements or having stored thereon a trained machine learning model is provided.

[0036] "User" refers to a person, such as a medical professional, who operates an imaging device or oversees an imaging procedure. In other words, the user is generally not the patient.

[0037] In general, the term "machine learning" includes computerized devices (or modules) that implement machine learning ("ML") algorithms. Some such ML algorithms operate to adjust a machine learning model configured to perform ("learn") a task. Other ML algorithms operate directly on training data, without necessarily using a model. This adjustment or updating of the training data corpus is referred to as "training." In general, task performance by an ML module can measurably improve with training experience. Training experience can include appropriate training data and exposure of the model to such data. Task performance can improve as the data better represents the task to be learned. "Training experience helps improve performance when the training data well represents the distribution of examples against which the ultimate system performance will be measured." Performance may be measured by objective testing based on the output generated by the module in response to providing test data to the module. Performance may be defined in terms of a particular error rate to be achieved for given test data. See, for example, TM Mitchell, "Machine Learning", page 2, section 1.1, page 6 1.2.1, McGraw-Hill, 1997.

[0038] "Lag" refers to the delay in signal buildup in an X-ray detector when exposed to X-rays at a given intensity. Therefore, the value registered at a pixel is generally inaccurate and does not represent the true signal intensity. Therefore, lag results in temporal artifacts in the images recorded by such detectors. The reason for this lag can be found in imperfections in some of the components of such detectors, such as charge trapping in semiconductor crystals. Artifacts in projection images caused by such lag are transmitted to images derived from such projection images, such as images reconstructed by tomographic reconstruction algorithms from such projection images.

[0039] "Front layer detector," in the context of machine learning embodiments (particularly with respect to training machine learning models), may refer not only to multi-detector systems, but also to more traditional detector systems that are only a single such detector layer (single layer detectors / systems).

[0040] Exemplary embodiments of the present invention will now be described with reference to the following drawings. [Brief explanation of the drawings]

[0041] [Figure 1] 1 shows a schematic block diagram of a medical imaging device. [Figure 2] 1 shows a multi-layer X-ray detection system. [Figure 3] 1 shows a block diagram of a system for correcting projection data for lag artifacts. [Figure 4] 4 shows a training system for training the machine learning model used in the system of FIG. 3. [Figure 5] In connection with training a model for use in the system of FIG. 4 above, various training settings are shown that may be used as training data and during inference. [Figure 6] 1 shows a block diagram of a machine learning model that may be used in embodiments contemplated herein; [Figure 6A] 1 shows a block diagram of a machine learning model that may be used in embodiments contemplated herein; [Figure 7] 1 shows a flowchart of a method for reducing imaged artifacts, a method for training a machine learning model, and a method for generating training data. DETAILED DESCRIPTION OF THE INVENTION

[0042] Reference is now first made to the block diagram of FIG. 1, which shows the medical imaging apparatus IAR envisaged here in an embodiment.

[0043] The imaging device IAR may include an X-ray based imaging device IA operable to generate, in an acquisition operation, projection data Λ that can be processed by the computing system SYS, for example for therapeutic or diagnostic purposes.

[0044] The computing system SYS may be operable to process the projection data into tomographic images that may be stored, displayed, or otherwise processed to aid in treatment and / or diagnosis, for example, for statistical or educational purposes, or for other applications such as planning, control tasks, etc.

[0045] The imaging device IA (also referred to herein as "imager" for brevity) is X-ray based and configured for multi-energy X-ray imaging, also referred to herein as spectral imaging. Purely projection-based imaging, such as radiography, is not excluded here, but tomographic imaging is primarily envisioned here. Imagers configured for tomographic imaging envisioned here include CT scanners, as shown in FIG. 1, or interventional imagers of the C-arm or U-arm type, such as those sometimes used, for example, in catheterization labs. Tomosynthesis-based imaging, such as mammography or dentistry, is also envisioned here.

[0046] Briefly, and as more fully expanded, the system SYS contemplated herein is configured to reduce, if not eliminate, certain image artifacts that may be added to projection data acquired in certain types of spectral imaging setups.

[0047] Spectral imaging, as referred to herein primarily below, enables material-specific imaging. That is, the obtainable images have contrast that varies primarily or exclusively with the concentration of a particular material of interest. Specifically, the undesirable contrast contribution of intervening structures or other materials is avoided with spectral imaging, unlike conventional energy-integrated conventional projection-based imaging.

[0048] Such materials of interest present within the field of view of a particular distribution imager may be sought to include contrast agent material types (iodine, barium, or combinations thereof) that may be used in contrast agent-assisted protocols to image structures that are nearly naturally transparent to x-rays. The contrast agent accumulates in particular regions or organs of interest, such as specific vascular portions or other locations of the cardiovascular system for cardiac imaging, or portions of the intestine in abdominal imaging. This allows for more accurate images that may better aid in treatment, diagnosis, planning, control of medical robots, etc.

[0049] Spectral imaging requires an imaging setup in which the imager IA is configured for multi-energy projection data acquisition. That is, for each pixel in the projection data, there are at least two intensities recorded at different energy levels. Therefore, in effect, (at least) two sets of projection data are acquired, at least one for each energy range. In embodiments, this is achieved by exposing a region of interest, such as a specific body or organ part, to X-rays with different energy spectra. This can be achieved, in particular, by a detector-side solution or a source-side solution. The detector-side solution, which is of primary interest here, includes a multi-layer detector module. This allows the region of interest to be exposed to radiation of different spectra, thus enabling multi-energy spectral imaging.

[0050] The acquired multi-energy projection data may be processed by a spectral processor component SP to obtain a spectral image having desired material-specific contrast for the target material of interest.

[0051] It has been observed that in multi-energy projection data acquisition, such as that required in spectral imaging, instances of inaccurate detection of intensity signals can occur, which in turn result in image artifacts either in the projection images or as these are commonly amplified in the 3D reconstruction of such material-specific spectral images. Other types of spectral images that are equally affected by artifacts include virtual mono-energy images, virtual non-contrast images, etc.

[0052] One such cause of this type of artifact has been found to be the manner in which image acquisition, i.e., the multi-energy intensities are recorded by a particular multi-layer detector system XD, which is an example of the detector-side spectral imaging solution described above. A computer-implemented correction system CS is proposed here to correct such artifacts that introduce errors into the acquired multi-energy projection data. Such corrected projection data may be sent for further processing via an output port OUT. Such further processing may include spectral processing of the corrected input image and / or reconstruction of the corrected input image into a tomographic image in the image domain, as needed.

[0053] The correction system CS may be implemented by machine learning, as described in more detail below. The corrector system CS may be implemented on a single computing platform or across multiple computing platforms or systems, such as in a distributed manner (cloud configuration) on multiple, preferably at least partially interconnected, computing devices, servers, etc. in a data communication network. The artifact correction system CS may be integrated into the imaging device IA, such as in the operating console OC, or into the workstation of the imager IA, or into the overall computing system SYS that processes the projection data as described above. Alternatively, the correction system CS is integrated into the detector system DX of the imager.

[0054] Before describing the artifact correction system CS for multi-energy projection data in more detail, reference will first be made to the components of the imaging apparatus IA to better provide a functional context around the operation of the correction system CS.

[0055] The imager IA, preferably X-ray based, allows for the acquisition of cross-sectional, preferably 3D, images of the ROI. To this end, the imager IA is configured for multidirectional (schematically indicated as "α" in the figures) projection data acquisition around the ROI. The acquisition (path) or "scan" does not necessarily define a full 360° angular range, although 180° or less may be sufficient here. The acquisition path does not necessarily define a circular arc, although in most cases it may. A helical path is envisioned here. Thus, the present disclosure is not limited to axial paths, although "step-and-shoot" setups, such as those described above, are not excluded here.

[0056] In some embodiments, the imaging device IA may be rotatable to perform multidirectional projection data acquisition operations. Such a rotating imaging device allows for the collection of projection images λ along different projection directions α in a scan around a region of interest, such as around the aforementioned stenosis in a coronary vessel. The imaging device IA includes an X-ray source XS and an X-ray-sensitive detector system XD. In some embodiments, the X-ray source XS and its opposite X-ray-sensitive detector XD are preferably disposed on a rotatable gantry MG attached to an optional fixed gantry SG. The examination region may be defined by the gantry SG, MG, between the source XS and the detector system DX, where the imaged patient, or at least the region of interest, resides. Some imagers do not have a fixed gantry SG, and their rotatable gantry G (appropriately journaled) may have a recess in which the examination region ER is so defined, as shown in the CT scanner shown in FIG. 1. The rotatable gantry MG may therefore have a "C," "U," or similar shape, such as in the C-arm imaging device envisioned herein. Biplanar imaging setups with two detectors with intersecting (e.g., perpendicular) imaging axes are not excluded here. Instead of using a C-arm / U-arm system, the illustrated CT scanner or similar is envisioned here in some embodiments. They can be of any generation, but preferably at least third generation, and the X-ray source XS and detector D are mounted on a rotatable gantry MG so that they rotate together around the patient. However, even higher-generation CT setups are also envisioned, in which mechanical rotation of the source and detector is not necessarily required for multidirectional projection data acquisition. Instead, a fixed ring of detectors around the ROI is used, and / or multiple fixed X-ray sources arranged in a ring around the ROI are used. Cone-beam imaging or other diverging beam geometries are also envisioned here, such as imaging using a helical scan path caused by a circular path of the X-ray source with translational motion relative to the imaging device IA and the ROI along the patient's longitudinal axis Z.In most cases, this translational motion is caused by the patient bed PS, on which the patient PAT is translated during projection data acquisition and through the examination region ER. However, scan paths with geometries other than circular / helical are not excluded here, as mentioned previously.

[0057] In general, imaging / acquisition involves energizing an X-ray source XS so that an X-ray beam XB emerges from the focal point of the source XS, traverses the examination region (having a region of interest therein), and interacts with the patient's tissue material therein. The interaction of the X-ray beam XB with the tissue material causes the beam to be modified. The modified beam can then be detected at a detector system XD.

[0058] The images obtainable by the imager IA may be stored in a memory DB, such as an image database, displayed on a display device DD via a visualizer VIZ, or processed in any other manner as desired depending on the medical task at hand.

[0059] The spectral processing component SP (abbreviated as "spectral processor") implements spectral (image) processing algorithms in appropriate software or hardware (or partially both). The spectral processor component SP may be part of the computing system SYS, the console OC, or a workstation. The spectral processor SP processes the multi-energy projection data to calculate spectral imaging, i.e., tomographic or projection images specific to the material of interest. One such spectral imaging algorithm is similar to the 1976 approach of R. Alvarez & A. Macovski (published, for example, as "Energy-selective reconstructions in X-ray computerized tomography", Phys Med Biol, vol. 21(5), pp. 733-44), in which two (or more) sets of projection images at different energy ranges per pixel are fed into a simultaneous linear system, which allows solving for material-specific attenuation coefficients as unknowns. N (≧2) energy levels can be used to solve for N different materials. Other types of material decomposition techniques have been devised, either in conjunction with or unrelated to the Alvarez approach, and each is contemplated herein in an embodiment for the spectral processor SP.

[0060] If a spectral tomographic image is desired, the system SYS may further include a reconstructor RECON configured to implement a tomographic reconstruction algorithm (FBP, iterative, algebraic, etc.) that transforms the projection images from the projection domain into the image domain (the portion of 3D space in which the patient PAT or imaging subject resides during imaging). Spectral processing by the spectral processor SP can occur in the projection domain before reconstruction or in the image domain after reconstruction. Alternatively, spectral processing and reconstruction are combined.

[0061] A multi-layer X-ray detector system XD that may be used here for dual energy imaging is exemplarily shown in more detail in the cross-sectional view given by FIG. 2 to which reference is now made.

[0062] The multi-layer X-ray detector system XD may include a housing H that houses multiple detector modules, each capable of detecting X-rays. The multiple detector modules are arranged in layers shown in FIG. 1 , including a front layer detector module FLD (also referred to herein simply as the “front layer”) and a back layer detector module BLD (also referred to herein for brevity as the “back layer”). While this is a basic setup for dual-energy imaging and will be primarily referred to herein, those skilled in the art will understand that the principles described herein are equally applicable to multi-energy detector systems having more than two such layers formed from multiple such modules (e.g., three, four, or more) in a multi-energy imaging setup for more than two (N>2) energy ranges, one higher than the other. Such layers include a front layer and a back layer, as well as one or more intermediate layers between the two. However, for present purposes, all layers other than the front layer may be referred to herein as “a / the” back layer module, since they are all positioned behind the front layer when viewed from the X-ray source XS. Thus, the front layer detector FLD is spatially closer to the X-ray source XS, proximal to the X-ray source XS, compared to any one of the other back layers, such as the single back layer BLD, which is distal to the X-ray source, i.e., quantifiers such as "front", "back", and "middle" are understood herein as references to the relative proximity of each layer to the X-ray source XS.

[0063] Each of the layers FLD and BLD preferably includes an array of X-ray sensitive pixels in a 2D layout of rows and columns. Each layer functions as a transducer that converts incident X-ray intensity into a correspondingly varying electrical charge. Incident X-rays that escape the patient tissue PAT after interaction are detected in the first, front layer FLD to generate such charges. The charges are then converted by appropriate A / D circuitry into a numerical format, such as a matrix of numbers (intensity values), each number representing and varying with the respective X-ray intensity that caused the particular charge. X-rays after interaction with the first layer similarly pass through and interact with each of the follow-up layers, such as the back layer BLD, to generate sets of charges that are similarly converted by appropriate A / D circuitry into a second set of projection data / projection images.

[0064] Due to the passage of X-rays through the first layer FLD, the spectrum of the radiation exiting the front layer is changed, and it is this radiation with the changed spectrum that subsequently interacts with the back layer BLD, thus enabling spectral imaging, since the two sets of projection data obtained result from X-rays at different spectra. The two intensity values ​​per pixel contained in the at least two sets of projection images so acquired by the multi-energy detection system XD can then be processed by the spectral processor SP using the spectral imaging algorithm as described above. The above-mentioned artifacts in layered multi-energy X-ray detection system XD setups such as the one described are due to a certain acquisition lag caused by the way the layers BLD, FLD, register the incident radiation intensity and convert it into charge. Indirect or direct conversion modules are envisioned for the front and back layers FLD. In each such layer of any type of module, a layer of semiconductor crystal is used. Interacting radiation is converted into a cloud of electron-hole pairs in such crystals, and the signal chain in each detector module uses the pairs, i.e., the aforementioned charges, to generate the relevant measurement signal. In direct conversion detector modules, such pairs are diffused across the crystal to generate corresponding charges that can be sensed. In indirect conversion detector modules, such pairs are used to generate visible light, which is then sensed by a photosensor. However, in each case of direct or indirect conversion, operation relies on the creation and / or movement (diffusion) of such electron pairs in the respective crystal layers. The creation or movement of such electron-hole pairs and their interaction within the detector module setup can be hindered due to crystal imperfections. Such imperfections cause charge traps that undesirably interfere with such electron-hole pairs (EHPs). In particular, such EHPs, or portions thereof (electrons or holes), may become trapped in such traps and fail to react as intended to the incident X-rays to generate associated charges or light photons that are thought to represent the respective local radiation intensities. That is, the expected charges are not generated at all, or, if generated, are generated with a delay. This effect is referred to as acquisition lag, a time effect that causes artifacts. Thus, each pixel exposed to different radiation intensities may respond in a delayed manner, thus resulting in erroneously recorded intensity values ​​for some pixels in each layer. This effect is more prevalent in medical imaging, which requires considerable dynamics and fast image signal recording in rotational tomography systems such as the one described above in FIG. 1. High image quality gradients can exacerbate such lag artifacts.For example, during the rotation of the detector, a pixel at a certain moment may be located behind a highly attenuating object such as a bone, but the next moment, within a few milliseconds, as soon as the pixel moves out of the shadow of this highly attenuating object, it may be exposed to much higher radiation. Such high intensity gradients may be too much for the described type of detection module to handle, which suffers from charge trapping, which causes delayed signal enhancement. Therefore, the recorded projection images may suffer from artifacts, i.e., inaccurately recorded and resulting intensity values, which may even be fed into the reconstructed image.

[0065] The proposed corrector system CS is configured to eliminate or at least reduce, i.e., correct, such lag artifacts. The described effects of such lag and the temporal artifacts resulting from delayed signal acquisition can be mitigated to some extent using hardware solutions. One such solution may include a lag reduction circuit LRC attached to the back layer detector module BLD. Such a circuit, e.g., a backside reset light circuit, can be implemented as a matrix of visible light-emitting diodes arranged on the distal surface of the back layer detector module BLD. During operation of the lag reduction circuit LRC, the semiconductor (crystalline) layer of the back layer BLD is exposed to visible light before some or each X-ray exposure. It has been found that such exposure of the crystal to visible light reduces the disturbing effect of such charge traps by energetically "filling" deep potential wells in the semiconductor crystal lattice.

[0066] Due to the setup as described above, the light reduction circuit LRC cannot be arranged on any of the front layer detectors FLD, since the material of the components of such a circuit LRC may result in artifacts and reduced efficiency, especially due to the shadow that such a circuit casts on the underlying back layer detector BLD. Undesirably, the back layer detector module BLD will record such shadow and thus completely corrupt the projection data.

[0067] Therefore, what is proposed here is a corrector system CS configured to computationally remove such lag-related artifacts from any front layer projection image detected by any of the front layer detectors FLD. No intervening signal jamming hardware is required in the newly proposed correction system CS. The lag reduction circuit LRC is typically effective enough to reduce or completely avoid the described lag-related artifacts in the projection images recorded by the back layer BLD. For present purposes, the back layer detected projection images can be considered artifact-free. Therefore, the back layer detected projection images do not need to be processed separately by the correction system CS; only the front layer images detected by one or more front layer detector modules FLD need to be so processed.

[0068] The projection image detected by the front layer detector module FLD may be referred to herein, for simplicity, simply as the "front image / imagery," in contrast to the projection image recorded by the back layer detector module BLD, which is correspondingly referred to as the "back image / imagery."

[0069] Reference is now made to the flowchart of FIG. 3, which lays out more detailed aspects of the lag-induced artifact reducer or corrector system CS. In particular, a front image X recorded by one of the front layer detector modules FLD is received at an input port IN of the corrector system CS. The corrector component CC processes the received front projection image and generates a corrected version thereof X', which is output at an output port OUT. The corrected front image X' may then be stored by a visualizer VIZ on a display device DD, or sent together with the back image to a spectral processor SP for spectral processing and / or a reconstructor RECON for conversion to the image domain, as described above, to obtain a spectral cross-section / tomographic image of the 3D distribution of a particular material of interest. If more than one front layer is present, the system SYS may process the same together or separately, as needed. A separate constructor system CS may be present for each front layer if necessary to better handle the specific characteristics of each such front layer.

[0070] The corrector component may be implemented using machine learning (ML). Thus, the corrector component CC may include a trained machine learning model M.

[0071] Thus, conventional energy-integrating single-layer detector systems use detector modules similar to, and even of the same structure and type as, the detector modules used in dual-energy setups such as those described in FIG. 2. Thus, the front-layer detector FLD of the spectral detector XD can be used successfully in conventional single-detector layer imaging, and the single-layer detector image can be used in both rows as the described FLD module for dual / spectral imaging. This structural fact can be exploited here to easily obtain suitable training data for training the machine learning model M, as will be explained in more detail below.

[0072] However, the applicant has discovered a surprisingly simple relationship between an already artifact-corrected back layer image and the corresponding lag-artifact-corrupted front layer image, so implementing such a machine learning model may not necessarily be required here. Experiments have found that the two can be related by scaling with a constant scaling factor for the entire image, or by a scaling map with a scale entry for each individual pixel location. The scaling factor can be determined by linear fitting based on recording a series of such pairs of front and back images. Such a linear modeling setup can reveal good approximations of scale factors that can be used in simpler embodiments by the corrector component CC. Specifically, this scaling approach is based on scaling a "lag-free" back image and comparing it with the associated front image to estimate the lag component. Alternatively, it is the back image that can be scaled in this way. More specifically, the front layer lag artifact can be estimated by recording a sequence of 2D images both with and without the lag reducer LRC for the back layer BLD. By comparing the deviation of the front layer image relative to the back layer image with and without lag correction for such a 2D projection image sequence, a scaling factor for the correction can be estimated by linear fitting to account for the lag artifacts imparted to the front layer detector module FLD. The scaling factor can then be applied to the corrupted front image to reveal a corrected version.

[0073] However, depending on the particularities of the semiconductor crystal lattice, the distribution of charge traps that cause lag, or any other factors, the relationship may be more complex. Thus, a highly nonlinear, more complex relationship may exist between the back image and the corrupted front image. While analytical setups such as those described by linear fitting may be one-way forward, it has been found that more general modeling that does not rely on specific assumptions can yield better results. This has been found by using a neural network-type model that does not require a specific modeling setup but can be generally used to better learn this back-to-front image relationship in terms of lag behavior.

[0074] Due to the spatially correlated nature of images, convolutional neural networks ("CNNs") have been found to work particularly well. As explained in more detail below, some such neural network models allow for multi-scale processing, which has been found to yield better results and is specifically envisioned herein in some embodiments. Again, details are explained below.

[0075] Describing now some general aspects of machine learning embodiments in more detail, these generally require two phases: a pre-learning / training phase and a subsequent inference / deployment phase. Training of the ML model M is a mathematical procedure that can be performed as a one-off operation, or can be performed in multiple stages as new training data emerges. The training phase is based on training data that can either be artificially generated by a training data generator system TDGS, as actually envisioned in embodiments, or can be otherwise sourced from a stock of existing images, such as may be found in medical databases across hospitals, medical facilities, etc.

[0076] In general, training a machine learning model involves adapting the parameters of the model M in light of a suitably varying set of training data (x, y). Here, x represents a training data input that is processed by the model M to generate an output M(x), which is then compared to a ground truth or label y previously associated with the input x. Thus, x may be a historical or synthetic front image, and y may be an estimate of what a relevant artifact-free version of x should look like. This corresponds to a supervised learning setup envisioned in some embodiments herein; however, unsupervised learning setups based on clustering or other approaches are not excluded herein. Another feature of ML learning not typically found in analytical modeling approaches is the backpropagation feature, where mismatches between M(x) and y are "backpropagated," or more generally, used in adapting the current model parameters, preferably over an iterative cycle. This allows for very efficient learning and extraction of good approximations of sought-after relationships for "freer" modeling without explicit analytical modeling, which tends to tie unknown relationships to specific forms that are not necessarily appropriate.

[0077] Lowercase notation such as "x", "y", etc. is used herein for training data, while data in the post-training deployment is referred to with uppercase "X", "Y", etc.

[0078] 4, a training system TS for training an ML model M based on training data TD={(x, y)} is shown. The training data is optionally provided by a training data generator system TDGS.

[0079] In a supervised setting, the training data is a set of k(x k ,y k ) for each pair k (index k is unrelated to the index used above to specify the generation of the feature map), the training input data x k, and an associated target or ground truth y. Thus, the training data is organized in k pairs, particularly for the supervised learning schemes primarily envisaged here. Note, however, that unsupervised learning schemes are not excluded here.

[0080] In the training phase, the architecture of a machine learning model M, such as an artificial neural network (discussed more fully below in Figure 6), is pre-populated with an initial set of weights or other parameters. The weights θ of the model M are defined by the parameterization M θ and the training data (x k ,y k It is the goal of the training system TS to optimize, and therefore adapt, the parameters θ based on the pair F. In other words, learning can be mathematically formulated as an optimization scheme in which a cost function F is minimized, although a dual formulation that maximizes a utility function can be used instead.

[0081] Now, assuming the paradigm of a cost function F, this measures the aggregate residual, i.e., the error occurring between the data estimated by the neural network model M and the target for some or all of the training data pairs k. argmin θ F=Σ k ||M θ (x k ),y k || (1)

[0082] In equation (1) and below, the function M() denotes the result of the model M applied to the input x. The cost function may be pixel / voxel based, such as an L1 or L2 cost function. In training, the training input data x of the training pairs k is propagated through the initialized network M. Specifically, the training input x for the kth pair k is received at the input section IL, passed through the model, and then output training data M θ(x). An appropriate similarity measure ||||, such as the derived p-norm or squared difference, is then calculated based on the actual training output M generated by the model M, also referred to here as the residual. θ (x k ) and the desired target y.

[0083] Output training data M(x k ) is the applied input training image data x k Target y related to k In general, this output M(x k ) and the associated target y of the kth pair currently under consideration. k Then, an optimization scheme such as backpropagation or other gradient-based methods is used to find the error between the considered pair (x k ,y k ) or can be used to adapt the parameters θ of the model M to reduce the residuals for a subset of training pairs from the full training dataset.

[0084] The model parameter θ is the pair {(x k ,y k After one or more iterations in the first inner loop, where the training data pairs {x k+1 ,y k+1} is processed accordingly. While it is possible for the outer loop to loop over individual pairs until the training data set is exhausted, it is preferable here to instead loop over the set of pairs (i.e., over the batch of training pairs) at a time. The inner loop iterations over the parameters for all pairs that make up the batch. Such batch-wise learning has been shown to be more efficient than proceeding pairwise. In the following, the subscript "k" indicating individual instances of training data pairs is largely dropped to free up notation.

[0085] The structure of the updater UP depends on the optimization scheme used. For example, the inner loop as managed by the updater UP may be implemented by one or more forward and backward passes in a backpropagation algorithm. While adapting the parameters, the aggregated, e.g., summed, residuals of all training pairs in a given batch are considered up to the current pair to improve the objective function. The aggregated residuals can be formed by constructing the objective function F as a sum of squared residuals, such as in equation (1), of some or all of the considered residuals for each pair. Other algebraic combinations instead of sums of squares are also envisioned.

[0086] A GPU may be used to implement the training system TS for improved efficiency.

[0087] Reference is now made to Figure 5, which shows a schematic block diagram of the learning phase LP and the subsequent inference / deployment phase IP. The deployment phase IP may refer to a phase in which the model trained by the corrector component CC is used in clinical practice, where the mode processes new "not-before-seen" inputs X, i.e., input data X not drawn from the training dataset TD (e.g., new artifact-corrupted front images). There may also be a testing phase in which some training data is reserved for later use in a testing phase to test the performance of the system. In general, the testing phase and the deployment / inference are very similar.

[0088] As mentioned above, the training data (x, y) are either obtained from a stock of prime images or artificially generated, for example, by imaging a phantom PH. Labels can be generated using graphical correction software by experienced human experts, but this is likely to be tedious and expensive. Other methods, particularly those that provide relevant labels y, are described in more detail below. The training data (x, y) obtained by either method are fed to an initial model M0=M, such as a convolutional neural network with a default set of parameters. For reasons that will soon become apparent, such CNN parameters are often referred to as filter parameters due to the nature of the specific type of operation (convolution) used by such models. Thus, the initial model M0 is initialized using a default, or even arbitrary, set of initial parameters, which applies not only to CNN / NN models but also to other model types. The output of the initial model M0(x) is compared to the label y, and the parameters are iteratively adapted, preferably in one or more cycles during the learning phase LP. Backpropagation is preferably used in one form or another when adapting parameters for a given batch of training data to be processed. The two curved arrows schematically indicate a training process, most commonly iterative, in which model parameters are adapted in light of training data instances (x, y) in iterative cycles. Once sufficiently trained (which can be established based on various criteria, such as convergence or a preset number of iterative cycles), the now-trained model M is then deployed in an inference phase or testing, as described. In deployment, given a new lag-corrupted front image X, a desired lag / artifact-corrected version X' of the front image X is calculated by the model M, where artifacts otherwise caused by the described lag mechanism are corrected (at least reduced, or entirely eliminated). Thus, version M(X)=X' is an approximation of the pattern of intensity values ​​that would have been obtained in the absence of lag-induced corruption.

[0089] Generative models, such as a GAN model M, may also be used with interest here in some embodiments, where the model M is trained to "redraw" X as X' in a lag-free diction learned from a stock of training data (X, Y). Such generative models are sometimes used to artificially generate motive images for artists who did not originally draw such motives, but in their "style." It has been found that such a GAN setup can be beneficially used here to learn to draw front images in a lag-free style. In such a GAN-type setup, the cost function F (Equation (1)) that drives the training represents the adversarial interaction between the generator network and the discriminator network that together constitute the model M in the GAN. See "Generative Adversarial Networks" by Goodfellow et al., available on arXiv under the reference code arXiv:1406.2661 (2014).

[0090] In Figure 5 and subsequent figures, the subscript "f" denotes "front image" and the subscript "b" stands for "back image." The asterisk symbol "*" is used in superscript to indicate whether the back image was acquired with a hardware lag reducer circuit activated, e.g., whether a backside reset light circuit was activated prior to acquisition of a given back image frame, or whether any other such type of lag reduction circuit LRC was activated. Thus, the presence of such an asterisk symbol "*" indicates that the inverted image was actually recorded with prior activation of the circuit LRC, while the absence of the asterisk indicates that such an LRC was not activated, and thus, hardware-based temporal artifact correction was not used. The asterisk therefore relates to the use of hardware-induced correction in connection with back image acquisition. Preferably, either type of lag reduction circuit LRC can be deactivated here to allow acquisition of an uncorrected back image, which may be useful for facilitating acquisition of training data TD, as described in more detail below.

[0091] In some embodiments, it has been found that the above-described approximately linear relationship due to scaling between the lag-corrupted front image and the LRC-corrected back image can now be utilized to implement data-based regularization in the training algorithm via a regulator channel RL. Such data-based regularization allows for better and more robust learning and better performance in deployment or testing. In this data-based regularization embodiment, the corrected front image X is fed to the training model M for processing. f As well as backlight correction image X b * is also fed to the model M through the second regularizer channel RL. Thus, in the proposed training embodiment, the front image X f A two-channel setup is used by model M when processing as input, and the hardware-corrected back image X b *(Thanks to the LRC circuit) but the front image X fare jointly processed by the model to derive a corrected version X' of X. This two-channel processing allows the model M to provide context so that it can better learn the relationship between the artifacts corrupting the front image and the corrected back image in order to better learn how lag corruption affects the image.

[0092] With continued reference to FIG. 5, as shown therein, multiple training configurations based on different combinations of training data X are contemplated herein for use in training system TS (see below in FIG. 6 for more details) to train model M.

[0093] For example, in one embodiment, a single-layer detector front image xf without lag compensation is optionally used along with a back image xb* of a dual-layer detector system XD as training input. The optional use of the back image xb* implements the regularization channel RL described above. Thus, the training input data in this embodiment of the training data configuration is x=xf, or, optionally with regularization, x=(xf, xb*).

[0094] Without abuse of linguistic logic, a projection image that can be recorded by a conventional energy integrating imager with a single layer detector system (not shown) may also be referred to herein in uniform terms as a "front image."

[0095] The single-layer front image xf* with LRC circuit lag correction may be used as the ground truth y for the supervised training setup, where y=xf*.

[0096] Such a single front image xf* can be obtained by physically removing the front layer FLD in the dual imager IA and using the remaining back layer BLD, which itself includes a lag corrector circuit LRC, to essentially temporarily transform the dual-layer imager IA into a single-layer detector system. Thus, the previous back layer module BLD is now the front layer module. Alternatively, the ground truth front image xf* with lag correction may be obtained with a different native single-layer imager having a detector module of the same type and make as that used in the dual imager IA. This is often the case, for example, for a given manufacturer, as observed above.

[0097] Training data (x=x f ,y=x f *) or (x=(x f ,x b The above set of input training images x, y = x*) can be obtained in the training data generator setup TDGS as follows: the input training images x can be obtained by regular operation of a dual imager with a multi-layer detector XD, for example by imaging variations of a suitably constructed phantom PH in a sufficient number of instances. The ground truth part y = x of the training data TD f * can be obtained by using a dual-energy imager IA as described, preferably in the same imaging geometry for the training input x and the ground truth y, by acquiring the same number of images of the phantom PH, with the front layer FLD removed.

[0098] Alternatively, single-layer data without lag correction x f can be obtained from a native single-layer detector imager with the lag-corrected LRC turned off, while x f *Images are acquired with pre-exposure with lag correction LRC activated as described above.

[0099] Preferably, the training images are acquired in sequence, for example, imaging for a training input x followed by the associated ground truth y, or vice versa, to maintain correct pairwise association.

[0100] Now, using the above protocol and training data configuration (x, y) to explain training data generation in more detail, in general, single-layer front images and two-layer data (front and back images) for training should preferably be acquired in the same imaging geometry, preferably using the same acquisition-related imaging settings (such as rotation speed).

[0101] Thus, the training data generation system TDGS includes operating a multi-energy imager IA or a single-layer imager with the above-described protocol. The training data generation system TDGS further includes acquiring different sets of training images either by changing the configuration, such as the spatial configuration, of the phantom PH imaged with the protocol over time or by changing the imaging context (such as the settings of the imager IA) relative to a static phantom. In particular, various time-dependent instances of phantom images are envisioned here as being controlled and managed by the training system TDGS. For example, the phantom PH may be positioned as a rotating disk with the disk's normal vector parallel to the detector normal (2D imaging). Alternatively, the phantom PH may be placed on a rotating table for a rotating imaging geometry, such as cone-beam CT (CBCT) or any other tomographic imaging geometry, diverging or parallel.

[0102] Alternatively, a sequence of images can be acquired using a static phantom PH, but with a strongly varying dose. The two-layer data should be acquired at the same positions as the single-layer detector data.

[0103] In any of the above cases, the phantom PH may contain insets of different materials normally found in the human body, such as water, fat, calcium, etc.

[0104] To obtain the corrected version x', we use the front x as input to the model for data-based regularization. f and back image x b In the two-channel training described above, where * is sent together, this can also be done during deployment. Alternatively, the two-channel approach is used only in training, but not during deployment.

[0105] The above learning setup has primarily focused on supervised learning, however, an unsupervised setup for training ML models is also envisioned here.

[0106] For example, an unsupervised framework for network learning strategies can be implemented by different positioning of the phantom PH, resulting in different realizations of the lag artifact from only the two-layer data (xf, xb*). Based on multiple independent such artifact realizations of the same basic scan object PH, the unsupervised learning scheme can recover a basic ground truth image y = xf* and reduce / remove the lag artifact for the front layer detector FLD. Thus, in this manner, several images xf acquired by the front layer detector FLD with lag artifacts from the phantom are recorded at different positions along with corresponding images xb* for the back layer detector BLD with a lag compensation circuit LRC for applying lag compensation. The training system TS can then proceed to distinguish the artifact from the intended image content, which obviates the need for a priori knowledge of xf*, thus implementing the unsupervised learning scheme here. For example, since the training system TS has recorded several versions / realizations of the lag artifact, it can compare the images via registration to align the images in order to distinguish the artifact from the actual image content (represented here by the phantom). The essentially artifact-free images arrived at in this way can be used for training when no other means of generating ground truth data are available. The proposed phantom imaging therefore provides a bootstrap method for generating ground truth data, which can then be used in any training method to adapt the parameters of the model M.

[0107] Reference is now made to FIG. 6, which illustrates the components of a convolutional neural network CNN-type model as envisaged herein in an embodiment, including a convolution filter CV to better account for the above-mentioned spatial correlation in the pixel data (x,y).

[0108] Specifically, Figure 6 shows a convolutional neural network M in a feedforward architecture. Network M has multiple computational nodes arranged in layers in a cascaded fashion, with data flow from left to right, thus from layer to layer. Recurrent networks are not excluded here.

[0109] In deployment or training, the front image x f .X f Input data including (and optionally a back image xb*) is applied to the input layer IL. The input data is fed to the model M in the input layer IL, then propagates through a sequence of hidden layers L1-L3 (only three are shown, but there could be one, two, or more than three), and then emerges at the output layer OL as the training data input M(x) or the corrected front image X'. The network M may be said to have a deep architecture because it has more than one hidden layer. In a feedforward network, the "depth" is the number of hidden layers between the input layer IL and the output layer OL, while in a network, the depth is the number of hidden layers multiplied by the number of passes.

[0110] The layers of the network, indeed the input and output images, as well as the inputs and outputs between hidden layers (here called feature maps), can be represented as two or more dimensional matrices ("tensors") for computational and memory allocation efficiency.

[0111] Preferably, the hidden layer is layer L1-L N-k, k>1 The number of convolutional layers is at least one, such as 2, 3, 4, or 5, or any other number. The number may be two digits.

[0112] In embodiments, downstream of the sequence of convolutional layers there may be one or more fully connected layers, although this is not necessarily the case in all embodiments, and indeed, preferably, fully connected layers are not used in the architectures envisioned herein in some embodiments. mThe layer and the input layer IL implement one or more convolution operators CV. m may implement the same number of convolution operators CV, or the number may be different in some or all layers.

[0113] The convolution operator CV implements the convolution operation to be performed on its respective input. A convolution operator may be conceptualized as a convolution kernel. It may be implemented as a matrix with entries that form filter elements, referred to herein as weights θ. In particular, it is these weights that are adjusted during the training phase. The first layer IL processes the input data through its one or more convolution operators. The feature maps FM are the output of the convolution layer, one feature map for each convolution operator in the layer. The feature maps of the previous layer are then input to the next layer, which generates higher-generation feature maps FM. i , F.M. i+1 and so on until the final layer OL combines all feature maps into an output M(x),x'. Each feature map entry may be described as a node. The final combiner layer OL may also be implemented as a convolution to provide the corrected front image.

[0114] The convolution operator CV in a convolutional layer is distinguished from a fully connected layer in that the entries in the output feature map of the convolutional layer are not a combination of all nodes received as inputs to that layer. In other words, the convolution kernel is only applied to a subset of the input volume V, or feature maps received from the previous convolutional layer. The subset is for each entry in the output feature map. Therefore, the operation of the convolution operator can be conceptualized as a "slide" on the input, similar to the discrete filter kernel in classical convolution operations known from classical signal processing. Hence the name "convolutional layer." In a fully connected layer, the output node is generally obtained by processing all nodes in the input layer.

[0115] The stride of the convolution operator can be selected as 1 or greater than 1. The stride defines how the subset is selected. A stride greater than 1 reduces the dimensions of the feature map relative to the dimensions of the input at that layer. A stride of 1 is preferred here. To maintain the sizing of the feature map to correspond to the dimensions of the input image, a padding layer P of zeros is applied. This allows for convolving feature map entries located at the edge of the processed feature map.

[0116] The number of pixels of input data so processed for each output node in a given layer describes the size of the neural network's receiving field for a given unit in that layer and is determined by the kernel size, stride, and padding of all convolution operators CV preceding that layer in which the node is located. More specifically, the receiving field of the entire convolutional neural network M is the size of the region of action of the input data that is "seen" or processed to determine a given node in the output data. The region of action is the number of nodes across the layers of the network M that contributed to a given node in the output data X',M(x).

[0117] In a preferred embodiment (see schematic inset of FIG. 6A), a convolutional neural network CNN is used for the model M configured for multi-scale processing. Such a model may comprise, for example, a U-unit architecture or its ilk as described by O. Ronneberger et al. in "U-Net: Convolutional Networks for Biomedical Image Segmentation," available online on the arXiv repository under the citation code arXiv:1505.04597 (2015).

[0118] In a multiscalable NN model of this or a similar architecture, a deconvolution operator DV is used. More specifically, such a multiscalable architecture of model M may have, in series, a downscaling path of layers DPH and, downstream, an upscaling path of layers UPH. A convolution operator with varying stride length is used at layer L in the downscaling path. j The feature map size / dimension (m×n) is gradually reduced as the layers pass through the upscaling pathway UPH, down to the lowest dimension representation by the feature map, i.e., the latent representation κ. The latent representation feature map κ is k , returning to the dimensions of the original input data X. This dimension bottleneck structure has been found to result in better learning because it forces the system to distill its learning down to simpler structures such as those represented by the low-dimensional latent representation.

[0119] An additional regularization channel can be defined, in which intermediate outputs (intermediate feature maps) from scale levels in the downscaling path are fed to corresponding layers at corresponding scale levels in the upscaling path UPH. Such inter-scale cross-connections ISX have been found to promote robustness and efficiency of learning and performance.

[0120] Preferably, in any NN model M, the receiving field of the network is selected to cover the characteristic artifact sizes.

[0121] Other regression-type models may be used in place of or in combination with the NN model, and thus the present disclosure is not limited to NN models.

[0122] Reference is now made to the flowchart of FIG. 7 for the above-described process of deployment, ML training, and, if necessary, generation of training data TD.

[0123] Referring first to the flowchart of FIG. 7A, this illustrates the steps of a lag correction method for processing projection images acquired by a spectral imager IA. The data includes two sets of projection data, referred to as the front image and the back image / images, as described above. The input front image X is the result of projection data acquisition using a multi-energy detector system, such as might be used in an X-ray-based imager IA of the radiography or multidirectional tomography type. Thus, the multi-energy projection images, the front image and the back image, record the same FOV in the same imaging geometry, but in different energy ranges or spectra.

[0124] In step S710, a potentially lag-corrupted front image X acquired by the front layer FLD of the multi-energy detection system DX in the spectral imager IA is received.

[0125] The front image X thus received is processed in step S720 into a corrected image X'.

[0126] The corrected front image X' is then stored, displayed (if desired), or made available for further processing, for educational purposes, for statistical analysis, etc., in step S730.

[0127] For example, as an optional step S740, and primarily as envisaged herein, the so-corrected front image X′ is then sent for processing by a spectral imaging processing algorithm to obtain spectrally processed projection data together with a simultaneously acquired back image (acquired by the back layer BLD) that has been corrected with the native LRC circuitry.

[0128] In such spectral projection images, such as "virtual-contrast-only" images, the contrast is almost exclusively material-specific, corresponding to a preselected material type, e.g., selected by the user, as described above in connection with Figure 1. At this stage, the spectral projection image is no longer multi-energy, but a single set with a single value per pixel location for any given projection direction (not counting effects from cone-beam or other divergent imaging geometries).

[0129] In step S750, the spectral projection images acquired in step S740 may then be reconstructed into tomographic images in the image domain, such as a 3D image volume or one or more sections (slices) thereof.

[0130] Machine learning may be used to perform the corrective action S720. In such an approach, a trained ML model M, previously trained on training data, may be used.

[0131] If machine learning is used, the steps of the method according to FIG. 7A may be used in deployment (eg, in clinical routine practice) or in testing prior to deployment.

[0132] Once the model is sufficiently trained (as judged by certain criteria as described above), the model may be used in a correction step S720.

[0133] The corrective action in step S720 is in fact preferably based on a trained machine learning model as described above.

[0134] In a simpler, not necessarily ML, approach or a simpler ML approach involving more explicit modeling, the scaling coefficients can be determined by fitting a linear model to a series of acquired front images and their simultaneously acquired back image counterparts. A single such scalar coefficient may be used, or a matrix (map) of such scaling coefficients is used, where the scaling coefficients are pixel-specific. The scaling coefficients may then be stored in memory MEM′. Step S720 may then be performed by retrieving the scaling coefficient(s) from memory MEM′ and applying it to the front image X by pixel-wise multiplication or division (or other algebraic combination) to obtain a lag-compensated version X′ thereof.

[0135] If a machine learning model M is used in step S720, such a model can be obtained by a training method as shown in the flowchart of FIG. 7B, to which reference is now made.

[0136] In general, such methods proceed to step S810, where training data TD is received. The training data is either drawn from an existing stock of medical images, such as those found in medical databases, preferably including human-generated annotations, or is synthetically generated, as described in more detail below.

[0137] In step S820, parameters of a predefined architecture of a machine learning model M, such as a convolutional neural network, are adapted based on the training data.

[0138] Training may proceed in one or more iterations to generate a fully trained model, which is then made available in step S830 for deployment or testing, which can be used as described above in Figure 7A. Training S820 may be one-off, or may be repeated once new training data becomes available.

[0139] Training may be supervised or unsupervised, such as by clustering the training dataset or by processing with an autoencoder network (AE), including AEs of the variational type (VAE). As mentioned above, generative models such as I Goodfellow's GAN approach, or any of its ilk, may be used.

[0140] Rather than relying (solely) on existing historical images for training data, the training data may instead be synthetically generated in a training data generation setup such as that shown in the flowchart of FIG. 7C, referenced hereto.

[0141] In such a method, in step S910, various combinations of front and back images are generated by imaging a phantom PH, such as that described above in connection with FIG. 5, with and without the lag reduction circuitry activated.

[0142] Dual-layer and / or single-layer imagers may be used. Front images xf / xf* can be acquired by sequentially imaging the phantom PH with and without the lag reduction circuit LRC activated by using a single-layer imager instead of the dual imager IA. Imager settings (e.g., dose) may be changed for a static phantom, or the same settings may be used but the phantom is moved during imaging. The single-layer detector imager preferably uses detector modules of the same type and / or manufacture as those used in the multi-energy detector system XD (FIG. 2). If this is not possible, care should be taken to ensure that the detector module specifications are the same or at least comparable. Preferably, as a minimum, the same type (e.g., doping pattern) of semiconductor crystals should be used in both imagers.

[0143] Instead of using such a single-layer imager as an auxiliary, the dual-imager IA itself isb * may be operated to obtain x f ,x f * can be obtained by removing the front layer FLD of the dual imager and using only the back layer BLD with the LRC circuit. In this regard, a dual imager IA with a removable front layer detector module FLD may be beneficial. For example, the front layer FLD may be configured as a slide-in / out unit that can be slid out of the housing H and back in as needed, simplifying training data generation. Alternatively, a lab setup configured to mimic the characteristics of a real dual-layer detector may be used. Again, the phantom PH may be moved during such imaging, or the imager settings may be changed to generate different realizations of training data, particularly training data pairs. Other training data generation options are not excluded here.

[0144] In step S920, once a sufficient supply of appropriately modified training data images of the front and back images is obtained, it can be used as training data for a machine learning model, for example as shown in FIG. 7B.

[0145] It was observed that the method described above in Figures 7A and 7B performed particularly well when, at the input during inference / testing 7A and / or training 7B, not only the corrupted front image was provided for processing by the machine learning model, but also the associated natively corrected back image. The two may then be processed jointly during inference or training to provide better context for the method so that inference and / or training can be improved. It may be sufficient to use this two-channel approach in training but not during inference. Preferably, however, it is used in both training and inference.

[0146] The components of the corrector system CS may be implemented as one or more software modules and may run on one or more general-purpose processing units, such as workstations associated with the imager IA, or on a server computer associated with a group of imagers.

[0147] Alternatively, some or all components of the corrector system CS may be arranged in hardware, such as a suitably programmed microcontroller or microprocessor, such as an FPGA (Field Programmable Gate Array), or as a hardwired IC chip, an application specific integrated circuit (ASIC) integrated into the imaging system IAR. In further embodiments, the system SYS may be implemented partly in software and partly in hardware.

[0148] The various components of the corrector system CS of the overall calculation SYS may be implemented on a single data processing unit PU, or alternatively, some or several components are implemented on different processing units PU, possibly located remotely in a distributed architecture and connectable within a suitable communication network, such as in a cloud configuration or a client-server setup.

[0149] One or more features described herein may be configured or implemented as or using circuitry encoded in a computer-readable medium, and / or combinations thereof. Circuitry may include discrete and / or integrated circuits, systems-on-chips (SOCs), and combinations thereof, machines, computer systems, processors and memories, computer programs.

[0150] In another exemplary embodiment of the invention, a computer program or a computer program element is provided, characterized in that it is configured to perform, on a suitable system, the method steps of the method according to one of the previous embodiments.

[0151] Thus, a computer program element may be stored in a computing unit that may be part of an embodiment of the present invention. This computing unit may be configured to perform or direct the execution of the steps of the above-mentioned method. Furthermore, it may be adapted to operate each component of the above-mentioned apparatus. The computing unit may be configured to operate automatically and / or to execute a user's order. The computer program may be loaded into the working memory of a data processor. The data processor may thus be configured to perform the method of the present invention.

[0152] This exemplary embodiment of the present invention encompasses both computer programs that use the present invention from the beginning, and computer programs that convert existing programs into programs that use the present invention by means of an update.

[0153] Furthermore, the computer program element may be capable of providing all the steps necessary to fulfill the procedures of the exemplary embodiments of the methods described above.

[0154] According to a further exemplary embodiment of the present invention, a computer readable medium, such as a CD-ROM, is presented, the computer readable medium having stored thereon a computer program element, the computer program element being as described by the preceding sections.

[0155] The computer program may be stored and / or distributed on a suitable medium (particularly, but not necessarily, a non-transitory medium), such as an optical storage medium or a solid-state medium supplied together with or as part of other hardware, but may also be distributed in other forms, such as via the Internet or other wired or wireless telecommunications systems.

[0156] However, the computer program may also be presented over a network such as the World Wide Web and can be downloaded into the working memory of a data processor from such a network. According to a further exemplary embodiment of the present invention, a medium for making a computer program element available for downloading is provided, the computer program element being configured to perform a method according to one of the aforementioned embodiments of the present invention.

[0157] It should be noted that the embodiments of the present invention are described with reference to different subject matters. In particular, some embodiments are described with reference to method-type claims, and other embodiments are described with reference to apparatus-type claims. However, those skilled in the art will understand from the above and below description that, unless otherwise specified, any combination of features belonging to one type of subject matter, as well as any combination between features relating to different subject matters, is considered to be disclosed in the present application. However, all features can be combined to provide a synergistic effect greater than the simple sum of the features.

[0158] While the invention has been illustrated and described in detail in the drawings and foregoing description, such illustration and description are to be considered exemplary or explanatory and not restrictive. The invention is not limited to the disclosed embodiments. Other variations to the disclosed embodiments can be understood and effected by those skilled in the art in practicing the claimed invention, from a study of the drawings, the disclosure and the dependent claims.

[0159] In the claims, the word "comprise" does not exclude other elements or steps, and the indefinite article "a" or "an" does not exclude a plurality. A single processor or other unit may fulfill the functions of several items recited in the claims. The mere fact that certain means are recited in mutually different dependent claims does not indicate that a combination of these means cannot be used to advantage. Any reference signs in the claims should not be interpreted as limiting the scope. Such reference signs may consist of numbers, letters or any alphanumeric combination.

Claims

1. 1. A system for acquisition lag correction in x-ray imaging, comprising: an input interface for receiving an input including an input image acquired by a front layer detector of a multi-energy x-ray imaging device and an additional image acquired by a back layer detector of the multi-energy x-ray imaging device, the back layer detector having circuitry for reducing time lag effects, the additional image being acquired by the back layer detector while the circuitry is activated; and a corrector component configured to process the input image and the additional input image to generate a corrected input image representing an estimate of image information of the front layer detector at an acquisition lag lower than an acquisition lag experienced for the input image; A system having:

2. 2. The system of claim 1, wherein the processing by the corrector component includes scaling the input image by a scaling factor obtained based on a series of previous input images and additional input images previously acquired by the or other multi-energy x-ray imaging device.

3. The system of claim 1 or 2, wherein the corrector component is implemented based on a trained machine learning model.

4. 4. The system of claim 3, wherein the machine learning model is trained in a supervised setup based on training data including training input data and associated ground truth data, the training input data including previous images acquired by the or other front layer detector having a time lag effect, and the ground truth data including previous images acquired by the or other front layer detector having circuitry for reducing the time lag effect with such circuitry activated.

5. 5. The system of claim 4, wherein the training input data further includes additional previous images acquired by the or another back layer detector, the additional previous images being acquired by the back layer detector with circuitry for reducing time delay effects activated.

6. 4. The system of claim 3, wherein the machine learning model is trained in an unsupervised setup based on training data including previous images acquired by the or other front-layer detector having time lag effects and previous images acquired by the or other back-layer detector having circuitry for reducing time lag effects with such circuitry activated.

7. 7. The system of claim 1, further comprising an output port for sending the corrected input image for further processing by a spectral processing algorithm and / or a tomographic image reconstruction algorithm.

8. An imaging device comprising the system according to any one of claims 1 to 7 and the multi-energy X-ray imaging device having the front layer detector and the back layer detector.

9. The imaging device of claim 8 , wherein the front layer detector and the back layer detector are provided as modules of a multi-layer detector system.

10. 1. A computer-implemented method for acquisition lag correction in x-ray imaging, comprising: receiving an input including an input image acquired by a front layer detector in a multi-energy x-ray imaging device and an additional image acquired by a back layer detector of the multi-energy x-ray imaging device, the back layer detector having circuitry for reducing time lag effects, the additional image being acquired by the back layer detector while the circuitry is activated; processing the input image and the additional input image to generate a corrected input image representing an estimate of image information of the front layer detector at an acquisition lag lower than the acquisition lag experienced for the input image; A method comprising:

11. 6. A method for training the machine learning model in a system according to claim 4 or 5, comprising generating training data by imaging a phantom from a plurality of relative directions, each in the same imaging geometry, with and without activation of the lag reduction circuitry for the front layer detector.

12. 7. The method of training a machine learning model in the system of claim 6, comprising generating training data by imaging a phantom using the front layer detector and the back layer detector at different positions of the phantom.

13. A computer program configured, when executed by at least one processing unit, to cause the processing unit to carry out the method of any one of claims 10 to 12.

14. At least one computer readable medium storing the computer program of claim 13 or the trained machine learning model of the system of claims 3 to 6.