Image processing device, radiography system, image processing method for image processing device, and program

The image processing apparatus efficiently processes large and varying-sized medical X-ray video images in real-time using a trained model and neural network for noise reduction, addressing the challenge of maintaining performance in medical imaging.

JP2026122171APending Publication Date: 2026-07-28CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
CANON KK
Filing Date
2025-01-15
Publication Date
2026-07-28

AI Technical Summary

Technical Problem

Medical X-ray video imaging requires real-time processing of large video images with varying sizes, which existing technologies often compromise real-time performance due to increased processing times.

Method used

An image processing apparatus that acquires radiation images, generates partial images based on shooting mode, and uses a trained model to infer and process these images efficiently, utilizing a multi-layer neural network for noise reduction without compromising real-time performance.

Benefits of technology

Enables effective noise reduction processing in real-time, handling varying image sizes with improved diagnostic ability and efficient image handling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026122171000001_ABST
    Figure 2026122171000001_ABST
Patent Text Reader

Abstract

To provide an image processing device capable of noise reduction processing without compromising real-time performance. [Solution] The image processing apparatus disclosed herein is An acquisition unit that acquires a first radiation image obtained by video recording using a radiation detector, The system includes an inference processing unit that generates a plurality of partial images using the first radiation image based on processing conditions corresponding to the shooting mode of the radiation detector, infers the plurality of partially processed partial images by inputting the plurality of partially processed partial images into a trained model, and obtains a second radiation image obtained by image processing the first radiation image using the plurality of partially processed partial images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an image processing apparatus, a radiation imaging system, an image processing method of the image processing apparatus, and a program.

Background Art

[0002] In recent years, radiation imaging systems equipped with a detection unit for detecting radiation such as X-rays have been widely used in fields such as industry and medicine. In particular, in the field of X-ray video imaging, digital radiation imaging systems that convert incident X-rays into visible light by a phosphor and obtain a moving image using a semiconductor sensor have become widespread. Here, a moving image is a set of a plurality of still images collected continuously, and hereinafter, each still image in the moving image is referred to as a frame.

[0003] In such a radiation imaging system, various image processes are applied to the image acquired by the semiconductor sensor to enhance the diagnostic ability (an index indicating the value of an image for diagnosis). As an example, noise reduction processing can be mentioned. In a series of imaging processes, it is known that various noises such as quantum noise due to fluctuations of X-ray quanta and system noise generated from a detector and a circuit are generated and superimposed on the image. Due to this phenomenon, the granularity of the obtained moving image may deteriorate, and the diagnostic ability may decrease.

[0004] In particular, in medical X-ray video imaging, imaging with a small X-ray dose is recommended from the viewpoint of exposure to the subject. Therefore, it is important to improve the diagnostic ability by suitable noise reduction processing.

[0005] Furthermore, in medical X-ray video imaging, it is common to handle images (large-sized moving images) in real time, consisting of several million to tens of millions of pixels per frame, such as 2688 pixels wide x 2688 pixels high. For example, when operating in real time at 15 FPS, the processing time per frame must be at least 66 ms or less. Another characteristic is that the image size handled varies depending on the imaging procedure, and efficient, high-speed processing of input images of various sizes is required.

[0006] Patent Document 1 describes a more efficient noise reduction process that utilizes machine learning-based technologies such as deep learning. It also proposes a configuration in which an image is divided into computational ROIs of a fixed size (for example, 256 x 256 pixels) and processed using a neural network.

[0007] Furthermore, Non-Patent Document 1 proposes a machine learning-based noise reduction processing technique for moving images. [Prior art documents] [Patent Documents]

[0008] [Patent Document 1] Japanese Patent Publication No. 2023-77989 [Non-patent literature]

[0009] [Non-Patent Document 1] “FastDVDnet:Towards Real-Time Deep Video Denoising Without Flow Estimation”, M Tassano,et.al,IEEE / CVF Conference on Computer Vision and Pattern Recognition(CVPR),2020,pp.1354~1363 [Overview of the project] [Problems that the invention aims to solve]

[0010] Medical X-ray video imaging requires real-time processing of large video images. Non-patent document 1 suggests that, for example, depending on the image size, the processing time may be long, impairing real-time performance.

[0011] Furthermore, in medical X-ray video imaging, different imaging techniques may result in the capture of moving images with varying image sizes.

[0012] In Patent Document 1, for example, depending on the size of the image, the processing time may increase, which may impair real-time performance.

[0013] Therefore, one of the purposes of this disclosure is to provide an image processing device that enables noise reduction processing without compromising real-time performance by shortening the time required for machine learning-based noise reduction processing. [Means for solving the problem]

[0014] The image processing apparatus disclosed herein is An acquisition unit that acquires a first radiation image obtained by video recording using a radiation detector, The system includes an inference processing unit that generates a plurality of partial images using the first radiation image based on processing conditions corresponding to the shooting mode of the radiation detector, infers the plurality of partially processed partial images by inputting the plurality of partially processed partial images into a trained model, and obtains a second radiation image obtained by image processing the first radiation image using the plurality of partially processed partial images. [Effects of the Invention]

[0015] According to this disclosure, an image processing apparatus capable of noise reduction processing without compromising real-time performance can be provided. [Brief explanation of the drawing]

[0016] [Figure 1](a)Shows an example of the schematic configuration of the radiation imaging system and the radiation detector according to Embodiment 1. (b)Shows an example of the schematic configuration of the radiation detector according to Embodiment 1. [Figure 2] (a)Shows an example of the schematic configuration of the control unit and the noise reduction processing unit according to Embodiment 1. (b)Shows an example of the schematic configuration of the noise reduction processing unit according to Embodiment 1. [Figure 3] (a)Shows an example of the schematic configuration of the learned model according to Embodiment 1. (b)Shows an example of the schematic configuration of the CNN according to Embodiment 1. (c)Is a diagram for explaining an operation example of the learning process according to Embodiment 1. [Figure 4] Shows an example of the schematic configuration of the CNN according to Embodiment 1. [Figure 5] Is a flowchart of the operation of the radiation imaging system according to Embodiment 1. [Figure 6] Is a flowchart of the operation of the radiation imaging system according to Embodiment 1. [Figure 7] (a)Is a diagram showing an input image according to Embodiment 1 and an example of the division of the calculated ROI. (b)Is a diagram showing an example of the calculated ROI according to Embodiment 1. (c)Is a diagram showing an input image according to Embodiment 1 and an example of the division of the calculated ROI. [Figure 8] (a)Is a diagram showing the relationship between the size of the calculated ROI and the calculation amount of the CNN according to Embodiment 1. (b)Is a diagram showing the relationship between the size of the calculated ROI and the number of calculated ROIs according to Embodiment 1. (c)Is a diagram showing the relationship between the size of the calculated ROI, the number of calculated ROIs, and the calculation amount of the CNN according to Embodiment 1. [Figure 9] (a)Is a table showing the mode of the radiation detector according to Embodiment 1 and an example of the output image size. (b)Is a diagram showing the relationship between the size of the calculated ROI and the calculation amount of the CNN according to Embodiment 1. [Figure 10] (a)Is a table showing the mode of the radiation detector according to Embodiment 1 and an example of the output image size. (b)Is a diagram showing the relationship between the size of the calculated ROI and the calculation amount of the CNN according to Embodiment 1. [Figure 11]Table showing the modes of the radiation detector according to Embodiment 1 and examples of the image sizes to be output. [Figure 12] (a) Table showing an example of the relationship between the size of the calculated ROI and the filling efficiency of the calculated ROI according to Embodiment 2. (b) Diagram showing an example of the relationship between the size of the calculated ROI and the filling efficiency of the calculated ROI according to Embodiment 2. [Figure 13] (a) An example of a schematic configuration of an image processing unit according to Embodiment 3 is shown. (b) An example of a schematic configuration of a super-resolution processing unit according to Embodiment 3 is shown. [Figure 14] An example of a schematic configuration of a learned model according to Embodiment 3 is shown.

Mode for Carrying Out the Invention

[0017] Hereinafter, exemplary embodiments for implementing the present disclosure will be described in detail with reference to the drawings. However, the dimensions, materials, shapes, relative positions of the components, etc. described in the following embodiments are arbitrary and can be changed according to the configuration of the device to which the present disclosure is applied or various conditions. Also, in the drawings, the same reference numerals are used between the drawings to indicate elements that are the same or functionally similar.

[0018] Hereinafter, a radiation imaging system using X-rays as an example of radiation will be described. However, the radiation may be X-rays or other radiation. In the following embodiments, the term "radiation" can include, for example, electromagnetic radiation such as X-rays and γ-rays, and particle radiation such as α-rays, β-rays, particle beams, proton beams, heavy ion beams, and neutron beams.

[0019] In the following, "machine learning model" refers to a learning model based on a machine learning algorithm. Specific machine learning algorithms include nearest neighbors, naive Bayes, decision trees, and support vector machines. Neural networks and deep learning may also be used. Appropriately, any of the above algorithms that are available can be applied to the following embodiments and modifications. Furthermore, "training data" refers to a dataset used to train a machine learning model, and consists of pairs of input data that are input to the machine learning model and ground truth data (training data) that represent the correct output results of the machine learning model.

[0020] A pre-trained model refers to a machine learning model that has been trained using appropriate training data in advance, following any machine learning algorithm such as deep learning. However, while a pre-trained model is obtained through prior training with appropriate training data, it is not the case that further training is not possible; additional training can be performed. Additional training can be performed even after the device has been installed at the user's site.

[0021] (Embodiment 1) (Configuration of the radiography system) Hereinafter, with reference to Figures 1(a) and 1(b), a radiography system, an image processing device, and an image processing device method according to Embodiment 1 of this disclosure will be described.

[0022] Figure 1(a) shows a schematic example of the configuration of the radiography system 1 according to this embodiment. In the following description, the object to be inspected O will be described as a human body, but the object to be inspected O photographed by the radiography system according to this disclosure is not limited to a human body, but may be other animals, plants, or objects to be tested using non-destructive testing.

[0023] The radiography system 1 according to this embodiment includes a radiation detector 10, a control unit 20, a radiation generator 30, an input unit 40, and a display unit 50. The radiography system 1 may also include an external storage device 70, such as a server, connected to the control unit 20 via a network 60, such as the Internet or an intranet.

[0024] The radiation generator 30 is equipped with a radiation source, such as an X-ray tube, and can emit radiation. The radiation detector 10 can detect the radiation emitted from the radiation generator 30 and generate a radiation image corresponding to the detected radiation. Therefore, the radiation detector 10 can generate a radiation image of the object under inspection O by detecting the radiation emitted from the radiation generator 30 and passing through the object under inspection O.

[0025] Here, Figure 1(b) shows an example of a schematic configuration of the radiation detector 10 according to this embodiment. The radiation detector 10 is provided with a phosphor 11 and an imaging sensor 12. The phosphor 11 converts radiation incident on the radiation detector 10 into light of a wavelength detectable by the imaging sensor 12. The phosphor 11 may include, for example, CsI or GOS (Gd2O2S). The imaging sensor 12 includes a photoelectric conversion element composed of, for example, a-Si or crystalline Si, and can detect light corresponding to the radiation converted by the phosphor 11 and output a signal corresponding to the detected light. The radiation detector 10 can generate a radiation image by performing A / D conversion or the like on the signal output by the imaging sensor 12.

[0026] Although not shown in Figure 1(b), the radiation detector 10 may include a calculation unit, an A / D conversion unit, etc. Furthermore, a grid may be installed between the radiation detector 10 and the object under inspection O to reduce scattered radiation that is generated when radiation passes through the object under inspection O and reaches the radiation detector 10.

[0027] The control unit 20 is connected to the radiation detector 10, the radiation generator 30, the input unit 40, and the display unit 50. The control unit 20 can acquire radiation images output from the radiation detector 10, perform image processing on the radiation images, and control the operation of the radiation detector 10 and the radiation generator 30. As a result, the control unit 20 can control the radiation generator 30 to generate radiation under predetermined shooting conditions at the appropriate timing, enabling video recording at any frame rate. Furthermore, the control unit 20 can function as an example of an image processing device.

[0028] The control unit 20 may be connected to an external storage device 70 via any network 60 such as the Internet or an intranet, and may acquire radiation images, etc., from the external storage device 70. Furthermore, the control unit 20 may be connected to other radiation detectors, radiation generators, etc., via the network 60. The control unit 20 may be connected to the external storage device 70, etc., by wire or by wireless connection.

[0029] The input unit 40 is equipped with input devices such as a mouse, keyboard, trackball, or touch panel, and can receive instructions from the control unit 20 by being operated by the operator. The display unit 50 includes, for example, any monitor and can display information and images output from the control unit 20, as well as information input by the input unit 40.

[0030] In this embodiment, the control unit 20, input unit 40, display unit 50, etc., are configured as separate devices, but they may be configured as an integrated unit. For example, the input unit 40 and display unit 50 may be configured as a touch panel display. Also, in this embodiment, the control unit 20 constitutes the image processing device, but the image processing device only needs to be able to acquire radiation images and perform image processing on the radiation images, and does not need to control the driving of the radiation detector 10 or the radiation generator 30.

[0031] Furthermore, the control unit 20, radiation detector 10, radiation generator 30, etc. may be connected by wire or wirelessly. In addition, the external storage device 70 may constitute an image system such as PACS within the hospital, or it may be a server outside the hospital.

[0032] (Configuration of the control unit) Next, the more specific configuration of the control unit 20 will be described with reference to Figures 2(a), 2(b), and 2(c).

[0033] Figure 2(a) shows a schematic example of the configuration of the control unit 20 according to this embodiment, and Figure 2(b) shows a schematic example of the configuration of the noise reduction processing unit 26 according to this embodiment.

[0034] The control unit 20 includes an acquisition unit 21, an image processing unit 22, a display control unit 23, a drive control unit 24, and a storage unit 25. Figure 2(c) shows an example of a schematic configuration of the inference processing unit 262 according to this embodiment.

[0035] The acquisition unit 21 can acquire radiation images output by the radiation detector 10 and various information input by the input unit 40. The acquisition unit 21 can also acquire radiation images, patient information, etc. from an external storage device 70 or the like.

[0036] The image processing unit 22 is equipped with a noise reduction processing unit 26 and a diagnostic image processing unit 27, and can perform the image processing according to this disclosure on the radiographic image acquired by the acquisition unit 21. In this embodiment, noise reduction processing will be described as an example of image processing performed by the image processing unit 22.

[0037] The noise reduction processing unit 26 is equipped with a learning processing unit 261, as shown in Figure 2(b). In addition to the inference processing unit 262 and the trained model selection unit 263, the learning processing unit 261 is equipped with a training data generation unit 264, a parameter update unit 265, and a processing condition determination unit 268. Furthermore, the noise reduction processing unit 26 has a pre-processing unit 266 that converts the image input to the noise reduction processing unit 26 into a format suitable for processing by the learning processing unit 261, and a post-processing unit 267 that applies appropriate processing to the output result of the learning processing unit 261. With this configuration, the noise reduction processing unit 26 can train a machine learning model for noise reduction processing. The noise reduction processing unit 26 can also apply noise reduction processing suitable for radiographic images using the trained machine learning model. Note that the noise reduction processing unit 26 may also perform noise reduction processing using trained parameters learned by other learning devices. In other words, the noise reduction processing unit 26 does not have to be configured to perform both machine learning model training and noise reduction processing (inference processing using trained parameters).

[0038] Furthermore, the diagnostic image processing unit 27 can perform diagnostic image processing on the image that has undergone noise reduction by the noise reduction processing unit 26 to convert it into an image suitable for diagnosis. Diagnostic image processing includes, for example, tone processing to adjust the gradation of the image, enhancement processing to highlight specific pixels in the image, and grid fringe reduction processing to reduce grid fringes in the image. The diagnostic image processing unit 27 may, for example, perform tone processing, enhancement processing, grid fringe reduction processing, etc., according to a region of interest (ROI) set in the radiographic image. For example, tone processing may be performed to broaden the gradation of the region of interest, and enhancement processing may be performed to enhance the region of interest. Here, the region of interest may be set according to the operator's instructions, or it may be set based on the imaging site, disease name information, findings information, etc.

[0039] Next, the configuration of the learning processing unit 261 will be described. The learning processing unit 261 performs the learning process that is applied when training a machine learning model. The learning processing unit 261 includes an inference processing unit 262, a trained model selection unit 263, a training data generation unit 264, a parameter update unit 265, and a processing condition determination unit 268.

[0040] When performing the learning process, the learning processing unit 261 receives images that have been appropriately processed by the preprocessing unit 266, and the learning data generation unit 264 creates the learning data. Here, an example configuration is shown in which images with artificial noise added are used as input data and images without added noise are used as ground truth data, as a set of learning data for learning noise reduction processing. The learning data generation unit 264 creates a set of learning data by adding artificial noise, which is created by simulating the characteristics of radiation images, to the input images. Here, the noise added by the learning data generation unit 264 may reflect the amount of noise that may vary due to manufacturing variations of the radiation detector 10, as calculated by the learning data generation unit 264.

[0041] The parameter update unit 265 updates the parameters of the machine learning model held by the inference processing unit 262 based on the calculation results of the inference processing unit 262 on the input data and the ground truth data.

[0042] The processing condition determination unit 268 determines various processing conditions for the learning process and the inference process. The detailed processing of the processing condition determination unit will be described later.

[0043] When a radiation image is input to a trained model that has been trained using the training data described above, the inference processing unit 262 generates an image in which image processing has been applied to the radiation image through inference processing. The trained model selection unit 263 selects a trained model to be used by the inference processing unit 262. Here, the trained models obtained by the series of training processes of the training processing unit 261 may be multiple for each model of radiation detector 10, for example, multiple for each type of phosphor 11, or multiple for each type of imaging sensor 12, or multiple for each binning, sensitivity, image size, frame rate, and imaging procedure for a single model of radiation detector 10. The trained model selection unit 263 selects at least one trained model from among the multiple trained models to be used by the inference processing unit 262.

[0044] Here, a portion of the learning processing unit 261 does not need to be included in the control unit 20. For example, the components other than the inference processing unit 262 and the trained model selection unit 263 may be configured on hardware other than the control unit 20 (such as a server). This hardware creates a trained model by performing training in advance using appropriate training data. In this case, the control unit 20 may access this other hardware via the inference processing unit 262 to obtain the trained model and perform only processing using that trained model. Alternatively, the trained model may be provided in advance in the noise reduction processing unit 26, and the control unit 20 may perform only processing using that trained model.

[0045] Alternatively, the learning processing unit 261 may be included in the control unit 20, allowing for additional learning using the learning data acquired after installation (sale) to the customer.

[0046] The display control unit 23 can control the display of the display unit 50. For example, it can display radiographic images before and after image processing by the image processing unit 22, patient information, etc., on the display unit 50.

[0047] The drive control unit 24 can control the driving of the radiation detector 10 and the radiation generator 30, etc. Therefore, the control unit 20 can control the acquisition of radiation images by controlling the driving of the radiation detector 10 and the radiation generator 30 by the drive control unit 24.

[0048] The memory unit 25 can store programs for implementing various application software, including the operating system (OS), device drivers for peripheral devices, and programs for performing the processing described later. The memory unit 25 can also store information acquired by the acquisition unit 21 and radiation images processed by the image processing unit 22. For example, the memory unit 25 can store radiation images acquired by the acquisition unit 21, or radiation images that have undergone noise reduction processing, as described later.

[0049] The control unit 20 can be configured using a general-purpose computer including a processor and memory, but it may also be configured as a dedicated computer for the radiography system 1. Here, the control unit 20 functions as an example of an image processing device according to this embodiment, but the image processing device according to this embodiment may be a separate (external) computer that is communicatively connected to the control unit 20. Furthermore, the control unit 20 and the image processing device may be, for example, personal computers, and desktop PCs, notebook PCs, or tablet PCs (portable information terminals) may be used. The processor may be a CPU (Central Processing Unit). The processor may also be, for example, an MPU (Micro Processing Unit), a GPU (Graphical Processing Unit), or an FPGA (Field-Programmable Gate Array).

[0050] Each function of the control unit 20 may be realized by a processor such as a CPU or MPU executing software modules stored in the memory unit 25. The processor may be, for example, a GPU or FPGA. Furthermore, each function may be configured by a circuit that performs a specific function, such as an ASIC. For example, the image processing unit 22 may be realized by dedicated hardware such as an ASIC, and the display control unit 23 may be realized using a dedicated processor such as a GPU that is different from the CPU. The memory unit 25 may be configured by any storage medium such as an optical disk such as a hard disk or memory.

[0051] (Machine learning model configuration) Next, with reference to Figures 3(a) to 3(c), an example of a machine learning model that constitutes the trained model according to this embodiment will be described. An example of a machine learning model used by the inference processing unit 262 according to this embodiment is a multi-layer neural network.

[0052] Figure 3(a) shows a schematic example of the neural network model according to this embodiment. The neural network model configuration 33 shown in Figure 3(a) is designed to output noise-reduced inference data 32 according to a trained model in response to input data 31. The output noise-reduced inference data 32 is based on the learning content in the machine learning process. The neural network according to this embodiment learns features for distinguishing between signals and noise contained in the input radiation image. In the example shown in Figure 3(a), the input data 31 includes the current frame and one or more frames prior to the current frame. Alternatively, the input data 31 includes the current frame and one or more frames in the future. Alternatively, the input data 31 includes any set of frames including the current frame, one or more frames prior to the current frame, and one frame in the future. The noise-reduced inference data 32 is the current frame with reduced noise. It is also possible to configure a trained model with only the current frame (one image) as input data 31, where the number of input frames is one. The input data 31 is an example of a first radiation image. Furthermore, inference data 32 is an example of a second radiographic image.

[0053] Furthermore, at least a portion of the multi-layer neural network may be a convolutional neural network (CNN), for example. Additionally, at least a portion of the multi-layer neural network may utilize techniques related to autoencoders or vision transformers (ViT).

[0054] This section describes an example of using a Convolutional Neural Network (CNN) as a machine learning model for noise reduction processing of radiographic images. Figure 3(b) shows an example of a schematic configuration 33 of the CNN that constitutes the neural network model according to this embodiment. In the example of the trained model according to this embodiment, when input data 31, which is a radiographic image, is input, inference data 32 can be output as a radiographic image with reduced noise.

[0055] The CNN shown in Figure 3(b) is composed of multiple layers responsible for processing the input data set and producing an output. The types of layers included in the CNN configuration 33 are convolutional layers, downsampling layers, upsampling layers, and merge layers. Here, the CNN configuration 33 further includes an additive layer 34, and it is preferable to configure a shortcut that adds the input data before output. This allows the CNN to adopt a configuration that learns the difference between the input data and the output data, and can suitably handle systems that target noise.

[0056] A convolutional layer is a layer that performs convolution on an input set of values ​​according to parameters such as the kernel size of the set filter, the number of filters, the stride value, and the dilation value. The dimensionality of the filter kernel size may also be changed depending on the dimensionality of the input image.

[0057] A downsampling layer is a layer that performs a process to reduce the number of output values ​​to less than the number of input values ​​by decimating or combining input values. Specifically, one example of such a process is Max Pooling.

[0058] An upsampling layer is a layer that performs a process to increase the number of output values ​​to the number of input values ​​by duplicating the input values ​​or adding interpolated values ​​from the input values. Specifically, one example of such a process is upsampling by deconvolution.

[0059] A synthesis layer is a layer that takes a set of values, such as the output values ​​of a certain layer or the pixel values ​​that make up an image, as input from multiple sources and performs processing to combine them by concatenating or adding them together.

[0060] It should be noted that different parameter settings for the layers and nodes that make up a neural network may result in differences in the degree to which the features of the data trained from the training data can be reproduced during inference. In other words, the appropriate parameters often differ depending on the implementation method, so they can be changed as needed.

[0061] In addition to changing the parameters as described above, the CNN may also achieve better characteristics by changing its configuration. These better characteristics include, for example, outputting radiation images with better noise reduction, shorter processing times, and shorter training times for machine learning models.

[0062] The CNN configuration 33 used in this embodiment is a U-net type machine learning model having the functionality of an encoder consisting of multiple layers including multiple downsampling layers, and the functionality of a decoder consisting of multiple layers including multiple upsampling layers. In a U-net type machine learning model, for example, skip connections can be used. That is, positional information (spatial information) that has been obscured in the multiple layers configured as an encoder can be used in layers of the same dimension (layers corresponding to the dimensions of the encoder) in the multiple layers configured as a decoder.

[0063] Although not shown in the diagram, one example of modifying the CNN configuration is to incorporate layers with activation functions (e.g., ReLu: Rectifier Linear Unit) before and after the convolutional layers.

[0064] Through these steps in the CNN, noise features can be extracted from the input radiation images.

[0065] Here, the learning processing unit 261 includes a parameter update unit 265. As shown in Figure 3(c), the parameter update unit 265 calculates a loss function from the inference data 32 obtained by applying the neural network model of the inference processing unit 262 to the input data 31 in the learning data, and from the ground truth data 35 in the learning data. Subsequently, the parameter update unit 265 updates the parameters of the neural network model based on the calculated loss function. Here, the loss function represents the error between the inference data 32 and the ground truth data 35.

[0066] The parameter update unit 265 can update the filter coefficients of the convolutional layer, for example, using backpropagation, so as to reduce the error between the inference data 32 represented by the loss function and the ground truth data 35. Backpropagation is a method for adjusting the parameters between each node of the neural network so as to reduce the above error. In addition, a method of randomly deactivating the units (each neuron or each node) that make up the CNN (dropout) may be used for training.

[0067] Furthermore, the pre-trained model used by the inference processing unit 262 may be generated using transfer learning. In this case, for example, a pre-trained model used for noise reduction processing may be generated by performing transfer learning on a machine learning model trained on radiographic images of objects O of different types. By performing such transfer learning, it is possible to efficiently generate a pre-trained model even for objects O of which it is difficult to obtain a large amount of training data. Here, objects O of different types may be, for example, animals, plants, or objects used in non-destructive testing.

[0068] Here, the GPU can perform calculations efficiently by processing more data in parallel. Therefore, when performing training multiple times using a machine learning model that utilizes a CNN as described above, it is effective to perform the processing on the GPU. Accordingly, the learning processing unit 261 in this embodiment can use a GPU in addition to the CPU. Specifically, when executing a learning program that includes a machine learning model, the CPU and GPU work together to perform calculations to perform the training. Note that the training process may be performed by the CPU or the GPU alone. Furthermore, each process of the inference processing unit 262 may also be implemented using the GPU in the same way as the learning processing unit 261. Alternatively, it is also possible to use a dedicated computing unit specialized for inference processing.

[0069] The above describes the configuration of the machine learning model, but the machine learning model is not limited to the CNN-based model shown above. The machine learning model used in this embodiment can be any machine learning-like model that is capable of extracting (representing) the features of training data such as images through learning.

[0070] Here, the learning processing unit 261 according to this embodiment can use any set of training data for learning noise reduction processing. For example, the learning processing unit 261 can use training data in which images with artificial noise added are input data and images without added noise are ground truth data. In addition, for example, learning may be performed using images before averaging as input data and images after averaging as ground truth data, or images before statistical processing such as MAP (maximum posterior probability) estimation as input data and images after statistical processing as ground truth data. Furthermore, although examples of supervised learning have been shown so far, the learning method is not limited to this, and any unsupervised learning or semi-supervised learning method may be used.

[0071] (Operation of the noise reduction processing unit) Next, Figure 4 will be used to explain the detailed operation of the noise reduction processing unit 26 in video recording. In video recording, past frames in the vicinity of the current frame often capture similar structures. Therefore, when performing noise reduction on the target pixel of the current frame, not only spatial information such as similar structures in the surrounding area within the same frame, but also temporal information such as similar structures in past or future frames can be used in the noise reduction process.

[0072] From this perspective, the inference processing unit 262 can process data while utilizing more temporal information by inputting multiple frames. In this case, in order to perform real-time processing, the inference processing unit 262 needs to be configured to input a total of N frames, consisting of the current frame and a predetermined number of past or future frames. The number of past frames to be used will vary depending on the frame rate during shooting and the required noise reduction performance, but below, as a suitable example, we will explain the case where N=10 and the current frame and past frames are used as input.

[0073] Figure 4 is a schematic diagram of the neural network configuration for N=10. The current frame number is denoted as t, and the current frame t and past frames numbered t-1 to t-9 are sequentially input to a pre-trained CNN41. CNN41 is an example of a machine learning model used by the inference processing unit 262. Using such a CNN41, the inference processing unit 262 can obtain a noise-reduced image F(t) by performing noise reduction processing using the spatial information of the current frame t and the temporal information of the nine past frames.

[0074] (Operation of the inference processing unit) Next, Figure 5 will be used to describe the detailed operation of the noise reduction processing unit 26 and the inference processing unit 262 during video recording. Figure 5 shows an example of the flow of the noise reduction processing unit 26 and the inference processing unit 262.

[0075] This section describes the operation when a pre-trained model is already available, and inference processing is performed using the pre-trained model during video recording to obtain an image after noise reduction processing.

[0076] First, in step S501, when the video recording sequence starts, the noise reduction processing unit 26 sets the initial frame number to t=0.

[0077] In step S502, the noise reduction processing unit 26 acquires the image of frame t via the acquisition unit 21. Initially, the noise reduction processing unit 26 acquires the frame at t=0.

[0078] In step S503, the preprocessing unit 266 performs preprocessing on the image acquired in step S502 to perform appropriate inference processing, thereby obtaining a preprocessed image. The method of preprocessing is not particularly limited. For example, in noise reduction processing, preprocessing may include square root transformation, logarithmic transformation, Anscombe transformation, etc. These transformations make quantum noise following a Poisson distribution approximately constant regardless of the intensity of the irradiated radiation, so that the noise contained in the input image can be treated as additive noise. Furthermore, appropriate preprocessing can be performed depending on the content of the image processing. For example, centering can be performed to set the mean of the data to 0 in order to stabilize processing by the neural network. Alternatively, standardization can be performed to set the standard deviation of the data to 1. Alternatively, normalization can be performed to normalize the data to a range of 0 to 1. Alternatively, both centering to set the mean of the data to 0 and standardization to set the standard deviation of the data to 1 can be performed. The results of the above preprocessing can be temporarily stored in memory as needed for use in the inference processing of subsequent frames. Furthermore, it is desirable that the preprocessing performed in preprocessing step 266 be the same during both inference and training.

[0079] In step S504, appropriate image segmentation is performed according to the various processing conditions determined by the processing condition determination unit 268. Details of the processing by the processing condition determination unit will be described later.

[0080] In step S505, the inference processing unit 262 performs inference processing on the segmented images obtained in step S504 using the trained model. This makes it possible to obtain images to which noise reduction processing has been applied.

[0081] In step S506, the divided images are combined according to the various processing conditions determined by the processing condition determination unit 268.

[0082] In step S507, the post-processing unit 267 performs post-processing on the combined image obtained in step S506. The post-processing reverses the operations performed in the pre-processing in step S503, such as various normalization and leveling inverse transformations, removal of padding, and merging of multiple divided ROIs.

[0083] In step S508, the noise reduction processing unit 26 determines whether or not to terminate image acquisition. The noise reduction processing unit 26 may determine the termination of image acquisition based, for example, on set shooting conditions or instructions from the operator. If image acquisition is to continue, the process moves to step S509. In step S509, the noise reduction processing unit 26 adds 1 to the frame number t and then moves the process to step S502, where the noise reduction processing unit 26 repeats the processes from steps S502 to S508.

[0084] (Operation of the learning processing unit) Figure 6 will be used to explain the detailed operation of the noise reduction processing unit 26 and the learning processing unit 261 in video recording. Figure 6 shows an example of the flow of the noise reduction processing unit 26 and the learning processing unit 261. Hereafter, the learning processing unit 261 will be assumed to perform supervised learning, and the example will be explained using a set of training data for learning the noise reduction process, in which images with artificial noise added are used as input data and images without added noise are used as ground truth data. In this case, the CNN 41 to be trained will be assumed to be a system that takes input consisting of a total of N current frames and past frames, as shown in Figure 4, and outputs a current frame with noise reduced.

[0085] In step S601, the training data generation unit 264 randomly selects image data from the storage unit 25 which holds multiple image data. For data to be used to train the noise reduction process for video, it is preferable to use multiple video images, for example, each containing multiple frames. Furthermore, it is desirable that the video images have a good signal-to-noise ratio (SNR), and a configuration may be adopted in which the SNR is improved by performing a separate noise reduction process beforehand. The training data generation unit 264 randomly selects video images.

[0086] In step S602, the image data extraction process is performed according to the various processing conditions determined by the processing condition determination unit 268.

[0087] In step S603, the learning data generation unit 264 adds artificial noise, created by simulating the characteristics of radiation images, to the cropped images obtained in step S602. It is desirable to simulate artificial noise that mimics actual radiation images according to the sensitivity of the radiation detector 10, the noise characteristics of the readout circuit, the MTF (Modulation Transfer Function) of the phosphor 11, etc. Through the above operation, learning data can be constructed consisting of input data with artificial noise added and ground truth data without artificial noise added.

[0088] In step S604, the preprocessing unit 266 performs preprocessing on the training data obtained in step S603 to perform appropriate inference processing, thereby obtaining preprocessed training data. Details of the preprocessing are as described above.

[0089] In step S605, the inference processing unit 262 inputs the pre-processed input data into the CNN41, applies the parameters of the CNN41 during training, and outputs the inference result.

[0090] In step S606, the parameter update unit 265 updates the parameters of CNN41 based on the inference results obtained in step S605 and the ground truth data, in order to minimize an appropriate loss function.

[0091] In step S607, the learning processing unit 261 determines whether or not to terminate the learning process. The determination of whether to terminate the learning process can be made based on any criteria, such as the number of iterations, the value of the loss function, or whether or not overfitting has occurred. If the learning process is complete, the updating of the parameters of CNN41 is stopped, and the trained CNN can be used for the various inference processes described above. If the learning process is to continue, the process moves to step S601, and the learning processing unit 261 repeats the processes from steps S601 to S607.

[0092] (Operation of the processing condition determination unit) The detailed operation of the processing condition determination unit 268 will be explained using Figures 7 to 11. The processing condition determination unit 268 determines the appropriate processing conditions for domain partitioning by optimizing the cost function related to the calculations in the inference processing performed by the inference processing unit 262.

[0093] Figure 7(a) shows an example of an image input to the noise reduction processing unit 26. Let's assume the input image 71 has width Ix and height Iy. In the inference processing unit 262, a padding region of width px and height py is added to improve the image quality at the edges of the image, and an image 72 with width Ax and height Ay is obtained. Ax and Ay are given by Equation 1. px and py can take any value greater than or equal to 0. Note that the input image 71 is an example of a first radiation image. Also, image 72 is an example of a third radiation image.

[0094] Ax = Ix + 2 × px Ay = Iy + 2 × py formula 1 Here, if the size of image 72 exceeds the image size that can be processed at once by CNN41, it can be divided into multiple computational ROIs 73. Note that computational ROIs 73 are just one example of partial images.

[0095] Figure 7(b) shows the details of the calculated ROI73. Here, the positions of ROI73 in the X and Y directions are denoted as i and j, respectively, and R(i,j), R(i+1,j), and R(i,j+1) are schematically shown. ROI73 has a width Lx and a height Ly, and R(i,j) and R(i+1,j) overlap in the X direction by length Wxi, while R(i,j) and R(i,j+1) overlap in the Y direction by length Wyj. The overlapping regions defined by Wxi and Wyj are used in the region merging process described above, and the results of the calculations of the individual calculated ROIs are weighted and added together to make the seams between the calculated ROIs less noticeable.

[0096] Figure 7(c) shows the arrangement of the calculated ROI73 in image 72. As shown in Figure 7(c), ROI73 is arranged in a total of n × m locations, n in the X direction and m in the Y direction. n and m are given by Equation 2. i and j are integers, (i=1, 2..., n) (j=1, 2..., m). For simplicity, the following explanation will describe the case where the overlap length is equal except for the right edge and bottom edge of the image. That is, assume that Wx=Wx0=Wx1=...=Wxn-2≠Wxn-1 and Wy=Wy0=Wy1=...=Wym-2≠Wym-1. Here, Lx, Ly, Wx, Wxn-1, Wy, and Wym-1 can take any value. However, if each value is not above a certain level, it may negatively affect the image quality at the seams and the computational accuracy of the neural network. Therefore, it is desirable to set each setting to a certain value or higher, taking into account the image quality after processing.

[0097]

number

[0098] CNN41 receives a total of n × m inputs of calculated ROI73 with width Lx and height Ly.

[0099] Here, since CNN41 is a computational unit composed of addition, multiplication, branching, and other operations such as convolution and activation function operations that make up a neural network, the computational complexity associated with inference operations can be measured.

[0100] Figure 8(a) shows the relationship between the size of the computational ROI73 (Lx × Ly) and the computational cost of CNN41 per computational ROI.

[0101] Figure 8(b) is an example of a diagram showing the relationship between the size (Lx × Ly) of the calculated ROI73 and the number of calculated ROI73s. Note that in the following, appropriate values ​​are assumed for px, py, Wxi, and Wyj.

[0102] Figure 8(c) is an example of a diagram showing the relationship between the size of the calculated ROI73 (Lx × Ly), the number of calculated ROI73, and the total computational cost when all ROIs are calculated using CNN41.

[0103] As shown in Figure 8(a), the computational complexity of CNN41 is proportional to the size of the input computational ROI73, and the computational complexity of CNN41 decreases as Lx × Ly decreases. On the other hand, as shown in Figure 8(b), the number of computational ROI73 decreases as Lx × Ly increases.

[0104] Here, the total computational complexity of CNN41 when all computational ROIs are calculated is not proportional to the size of the computational ROI73. Furthermore, the total computational complexity is not proportional to the number of computational ROI73. As shown in Figure 8(c), the total computational complexity varies in a complex way depending on the size of the computational ROI73 (Lx × Ly) and the number of ROIs.

[0105] Based on the above, the processing condition determination unit 268 can determine the set of Lx, Ly, Wx, Wy, px, and py according to Equation 3. The computational complexity of CNN41 per computation ROI 73 is denoted as Comp. Note that Comp varies depending on the network structure, such as the configuration of the intermediate layers and the number of layers in CNN41.

[0106]

number

[0107] In Embodiment 1, the computational complexity of CNN41 is used as the first cost function used by the processing condition determination unit 268, and the optimization problem of minimizing this complexity under various conditions is addressed.

[0108] In the example shown in Figure 8(c), assume that the total computational complexity of CNN41 is minimized with parameter set 81, based on Equation 3. In that case, the processing condition determination unit 268 can determine that parameter set 81 is the condition that allows for the fastest inference processing.

[0109] Here, it is common for radiography systems to handle different image sizes depending on the imaging procedure. Figure 9(a) is a table listing examples of imaging modes for the radiation detector 10 and examples of output image sizes.

[0110] The radiation detector 10 has a binning mode that combines multiple pixels into a single pixel for signal reading. For example, in binning mode 2x2, four pixels are treated as one pixel, which reduces the number of pixels but increases the signal-to-noise ratio (SNR). Binning mode is particularly effective for low-dose imaging and video recording.

[0111] Furthermore, the radiation detector 10 typically has multiple sensitivity modes even within the same binning mode. For example, there is a high-sensitivity mode for obtaining highly sensitive images at low doses, which is used in pediatric and gynecological examinations where low radiation exposure is required. Another mode is a standard sensitivity mode that provides a balanced level of sensitivity and resolution. Yet another mode is a low-sensitivity mode used when high resolution is required for detailed observation of fractures or detection of fine structures.

[0112] For example, when photographing swallowing, a 1x1 binning mode and a 17-inch field of view in standard sensitivity mode are used. When photographing nasogastric sphincter function, a 2x2 binning mode and a 12-inch field of view in high sensitivity mode are used. Thus, it is common to use different modes and field of view depending on the application.

[0113] The noise characteristics of the radiation detector 10 may differ depending on the binning mode and sensitivity mode. Therefore, to perform noise reduction processing effectively, it is preferable for CNN 41 to have a pre-trained parameter for each mode. For example, it is preferable to have one pre-trained parameter for each binning mode and sensitivity mode shown in Figure 9(a). Figure 9(a) shows an example where one of type 1 to type 4 pre-trained parameters is used depending on the binning mode and sensitivity mode.

[0114] Here, noise reduction is a process where spatial frequency information contained in the image is crucial. For example, scaling the input image will cause spatial information to be lost, significantly degrading the performance of the noise reduction process. Therefore, for each trained parameter, the image sizes Lx and Ly of the ROI73 input to CNN41 must be the same during both training and inference.

[0115] The processing condition determination unit 268 can determine a set of Lx, Ly, Wx, Wy, px, and py to optimize a single learned parameter for inputs of multiple field angles, as described below.

[0116] One example of the configuration of the processing condition determination unit 268 is to determine the minimum total computational cost for all fields of view according to Equation 4.

[0117]

number

[0118] Figure 9(b) shows the relationship between the size of the calculated ROI 73 (Lx × Ly) and the total computational cost when calculating all ROIs using CNN41 for each field of view. Here, s=5, and as an example of the field of view shown in Figure 9(a), k=1 is 17 inches, k=2 is 14 inches, k=3 is 12 inches, k=4 is 9 inches, and k=5 is 6 inches. As shown in Figure 9(b), the processing condition determination unit 268 calculates the computational cost Compk of CNN41 for each k and determines the parameter set 91 that minimizes the sum of these costs as the processing condition. This makes it possible to calculate the processing condition that minimizes the average computational processing cost for all fields of view.

[0119] Another example of the configuration of the processing condition determination unit 268 is that when calculating the total computation time for all fields of view, it is possible to assign arbitrary weights to each field of view.

[0120] As an example, we will describe a configuration in which processing conditions are determined by weighting and adding information on the frequency of use of each field of view in each mode.

[0121] Figure 10(a) is a table listing examples of modes for the radiation detector 10 and other examples of output image sizes. Here, the frequency of use of each condition is shown for each mode and field of view.

[0122] The processing condition determination unit 268 can be configured to weight the calculations based on the frequency of use of each field of view in each mode, and to find the minimum weighted sum of the computational complexity for all fields of view according to Equation 5.

[0123]

number

[0124] Let weightk be the weight used in the weighted sum. In the example of trained parameter type1 shown in Figure 10(a), s=5, where k=1 corresponds to a field of view of 17 inches, k=2 to 14 inches, k=3 to 12 inches, k=4 to 9 inches, and k=5 to 6 inches.

[0125] Figure 10(b) shows the relationship between the size of the calculated ROI 73 (Lx × Ly) and the weighted sum of the computational complexity when all ROIs are calculated using CNN41 for each field of view. As shown in Figure 10(b), the processing condition determination unit 268 calculates Compk for each and determines the parameter set 101 that minimizes the weighted sum as the processing condition.

[0126] The weight 'weightk' can be calculated using the frequency of use 'Uratek' (0 ≤ Uratek ≤ 100) for each field of view, as shown in Equation 6.

[0127]

number

[0128] Alternatively, following Equation 7, it is possible to adopt a configuration that prioritizes only the most frequently used field of view.

[0129]

number

[0130] According to Equation 7, in the example of trained parameter type 1 shown in Figure 10(a), k=2 (field of view 14 inches) is the most frequently used, so kmax=2. Therefore, the weight weight2 for k=2 is 1, and the weights for k=1, 3, 4, and 5 are 0. As a result, in Equation 5, the processing conditions are determined so that the computational cost for k=2 (field of view 14 inches) is minimized.

[0131] With the above configuration, the processing condition determination unit 268 can determine processing conditions considering the frequency of use of the radiography system. The computational cost of frequently used modes accounts for a large portion of the overall computational cost of the radiography system. With the above configuration, the computational cost of frequently used modes is reduced, thus reducing the overall computational cost of the radiography system. As a result, the processing condition determination unit 268 can determine processing conditions that are optimal for the entire radiography system and enable high-speed inference processing.

[0132] As yet another example of the configuration of the processing condition determination unit 268, it can be configured to determine the processing conditions according to the amount of image data transferred from the radiation detector 10 to the control unit 20.

[0133] Figure 11 is a table listing examples of modes for the radiation detector 10 and other examples of output image sizes. Here, the maximum frame rate of the video and the maximum data transfer rate (amount of image data transferred per second) are shown for each mode and field of view. When a radiation imaging system performs real-time processing to display images immediately after acquisition, a higher data transfer rate means that a larger amount of image data must be processed per unit of time. Therefore, the conditions can be determined by prioritizing the mode with the highest data transfer rate.

[0134] Here, assuming that pixel information is stored in 16-bit unsigned short format (2 bytes), the data transfer rate Drate can be calculated using Equation 8 with the maximum frame rate f, Ix, and Iy in each mode.

[0135] Drate=Ix×Iy×2×f Formula 8 The processing condition determination unit 268 calculates the field of view conditions Ix and Iy that minimize the Drate under the conditions corresponding to each learned parameter.

[0136] In the example shown in Figure 11, for trained parameter type 1, condition 111 with a 12-inch field of view is selected. For type 2, condition 112 with a 14-inch field of view is selected. For type 3, condition 113 with a 14-inch field of view is selected. For type 4, condition 114 with a 17-inch field of view is selected. In this way, the field of view condition with the highest data transfer rate is selected.

[0137] The processing condition determination unit 268 applies Equation 3 to each of the field-of-view conditions obtained above and selects the parameter set that minimizes the total computation time of CNN41.

[0138] The above configuration allows us to determine the processing conditions that maximize the data transfer rate, minimize computational costs, and enable the fastest inference processing under the most computationally intensive situations.

[0139] In addition to the configuration described above, it is also possible to use a configuration with multiple pre-trained parameters for each binning mode and sensitivity mode, depending on the difference in field of view.

[0140] According to the configuration described above, the processing condition determination unit 268 can appropriately determine the parameter sets Lx, Ly, Wx, Wy, px, and py related to the calculated ROI 73. This makes it possible to provide an image processing device that performs processing at high speed by defining efficient processing conditions according to the processing system in machine learning-based noise reduction processing.

[0141] (Embodiment 2) The image processing unit according to Embodiment 2 of this disclosure will be described with reference to Figures 7 and 12. In Embodiment 1, the operation of the processing condition determination unit, in which the computational complexity of the CNN is used as the first cost function, was described. In Embodiment 2, as another operation of the radiography system, when it is difficult to uniquely calculate the computational complexity of the CNN, for example, when it changes depending on the situation, the processing conditions can be determined based on the efficiency of the array of computational ROIs (referred to as the packing efficiency). An example of a CNN whose computational complexity changes depending on the situation is a CNN with an architecture whose structure changes dynamically. Dynamically changing structure means, for example, dynamically activating or deactivating a part of the network according to the features of the input image. Note that the configuration of the radiography system according to this embodiment, other than the processing condition determination unit 268, is the same as the configuration of the radiography system 1 according to Embodiment 1, so the same reference numerals are used and the description is omitted.

[0142] As shown in Figure 7(c), the calculated ROI 73 is arranged with n units in the X direction and m units in the Y direction, for a total of n × m units. m and n are integers. If the filling efficiency of the calculated ROI 73 in the X direction in Image 72 is Ex and the filling efficiency in the Y direction is Ey, then the filling efficiency E of the calculated ROI 73 in Image 72 can be expressed by Equation 9. When E = 1, it indicates that the calculated ROI 73 can be placed in Image 72 without any waste.

[0143] In Embodiment 2, the filling efficiency E of the calculated ROI 73 in Image 72 is used as the second cost function by the processing condition determination unit 268, and the optimization problem of maximizing this efficiency under various conditions is addressed.

[0144]

number

[0145] Figure 12(a) is a table showing an example of the relationship between Lx, Wx, px, n, and Ex.

[0146] Figure 12(b) is an example of a diagram showing the relationship between the width Lx of the calculated ROI73 and the packing efficiency Ex in the X direction. While unique values ​​are shown for px and Wx here, these values ​​can be changed as needed.

[0147] As shown in Figure 12(a), the filling efficiency Ex in the X direction varies depending on the value of Lx. The processing condition determination unit 268 determines the parameter set 121 that maximizes the filling efficiency Ex in the X direction as the processing condition, as shown in Figure 12(b). Specifically, in Figure 12(a), the parameter set 121 of the row containing Lx=384 is determined as the processing condition. Note that the above discussion has been limited to parameters related to the X direction of the image, but the Y direction is the same as the X direction. The processing condition determination unit 268 can appropriately determine the parameter sets Lx, Ly, Wx, Wy, px, and py that govern the division of the calculated ROI 73 at a specific field of view, according to equation 10.

[0148]

number

[0149] Furthermore, when determining the optimal processing conditions for all angles of view, the processing condition determination unit 268 can adopt any configuration described in Embodiment 1, with each binning mode and sensitivity mode having a single learned parameter.

[0150] For example, the processing condition determination unit 268 can select a parameter set that maximizes the sum of the filling efficiency E across all fields of view. Alternatively, it can select a parameter set that maximizes the weighted sum across all fields of view. If the weights used in the weighted sum are denoted by eweightk, it is expressed by equation 11. Note that s in equation 11 is the same as s in equation 5.

[0151]

number

[0152] In addition, the processing condition determination unit 268 can also select a parameter set in which the filling efficiency E is maximized at the field of view where the data transfer rate is maximum.

[0153] With the above configuration, for example, when the computational complexity of a CNN varies depending on the situation, making it difficult to calculate uniquely, processing conditions can be determined based on the efficiency of filling the computational ROI.

[0154] (Embodiment 3) The image processing unit according to Embodiment 3 of this disclosure will now be described. In Embodiment 3, as another operation of the radiography system, it is assumed that a dedicated computing device is electrically connected to the radiography system and that CNN calculations are performed therein. The dedicated computing device is, for example, a computing device separate from the control unit 20 (such as an external GPU connected to the control unit). Depending on the dedicated computing device, some operations such as convolution and activation function calculations that constitute the neural network, such as addition, multiplication, and branching operations, may be performed particularly quickly, while others may be performed slowly. In this case, the amount of computation associated with the inference calculation calculated from the CNN structure described in Embodiment 1 may not necessarily correlate with the computing speed of the dedicated computing device. The dedicated computing device is an example of a second inference processing unit.

[0155] In Embodiment 3, processing conditions can be determined based on the calculation speed when inference calculations are performed by a dedicated arithmetic unit. Note that, apart from the processing condition determination unit 268 of the radiography system according to this embodiment, the configurations of the radiography system 1 according to Embodiment 1 are the same, and therefore, the same reference numerals are used and their description is omitted.

[0156] When determining the set of Lx, Ly, Wx, Wy, px, and py, the processing condition determination unit 268 can determine the processing conditions according to equation 12, based on the computation time PT of CNN41 per computation ROI when inference calculations are performed by a dedicated computing device.

[0157]

number

[0158] In Embodiment 3, the processing condition determination unit 268 uses the computation time PT of the CNN41 when inference calculations are performed by a dedicated computing unit as the third cost function, and deals with an optimization problem that minimizes this time under various conditions.

[0159] Furthermore, when determining the optimal processing conditions for all angles of view, the processing condition determination unit 268 can adopt any configuration described in Embodiment 1, with each binning mode and sensitivity mode having a single learned parameter.

[0160] For example, the processing condition determination unit 268 can select a parameter set that minimizes the total computation time of CNN41 or the weighted sum across all fields of view. Alternatively, it can select a parameter set that minimizes the computation time at the field of view where the data transfer rate is highest.

[0161] The above configuration can be applied when performing inference using a dedicated computing device provided in a radiography system. In other words, even when the computational complexity and computation speed associated with the inference calculation calculated from the CNN structure are not correlated, the above configuration allows for determining processing conditions based on the processing time of the computational ROI array.

[0162] (Embodiment 4) Up to this point, as an example of image processing performed by the image processing unit 22, we have described noise reduction processing by the noise reduction processing unit 26 as an example. However, this disclosure is not limited to this, and the above configuration can be adopted for any image processing using a machine learning model on moving images. Here, we will describe an example in which the image processing unit according to Embodiment 4 of this disclosure performs super-resolution processing to improve the resolution of an image as image processing on a moving image. Note that the configuration of the radiography system according to this embodiment, other than the image processing unit 22, is the same as the configuration of the radiography system 1 according to Embodiment 1, so the same reference numerals are used and the explanation is omitted.

[0163] Figure 13(a) shows an example of a schematic configuration of the image processing unit 22 according to this embodiment. In this embodiment, the image processing unit 22 is provided with a super-resolution processing unit 136 instead of a noise reduction processing unit 26.

[0164] Figure 13(b) shows an example of the schematic configuration of the super-resolution processing unit 136. The configuration of the super-resolution processing unit 316 is the same as that of the noise reduction processing unit 26 according to Embodiment 1, and the super-resolution processing unit 136 is provided with a learning processing unit 261, an inference processing unit 262, a pre-processing unit 266, and a post-processing unit 268. In addition to the configuration of the inference processing unit 262 and the trained model selection unit 263, the learning processing unit 261 is provided with a training data generation unit 264, a parameter update unit 265, and a processing condition determination unit 268.

[0165] Figure 14 shows a schematic example of the neural network model used for super-resolution processing in this embodiment. The neural network model configuration 143 shown in Figure 14 is designed to output inference data 142 with improved resolution according to pre-trained parameters for input data 141. The input data 141 includes one or more frames with a resolution lower than the desired resolution, and may be, for example, the current frame and one or more past or future frames. The neural network model has a configuration 143 that has been trained with a set of training data where low-resolution images are input data and images of the desired resolution are ground truth data. The neural network model is configured to output inference data 142, which is an image after super-resolution processing, when input data 141, which is a radiographic image, is input, according to the configuration 143. It is also possible to configure a trained model where the input data 141 is a single current frame, and the number of input frames is one.

[0166] The training data can be generated, for example, by using a low-resolution image as input data and a ground truth image generated by applying a known super-resolution process to the input data. Alternatively, the training data may be generated by using an image obtained using a radiation detector capable of acquiring images of the desired resolution as ground truth data and an image with reduced resolution as input image. Furthermore, the training data may be generated by using an image obtained with a low resolution set as the shooting condition as input data and an image obtained with a high resolution (desired resolution) set as the shooting condition as ground truth data. In addition, if the input data consists of multiple frames, similar to Embodiment 1, the frame to be processed and frames acquired in the past or future than that frame can be used as input data for the training data.

[0167] An example of a machine learning model used by the inference processing unit 262 according to this embodiment may be a multilayer neural network, and at least a part of the multilayer neural network may be, for example, a CNN. Furthermore, at least a part of the multilayer neural network may utilize techniques related to autoencoders. In addition, the trained model used by the inference processing unit 262 may be generated using transfer learning. In this case, for example, a trained model used for super-resolution processing may be generated by performing transfer learning on a machine learning model trained on radiographic images of objects O of different types. By performing such transfer learning, a trained model can be efficiently generated even for objects O of which it is difficult to obtain a large amount of training data. The objects O of different types referred to here may be, for example, animals, plants, objects used for non-destructive testing, etc.

[0168] In such systems, for example, as shown in Figure 7(a), if the size of image 72 exceeds the image size that can be processed at once by CNN41, it will be divided into multiple computational ROIs 73. Similar to noise reduction processing, super-resolution processing is a task in which spatial frequency information contained in the image is important, and scaling the input image will cause spatial information to be lost, resulting in a significant decrease in performance. Therefore, for each trained parameter, the image sizes Lx and Ly of the ROI 73 input to CNN41 must be the same during both training and inference. Also, as with noise reduction processing, it is desirable for the radiography system equipped with the super-resolution processing unit 136 to have a configuration with one trained parameter for each binning mode and sensitivity mode, as shown in Figure 9(a).

[0169] Therefore, it can be said that the super-resolution processing unit 136 in Embodiment 4 is a system that has the same problems as the noise reduction processing unit 26 described in Embodiments 1 to 3.

[0170] The processing condition determination unit 268 can operate with the configuration described in Embodiments 1 to 3, and the super-resolution processing unit 136 can also determine a set of Lx, Ly, Wx, Wy, px, and py to optimize for multiple field-of-view inputs with a single learned parameter.

[0171] This makes it possible to provide an image processing device that can perform high-speed processing even in machine learning-based super-resolution processing by defining efficient processing conditions according to the processing system.

[0172] In the control unit 20 according to this embodiment, image processing may include super-resolution processing to improve the resolution of the image. In such a configuration, by performing processing using a trained model by the super-resolution processing unit 136, real-time processing can be performed to obtain an image with suitably improved resolution, even immediately after capture.

[0173] Furthermore, the inference processing unit 262 may perform image processing using a trained model that combines not only the noise reduction processing described in Embodiments 1 to 3 and the super-resolution processing described in Embodiment 4, but also other diagnostic image processing techniques.

[0174] Regarding image processing, there are no restrictions on the type of processing, as long as the spatial frequency information contained in the image is important. Examples of types include enhancement, style transfer, grayscale processing, segmentation (especially those that do not involve scaling the input image), scattered radiation reduction, and grid fringe reduction.

[0175] In this case, the training data may include a set of data in which one or more images before various combined image processing are applied are used as input data, and one image after various processing is applied is used as the ground truth data. In this configuration as well, if the input data consists of multiple frames, the frame to be processed and frames acquired in the past or future than that frame can be used as input data for the training data, as described above. Therefore, the image processing performed by the inference processing unit 262 may include at least one of the following: noise reduction processing, super-resolution processing, enhancement processing, style transfer, grayscale processing, segmentation processing, scattered radiation reduction processing, and grid fringe reduction processing. With such a configuration, the control unit 20 can apply the desired image processing to the moving image suitably and in real time using the trained model.

[0176] (Variation 1) The machine learning model used by the inference processing unit 262 can be a combination of any layer configuration, such as a Variational Auto-Encoder (VAE), a Fully Convolutional Network (FCN), SegNet, or DenseNet, as part of the CNN configuration. Alternatively, the machine learning model may also use a configuration such as a Vision Transformer (ViT).

[0177] (Modification 2) Furthermore, the training data for various pre-trained models is not limited to data obtained using the radiation detector itself that actually performs the imaging; depending on the desired configuration, it may also be data obtained using the same type of radiation detector or data obtained using the same kind of radiation detector. The pre-trained models according to the above embodiments and modifications are thought to be used in estimation processes related to the generation of radiation images that have undergone various image processing, by extracting features such as the magnitude of the brightness values ​​of the radiation image, the order and slope of the bright and dark areas, their positions, distributions, and continuity.

[0178] Furthermore, the trained models according to the embodiments and modifications described above can be provided in the control unit 20. The trained models may consist of, for example, software modules executed by a processor such as a CPU, MPU, GPU, or FPGA, or they may consist of circuits that perform specific functions such as ASICs. These trained models may also be provided in a separate server device connected to the control unit 20. In this case, the control unit 20 can use the trained models by connecting to the server equipped with the trained models via any network such as the Internet. Here, the server equipped with the trained models may be, for example, a cloud server, a fog server, an edge server, etc.

[0179] (Variation 3) Furthermore, in the embodiments and modifications described above, the radiation detector 10 is an indirect conversion type detector that first converts radiation into visible light using a phosphor 11 and then converts the visible light into an electrical signal using a photoelectric conversion element. In contrast, the radiation detector 10 may also be a direct conversion type detector that directly converts incident radiation into an electrical signal.

[0180] (Other embodiments) Furthermore, the disclosed technology can also be realized by performing the following process: that is, the disclosed technology can also be realized by supplying software (programs) that implement one or more functions of the various embodiments described above to a system or device via a network or storage medium, and the computer (or CPU, MPU, etc.) of that system or device reads and executes the program. The computer may have one or more processors or circuits and may include a network of separate computers or separate processors or circuits for reading and executing computer executable instructions. In this case, the processor or circuit may include a central processing unit (CPU), a microprocessing unit (MPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), or a field-programmable gateway (FPGA). The processor or circuit may also include a digital signal processor (DSP), a dataflow processor (DFP), or a neural processing unit (NPU).

[0181] (Composition 1) An acquisition unit that acquires a first radiation image obtained by video recording using a radiation detector, An inference processing unit generates multiple partial images using the first radiation image based on processing conditions corresponding to the imaging mode of the radiation detector, infers the multiple partial images that have undergone image processing by inputting the multiple partial images into a trained model, and obtains a second radiation image obtained by image processing the first radiation image using the multiple partial images that have undergone image processing. An image processing device equipped with the following features.

[0182] (Configuration 2) The image processing apparatus according to configuration 1, wherein the shooting mode is at least one of the binning mode and sensitivity mode in the video shooting.

[0183] (Composition 3) The image processing apparatus according to configuration 1 or 2, wherein the processing conditions include the size of the partial image and the size of the overlapping region between the plurality of partial images.

[0184] (Composition 4) The inference processing unit, Multiple partial images are generated by performing a division process on a third radiation image obtained by padding the first radiation image, By inputting the aforementioned multiple subimages into the trained model, the multiple subimages that have undergone image processing are inferred. An image processing apparatus according to any one of configurations 1 to 3, which obtains a second radiation image obtained by combining a plurality of partial images that have undergone the aforementioned image processing into a first radiation image.

[0185] (Composition 5) The image processing apparatus according to configuration 4, wherein the processing conditions further include the amount of padding.

[0186] (Composition 6) The image processing apparatus according to any one of configurations 1 to 5, wherein the processing conditions are determined based on a first cost function relating to the processing time of the inference processing unit.

[0187] (Composition 7) The image processing apparatus according to configuration 6, wherein the processing condition is the processing condition in which the first cost function is smallest among a plurality of processing conditions in which at least one of the size of the partial image and the size of the overlapping region between the plurality of partial images is different.

[0188] (Composition 8) The image processing apparatus according to configuration 6 or 7, wherein the first cost function is the time required to infer the multiple partial images on which the image processing has been performed.

[0189] (Composition 9) The image processing apparatus according to any one of configurations 1 to 8, wherein the trained model is configured with parameters corresponding to the imaging mode of the radiation detector.

[0190] (Composition 10) The image processing apparatus according to any one of configurations 1 to 9, wherein the processing conditions are determined based on a second cost function relating to the sequence of the plurality of partial images for the third radiographic image.

[0191] (Composition 11) The image processing apparatus according to configuration 10, wherein the processing condition is the processing condition in which the second cost function is largest among a plurality of processing conditions in which at least one of the size of the partial image and the size of the overlapping region between the plurality of partial images is different.

[0192] (Composition 12) The image processing apparatus according to any one of configurations 1 to 11, wherein the processing conditions are determined based on a third cost function relating to the processing time of a second inference processing unit connected to the outside of the image processing apparatus.

[0193] (Composition 13) The image processing apparatus according to configuration 12, wherein the processing condition is the processing condition in which the third cost function is smallest among a plurality of processing conditions in which at least one of the size of the partial image and the size of the overlapping region between the plurality of partial images is different.

[0194] (Composition 14) The image processing apparatus according to any one of configurations 1 to 13, wherein the processing performed by the inference processing unit is one of noise reduction processing, super-resolution processing, enhancement processing, style conversion, grayscale processing, segmentation processing, scattered radiation reduction processing, or grid fringe reduction processing.

[0195] (Composition 15) A radiation detector that detects radiation, A radiography system comprising an image processing device according to any one of configurations 1 to 14, which is connected to the aforementioned radiation detector in a manner that enables communication.

[0196] (Method 1) A step of acquiring a first radiation image obtained by video recording using a radiation detector, A step of generating a plurality of partial images using the first radiation image based on processing conditions corresponding to the imaging mode of the radiation detector, The process involves inputting the aforementioned multiple subimages into a trained model to perform image processing on the multiple subimages and then inferring from those multiple subimages, A step of obtaining a second radiation image obtained by applying image processing to a first radiation image using a plurality of partial images on which the aforementioned image processing has been performed, An image processing method for an image processing apparatus comprising an image processing apparatus.

[0197] (Program 1) A program that causes a computer to execute the image processing method described in Method 1. [Explanation of Symbols]

[0198] 10. Radiation detectors 20 Control Unit (Image Processing Unit) 22 Image Processing Unit

Claims

1. An acquisition unit that acquires a first radiation image obtained by video recording using a radiation detector, An inference processing unit generates multiple partial images using the first radiation image based on processing conditions corresponding to the imaging mode of the radiation detector, infers the multiple partial images that have undergone image processing by inputting the multiple partial images into a trained model, and obtains a second radiation image obtained by image processing the first radiation image using the multiple partial images that have undergone image processing. An image processing device equipped with the following features.

2. The image processing apparatus according to claim 1, wherein the shooting mode is at least one of the binning mode and sensitivity mode in the video shooting.

3. The image processing apparatus according to claim 1, wherein the processing conditions include the size of the partial image and the size of the overlapping region between the plurality of partial images.

4. The inference processing unit, Multiple partial images are generated by performing a division process on a third radiation image obtained by padding the first radiation image, By inputting the aforementioned multiple subimages into the trained model, the multiple subimages that have undergone image processing are inferred. The image processing apparatus according to claim 1, wherein a second radiation image obtained by combining a plurality of partial images that have undergone the aforementioned image processing is obtained by combining the first radiation image.

5. The image processing apparatus according to claim 4, wherein the processing conditions further include the amount of padding.

6. The image processing apparatus according to claim 1, wherein the processing conditions are determined based on a first cost function relating to the processing time of the inference processing unit.

7. The image processing apparatus according to claim 6, wherein the processing condition is the processing condition in which the first cost function is smallest among a plurality of processing conditions in which at least one of the size of the partial image and the size of the overlapping region between the plurality of partial images is different.

8. The image processing apparatus according to claim 6, wherein the first cost function is the time required to infer the multiple partial images on which the image processing has been performed.

9. The image processing apparatus according to claim 1, wherein the trained model is configured with parameters corresponding to the imaging mode of the radiation detector.

10. The image processing apparatus according to claim 1, wherein the processing conditions are determined based on a second cost function relating to the sequence of the plurality of partial images for the third radiographic image.

11. The image processing apparatus according to claim 10, wherein the processing condition is the processing condition in which the second cost function is largest among a plurality of processing conditions in which at least one of the size of the partial image and the size of the overlapping region between the plurality of partial images is different.

12. The image processing apparatus according to claim 1, wherein the processing conditions are determined based on a third cost function relating to the processing time of a second inference processing unit connected to the outside of the image processing apparatus.

13. The image processing apparatus according to claim 12, wherein the processing condition is the processing condition in which the third cost function is smallest among a plurality of processing conditions in which at least one of the size of the partial image and the size of the overlapping region between the plurality of partial images is different.

14. The image processing apparatus according to claim 1, wherein the processing performed by the inference processing unit is any of noise reduction processing, super-resolution processing, enhancement processing, style conversion, grayscale processing, segmentation processing, scattered radiation reduction processing, or grid fringe reduction processing.

15. A radiation detector that detects radiation, A radiation imaging system comprising an image processing device according to any one of claims 1 to 14, which is connected to the radiation detector in a manner that enables communication.

16. A step of acquiring a first radiation image obtained by video recording using a radiation detector, A step of generating a plurality of partial images using the first radiation image based on processing conditions corresponding to the imaging mode of the radiation detector, The process involves inputting the aforementioned multiple subimages into a trained model to perform image processing on the multiple subimages and then inferring from those multiple subimages, A step of obtaining a second radiation image obtained by applying image processing to a first radiation image using a plurality of partial images on which the aforementioned image processing has been performed, An image processing method for an image processing apparatus comprising an image processing apparatus.

17. A program that causes a computer to execute the image processing method described in claim 16.