Image processing method, readable storage medium and medical device

By using generative adversarial networks to process low-dose CT images, the problems of image noise and artifacts in CT equipment when reducing radiation dose are solved, and high-quality image generation and radiation dose reduction are achieved.

CN116309511BActive Publication Date: 2026-05-22MIDEA GRP (SHANGHAI) CO LTD +1
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
MIDEA GRP (SHANGHAI) CO LTD
Filing Date
2023-03-30
Publication Date
2026-05-22

Smart Images

  • Figure CN116309511B_ABST
    Figure CN116309511B_ABST
Patent Text Reader

Abstract

The application provides an image processing method, a readable storage medium and medical equipment, and the image processing method comprises the following steps: acquiring a first image; inputting the first image into a preset generative adversarial network, so that the preset generative adversarial network processes the first image to obtain a second image, and the preset generative adversarial network is an adversarial network which is trained based on image blocks of different sizes in sequence; and outputting the second image; wherein the first image is a computed tomography image which is photographed under a first radiation dose, the second image is a computed tomography image which is photographed under a second radiation dose, and the first radiation dose is less than the second radiation dose.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and more specifically, to an image processing method, a readable storage medium, and a medical device. Background Technology

[0002] CT scanners, or CT equipment, have become an indispensable medical diagnostic tool for doctors. They use X-rays to scan a part of the human body at a certain thickness. However, due to their radioactivity, they have become an increasingly important focus of public attention.

[0003] However, reducing radiation dose increases image noise and generates artifacts, failing to ensure image quality and meet the requirements of clinicians. Summary of the Invention

[0004] The present invention aims to solve at least one of the technical problems existing in the prior art or related art.

[0005] Therefore, a first aspect of the present invention is to provide an image processing method.

[0006] A second aspect of the present invention is that a readable storage medium is provided.

[0007] A third aspect of the invention is that a medical device is provided.

[0008] In view of this, according to a first aspect of the present invention, the present invention provides an image processing method, comprising: acquiring a first image; inputting the first image into a preset generative adversarial network (GAN) to enable the preset GAN to process the first image to obtain a second image, wherein the preset GAN is an adversarial network trained sequentially based on image patches of different sizes; and outputting the second image; wherein the first image is a computed tomography image obtained under a first radiation dose, and the second image is a computed tomography image obtained under a second radiation dose, wherein the first radiation dose is less than the second radiation dose.

[0009] According to a second aspect of the present invention, a readable storage medium is provided on which a program or instructions are stored, which, when executed by a processor, implement the steps of any of the methods described above.

[0010] According to a third aspect of the present invention, a medical device is provided, comprising: a readable storage medium as described above.

[0011] This application enables the processing of a first image obtained from low-dose radiography to obtain a second image obtained from radiography at a higher dose.

[0012] In this process, by implementing the above technical solution, the image quality of computed tomography images obtained by low-dose radiography can be improved. Therefore, images that meet the requirements of clinicians can be captured using low radiation doses.

[0013] Furthermore, since the image quality of computed tomography images is positively correlated with the radiation dose, this method reduces the radiation dose for computed tomography images and improves the safety of the imaging process compared to related technical solutions, namely, those that use a second radiation dose to capture computed tomography images.

[0014] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0015] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the description of the embodiments taken in conjunction with the following drawings, in which:

[0016] Figure 1 A flowchart illustrating the image processing method in an embodiment of the present invention is shown.

[0017] Figure 2 One of the schematic diagrams showing the effects before and after application of an embodiment of the present invention is shown;

[0018] Figure 3 The second illustration shows the effect before and after the application of the embodiment of the present invention;

[0019] Figure 4 A schematic diagram of the third image in an embodiment of the present invention is shown;

[0020] Figure 5 A schematic diagram of noise in an embodiment of the present invention is shown;

[0021] Figure 6 A schematic diagram of the fourth image in an embodiment of the present invention is shown;

[0022] Figure 7 A schematic diagram of the preprocessing process in an embodiment of the present invention is shown;

[0023] Figure 8 A logical block diagram of a generative adversarial network in an embodiment of the present invention is shown;

[0024] Figure 9 A schematic diagram of the short-circuit structure in the generator in an embodiment of the present invention is shown;

[0025] Figure 10 This diagram illustrates the image size change during image processing by the discriminator in an embodiment of the present invention.

[0026] Figure 11 A schematic diagram of the noise power spectrum curves before and after application in an embodiment of the present invention is shown;

[0027] Figure 12 A schematic block diagram of an image processing apparatus according to an embodiment of the present invention is shown. Detailed Implementation

[0028] To better understand the above aspects, features, and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in these embodiments can be combined with each other.

[0029] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and therefore the scope of protection of the invention is not limited to the specific embodiments disclosed below.

[0030] In one embodiment of this application, such as Figure 1 As shown, an image processing method is provided, including:

[0031] Step 102, acquire the first image;

[0032] Step 104: Input the first image into a preset generative adversarial network so that the preset generative adversarial network processes the first image to obtain the second image. The preset generative adversarial network is an adversarial network trained sequentially based on image blocks of different sizes.

[0033] Step 106: Output the target image;

[0034] The first image is a computed tomography image taken at a first radiation dose, and the second image is a computed tomography image taken at a second radiation dose, wherein the first radiation dose is less than the second radiation dose.

[0035] An embodiment of this application proposes an image processing method. By running the above image processing method, a first image obtained from low-radiation-dose photography can be processed to obtain a second image obtained from photography at a higher radiation dose.

[0036] In this process, by running the above embodiments, the image quality of computed tomography images obtained by low-dose radiography can be improved. Therefore, images that meet the requirements of clinicians can be captured using low radiation doses.

[0037] Furthermore, since the image quality of computed tomography images is positively correlated with the radiation dose, the radiation dose for taking computed tomography images is reduced and the safety of taking images is improved compared to the relevant embodiments, i.e., the scheme of taking computed tomography images using a second radiation dose.

[0038] It should be noted that the preset generative adversarial network proposed in this application embodiment is an adversarial network trained sequentially on image blocks of different sizes. Therefore, it can integrate local and global information of the image, effectively improve the processing speed of the preset generative adversarial network, and thus reduce the computational cost of the preset generative adversarial network.

[0039] For example, such as Figure 2 As shown, in Figure 2 The image shown in (a) is input into a preset generative adversarial network, resulting in the following: Figure 2 The second image shown in (b) contains, as Figure 2 The second image shown in (b) is compared with the computed tomography image obtained at the second radiation dose, i.e. Figure 2 (c) Almost identical, therefore, as Figure 2 The second image shown in (b) can serve as an alternative image to a computed tomography image obtained at the second radiation dose.

[0040] For example Figure 3 As shown, in Figure 3 The image shown in (a) is input into a preset generative adversarial network, resulting in the following: Figure 3 The second image shown in (b) contains, as Figure 3 The second image shown in (b) is compared with the computed tomography image obtained at the second radiation dose, i.e. Figure 3 (c) Almost identical, therefore, as Figure 3 The second image shown in (b) can serve as an alternative image to a computed tomography image obtained at the second radiation dose.

[0041] In the above embodiments, computed tomography (CT) images, also known as electronic computed tomography scans, utilize precisely collimated X-ray beams, ultrasound waves, etc., along with highly sensitive detectors to perform a series of cross-sectional scans around a certain part of the human body.

[0042] In the above embodiments, the method further includes: acquiring a third image obtained by photography under a second radiation dose; superimposing noise on the third image to obtain a fourth image corresponding to the third image, wherein the third image and the fourth image form an image pair; processing the image pair according to a preset image block size to obtain training data of a first size and training data of a second size, wherein the first size is smaller than the second size; training the adversarial network according to the training data of the first size to obtain a converged first adversarial network; and training the first adversarial network according to the training data of the second size to obtain a preset generative adversarial network.

[0043] In this embodiment, the pre-defined generative adversarial network typically requires training and learning using existing empirical data before data processing can be performed. However, as for images such as computer tomographic images, it is difficult to obtain image pairs that can be used as training samples using the equipment. In other words, there is a problem of difficulty in obtaining training samples and a shortage of samples.

[0044] Considering that the image obtained by the camera is image noise caused by differences in radiation dose, which is additive noise, the embodiments of this application, after obtaining the third image taken when working at the second radiation dose, superimpose noise on the third image to obtain an image similar to that taken at a low radiation dose, i.e., the fourth image mentioned above. In this process, it can be ensured that the information contained in the tomographic scan of the third image and the fourth image is the same, that is, the image information contained in the tomographic scan is the same. Therefore, it can be used as a training sample for generative adversarial networks. This reduces the difficulty of obtaining training samples.

[0045] by Figure 4 , Figure 5 and Figure 6 For example, as the third image Figure 4 In superposition, such as Figure 5 After the noise shown, the following can be obtained: Figure 6 The fourth image shown.

[0046] Specifically, the pixel sizes in region A of the third image and region A' of the fourth image are 30.6486±3.8959 and 30.9460±5.1248, respectively. The pixel sizes in region B of the third image and region B' of the fourth image are 5.5120±4.0393 and 5.6171±5.2242, respectively. The average pixel value in the selected region (such as A or B) of the fourth image is basically unchanged compared to the third image, but the variance increases. That is, the image quality of the fourth image is better than that of the third image.

[0047] In the above embodiments, the noise superimposed on the third image can be randomly generated or superimposed according to a preset noise. It can be selected according to actual needs, and will not be elaborated here.

[0048] In the above embodiments, since deep learning requires a large number of samples for training, the technical solution of this application can use a sliding window to randomly crop from the image pair when the preset image patch size is known, thereby obtaining samples for training the adversarial network.

[0049] In this embodiment, the use of a sliding window increases the number of training samples while also enhancing the ability to detect differences in local regions within image pairs.

[0050] Furthermore, during the training process of adversarial networks, the inductive bias problem limits the network structure's ability to learn effective information between image patches. By selecting image patches of different sizes for training, that is, training the adversarial network with training data of the first size and then training it with training data of the second size, a local-to-global training strategy is adopted. This gradually increases the size of the cropped image patches for training, thereby improving the overall understanding of CT images and enhancing the ability to extract high-level information from the images.

[0051] In one embodiment, in addition to the first and second sizes mentioned above, the preset image block size can also be set to a third size, a fourth size, and so on, according to actual usage needs.

[0052] In one embodiment, the first size patch = 128 and the second size patch = 512.

[0053] In one embodiment, when a third size exists, the training process described above includes training the adversarial network with training data of the first size to obtain a converged first adversarial network; training the first adversarial network with training data of the third size to obtain a converged second adversarial network; and training the second adversarial network with training data of the second size to obtain a preset generative adversarial network.

[0054] Specifically, for example, the input image patch size is 128, and the number of training images per iteration is 16. After convergence, adjustments are made: patch = 256, batch = 8, the best-performing model from the previous training is loaded as the initial model and training continues. After convergence, a final adjustment is made: patch = 512, batch = 4, and training is completed.

[0055] In one embodiment, the method further includes: converting the pixel values ​​of each image in the image pair into CT values; normalizing the CT values ​​to obtain CT image data; and processing the CT image data according to a preset image block size to obtain training data of a first size and training data of a second size.

[0056] In this embodiment, after obtaining the image pairs, they need to be preprocessed to obtain samples for training the adversarial network. The preprocessing includes intensity normalization and sliding cropping.

[0057] CT value is a unit of measurement used to determine the density of a local tissue or organ in the human body. It is usually called a Hounsfield unit (HU). Because CT values ​​are independent of the equipment, different ranges of values ​​can represent different organs. Therefore, it is usually necessary to convert the pixel values ​​of CT images.

[0058] Specifically, the conversion formula is as follows:

[0059] CT value = CT image pixel value × slope + intercept

[0060] The slope and intercept are information from the CTdicom file.

[0061] Among them, such as Figure 7 As shown, the preprocessing steps specifically include:

[0062] Step 702: Convert image pixel values ​​to CT values;

[0063] Step 704, CT value intensity normalization;

[0064] Step 706, slide to cut.

[0065] During the normalization process of CT values, the CT values ​​are normalized to a range of 0 to 1.

[0066] In one embodiment, the method further includes: obtaining a pixel threshold; and processing the training data of the first size and the training data of the second size based on the pixel threshold to filter out the background image in the image pair.

[0067] In this embodiment, setting a pixel threshold can filter out areas of the image block that contain more background. By using simple threshold segmentation to separate the background from the foreground, only the pixel values ​​of the foreground area in the image block are calculated, and the background part is not included in the calculation, thereby speeding up the convergence speed of the generator.

[0068] In one embodiment, the pixel threshold can be determined according to the application scenario in which the embodiments of this application are applied, and its specific value range will not be elaborated here.

[0069] In one embodiment, the adversarial network includes a generator; the generator includes an encoder network structure and a decoder network structure corresponding to the encoder network structure, wherein the feature output map of the downsampled layer of the encoder network structure is concatenated with the feature output map of the upsampled layer of the corresponding decoder network structure and then input to the convolutional layer of the decoder network structure adjacent to the corresponding decoder network structure; and / or one or more encoder network structures include convolutional layers, hidden layer neuron outputs and fast attention modules arranged in sequence; and / or one or more decoder network structures include convolutional layers, hidden layer neuron outputs and fast attention modules arranged in sequence.

[0070] Generative Adversarial Networks (GANs) are a commonly used type of neural network. They consist of two modules: a generative model G (the generator in this application) and a discriminative model D (the discriminator in this application). The generative model G is used to generate realistic images as much as possible to deceive the discriminative model D, while D tries to distinguish the images generated by the generative model G from the real images. This leads to the generative model G and the discriminative model D learning from each other through game theory, ultimately generating a more realistic result.

[0071] The generator training principle is as follows: The generator processes the third image to obtain a fake image; the fourth image and the fake image are input into the discriminator, so that if the fourth image and the fake image are inconsistent, the generator is updated until the fourth image matches the fake image output by the updated generator. Figure 1 To.

[0072] Specifically, in the embodiments of this application, such as Figure 8 As shown, a low-dose CT head image is processed by a generator to generate a fake image. The discriminator compares the fake image with a normal-dose CT head image to determine whether it can distinguish between the real and fake images. If the fake image is identified, the generator is updated.

[0073] In this embodiment, the generator employs an encoder-decoder network structure, which can effectively suppress noise and artifacts, retain valid information, and restore tissue structure details.

[0074] Considering that the convolutional layers used in the aforementioned encoder-decoder network structure can lead to image detail loss, embodiments of this application propose a compensation mechanism. This compensation mechanism can be understood as a short-circuit structure, that is, the feature output map of the downsampling layer of the encoder network structure is concatenated with the feature output map of the upsampling layer of the corresponding decoder network structure, and then input to the convolutional layer of the decoder network structure adjacent to the corresponding decoder network structure. For example... Figure 9 As shown, the feature output maps of the second, third, fourth, and fifth downsampling layers D2, D3, D4, and D5 are concatenated with the corresponding feature output maps of the upsampling layers U2, U3, U4, and U5.

[0075] The aforementioned short-circuit structure can obtain more high-level semantic features and low-level features, thereby preserving more structural information and contrast details, and effectively improving image quality.

[0076] In the above embodiments, considering the relatively large computational overhead of generative adversarial networks, in the embodiments of this application, a fast attention module is added to the encoder-decoder network structure so that the network can be forced to learn local and non-local features of the image, thereby reducing the computational overhead and achieving a win-win situation for both speed and image quality.

[0077] In the above embodiments, the fast attention module is an improved self-attention mechanism module. By replacing the softmax operation with cosine similarity and changing the order of matrix multiplication, it can effectively reduce the amount of computation and improve the operation speed, while maintaining the ability of the self-attention mechanism to provide a larger receptive field and more refined spatial features.

[0078] The fast attention module borrows ideas from the traditional nonlocal module. Specifically, the nonlocal module calculates the product QK of the query value Q and the address K during computation. T The attention matrix A is obtained through the softmax exponentiation operation. Attention matrix A is then multiplied by V (value) to obtain the output Y = Softmax(QK). T The computational cost of V is very high.

[0079] In one embodiment, the encoder network structure includes a fast attention module or the decoder network structure includes a fast attention module, wherein the fast attention module employs a normalized cosine similarity operation and / or an attention matrix with altered matrix multiplication order.

[0080] As can be seen from the above, the nonlocal module approach suffers from instruction overhead. Therefore, the embodiments of this application use normalized cosine similarity instead of the Softmax approach in the traditional nonlocal module approach, i.e. in These are the L2 regularized values ​​of Q and K, respectively, with the multiplication order changed. Where n is the spatial size of the feature map (n = height × width).

[0081] Specifically, such as Figure 10 As shown, changing the order of multiplication reduces the computational complexity from O(n^2) to O(n^2). 2 c) becomes O(nc) 2 ), where c is the number of channels, effectively solving the problem of excessive computational load in traditional non-local modules.

[0082] In one embodiment, the discriminator of the preset generative adversarial network restricts the weights of the discriminator to a range of real numbers based on a weight pruning and truncation strategy.

[0083] Specifically, in one embodiment, the preset generative adversarial network mentioned above is a WGAN network. In the WGAN network, the discriminator eliminates the logarithmic operation performed on the output by the traditional GAN ​​network, and its value range is no longer limited to 0-1, but rather falls within the real number range. By employing a weight pruning and truncation strategy, the weights of the discriminator are restricted to a certain range, preventing large oscillations and avoiding the problem of discriminator training oscillations and non-convergence.

[0084] Weight pruning and truncation can be understood as restricting the weights from the range of real numbers to a certain interval in order to solve the problem of unstable training of GAN networks and achieve higher quality generation results.

[0085] In any of the above embodiments, the generator's loss function includes: the sum of the mean squared error loss function, the generator loss function, and the perceptual loss function; wherein, the mean squared error loss function is used to minimize the distance between pixels in the image generated by the generator and the second image, the generator loss function is used to characterize the consistency between the generator output image and the second image, and the perceptual loss function is used to compare the features obtained by convolving the image generated by the generator with the features obtained by convolving the second image.

[0086] The loss function of the generator in this application consists of the above three parts, and its expression is as follows:

[0087] L G =L MSF +L GA +L perp

[0088] Among them, L G L represents the loss function of the generator in this application. MSE The mean squared error loss function is used primarily to minimize the pixel distance between the image generated by the generator and the real standard image; L GA The generator loss function is primarily designed to make the generator's output data as realistic as possible, ideally so that the discriminator cannot detect it at all; L perp The perceptual loss function compares the features obtained by convolving real images with the features obtained by convolving generated images, making the high-level information (content and global structure) close to the feature maps extracted using a pre-trained VGG network.

[0089] Specifically, L MSE The expression is as follows:

[0090]

[0091] L GA The expression is as follows:

[0092]

[0093] L perp The expression is as follows:

[0094]

[0095] Where I(x,y) represents the pixel coordinates of the generated image, R(x,y) represents the pixel coordinates of the real image, and N represents the number of pixels; G(z) represents the data generated by the generator, and D(G(z)) represents the probability of judging the data generated by the generator as real; VGG(I(x,y)) is the feature obtained by convolving the image generated by the generator, and VGG(R(x,y)) is the feature obtained by convolving the real image.

[0096] In any of the above embodiments, an autoencoder (AE) is used as the generator.

[0097] In any of the above embodiments, the adversarial network includes a discriminator; the discriminator includes a self-attention mechanism module and an output layer arranged in sequence.

[0098] The self-attention mechanism module is a non-local information statistical attention mechanism that can capture the internal correlations of data or features. It can be understood as the self-attention mechanism selectively filtering out a small amount of important information from a large amount of data and focusing on this information. It typically adopts a Query-Key-Value format, and the focusing process is reflected in the calculation of the weights (Values). Specifically, the process involves first calculating weight coefficients based on the Query and Key, which are generally calculated based on the similarity between the Query and Key. Then, all similarities are normalized, and finally, the Values ​​are weighted and summed according to the weight coefficients. In the self-attention module, the Query and Key are the same.

[0099] In this embodiment, the discriminator employs a Markov discriminator based on a self-attention mechanism. Structurally, the Markov discriminator is entirely composed of convolutional layers. When the input image size is W×H×C, the output is a... The matrix is ​​used to calculate the true / false output, and the mean of the output matrix is ​​taken as the true / false output. The matrix size is [value missing]. Each output in the output matrix represents a receptive field in the original image, corresponding to a patch in the original image.

[0100] Compared to traditional discriminators that capture global information in an image, this discriminator focuses on the evaluation results of a small region of the image, paying more attention to image details and generating more realistic CT images.

[0101] Specifically, in the discriminator, the input image size changes in the following order: W×H×C, W / 2×H / 2×C, W / 2×H / 2×2C, W / 2×H / 2×4C, W / 4×H / 4×8C.

[0102] In one embodiment, such as Figure 10 As shown, the input image sequentially goes through Conv+Spectral_norm+LeakyReLU (i.e., Markov discriminator), Self attention layer, and Output layer.

[0103] Where C is an abbreviation for channel, H is an abbreviation for height, and W is an abbreviation for width.

[0104] Specifically, to address the issue of Markov discriminators lacking global consistency, embodiments of this application incorporate a self-attention mechanism module into the discriminator. This self-attention module effectively increases the receptive field and enhances the globality of image features.

[0105] Wherein, the loss function L of the discriminator D The expression is as follows:

[0106]

[0107] in, Let D(x) be the loss function of the Markov discriminator, and let D(x) be the probability of the output being true or false when the input is x.

[0108] In one embodiment, such as Figure 11 As shown in the figure, the noise power spectrum curves of the images obtained before and after processing the low radiation dose image using the image processing method proposed in this embodiment are compared with the noise power spectrum curves of the image under normal radiation dose. As can be seen from the figure, the noise power spectrum curve of the processed image is almost the same as that of the image under normal radiation dose.

[0109] In one embodiment, the present invention provides a readable storage medium on which a program or instructions are stored, which, when executed by a processor, implement the steps of the method as described above.

[0110] In one embodiment, the present invention provides a computer program product stored in a storage medium, which is executed by at least one processor to perform the steps of any of the methods described above.

[0111] In one embodiment, such as Figure 12 As shown, the present invention provides an image processing apparatus 1200, including a processor 1202 and a memory 1204. The memory 1204 stores programs or instructions that can be executed on the processor 1202. When the program or instructions are executed by the processor 1202, they implement the steps of any of the methods described above.

[0112] The memory 1204 can be used to store software programs and various data. The memory 1204 may primarily include a first storage area for storing programs or instructions and a second storage area for storing data. The first storage area may store the operating system, application programs or instructions required for at least one function (such as sound playback function, image playback function, etc.). Furthermore, the memory may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory can be random access memory (RAM), static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus RAM (DRRAM). The memory in the embodiments of this application includes, but is not limited to, these and any other suitable types of memory.

[0113] In one embodiment, the present invention provides a medical device comprising: any of the image processing devices described above; and / or a readable storage medium as described above; and / or a computer program product as described above.

[0114] In one embodiment, the medical device includes a CT scanner.

[0115] In one embodiment, the CT scanner is a brain CT scanner.

[0116] The terms "first" and "second" in the specification and claims of this application may explicitly or implicitly include one or more of the features. In the textual description of this invention, unless otherwise stated, "a plurality of" means two or more. Furthermore, in the specification and claims, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.

[0117] In the textual description of this invention, it is understood that, unless explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal communication between two components. Those skilled in the art will understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0118] In the claims, description, and accompanying drawings of this invention, the term "plural" refers to two or more. Unless otherwise explicitly defined, the terms "upper," "lower," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the invention and simplifying the description process, not to indicate or imply that the device or element referred to must have the described specific orientation, or be constructed and operated in a specific orientation. Therefore, these descriptions should not be construed as limiting the invention. The terms "connect," "install," "fix," etc., should be interpreted broadly. For example, "connect" can be a fixed connection between multiple objects, a detachable connection between multiple objects, or an integral connection; it can be a direct connection between multiple objects or an indirect connection between multiple objects through an intermediate medium. For those skilled in the art, the specific meaning of the above terms in this invention can be understood based on the specific circumstances described above.

[0119] In the claims, description, and accompanying drawings of this invention, the terms "one embodiment," "some embodiments," "specific embodiment," etc., refer to a specific feature, structure, material, or characteristic described in connection with that embodiment or example, which is included in at least one embodiment or example of the invention. In the claims, description, and accompanying drawings of this invention, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0120] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. An image processing method, characterized in that, include: Get the first image; The first image is input into a preset generative adversarial network (GAN) so that the preset GAN processes the first image to obtain a second image. The preset GAN is an adversarial network trained sequentially based on image blocks of different sizes. Output the second image; Wherein, the first image is a computed tomography image obtained under a first radiation dose, the second image is a computed tomography image obtained under a second radiation dose, and the first radiation dose is less than the second radiation dose; The image processing method further includes: Acquire a third image obtained by photography at the second radiation dose; Noise is superimposed on the third image to obtain a fourth image corresponding to the third image, and the third image and the fourth image form an image pair; The image pairs are processed according to a preset image block size to obtain training data of a first size and training data of a second size, wherein the first size is smaller than the second size; The adversarial network is trained based on the training data of the first size to obtain a converged first adversarial network; The first adversarial network is trained based on the training data of the second size to obtain the preset generative adversarial network; the preset generative adversarial network is a WGAN network.

2. The image processing method according to claim 1, characterized in that, Also includes: Convert the pixel values ​​of each image in the image pair to CT values; The CT values ​​are normalized to obtain CT image data; The CT image data is processed according to a preset image block size to obtain training data of the first size and training data of the second size.

3. The image processing method according to claim 2, characterized in that, Also includes: Obtain the pixel threshold; The training data of the first size and the training data of the second size are processed based on the pixel threshold to filter out the background image in the image pair.

4. The image processing method according to any one of claims 1 to 3, characterized in that, The adversarial network includes a generator; The generator includes an encoder network structure and a decoder network structure corresponding to the encoder network structure. The feature output map of the downsampling layer of the encoder network structure is concatenated with the feature output map of the upsampling layer of the corresponding decoder network structure, and then input to the convolutional layer of the decoder network structure adjacent to the corresponding decoder network structure; and / or One or more of the encoder network structures include sequentially ordered convolutional layers, hidden neuron outputs, and fast attention modules; and / or One or more of the decoder network structures include sequentially ordered convolutional layers, hidden neuron outputs, and fast attention modules.

5. The image processing method according to claim 4, characterized in that, The encoder network structure includes a fast attention module or the decoder network structure includes a fast attention module, wherein the fast attention module employs a normalized cosine similarity operation and / or an attention matrix with a changed matrix multiplication order.

6. The image processing method according to claim 4, characterized in that, The adversarial network includes a discriminator; The discriminator includes a self-attention mechanism module and an output layer arranged in sequence.

7. The image processing method according to claim 1, characterized in that, The discriminator of the pre-defined generative adversarial network restricts the weights of the discriminator to a range of real numbers based on a weight pruning and truncation strategy.

8. The image processing method according to claim 4, characterized in that, The loss function of the generator includes: The sum of the mean squared error loss function, the generator loss function, and the perceptual loss function; Wherein, the mean squared difference loss function is used to minimize the distance between pixels in the image generated by the generator and the second image, the generator loss function is used to characterize the consistency between the generator output image and the second image, and the perceptual loss function is used to compare the features obtained by convolving the image generated by the generator with the features obtained by convolving the second image.

9. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the method as described in any one of claims 1 to 8.

10. A medical device, characterized in that, include: The readable storage medium as described in claim 9.