Medical image reconstruction method and device, equipment, storage medium and product

Through the iterative reconstruction and multi-network model fusion method, low-quality medical images are enhanced, which solves the problem of poor compatibility of multimodal reconstruction solutions, improves the reconstruction quality of medical images, and completes the reconstruction without increasing scanning time or radiation dose.

CN120125697APending Publication Date: 2025-06-10SHENZHEN INST OF ADVANCED TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510326234.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-19
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The image compatibility of the multimodal reconstruction scheme is poor, resulting in low quality of medical image reconstruction and long scanning time.

Method used

By iterative reconstruction processing on low-quality medical images, the first network model is used to generate image prompt word sets, and a priori feature images are extracted in combination with the second network model to perform image enhancement, so as to realize single-modal medical image reconstruction.

Benefits of technology

The reconstruction quality of medical images is improved, the poor image compatibility problem of multimodal reconstruction scheme is solved, and the reconstruction is completed without increasing scanning time or radiation dose.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125697A_ABST
    Figure CN120125697A_ABST
Patent Text Reader

Abstract

The invention discloses a medical image reconstruction method and device, equipment, a storage medium and a product. The method comprises the following steps: performing reconstruction processing on a low-quality first medical image to obtain a second medical image; inputting the second medical image into the first network model for text generation to obtain an output image prompt word set; inputting the second medical image and the image cue word set into a second network model for feature extraction to obtain an output prior feature image; according to the prior feature image, performing enhancement processing on the second medical image to obtain a third medical image; taking the third medical image as a low-quality first medical image in the next iterative reconstruction process, and returning to the step of performing reconstruction processing on the low-quality first medical image to obtain a second medical image; and taking the second medical image in the current iterative reconstruction process as the reconstructed medical image until the reconstruction iteration ending condition is met, so that the reconstruction quality of the medical image is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image reconstruction, and particularly to a method, device, equipment, storage medium and product for reconstructing medical images. Background Art

[0002] Medical images collected by medical imaging devices often have low-quality problems such as noise interference, insufficient resolution, and blurred details. Medical image reconstruction technology can effectively improve the resolution of low-quality medical images and clearly present the fine structure of tissues.

[0003] Currently, medical image reconstruction technology mainly relies on prior knowledge provided by other modality images to reconstruct the target modality image. For example, when reconstructing PET (Positron Emission Tomography) images, structural prior information provided by CT (Computed Tomography) images or MRI (Magnetic Resonance Imaging) images needs to be introduced.

[0004] However, the multi-modal reconstruction scheme will bring a long scanning time, and the compatibility between different modality medical images is poor. The image registration error will seriously affect the reconstruction quality of medical images. Summary of the Invention

[0005] Embodiments of the present invention provide a method, device, equipment, storage medium and product for reconstructing medical images to solve the problem of poor image compatibility of the multi-modal reconstruction scheme and improve the reconstruction quality of medical images.

[0006] According to an embodiment of the present invention, a method for reconstructing a medical image is provided. The method includes:

[0007] Performing reconstruction processing on a low-quality first medical image to obtain a second medical image;

[0008] Inputting the second medical image into a first network model for text generation to obtain an output image prompt word set;

[0009] Inputting the second medical image and the image prompt word set into a second network model for feature extraction to obtain an output prior feature image;

[0010] Performing enhancement processing on the second medical image according to the prior feature image to obtain a third medical image;

[0011] Use the third medical image as the low-quality first medical image in the next iterative reconstruction process, and return to execute the step of reconstructing the low-quality first medical image to obtain a second medical image;

[0012] Until the reconstruction iteration end condition is met, use the second medical image in the current iterative reconstruction process as the reconstructed medical image.

[0013] According to another embodiment of the present invention, there is provided a medical image reconstruction device, which includes:

[0014] A second medical image determination module, configured to perform reconstruction processing on a low-quality first medical image to obtain a second medical image;

[0015] An image prompt word set output module, configured to input the second medical image into a first network model for text generation to obtain an output image prompt word set;

[0016] A prior feature image output module, configured to input the second medical image and the image prompt word set into a second network model for feature extraction to obtain an output prior feature image;

[0017] A third medical image determination module, configured to perform enhancement processing on the second medical image according to the prior feature image to obtain a third medical image;

[0018] An iterative execution module, configured to use the third medical image as the low-quality first medical image in the next iterative reconstruction process, and return to execute the step of performing reconstruction processing on the low-quality first medical image to obtain a second medical image;

[0019] A reconstructed medical image determination module, configured to, until the reconstruction iteration end condition is met, use the second medical image in the current iterative reconstruction process as the reconstructed medical image.

[0020] According to another embodiment of the present invention, there is provided an electronic device, which includes:

[0021] At least one processor; and

[0022] A memory communicatively connected to the at least one processor; wherein,

[0023] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the medical image reconstruction method according to any embodiment of the present invention.

[0024] According to another embodiment of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the medical image reconstruction method according to any embodiment of the present invention when executed.

[0025] According to another embodiment of the present invention, there is provided a computer program product including a computer program which implements the medical image reconstruction method according to any embodiment of the present invention when executed by a processor.

[0026] The technical solution of this embodiment performs iterative reconstruction processing on the low-quality first medical image. For the second medical image reconstructed in each iterative reconstruction process, a first network model is used to generate text according to the second medical image to determine an image prompt word set. A second network model is used to fuse the global features represented by the second medical image and the local features represented by the image prompt word set to obtain a prior feature image capable of representing deeper features of the second medical image. According to the prior feature image, the second medical image is enhanced to obtain a third medical image, and the third medical image is used as the low-quality first medical image in the next iterative reconstruction process, achieving the purpose of single-modal medical image reconstruction, solving the problem of poor image compatibility of multi-modal reconstruction schemes, and ensuring the reconstruction quality of medical images without additionally increasing the scanning time or radiation dose.

[0027] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0029] Figure 1 It is a flowchart of a medical image reconstruction method provided by an embodiment of the present invention;

[0030] Figure 2 It is a schematic diagram of a PET image reconstruction method provided by an embodiment of the present invention;

[0031] Figure 3 It is a flowchart of another medical image reconstruction method provided by an embodiment of the present invention;

[0032] Figure 4Schematic diagram of a specific example of an image enhancement framework provided by an embodiment of the present invention;

[0033] Figure 5 Comparison schematic diagram of a PET image provided by an embodiment of the present invention;

[0034] Figure 6 Schematic diagram of the structure of a medical image reconstruction device provided by an embodiment of the present invention;

[0035] Figure 7 Schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0036] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0037] It should be noted that the terms "first", "second", "third", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0038] Figure 1 Flowchart of a medical image reconstruction method provided by an embodiment of the present invention. This embodiment is applicable to the situation of reconstructing medical images. This method can be executed by a medical image reconstruction device, which can be implemented in the form of hardware and / or software, and the medical image reconstruction device can be configured in a terminal device. As Figure 1 shown, the method includes:

[0039] S110. Reconstruct the low-quality first medical image to obtain a second medical image.

[0040] Specifically, in the first iterative reconstruction process, the first medical image is a medical image acquired by a medical imaging device. Exemplarily, the medical imaging device includes, but is not limited to, a CT device, an MRI device, a PET device, or a SPECT device (Single-Photon Emission Computed Tomography).

[0041] In a specific embodiment, in the PET image reconstruction scenario, the first medical image in the first iterative reconstruction process is a sinogram. Specifically, the imaging principle of a PET device is to inject a tracer labeled with a radionuclide into an organism. The radionuclide decays to produce positrons, and the positrons annihilate with electrons in the organism tissue and emit a pair of annihilation photons in opposite directions. The detector combined with the electronics system uses coincidence measurement technology to detect a pair of photons generated by the same annihilation event. The line connecting the positions of each pair of photons is called a line of response. The sinogram takes the projection coordinate data of a large number of lines of response as the storage address and the count as the storage data of a two-dimensional statistical histogram, recording the number of annihilation events in the radial coordinate and the angular coordinate.

[0042] In the subsequent iterative reconstruction process, the first medical image is a third medical image obtained by enhancing the second medical image in the previous iterative reconstruction process.

[0043] Specifically, the image reconstruction algorithm used for the reconstruction process is an iterative reconstruction algorithm. Exemplarily, the iterative reconstruction algorithm can be the Maximum Likelihood Expectation-Maximization (MLEM) algorithm, the Ordered Subset Expectation-Maximization (OSEM) algorithm, the Penalized Weighted Least Squares (PWLS) algorithm, or the Alternating Direction Method of Multipliers (ADMM) algorithm, but is not limited to the exemplary cases.

[0044] Among them, the MLEM algorithm, based on the principle of maximum likelihood estimation, starts from the initial image and continuously updates the estimated value of the initial image to maximize the likelihood between the estimated image and the actual measurement data. The OSEM algorithm divides the initial image into multiple subsets. During each iterative reconstruction process, also based on the principle of maximum likelihood estimation, the image data of each subset is used in turn to update the reconstructed image. One round of all subsets being used is one iteration. The ADMM algorithm is an algorithm for solving large-scale convex optimization problems. It decomposes the complex image reconstruction problem into multiple sub-problems and gradually approaches the optimal solution by alternately updating variables and multipliers.

[0045] Exemplarily, when the iterative reconstruction algorithm is the OSEM algorithm, the number of subsets can be 8. When the iterative reconstruction algorithm is the ADMM algorithm, the regularization parameter is 0.07, the splitting variable is the zero matrix, and the step size parameter is 0.01.

[0046] Specifically, the second medical image represents the medical image obtained by performing one iterative reconstruction on the first medical image.

[0047] S120. Determine whether the reconstruction iteration end condition is satisfied. If so, execute S170. If not, execute S130.

[0048] Specifically, the reconstruction iteration end condition includes at least one of the reconstruction iteration times reaching the number threshold, the image error reaching the error threshold, and the objective function converging. Among them, the image error characterizes the change amount of the second medical image in the current iterative reconstruction process compared to the second medical image in the previous iterative reconstruction process. Exemplarily, the image error can be the mean square error or the absolute error, but is not limited to the example situation.

[0049] Exemplarily, when the iterative reconstruction algorithm is the MLEM algorithm, the number threshold can be 100 times, the error threshold can be 0.0001, and the objective function is used to measure the matching degree between the estimated image and the actual measurement data. When the iterative reconstruction algorithm is the OSEM algorithm, the number threshold can be 20 times, the error threshold can be 0.0001. When the iterative reconstruction algorithm is the ADMM algorithm, the number threshold can be 100 times, and the error threshold can be 0.0001.

[0050] Here, there is no limitation on the reconstruction iteration end condition corresponding to different iterative reconstruction algorithms, and it can be specifically customized according to actual needs.

[0051] S130. Input the second medical image into the first network model for text generation to obtain the output image prompt word set.

[0052] Exemplarily, the first network model can be a Transformer model or a convolutional neural network, but is not limited to the example situation.

[0053] Specifically, the image prompt set contains at least one image prompt, and the image prompt represents the descriptive information of the attribute features of the second medical image. Exemplarily, the attribute features can be geometric features, texture features, color features, semantic features, and so on.

[0054] S140. Input the second medical image and the image prompt set into the second network model for feature extraction to obtain the output prior feature image.

[0055] Exemplarily, the second network model can be a multi-modal Transformer model or a multi-modal ResNet, but is not limited to the exemplary cases.

[0056] Among them, the prior feature image represents the fusion information of the local prior knowledge expressed by the image prompt set and the global prior knowledge expressed by the second medical image.

[0057] S150. Perform enhancement processing on the second medical image according to the prior feature image to obtain the third medical image.

[0058] In an alternative embodiment, performing enhancement processing on the second medical image according to the prior feature image to obtain the third medical image includes: determining the weight value corresponding to each pixel point in the second medical image according to the prior feature image, and performing linear enhancement on the second medical image according to each weight value to obtain the third medical image.

[0059] Exemplarily, perform normalization processing on the prior feature image to obtain the weight value corresponding to each pixel point in the second medical image, or perform non-linear transformation on the prior feature image using a non-linear function to obtain the weight value corresponding to each pixel point in the second medical image. Among them, the non-linear function can be a Sigmoid function, a power-law transformation function, a logarithmic transformation function, etc.

[0060] In another alternative embodiment, performing enhancement processing on the second medical image according to the prior feature image to obtain the third medical image includes: inputting the second medical image and the prior feature image into the third network model for enhancement processing to obtain the output third medical image.

[0061] In an alternative embodiment, the third network model is a U-net network model, a residual network model, or an Inception network model.

[0062] In another alternative embodiment, the third network model is composed of a concatenation layer and at least two convolutional networks connected in series, and the number of convolutional kernels corresponding to the at least two convolutional networks decreases in sequence.

[0063] In a specific embodiment, the number of convolutional networks in the third network model is 3. Among them, the first convolutional network corresponds to 128 3×3 convolutional kernels with a stride of 1, and is used to extract features from the feature data obtained by splicing the second medical image and the prior feature image. The second convolutional network corresponds to 64 3×3 convolutional kernels with a stride of 1, and is used to further compress and extract features from the image features output by the first convolutional network. The third convolutional network corresponds to 32 1×1 convolutional kernels, and is used to output the third medical image according to the image features output by the second convolutional network.

[0064] Exemplarily, each convolutional network is composed of a cascaded convolutional layer, ReLU activation function, and batch normalization.

[0065] S160: Use the third medical image as the low-quality first medical image in the next iterative reconstruction process, and return to execute S110.

[0066] S170: Use the second medical image in the current iterative reconstruction process as the reconstructed medical image.

[0067] Figure 2 This is a schematic diagram of a method for reconstructing a PET image provided by an embodiment of the present invention. Specifically, Figure 2 the x in 0 represents the sinogram, x 1 represents the first second medical image in the iterative reconstruction framework, x i represents the i-th second medical image in the iterative reconstruction framework, x n represents the finally reconstructed PET image in the iterative reconstruction framework, x i ' represents the second medical image x in the iterative reconstruction framework i the third medical image obtained after enhancement processing, n represents the number of iterative reconstructions, Figure 2 Taking the enhancement process of the second medical image x as an example, the enhancement processing processes of the other second medical images except the second medical image x i are not shown.

[0068] In the image enhancement framework, the second medical image x i is processed by the first network model to obtain an image prompt word set. The second medical image x i and the image prompt word set are processed by the second network model to obtain a prior feature image. The second medical image x i and the prior feature image are processed by the third network model to obtain the third medical image x i '.

[0069] Based on the above embodiments, optionally, the method further includes: for each training medical image in the low-quality training image set, performing iterative reconstruction processing on the training medical image; in each iterative reconstruction process, inputting the intermediate medical image obtained by reconstruction into the first network model that has not been trained yet to generate text, obtaining the output predicted prompt word set, and inputting the intermediate medical image and the predicted prompt word set into the second network model that has not been trained yet to perform feature extraction, obtaining the output predicted prior image, and inputting the intermediate medical image and the predicted prior image into the third network model that has not been trained yet to perform enhancement processing, obtaining the output enhanced medical image, and using the enhanced medical image as the medical image to be reconstructed in the next iterative reconstruction process; until the reconstruction iteration end condition is satisfied, determining the predicted image set according to at least one intermediate medical image in the current iterative reconstruction process; determining the loss function value according to the predicted image set and the high-quality standard image set, and training the first network model, the second network model, and the third network model according to the loss function value; until the loss function value converges, obtaining the trained first network model, second network model, and third network model.

[0070] Exemplarily, the loss function corresponding to the loss function value can be a square loss function, a logarithmic loss function, an exponential loss function, a mean squared error loss function, a logistic regression loss function, a Huber loss function, a cross-entropy loss function, a Kullback-Leibler divergence loss function, etc., but is not limited to the example cases.

[0071] Exemplarily, the training image set X is represented as X = {x 1 , x 2 ,..., x n}, the standard image set Y is represented as Y = {y 1 , y 2 ,..., y n}, and n represents the number of samples of the training medical images. When the loss function is the mean squared error loss function, the loss function value L MSE satisfies the following formula:

[0072]

[0073] Wherein, x i represents the i-th training medical image in the training image set X, y i represents the i-th high-quality standard medical image in the standard image set Y, and Θ represents the model parameters corresponding to the first network model, the second network model, and the third network model in the above embodiments.

[0074] Exemplarily, the Adam optimizer is adopted to determine the updated gradient corresponding to the loss function value, and the model parameters of the first network model, the second network model, and the third network model are adjusted according to the updated gradient.

[0075] In the technical solution of this embodiment, by performing iterative reconstruction processing on the low-quality first medical image, for the second medical image reconstructed in each iterative reconstruction process, the first network model is used to generate text according to the second medical image to determine the image prompt word set. The second network model is used to fuse the global features represented by the second medical image and the local features represented by the image prompt word set to obtain a prior feature image that can represent deeper features of the second medical image. According to the prior feature image, the second medical image is enhanced to obtain a third medical image, and the third medical image is used as the low-quality first medical image in the next iterative reconstruction process, achieving the purpose of single-modal medical image reconstruction, solving the problem of poor image compatibility of the multi-modal reconstruction scheme, and ensuring the reconstruction quality of the medical image without additionally increasing the scanning time or radiation dose.

[0076] Figure 3 It is a flowchart of another medical image reconstruction method provided by an embodiment of the present invention. In this embodiment, the "first network model" in the above embodiment is further refined. In this embodiment, the first network model includes at least two parallel proxy robot networks. Correspondingly, when the second medical image is input into the first network model for text generation to obtain the output image prompt word set, it includes: generating image prompt words for the second medical image through each proxy robot network; determining the image prompt word set according to at least two image prompt words. As Figure 3 shown, the method includes:

[0077] S210. Perform reconstruction processing on the low-quality first medical image to obtain a second medical image.

[0078] S220. Determine whether the reconstruction iteration end condition is satisfied. If so, execute S280; if not, execute S230.

[0079] S210 - S220 in this embodiment are the same or similar to Figure 1 S110 - S120 shown above, and will not be elaborated herein.

[0080] S230. Generate image prompt words for the second medical image through each proxy robot network.

[0081] In this embodiment, the first network model includes at least two proxy robot networks in parallel. Specifically, the network architectures corresponding to the at least two proxy robot networks may be the same, partially the same, or different.

[0082] Exemplarily, the number of proxy robot networks may be 3, but is not limited to the exemplary situation.

[0083] In an alternative embodiment, the proxy robot network is a Transformer model or a convolutional neural network.

[0084] In another alternative embodiment, the proxy robot network is composed of at least two convolutional layers in series, a global average pooling layer, and two fully connected layers; wherein, the number of convolutional kernels corresponding to the at least two convolutional layers increases in sequence, and the number of neurons corresponding to the first fully connected layer is more than the number of neurons corresponding to the second fully connected layer.

[0085] In a specific embodiment, the number of convolutional layers in the proxy robot network is 3. Among them, the first convolutional layer corresponds to 32 3×3 convolutional kernels with a stride of 1, which is used to extract the primary image features expressed by the second medical image. The second convolutional layer corresponds to 64 3×3 convolutional kernels with a stride of 2, which is used to extract the intermediate image features expressed by the second medical image. The third convolutional layer corresponds to 128 3×3 convolutional kernels with a stride of 2, which is used to extract the high-level image features expressed by the second medical image.

[0086] Exemplarily, each convolutional layer is composed of a convolutional module, a ReLU activation function, and batch normalization in series. Among them, the ReLU activation function is used to enhance the non-linear expression ability of the proxy robot network, and batch normalization is used to accelerate the convergence speed of the first network model.

[0087] Among them, the global average pooling layer is used to convert the image features output by the last convolutional layer into a one-dimensional feature vector.

[0088] In a specific embodiment, the number of neurons corresponding to the first fully connected layer is 128, and the number of neurons corresponding to the second fully connected layer is 64.

[0089] S240. Determine an image prompt word set according to at least two image prompt words.

[0090] Specifically, the at least two image prompt words together constitute the image prompt word set, and the image prompt word set may contain duplicate image prompt words.

[0091] S250. Input the second medical image and the image prompt word set into the second network model for feature extraction to obtain the output prior feature image.

[0092] In an alternative embodiment, the second network model includes an image encoder, a mask encoder, and a prior decoder; the second medical image and the image prompt set are input into the second network model for feature extraction to obtain the output prior feature image, including: inputting the second medical image into the image encoder for image encoding to obtain the output image encoding feature; inputting the image prompt set into the mask encoder for mask encoding to obtain the output mask encoding feature; inputting the image encoding feature and the mask encoding feature into the prior decoder for decoding processing to obtain the output prior feature image.

[0093] In an alternative embodiment, the image encoder is a convolutional neural network, an autoencoder, or a variational autoencoder.

[0094] In another alternative embodiment, the image encoder is composed of at least two cascaded Transformer networks. The Transformer network adopts a multi-head self-attention mechanism, and adjacent two Transformer networks are connected through a residual connection and a layer normalization connection.

[0095] In a specific embodiment, the image encoder contains 12 cascaded Transformer networks, the number of attention heads in the multi-head self-attention mechanism is 16, and the feature dimension is 1024.

[0096] The advantage of setting the residual connection is to enhance the feature representation ability of the image encoder. The advantage of setting the layer normalization connection is to accelerate the convergence speed of the graph encoder, reduce gradient fluctuations, and thus improve the stability and generalization ability of the graph encoder.

[0097] In an alternative embodiment, the mask encoder is a bag-of-words model or a Transformer model.

[0098] In another alternative embodiment, the mask encoder is composed of at least two cascaded convolutional modules. Adjacent two convolutional modules are connected through an upsampling connection, and the number of convolutional kernels corresponding to at least two convolutional modules is the same.

[0099] In a specific embodiment, the mask encoder contains 4 cascaded convolutional modules, and each convolutional module corresponds to 256 3×3 convolutional kernels.

[0100] The advantage of setting the upsampling connection is to restore the feature dimension reduced by the convolutional processing and ensure the richness of the feature representation.

[0101] In an alternative embodiment, the prior decoder is a cross-attention Transformer structure.

[0102] S260. According to the prior feature image, perform enhancement processing on the second medical image to obtain a third medical image.

[0103] S270. Use the third medical image as the low-quality first medical image in the next iterative reconstruction process, and execute S210.

[0104] S280. Use the second medical image in the current iterative reconstruction process as the reconstructed medical image.

[0105] Figure 4 Schematic diagram of a specific example of an image enhancement framework provided by an embodiment of the present invention. Specifically, the image enhancement framework includes a first network model, a second network model, and a third network model. Among them, the first network model is composed of K proxy robot networks, and each proxy robot network is composed of 3 convolutional layers in series, a global average pooling layer, and two fully connected layers. The image prompt words respectively output by the K proxy robot networks together constitute an image prompt word set.

[0106] Among them, the second network model is composed of an image encoder, a mask encoder, and a prior decoder. Specifically, the image encoder is composed of multiple Transformer networks in series, and the mask encoder is composed of 4 upsampling-connected convolutional modules. The third network model is composed of a concatenation layer in series and 3 convolutional networks.

[0107] The technical solution of this embodiment, by setting the first network model to include at least two parallel proxy robot networks, on the one hand, enables each proxy robot network to focus on one image prompt word, so as to more deeply extract the deeper detail descriptions or semantic descriptions of the second medical image. On the other hand, the number of repetitions of the image prompt words can represent the importance of the image prompt words, thus enriching the feature dimensions in the reconstruction process of medical images and further improving the reconstruction quality of medical images.

[0108] Figure 5 Schematic diagram for comparing PET images provided by an embodiment of the present invention. Specifically, Figure 5 The PET images shown from left to right are a high-quality standard PET image, a reconstructed PET image A obtained by using the OSEM algorithm, and a reconstructed PET image B obtained by using the medical image reconstruction method provided by this embodiment.

[0109] From Figure 5 It can be seen that the medical image reconstruction method provided by this embodiment significantly improves the reconstruction quality of medical images compared with the traditional iterative reconstruction algorithm by adding an image enhancement mechanism to the traditional iterative reconstruction algorithm.

[0110] The following are embodiments of a medical image reconstruction device provided by embodiments of the present invention. This device and the medical image reconstruction method in the above embodiments belong to the same inventive concept. For details not described in detail in the embodiments of the medical image reconstruction device, reference may be made to the content of the medical image reconstruction method in the above embodiments.

[0111] Figure 6 The following is a schematic structural diagram of a medical image reconstruction device provided by an embodiment of the present invention. As Figure 6 shown, the device includes: a second medical image determination module 310, an image prompt word set output module 320, a prior feature image output module 330, a third medical image determination module 340, an iterative execution module 350, and a reconstructed medical image determination module 360.

[0112] Among them, the second medical image determination module 310 is configured to perform reconstruction processing on a low-quality first medical image to obtain a second medical image;

[0113] The image prompt word set output module 320 is configured to input the second medical image into a first network model for text generation to obtain an output image prompt word set;

[0114] The prior feature image output module 330 is configured to input the second medical image and the image prompt word set into a second network model for feature extraction to obtain an output prior feature image;

[0115] The third medical image determination module 340 is configured to perform enhancement processing on the second medical image according to the prior feature image to obtain a third medical image;

[0116] The iterative execution module 350 is configured to use the third medical image as the low-quality first medical image in the next iterative reconstruction process, and return to execute the step of performing reconstruction processing on the low-quality first medical image to obtain a second medical image;

[0117] The reconstructed medical image determination module 360 is configured to use the second medical image in the current iterative reconstruction process as the reconstructed medical image until the reconstruction iteration end condition is met.

[0118] The technical solution of this embodiment performs iterative reconstruction processing on the low-quality first medical image. For the second medical image obtained by reconstruction in each iterative reconstruction process, a first network model is used to generate text based on the second medical image to determine an image prompt word set. A second network model is used to fuse the global features represented by the second medical image and the local features represented by the image prompt word set to obtain a prior feature image that can represent deeper features of the second medical image. According to the prior feature image, the second medical image is enhanced to obtain a third medical image, and the third medical image is used as the low-quality first medical image in the next iterative reconstruction process, achieving the purpose of single-modal medical image reconstruction, solving the problem of poor image compatibility of the multi-modal reconstruction solution, and ensuring the reconstruction quality of the medical image without additionally increasing the scanning time or radiation dose.

[0119] In an alternative embodiment, the first network model includes at least two parallel proxy robot networks. Correspondingly, the image prompt word set output module 320 is specifically configured to:

[0120] Generate image prompt words for the second medical image through each proxy robot network;

[0121] Determine the image prompt word set according to at least two image prompt words.

[0122] In an alternative embodiment, the proxy robot network is composed of at least two serially connected convolutional layers, a global average pooling layer, and two fully connected layers;

[0123] Among them, the number of convolution kernels corresponding to at least two convolutional layers increases in sequence, and the number of neurons corresponding to the first fully connected layer is more than the number of neurons corresponding to the second fully connected layer.

[0124] In an alternative embodiment, the second network model includes an image encoder, a mask encoder, and a prior decoder; the prior feature image output module 330 is specifically configured to:

[0125] Input the second medical image into the image encoder for image encoding to obtain the output image encoding features;

[0126] Input the image prompt word set into the mask encoder for mask encoding to obtain the output mask encoding features;

[0127] Input the image encoding features and the mask encoding features into the prior decoder for decoding processing to obtain the output prior feature image.

[0128] In an alternative embodiment, the image encoder is composed of at least two Transformer networks connected in series. The Transformer networks adopt a multi-head self-attention mechanism, and adjacent Transformer networks are connected by residual connections and layer normalization;

[0129] The mask encoder is composed of at least two convolutional modules connected in series. Adjacent convolutional modules are connected by upsampling connections, and the number of convolutional kernels corresponding to at least two convolutional modules is the same;

[0130] The prior decoder is a cross-attention Transformer structure.

[0131] In an alternative embodiment, the third medical image determination module 340 is specifically configured to:

[0132] Input the second medical image and the prior feature image into a third network model for enhancement processing to obtain the output third medical image;

[0133] Wherein, the third network model is composed of a concatenation layer and at least two convolutional networks connected in series, and the number of convolutional kernels corresponding to at least two convolutional networks decreases in sequence.

[0134] The medical image reconstruction device provided by the embodiments of the present invention can execute the medical image reconstruction method provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method.

[0135] Figure 7 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. The electronic device 10 is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital assistants, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0136] Such as Figure 7As shown, the electronic device 10 includes at least one processor 11 and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. The memory stores computer programs executable by the at least one processor 11. The processor 11 can execute various appropriate actions and processes according to the computer programs stored in the read-only memory (ROM) 12 or the computer programs loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0137] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information or data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0138] The processor 11 can be various general and / or special processing components with processing and computing capabilities. Some examples of the processor 11 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the medical image reconstruction method provided in the above embodiments.

[0139] In some embodiments, the method for reconstructing a medical image provided in the above embodiments may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps in the method for reconstructing a medical image described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to execute the method for reconstructing a medical image by any other suitable means (e.g., by means of firmware).

[0140] In particular, according to an embodiment of the present invention, the process described above with reference to the flowchart may be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product that includes a computer program carried on a non-transitory computer-readable medium, the computer program including program code for performing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from a network via the communication unit 19, or installed from the storage unit 18, or installed from the ROM 12. When the computer program is executed by the processor 11, the above-described functions defined in the method of the embodiment of the present invention are performed.

[0141] The various embodiments of the systems and techniques described above herein may be implemented in the following systems or combinations thereof: digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), system on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: implemented in one or more computer programs that may be executed and / or interpreted on a programmable system including at least one programmable processor, the programmable processor may be a dedicated or general-purpose programmable processor that may receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0142] A computer program for implementing the medical image reconstruction method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the computer programs are executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer programs can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.

[0143] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable storage medium. Examples of machine-readable storage media would include electrical connections based on at least one wire, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0144] To provide interaction with a user, the systems and techniques described herein can be implemented on a terminal device having: a display device for displaying information to the user (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the terminal device. Other kinds of devices can also provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0145] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected with each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: Local Area Network (LAN), Wide Area Network (WAN), blockchain network, and the Internet.

[0146] A computing system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and Virtual Private Server (VPS) services.

[0147] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is imposed herein.

[0148] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A method for reconstructing a medical image, characterized in that: include: Reconstructing the low-quality first medical image to obtain a second medical image; Inputting the second medical image into the first network model for text generation to obtain an output image prompt word set; Inputting the second medical image and the image prompt word set into a second network model for feature extraction to obtain an output prior feature image; According to the prior feature image, enhancing the second medical image to obtain a third medical image; Using the third medical image as the low-quality first medical image in the next iterative reconstruction process, and returning to the step of reconstructing the low-quality first medical image to obtain the second medical image; Until the reconstruction iteration end condition is met, the second medical image in the current iterative reconstruction process is used as the reconstructed medical image.

2. The method according to claim 1, characterized in that The first network model includes at least two parallel proxy robot networks. Accordingly, the second medical image is input into the first network model for text generation to obtain an output image prompt word set, including: Through each agent robot network, text generation is performed on the second medical image to obtain image prompt words; An image prompt word set is determined according to at least two image prompt words.

3. The method according to claim 2, characterized in that The agent robot network is composed of at least two convolutional layers, a global average pooling layer and two fully connected layers connected in series; Among them, the numbers of convolution kernels corresponding to at least two convolutional layers increase successively, and the number of neurons corresponding to the first fully connected layer is greater than the number of neurons corresponding to the second fully connected layer.

4. The method according to claim 1, characterized in that: The second network model includes an image encoder, a mask encoder and a priori decoder; Inputting the second medical image and the image prompt word set into a second network model for feature extraction to obtain an output prior feature image, including: Inputting the second medical image into the image encoder for image encoding to obtain output image encoding features; Inputting the image prompt word set into the mask encoder for mask encoding to obtain output mask encoding features; The image coding features and the mask coding features are input into the priori decoder for decoding processing to obtain an output priori feature image.

5. The method according to claim 4, characterized in that The image encoder is composed of at least two transformer networks connected in series, wherein the transformer network adopts a multi-head self-attention mechanism, and two adjacent transformer networks are connected through residual connection and layer normalization; The mask encoder is composed of at least two convolution modules connected in series, two adjacent convolution modules are connected by upsampling, and the number of convolution kernels corresponding to the at least two convolution modules is the same; The prior decoder is a cross-attention Transformer structure.

6. The method according to claim 1, characterized in that The step of performing enhancement processing on the second medical image according to the prior feature image to obtain a third medical image includes: Inputting the second medical image and the prior feature image into a third network model for enhancement processing to obtain an output third medical image; Among them, the third network model is composed of a splicing layer and at least two convolutional networks connected in series, and the number of convolution kernels corresponding to the at least two convolutional networks respectively decreases successively.

7. A medical image reconstruction device, characterized in that: include: A second medical image determination module, used for reconstructing the low-quality first medical image to obtain a second medical image; An image prompt word set output module, used for inputting the second medical image into the first network model for text generation to obtain an output image prompt word set; A priori feature image output module, used for inputting the second medical image and the image prompt word set into a second network model for feature extraction to obtain an output priori feature image; A third medical image determination module, configured to perform enhancement processing on the second medical image according to the prior feature image to obtain a third medical image; an iterative execution module, configured to use the third medical image as the low-quality first medical image in the next iterative reconstruction process, and return to execute the step of reconstructing the low-quality first medical image to obtain the second medical image; The reconstructed medical image determination module is used to use the second medical image in the current iterative reconstruction process as the reconstructed medical image until the reconstruction iteration end condition is met.

8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can perform the medical image reconstruction method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the medical image reconstruction method according to any one of claims 1 to 6 when executed.

10. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the medical image reconstruction method according to any one of claims 1 to 6.