Image Reconstruction Method, Apparatus, Storage Medium, and Electronic Device

By using asymmetric encoder-decoder structure and self-attention mechanism in CT image reconstruction, the shortcomings of convolutional neural networks in adapting to the differences in human anatomical structure and modeling long-distance dependence are solved, high-quality image reconstruction is achieved, and the requirements of different resolutions are supported.

CN114742916BActive Publication Date: 2025-05-27INFERVISION MEDICAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210415984.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-20
Publication Date
2025-05-27
Estimated Expiration
2042-04-20

AI Technical Summary

Technical Problem

When the prior art reconstructs thick-layer CT images into thin-layer CT images, the inherent defects of convolutional neural networks make it difficult to adapt to the differences in human anatomy structure, and the long-distance dependence in the image cannot be effectively modeled, affecting the reconstruction quality.

Method used

Using an asymmetric encoder-decoder structure, M mask vector matrices are added to the first feature diagram sequence, and the self-attention mechanism in the decoder module is used to determine the learning vector matrix corresponding to each of the M mask vector matrices, thereby generating a second image sequence with higher spatial resolution.

Benefits of technology

Through the use of the self-attention mechanism, different regional features in the image can be extracted more effectively, the image reconstruction quality can be improved, and the number of mask vector matrix can be adjusted as needed, and image sequences of any resolution can be generated, simplifying the process and saving resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114742916B_ABST
    Figure CN114742916B_ABST
Patent Text Reader

Abstract

The present application provides an image reconstruction method, apparatus, storage medium, and electronic device, relating to the technical field of image processing. The method includes: adding M mask vector matrices to the first feature map sequence output by an encoder module, where the first feature map sequence is the feature map sequence of a first image sequence, and M is a positive integer; using the self-attention mechanism in a decoder module to determine, based on the first feature map sequence, M learning vector matrices respectively corresponding to the M mask vector matrices; and generating a second image sequence based on the first feature map sequence and the M learning vector matrices. Based on an asymmetric encoder-decoder structure and the self-attention mechanism, the present application can perform adaptive feature extraction on different regions in the first feature map sequence, can more accurately reconstruct the first image sequence of different tissue structures, and can generate a second image sequence with a required slice thickness according to the needs of doctor diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and particularly relates to an image reconstruction method, apparatus, storage medium, and electronic device. Background Art

[0002] The slice thickness of Computed Tomography (CT) images is generally below 10 mm. Among them, the slice thickness of thin-slice CT images is generally below 5 mm, and the slice thickness of thick-slice CT images is generally between 5 - 10 mm. The smaller the slice thickness of CT images, the smaller the slice interval and the higher the spatial resolution it has. Therefore, thin-slice CT images have a smaller slice interval and a higher spatial resolution, which can help doctors achieve more accurate diagnoses. However, compared with thick-slice CT images, thin-slice CT images have increased the data volume by several times, which poses great challenges to data transmission and storage. In addition, thin-slice CT images require a longer scanning time, which means that patients need to receive a longer radiation scan. Therefore, in actual clinical practice, thick-slice CT images are still one of the conventional choices for imaging examinations. However, there are significant differences in the inter-slice and intra-slice resolutions of thick-slice CT images, and each voxel in thick-slice CT images is anisotropic.

[0003] Reconstructing thick-slice CT images into thin-slice CT images can, to some extent, solve the above problems. In related technologies, methods for image reconstruction based on deep learning are generally developed based on the convolutional neural network structure. However, after training, the convolutional parameters of the convolutional operator are fixed, making it difficult to adapt to the differences in human anatomical structures, and the inherent inductive bias characteristics of the convolutional operator cannot effectively model the long-range dependencies in images. Summary of the Invention

[0004] To solve the above technical problems, this application is proposed. Embodiments of this application provide an image reconstruction method, apparatus, storage medium, and electronic device.

[0005] In a first aspect, an embodiment of this application provides an image reconstruction method for reconstructing a first image sequence into a second image sequence, where the slice thickness of the second image sequence is less than that of the first image sequence. The method includes: adding M mask vector matrices to the first feature map sequence output by the encoder module, where the first feature map sequence is the feature map sequence of the first image sequence, and M is a positive integer; using the self-attention mechanism in the decoder module to determine M learning vector matrices corresponding to the M mask vector matrices respectively based on the first feature map sequence; generating the second image sequence based on the first feature map sequence and the M learning vector matrices.

[0006] In combination with the first aspect, in some implementations of the first aspect, the decoder module includes a cross-plane self-attention block unit and a first Transformer calculation unit. The cross-plane self-attention block unit and the first Transformer calculation unit incorporate a self-attention mechanism. The first feature map sequence includes first plane feature data, second plane feature data, and third plane feature data corresponding to three coordinate planes of a spatial rectangular coordinate system respectively. Using the self-attention mechanism in the decoder module, based on the first feature map sequence, M learning vector matrices corresponding to M mask vector matrices are determined, including: using the cross-plane self-attention block unit to perform feature data enhancement processing on the first plane feature data and the second plane feature data in the first feature map sequence respectively to generate a second feature map sequence; using the first Transformer calculation unit to perform feature data enhancement processing on the third plane feature data in the second feature map sequence to generate a third feature map sequence; based on the third feature map sequence, determining M learning vector matrices corresponding to M mask vector matrices respectively.

[0007] In combination with the first aspect, in some implementations of the first aspect, using the cross-plane self-attention block unit to perform feature data enhancement processing on the first plane feature data and the second plane feature data in the first feature map sequence respectively to generate a second feature map sequence includes: using the cross-plane self-attention block unit to perform feature position annotation on the first plane feature data and the second plane feature data respectively to generate first plane feature annotation data and second plane feature annotation data; respectively performing extraction and calculation of the feature annotation data on the first plane feature annotation data and the second plane feature annotation data to generate first plane feature extraction data and second plane feature extraction data; adding the first plane feature data and the first plane feature extraction data to obtain first plane feature enhancement data; adding the second plane feature data and the second plane feature extraction data to obtain second plane feature enhancement data; based on the first plane feature enhancement data, the second plane feature enhancement data, and the third plane feature data, generating a second feature map sequence.

[0008] In combination with the first aspect, in some implementations of the first aspect, the third feature map sequence includes P third feature maps, the size of the third feature maps is the same as the size of the learning vector matrices, and P is a positive integer. Based on the third feature map sequence, determining M learning vector matrices corresponding to M mask vector matrices respectively includes: based on the correlation relationship of the feature data between the P third feature maps, determining the eigenvector data of each of the M mask vector matrices; based on the eigenvector data of each of the M mask vector matrices, determining M learning vector matrices corresponding to M mask vector matrices respectively.

[0009] In connection with the first aspect, in some implementations of the first aspect, the first feature map sequence includes N feature maps, where N is a positive integer greater than or equal to 2. Adding M mask vector matrices to the first feature map sequence output by the encoder module includes: for the N feature maps, adding S mask vector matrices between every two adjacent feature maps so as to add M mask vector matrices, where the value of S is determined based on the difference between the slice thickness of the second image sequence and the slice thickness of the first image sequence.

[0010] In connection with the first aspect, in some implementations of the first aspect, the encoder module includes a feature mapping unit and a second Transformer calculation unit. Before adding M mask vector matrices to the first feature map sequence output by the encoder module, the method further includes: using the feature mapping unit to perform feature linear mapping on the first image sequence to obtain a feature mapped image sequence; using the second Transformer calculation unit to perform feature extraction on the feature mapped image sequence to obtain the first feature map sequence.

[0011] In connection with the first aspect, in some implementations of the first aspect, the mask vector matrix includes a zero vector matrix.

[0012] In connection with the first aspect, in some implementations of the first aspect, generating the second image sequence based on the first feature map sequence and M learning vector matrices includes: performing an identical feature subtraction operation on the first feature map sequence based on the M learning vector matrices to obtain a fourth feature map sequence; performing linear mapping on the fourth feature map sequence and the M learning vector matrices to generate the second image sequence.

[0013] In a second aspect, an embodiment of the present application provides an image reconstruction device for reconstructing a first image sequence into a second image sequence, where the slice thickness of the second image sequence is less than the slice thickness of the first image sequence. The device includes: an adding module for adding M mask vector matrices to the first feature map sequence output by the encoder module, where the first feature map sequence is the feature map sequence of the first image sequence and M is a positive integer; a determining module for using the self-attention mechanism in the decoder module to determine M learning vector matrices corresponding to the M mask vector matrices based on the first feature map sequence; and a generating module for generating the second image sequence based on the first feature map sequence and the M learning vector matrices.

[0014] In a third aspect, an embodiment of the present application provides a computer-readable storage medium storing a computer program for executing the method described in the first aspect above.

[0015] Fourthly, an embodiment of the present application provides an electronic device, which includes: a processor; a memory for storing instructions executable by the processor; and the processor is configured to execute the method described in the first aspect above.

[0016] The technical solution of the present application is based on an asymmetric encoder-decoder structure. M mask vector matrices are added to the first feature map sequence, and then the image reconstruction problem is transformed into the problem of recovering the M mask vector matrices. In addition, the decoder module contains a self-attention mechanism, which can adaptively extract the feature data of different regions in the first feature map sequence, improving the image reconstruction quality. By adjusting the number of added mask vector matrices, a second image sequence with any resolution can be generated to meet the different needs of doctor diagnosis. This method is simple, convenient, and widely applicable, and can be implemented without training multiple models. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] By describing the embodiments of the present application in more detail with reference to the accompanying drawings, the above and other objects, features, and advantages of the present application will become more obvious. The accompanying drawings are used to provide a further understanding of the embodiments of the present application, and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application, but do not constitute a limitation to the present application. In the accompanying drawings, the same reference numerals generally represent the same components or steps.

[0018] Figure 1 Shown is a schematic diagram of a scenario applicable to the embodiments of the present application.

[0019] Figure 2 Shown is a schematic flowchart of an image reconstruction method provided by an exemplary embodiment of the present application.

[0020] Figure 3 Shown is a schematic flowchart of determining M learning vector matrices corresponding to the M mask vector matrices respectively based on the first feature map sequence provided by an exemplary embodiment of the present application.

[0021] Figure 4 Shown is a schematic flowchart of generating a second feature map sequence provided by an exemplary embodiment of the present application.

[0022] Figure 5 Shown is a schematic flowchart of determining M learning vector matrices corresponding to the M mask vector matrices respectively based on the third feature map sequence provided by an exemplary embodiment of the present application.

[0023] Figure 6 Shown is a schematic flowchart of an image reconstruction method provided by another exemplary embodiment of the present application.

[0024] Figure 7The figure shows a schematic flowchart of generating a second image sequence provided by an exemplary embodiment of the present application.

[0025] Figure 8 The figure shows a schematic diagram of the network structure of an image reconstruction algorithm.

[0026] Figure 9 The figure shows a schematic diagram of the structure of an image reconstruction device provided by an exemplary embodiment of the present application.

[0027] Figure 10 The figure shows a schematic diagram of the structure of an electronic device provided by an exemplary embodiment of the present application. Detailed implementation manners

[0028] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0029] The slice thickness of thin-slice CT images is generally below 5 mm, and the slice thickness of thick-slice CT images is generally 5 - 10 mm. Currently, thick-slice CT is one of the conventional means of imaging examinations. However, there are significant differences in the resolution between and within slices of thick-slice CT. The in-slice resolution of CT scans is usually about 1 mm, while the slice thickness of thick-slice CT can reach 5 - 10 mm. This means that each voxel in thick-slice CT is anisotropic and the gap is large, which is very disadvantageous for tasks based on the analysis of regions of interest.

[0030] Reconstructing thick-slice CT into thin-slice CT can, to some extent, solve the above problems. Thick-slice CT reconstruction can be regarded as super-resolution of images in the depth dimension. Methods for image super-resolution include interpolation based on filter operators, reconstruction based on edge gradient information, etc. In recent years, deep learning methods have made great progress in the field of image super-resolution and have become a better implementation method in the current super-resolution field. Existing methods for reconstructing thick-slice CT into thin-slice CT based on deep learning are all developed based on the convolutional neural network structure, but there are still the following problems.

[0031] First, after the convolution operator is trained, the convolution parameters are fixed. For any region input into the convolutional neural network, the same parameters are used for calculation in the same network stage. However, this method does not take into account the differences in anatomical structures such as the human body or animals. Since different anatomical structures have different absorption rates for radiation, the density distribution patterns of different anatomical regions are also different. For example, the density of the lung region is usually small, while the density of the abdominal region is usually large. Using the same convolution operator to calculate the lung and abdominal regions may lead to sub-optimal solutions.

[0032] Second, the bias characteristic of the convolution operator assumes that the relationship between adjacent regions in the image is closer. However, in the field of super-resolution, it has been found that the non-local similarity of the image plays an important role in improving the image reconstruction quality. Constrained by its own calculation method, the convolution operator cannot effectively model the long-range dependencies in the image and cannot utilize this non-local similarity.

[0033] Finally, the algorithms based on the convolutional neural network structure usually adopt a fixed multiple of the upsampling ratio. For different resolution requirements, multiple models need to be trained, which will consume more resources.

[0034] Figure 1 The following is a schematic diagram of a scenario applicable to the embodiments of the present application. This scenario includes an image acquisition device 110 and a computer device 120. There is a communication connection relationship between the computer device 120 and the image acquisition device 110. An encoder module and a decoder module are deployed on the computing device 120.

[0035] Specifically, the image acquisition device 110 is used to acquire the first image sequence. The image acquisition device 110 can be a CT scanner, an X-ray machine, or other devices with image acquisition functions. The present application does not specifically limit the structure of the image acquisition device 110.

[0036] The computer device 120 is used to receive the first image sequence acquired by the image acquisition device 110, add M mask vector matrices to the first feature map sequence output by the encoder module, and then use the self-attention mechanism in the decoder module to determine M learning vector matrices corresponding to the M mask vectors according to the first feature map sequence. Then, according to the first feature map sequence and the M learning vector matrices, a second image sequence is generated. The computer device 120 can be a general-purpose computer or a computer device composed of dedicated integrated circuits. The present application does not make a limit on this. For example, the computer device 120 can be a mobile terminal device such as a tablet computer or a personal computer. Moreover, the number of computer devices 120 can be one or more, and their types can be the same or different. The embodiments of the present application do not limit the number and type of the computer devices 120.

[0037] Figure 2 The figure shows a schematic flowchart of an image reconstruction method provided by an exemplary embodiment of the present application. It is used to reconstruct a first image sequence into a second image sequence, where the slice thickness of the second image sequence is smaller than that of the first image sequence. As Figure 2 shown, the image reconstruction method provided by the embodiment of the present application includes the following steps.

[0038] Step 30: Add M mask vector matrices to the first feature map sequence output by the encoder module.

[0039] The first feature map sequence is the feature map sequence of the first image sequence, and M is a positive integer. The mask vector matrix is used to represent the to-be-learned feature information of the missing layers of the first image sequence relative to the second image sequence.

[0040] The first image sequence can be a CT image or a Magnetic Resonance Imaging (MRI) image. The embodiment of the present application does not make specific limitations on the type of the first image sequence. Preferably, the CT image with fast scanning time and clear image is used as the first image sequence in the embodiment of the present application.

[0041] Step 40: Use the self-attention mechanism in the decoder module to determine M learning vector matrices corresponding to the M mask vector matrices based on the first feature map sequence.

[0042] The learning vector matrix is used to represent the learned feature information of the missing layers of the first image sequence relative to the second image sequence.

[0043] Specifically, by using the self-attention mechanism in the decoder module, adaptive feature extraction can be performed on different regions in the first feature map sequence, and then M learning vector matrices corresponding to the M mask vector matrices can be determined.

[0044] The structures of the encoder module and the decoder module in this embodiment are asymmetric, which is mainly manifested in two aspects. First, the number of input layers of the encoder module is different from the number of output layers of the decoder module, because the input of the encoder module is the first image sequence, while the input of the decoder module is the first feature map sequence corresponding to the first image sequence plus M mask vector matrices. For example, if the first image sequence includes K first images, the number of input layers of the encoder is K, while the number of input layers of the decoder is (K + M). Second, the network structures of the encoder module and the decoder module are different, because M mask vector matrices are input into the decoder module, and when these mask vector matrices are input into the decoder module, they are the to-be-learned feature information of the missing layers of the first image sequence relative to the second image sequence. Inside the decoder module, through a large number of interactive calculations and learning on the feature data of the first feature map sequence, M learning vector matrices corresponding to the M mask vector matrices are finally obtained.

[0045] Step 50: Generate a second image sequence based on the first feature map sequence and M learning vector matrices.

[0046] The technical solution in this embodiment is based on an asymmetric encoder-decoder structure. M mask vector matrices are added to the first feature map sequence, and the image reconstruction problem is transformed into the problem of recovering M unknown mask vector matrices from the feature data of the known first image sequence. In addition, by using the self-attention mechanism in the decoder module, the relevant feature data in different regions of the first feature map sequence can be adaptively extracted, improving the image reconstruction quality. Finally, the first image sequence may be a thick-slice CT image sequence or a thin-slice CT image sequence, and the generated second image sequence may be a thick-slice CT image sequence or a thin-slice CT image sequence. The conversion relationship between the two can be achieved by adaptively adjusting the number of added mask vector matrices according to the doctor's diagnosis needs, that is, the different resolution requirements of the second image sequence can be met by adjusting the number of added mask vector matrices. This solution is simple, convenient, resource-saving, and does not require training multiple models to achieve.

[0047] Figure 3 The following is a schematic flowchart of a process for determining M learning vector matrices corresponding to M mask vector matrices respectively based on a first feature map sequence provided by an exemplary embodiment of the present application. Figure 2 Based on the embodiment shown, Figure 3 the following embodiment is extended. Figure 3 The differences between the embodiment shown Figure 2 and the embodiment shown are emphasized below, and the same parts will not be elaborated.

[0048] As Figure 3 shown, using the self-attention mechanism in the decoder module to determine M learning vector matrices corresponding to M mask vector matrices respectively based on the first feature map sequence includes the following steps.

[0049] Step 41: Use a Through-plane Attention Block (TAB) unit to perform feature data enhancement processing on the first plane feature data and the second plane feature data in the first feature map sequence respectively to generate a second feature map sequence.

[0050] First, the decoder module includes a TAB unit and a first Transformer calculation unit. The TAB unit and the first Transformer calculation unit contain a self-attention mechanism. The first feature map sequence includes first plane feature data, second plane feature data, and third plane feature data corresponding to the three coordinate planes of the spatial rectangular coordinate system respectively.

[0051] The spatial rectangular coordinate system is the O-xyz rectangular coordinate system, that is, this spatial rectangular coordinate system is composed of the origin O, and the x-axis, y-axis, and z-axis that pass through the origin O and are perpendicular to each other. Through this spatial rectangular coordinate system, three mutually perpendicular coordinate planes can be determined, namely the xOy plane, the xOz plane, and the yOz plane. Placing the first feature map sequence in the spatial rectangular coordinate system, the first plane in the first feature map sequence is any one of the xOy plane, the xOz plane, and the yOz plane, the second plane is one of the xOy plane, the xOz plane, and the yOz plane that is different from the first plane, and the third plane is one of the xOy plane, the xOz plane, and the yOz plane that is different from the first plane and the second plane.

[0052] Exemplarily, the first plane feature data is the xOz plane feature data, the second plane feature data is the yOz plane feature data, and the third plane feature data is the xOy plane feature data. Then, the first feature map sequence and the M mask vector matrices are input into the TAB unit in the decoder module. The TAB unit first performs self-attention feature extraction interaction on the xOz plane feature data and the yOz plane feature data of the input first feature map sequence to obtain a second feature map sequence. Among them, the xOy plane feature data in the second feature map sequence is not processed. The TAB unit adjusts the size of the second feature map sequence and inputs the second feature map sequence with adjusted size into the first Transformer calculation unit.

[0053] Step 42: Use the first Transformer calculation unit to perform feature data enhancement processing on the third plane feature data in the second feature map sequence to generate a third feature map sequence.

[0054] Continuing with the example in step 41, use the first Transformer calculation unit to perform self-attention feature extraction interaction on the xOy plane feature data in the second feature map sequence to obtain a third feature map sequence, and adjust the size of the third feature map sequence to meet the needs of the next-stage calculation.

[0055] Step 43: Based on the third feature map sequence, determine the M learning vector matrices corresponding to the M mask vector matrices respectively.

[0056] Specifically, the xOz plane feature data, the yOz plane feature data, and the xOy plane feature data of the third feature map sequence are all enhanced. According to the feature data of the second feature map sequence after enhancement, determine the M learning vector matrices corresponding to the M mask vector matrices respectively.

[0057] In the embodiments of the present application, both the TAB unit and the first Transformer computing unit include a self-attention mechanism, which can adaptively extract features from different regions of the first feature map sequence and the second feature map sequence input thereto, improving the image reconstruction quality. Secondly, the first Transformer computing unit can effectively model the long-range dependencies of the second feature map sequence and utilize the non-local similarity of the second feature map sequence to improve the image reconstruction quality. By the TAB unit and the first Transformer computing unit, the feature data of different planes of the first feature map sequence is enhanced, which can further improve the image reconstruction quality.

[0058] Figure 4 The following is a schematic flowchart of generating a second feature map sequence provided by an exemplary embodiment of the present application. In Figure 3 Based on the embodiment shown, Figure 4 the following embodiment is extended. Figure 4 The differences between the embodiment shown below and Figure 3 the embodiment shown will be emphasized below, and the same parts will not be elaborated.

[0059] As Figure 4 shown, by using the cross-plane self-attention block unit, feature data enhancement processing is respectively performed on the first plane feature data and the second plane feature data in the first feature map sequence to generate a second feature map sequence, including the following steps.

[0060] Step 411, using the cross-plane self-attention block unit, respectively perform feature position annotation on the first plane feature data and the second plane feature data to generate first plane feature annotation data and second plane feature annotation data.

[0061] Specifically, using the TAB unit, relative position encoding is performed on the input first plane feature data and second plane feature data to avoid loss of the position information of the first plane feature data and the second plane feature data, introducing the relative position of the mask vector matrix and the feature data of the first feature map sequence into the feature interaction process. On the other hand, the non-local similarity of the first feature map sequence can be introduced into the feature interaction process. The generated first plane feature annotation data and second plane feature annotation data contain the position encoding information of the feature data.

[0062] In addition, using the TAB unit, relative position encoding is also performed on the third plane feature data of the first feature map sequence to generate third plane feature annotation data, which also contains the position encoding information of the feature data, so that when the first Transformer computing unit enhances the third plane feature data, the position information of the third plane feature data will not be lost.

[0063] Step 412: Extract and calculate the feature annotation data from the first planar feature annotation data and the second planar feature annotation data respectively to generate the first planar feature extraction data and the second planar feature extraction data.

[0064] Step 413: Add the first planar feature data and the first planar feature extraction data to obtain the first planar feature enhancement data.

[0065] Step 414: Add the second planar feature data and the second planar feature extraction data to obtain the second planar feature enhancement data.

[0066] Step 415: Generate the second feature map sequence based on the first planar feature enhancement data, the second planar feature enhancement data, and the third planar feature data.

[0067] In the embodiment of the present application, by performing feature position annotation on the first planar feature data, the second planar feature data, and the third planar feature data, the non-local self-similarity of the first feature map sequence can be introduced into the feature interaction process, and the loss of the position information of the first planar feature data, the second planar feature data, and the third planar feature data can also be avoided. In addition, through the technical solution in this embodiment, the first planar feature data and the second planar feature data are enhanced, which helps to improve the image reconstruction quality.

[0068] Figure 5 The following is a schematic flowchart of determining the M learning vector matrices corresponding to the M mask vector matrices based on the third feature map sequence provided by an exemplary embodiment of the present application. On the basis of Figure 3 the embodiment shown, an embodiment extended therefrom is Figure 5 shown. Next, the differences between the embodiment Figure 5 shown and the embodiment Figure 3 shown will be emphasized, and the same parts will not be elaborated again.

[0069] As Figure 5 shown, determining the M learning vector matrices corresponding to the M mask vector matrices based on the third feature map sequence includes the following steps.

[0070] Step 431: Determine the eigenvector data of each of the M mask vector matrices based on the correlation relationship of the feature data between the P third feature maps.

[0071] Among them, the third feature map sequence includes P third feature maps, the size of the third feature map is the same as the size of the learning vector matrix, and P is a positive integer.

[0072] Step 432: Determine the M learning vector matrices corresponding to the M mask vector matrices based on the eigenvector data of each of the M mask vector matrices.

[0073] In this embodiment, by using the correlation relationship of the feature data between P third feature maps, the non-local similarity between the P third feature maps is introduced, and then M learning vector matrices are determined, so as to convert the common super-resolution problem into the problem of restoring the learning vector matrices, and finally image reconstruction with any magnification factor is realized, and the applicability of this solution is wider.

[0074] In an embodiment of the present application, the first feature map sequence includes N feature maps, where N is a positive integer greater than or equal to 2. M mask vector matrices are added to the first feature map sequence output by the encoder module, including: for N feature maps, S mask vector matrices are added between every two adjacent feature maps, so as to add M mask vector matrices, where the value of S is determined based on the difference between the slice thickness of the second image sequence and the slice thickness of the first image sequence.

[0075] Exemplarily, the first image sequence includes 10 first images, and the slice thickness of each image is 5 mm. Correspondingly, the first feature map sequence includes 10 feature maps, and the slice thickness of each feature map sequence is 5 mm. Two mask vector matrices are added in each of the upper and lower directions of each feature map. That is, 4 mask vector matrices are added between every two adjacent feature maps, and finally the slice thickness of the obtained second image sequence is 1 mm.

[0076] Through the technical solution in this embodiment, it is possible to adjust the number of added mask vector matrices according to the doctor's diagnosis needs to obtain a second image sequence with any slice thickness. This method is simpler, more convenient, and has a wider applicability.

[0077] Figure 6 The following shows a schematic flowchart of an image reconstruction method provided by another exemplary embodiment of the present application. Figure 2 Based on the embodiment shown, Figure 6 the following embodiment is extended. Figure 6 The differences between the embodiment shown Figure 2 and the embodiment shown are described below, and the same parts will not be repeated.

[0078] As Figure 6 shown, before adding M mask vector matrices to the first feature map sequence output by the encoder module, the following steps are further included.

[0079] Step 10: Use the feature mapping unit to perform feature linear mapping on the first image sequence to obtain a feature mapping image sequence.

[0080] Among them, the encoder module includes a feature mapping unit and a second Transformer calculation unit.

[0081] Step 20: Use the second Transformer calculation unit to perform feature extraction on the feature mapping image sequence to obtain a first feature map sequence.

[0082] The second Transformer computing unit includes a self-attention mechanism, which can perform adaptive feature extraction on different regions in the feature map image sequence.

[0083] Through the technical solution in this embodiment, the image information of the first image sequence can be converted into a vector representation form, and the features of the first image sequence are fully extracted for the decoder module to learn.

[0084] In an embodiment of the present application, the mask vector matrix includes a zero vector matrix.

[0085] Specifically, the mask vector matrix includes a zero vector matrix, and its corresponding learning vector matrix needs to be obtained through a large amount of calculation and learning in the decoder module.

[0086] Figure 7 The following is a schematic flowchart of generating a second image sequence provided by an exemplary embodiment of the present application. Figure 2 Based on the embodiment shown, Figure 7 the following embodiment is extended. Figure 7 The following focuses on Figure 2 the differences between the embodiment shown and

[0087] As Figure 7 shown, based on the first feature map sequence and M learning vector matrices, generating a second image sequence includes the following steps.

[0088] Step 51, based on the M learning vector matrices, perform the same feature subtraction operation on the first feature map sequence to obtain a fourth feature map sequence.

[0089] The features in the first feature map sequence are represented in the form of vectors. Subtract the features in the first feature map sequence that are the same as the vectors in the M learning vector matrices to obtain a fourth feature map sequence.

[0090] Step 52, perform a linear mapping on the fourth feature map sequence and the M learning vector matrices to generate a second image sequence.

[0091] Perform a linear mapping of the features of the fourth feature map sequence and the M learning vector matrices to generate a second image sequence.

[0092] Through the technical solution in this embodiment, the fourth feature sequence output by the decoder and the M learning vector matrices can be mapped into an image, and finally a second image sequence with the desired slice thickness is obtained.

[0093] Figure 8The figure shows a schematic diagram of the network structure of the image reconstruction algorithm. Taking the generation of the second CT image sequence with a slice thickness of 1 mm from the first CT image sequence with a slice thickness of 5 mm as an example, the relevant algorithm is exemplarily described according to Figure 8 the asymmetric encoder-decoder network structure shown.

[0094] The first CT image sequence contains G consecutive CT images. Assuming that the size of each CT image is 512x512, taking the first CT image sequence as the input of the encoder module, the input scale of the encoder module is Gx512x512.

[0095] The input first image sequence is divided into blocks. Assuming that the size of each block is 4x4, that is, 4x4 pixels are used as an overall pixel vector, then the size of the input first image sequence is converted to (16G)x(512 / 4)x(512 / 4). Then, the feature mapping unit in the encoder module is used for feature extraction, and the size of the obtained feature mapping image sequence is Dx(512 / 4)x(512 / 4). Among them, D is a preset mapping relationship, and D can take 128 or any other value, which is not affected by the 16G variable.

[0096] The obtained feature mapping image sequence is input into the first size adjustment unit to adjust the size of the feature mapping image sequence to (D / 16)x512x512.

[0097] The feature mapping image sequence with adjusted size is input into the second Transformer calculation unit. Through E consecutive second Transformer calculation units, the self-attention mechanism is used to fully extract the features of the information of the input image sequence, and the first feature map sequence is output. The first feature map sequence is input into the second size adjustment unit, and after size adjustment, the first feature map sequence with a size of GxCx512x512 output by the encoder module is obtained. Among them, C represents the number of channels at each voxel position. At this time, each feature map in the first feature map sequence is converted into a feature vector representation.

[0098] A learnable mask vector matrix is added between the encoder module and the decoder module to represent the missing layers of the first CT image sequence relative to the second CT image sequence. The first CT image sequence with a slice thickness of 5 mm is 4 layers different from the second CT image sequence with a slice thickness of 1 mm. Therefore, two mask vector matrices can be added in the up and down directions of each first CT image in the first CT image sequence. Therefore, 4 mask vector matrices need to be added between every two adjacent first CT images, as shown in Figure 8As shown in the figure, 1 is the added mask vector matrix. It should be noted that for the first CT images 2 and 3 at both ends of the first CT image sequence, only two mask vector matrices need to be added below the first CT image 2 and two mask vector matrices need to be added above the first CT image 3. All mask vector matrices are initialized as learnable vectors, and the size of each mask vector matrix is Cx512x512.

[0099] There are many ways to insert the mask vector matrix. For example, an empty vector can be inserted as the mask vector matrix, or the first feature map sequence can be interpolated. According to the different definition methods of the used mask vector matrix, super-resolution with any magnification can be achieved in the same network. Taking the 5-fold super-resolution of this embodiment as an example, after adding the mask vector matrix, the output size of the decoder becomes (5G - 4)xCx512x512.

[0100] The decoder module includes L feature interaction modules. In each feature interaction module, first, the cross-plane self-attention block is used to perform self-attention feature extraction and interaction between the first plane feature data and the second plane feature data of the input first feature map sequence. Continuing with the example in step 41, that is, first, the cross-plane self-attention block is used to perform self-attention feature extraction and interaction between the xOz plane feature data and the yOz plane feature data, and the generated second feature map sequence is input to the third size adjustment unit for size adjustment. Then, the size-adjusted second feature map sequence is input to 4 consecutive first Transformer calculation units to perform self-attention feature extraction and interaction of the xOy plane feature data, generating a third feature map sequence. Based on the third feature map sequence, a fourth feature map sequence and M learning vector matrices are obtained.

[0101] The fourth size adjustment unit adjusts the size of the input fourth feature map sequence and M learning vector matrices, and outputs the size-adjusted fourth feature map sequence and M learning vector matrices. Among them, in the cross-plane self-attention block unit, the calculation processes of the xOz plane feature data and the yOz plane feature data are synchronous and share the model parameters of the self-attention block.

[0102] The output and input of the decoder module have the same size. Only after passing through the decoder module, the M mask vector matrices have specific representations, and M learning vector matrices are obtained. The learning vector matrices are equivalent to the image information of the missing layers of the first image sequence relative to the second image sequence.

[0103] Finally, through a linear mapping, the fourth feature map sequence output by the decoder module and the M learning vector matrices are mapped into an image. The output size of the decoder module is (5G - 4) x C x 512 x 512. After the linear mapping, the size becomes (5G - 4) x 1 x 512 x 512, that is, the number of channels is equal to 1, indicating that there is only a unique predicted value at each spatial position of the second image sequence.

[0104] Among them, during the training phase of the decoder module, the difference between the output of the linear mapping and the calculated value of the real second image sequence is used as the loss function to guide the training of the decoder module. During the test phase, the output of the linear mapping is the reconstruction prediction result of the second image sequence.

[0105] As described above in conjunction with Figures 1 to 8 , the embodiments of the image reconstruction method of the present application have been described in detail. Below, in conjunction with Figure 9 , the embodiments of the image reconstruction device of the present application will be described in detail. It should be understood that the description of the embodiments of the image reconstruction method corresponds to the description of the embodiments of the image reconstruction device. Therefore, for the parts not described in detail, reference can be made to the previous method embodiments.

[0106] Figure 9 The following shows a schematic structural diagram of an image reconstruction device provided by an exemplary embodiment of the present application. As Figure 9 shown, the image reconstruction device provided by the embodiments of the present application includes:

[0107] An adding module 910, configured to add M mask vector matrices to the first feature map sequence output by the encoder module, where the first feature map sequence is the feature map sequence of the first image sequence, and M is a positive integer;

[0108] A determining module 920, configured to use the self-attention mechanism in the decoder module to determine M learning vector matrices corresponding to the M mask vector matrices respectively based on the first feature map sequence;

[0109] A generating module 930, configured to generate a second image sequence based on the first feature map sequence and the M learning vector matrices.

[0110] In an embodiment of the present application, the determining module 920 is further configured to use a cross-plane self-attention block unit to perform feature data enhancement processing on the first plane feature data and the second plane feature data in the first feature map sequence respectively to generate a second feature map sequence; use a first Transformer calculation unit to perform feature data enhancement processing on the third plane feature data in the second feature map sequence to generate a third feature map sequence; and determine M learning vector matrices corresponding to the M mask vector matrices respectively based on the third feature map sequence.

[0111] In one embodiment of the present application, the determination module 920 is further configured to use a cross-plane self-attention block unit to perform feature position annotation on the first plane feature data and the second plane feature data respectively, generating first plane feature annotation data and second plane feature annotation data; perform extraction and calculation of the feature annotation data on the first plane feature annotation data and the second plane feature annotation data respectively, generating first plane feature extraction data and second plane feature extraction data; add the first plane feature data and the first plane feature extraction data to obtain first plane feature enhancement data; add the second plane feature data and the second plane feature extraction data to obtain second plane feature enhancement data; generate a second feature map sequence based on the first plane feature enhancement data, the second plane feature enhancement data, and the third plane feature data.

[0112] In one embodiment of the present application, the determination module 920 is further configured to determine the eigenvector data of each of the M mask vector matrices based on the correlation relationship of the feature data between the P third feature maps; determine the M learning vector matrices corresponding to each of the M mask vector matrices based on the eigenvector data of each of the M mask vector matrices.

[0113] In one embodiment of the present application, the addition module 910 is further configured to, for N feature maps, add S mask vector matrices between every two adjacent feature maps, so as to add M mask vector matrices, where the value of S is determined based on the difference between the slice thickness of the second image sequence and the slice thickness of the first image sequence.

[0114] In one embodiment of the present application, the addition module 910 is further configured to use a feature mapping unit to perform feature linear mapping on the first image sequence to obtain a feature mapping image sequence; use a second Transformer calculation unit to perform feature extraction on the feature mapping image sequence to obtain a first feature map sequence.

[0115] In one embodiment of the present application, the mask vector matrix includes a zero vector matrix.

[0116] In one embodiment of the present application, the generation module 930 is further configured to perform an identical feature subtraction operation on the first feature map sequence based on the M learning vector matrices to obtain a fourth feature map sequence; perform linear mapping on the fourth feature map sequence and the M learning vector matrices to generate a second image sequence.

[0117] Next, refer to Figure 10 to describe the electronic device according to the embodiments of the present application. Figure 10 The following shows a schematic structural diagram of an electronic device provided by an exemplary embodiment of the present application.

[0118] As Figure 10 shown, the electronic device 100 includes one or more processors 1001 and a memory 1002.

[0119] The processor 1001 can be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and can control other components in the electronic device 100 to perform desired functions.

[0120] The memory 1002 can include one or more computer program products, and the computer program products can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory can include, for example, random access memory (RAM) and / or cache memory, etc. The non-volatile memory can include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions can be stored on the computer-readable storage media, and the processor 1001 can run the program instructions to implement the image reconstruction methods of the various embodiments of the present application described above and / or other desired functions. Various contents such as a first feature map sequence, M mask vector matrices, M learning vector matrices, a second image sequence, etc. can also be stored in the computer-readable storage media.

[0121] In one example, the electronic device 100 can further include: an input device 1003 and an output device 1004, and these components are interconnected through a bus system and / or other forms of connection mechanisms (not shown).

[0122] The input device 1003 can include, for example, a keyboard, a mouse, etc.

[0123] The output device 1004 can output various information to the outside, including a first feature map sequence, M mask vector matrices, M learning vector matrices, a second image sequence, etc. The output device 1004 can include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.

[0124] Of course, for simplicity, Figure 10 only some of the components related to the present application in the electronic device 100 are shown, and components such as buses, input / output interfaces, etc. are omitted. In addition, according to specific application scenarios, the electronic device 100 can further include any other appropriate components.

[0125] In addition to the above methods and devices, the embodiments of the present application can also be computer program products, which include computer program instructions that, when run by a processor, cause the processor to execute the steps in the image reconstruction methods according to the various embodiments of the present application described above in this specification.

[0126] The computer program product may be written in any combination of one or more programming languages for programming code to perform the operations of the embodiments of the present application. The programming languages include object-oriented programming languages such as Java, C++, etc., and also include conventional procedural programming languages such as the "C" language or similar programming languages. The programming code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0127] In addition, an embodiment of the present application may also be a computer-readable storage medium having computer program instructions stored thereon, and when the computer program instructions are run by a processor, the processor is caused to execute the steps in the image reconstruction method according to various embodiments of the present application described above in this specification.

[0128] The computer-readable storage medium may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may include, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0129] The basic principles of the present application have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, benefits, effects, etc. mentioned in the present application are only examples and not limitations, and it cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present application. In addition, the above-disclosed specific details are only for illustrative and easy-to-understand purposes and are not limitations. The above details do not limit the present application to necessarily implement using the above specific details.

[0130] The block diagrams of the devices, apparatuses, equipment, and systems involved in this application are only illustrative examples and are not intended to require or imply that they must be connected, arranged, and configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, equipment, and systems can be connected, arranged, and configured in any way. Words such as "including," "comprising," "having," etc. are open-ended terms that mean "including but not limited to" and can be used interchangeably with each other. The word "or" and "and" used herein refer to the phrase "and / or" and can be used interchangeably with it, unless the context clearly indicates otherwise. The phrase "such as" used herein refers to the phrase "such as but not limited to" and can be used interchangeably with it.

[0131] It should also be noted that in the devices, equipment, and methods of this application, each component or each step can be decomposed and / or recombined. These decompositions and / or recombinations should be regarded as equivalent solutions of this application.

[0132] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this application. Therefore, this application is not intended to be limited to the aspects shown herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.

[0133] The above description is not intended to limit the embodiments of this application to the forms disclosed herein. Although multiple example aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, changes, additions, and sub-combinations thereof.

Claims

1. An image reconstruction method, characterized in that, it is used to reconstruct a first image sequence into a second image sequence, wherein the slice thickness of the second image sequence is smaller than that of the first image sequence, the method includes: adding M mask vector matrices to the first feature map sequence output by the encoder module, wherein the first feature map sequence is the feature map sequence of the first image sequence, and M is a positive integer; using the self-attention mechanism in the decoder module, based on the first feature map sequence, determining M learning vector matrices respectively corresponding to the M mask vector matrices; generating the second image sequence based on the first feature map sequence and the M learning vector matrices.

2. The image reconstruction method according to claim 1, characterized in that, the decoder module includes a cross-plane self-attention block unit and a first Transformer calculation unit. The cross-plane self-attention block unit and the first Transformer calculation unit contain the self-attention mechanism. The first feature map sequence includes first plane feature data, second plane feature data, and third plane feature data respectively corresponding to the three coordinate planes of the spatial rectangular coordinate system, the step of using the self-attention mechanism in the decoder module, based on the first feature map sequence, determining M learning vector matrices respectively corresponding to the M mask vector matrices includes: using the cross-plane self-attention block unit to perform feature data enhancement processing on the first plane feature data and the second plane feature data in the first feature map sequence respectively, to generate a second feature map sequence; using the first Transformer calculation unit to perform feature data enhancement processing on the third plane feature data in the second feature map sequence, to generate a third feature map sequence; determining M learning vector matrices respectively corresponding to the M mask vector matrices based on the third feature map sequence.

3. The image reconstruction method according to claim 2, characterized in that, the step of using the cross-plane self-attention block unit to perform feature data enhancement processing on the first plane feature data and the second plane feature data in the first feature map sequence respectively, to generate a second feature map sequence includes: using the cross-plane self-attention block unit, performing feature position annotation on the first plane feature data and the second plane feature data respectively, to generate first plane feature annotation data and second plane feature annotation data; extracting and calculating the feature annotation data of the first plane feature annotation data and the second plane feature annotation data respectively, to generate first plane feature extraction data and second plane feature extraction data; adding the first plane feature data and the first plane feature extraction data to obtain first plane feature enhancement data; adding the second plane feature data and the second plane feature extraction data to obtain second plane feature enhancement data; generating the second feature map sequence based on the first plane feature enhancement data, the second plane feature enhancement data, and the third plane feature data.

4. The image reconstruction method according to claim 2, wherein, the third feature map sequence includes P third feature maps, the size of the third feature map is the same as the size of the learning vector matrix, and P is a positive integer. The determining the M learning vector matrices respectively corresponding to the M mask vector matrices based on the third feature map sequence includes: determining the eigenvector data of each of the M mask vector matrices based on the correlation relationship of the feature data between the P third feature maps; determining the M learning vector matrices respectively corresponding to the M mask vector matrices based on the eigenvector data of each of the M mask vector matrices.

5. The image reconstruction method according to any one of claims 1 to 4, wherein, the first feature map sequence includes N feature maps, N is a positive integer greater than or equal to 2, and adding the M mask vector matrices to the first feature map sequence output by the encoder module includes: for the N feature maps, adding S mask vector matrices between every two adjacent feature maps so as to add the M mask vector matrices, wherein the value of S is determined based on the difference between the slice thickness of the second image sequence and the slice thickness of the first image sequence.

6. The image reconstruction method according to any one of claims 1 to 4, wherein, the encoder module includes a feature mapping unit and a second Transformer calculation unit. Before adding the M mask vector matrices to the first feature map sequence output by the encoder module, it further includes: using the feature mapping unit to perform feature linear mapping on the first image sequence to obtain a feature mapping image sequence; using the second Transformer calculation unit to perform feature extraction on the feature mapping image sequence to obtain the first feature map sequence.

7. The image reconstruction method according to any one of claims 1 to 4, wherein, the mask vector matrix includes a zero vector matrix.

8. The image reconstruction method according to any one of claims 1 to 4, wherein, generating the second image sequence based on the first feature map sequence and the M learning vector matrices includes: performing the same feature subtraction operation on the first feature map sequence based on the M learning vector matrices to obtain a fourth feature map sequence; performing linear mapping on the fourth feature map sequence and the M learning vector matrices to generate the second image sequence.

9. An image reconstruction device, wherein, for reconstructing a first image sequence into a second image sequence, wherein the slice thickness of the second image sequence is less than the slice thickness of the first image sequence, and the device includes: an adding module, configured to add M mask vector matrices to the first feature map sequence output by the encoder module, wherein the first feature map sequence is the feature map sequence of the first image sequence, and M is a positive integer; a determining module, configured to use the self-attention mechanism in the decoder module to determine the M learning vector matrices respectively corresponding to the M mask vector matrices based on the first feature map sequence; A generation module, configured to generate the second image sequence based on the first feature map sequence and the M learning vector matrices.

10. A computer-readable storage medium, characterized in that the storage medium stores a computer program, and the computer program is used to execute the image reconstruction method according to any one of claims 1 to 8 above.

11. An electronic device, characterized in that the electronic device includes: a processor; a memory for storing instructions executable by the processor; the processor is configured to execute the image reconstruction method according to any one of claims 1 to 8 above.

Citation Information

Patent Citations

  • A medical image super-resolution three-dimensional reconstruction method based on two-dimensional image transfer learning

    CN109584164A

  • Systems and methods for positron emission tomography image reconstruction

    CN110298897A