Super-resolution interpolation reconstruction method, device, equipment and medium
Through the frequency domain spatial domain collaborative learning module, the problem of insufficient details on the sagittal and coronal planes of CT images is solved, and effective learning of high-frequency information and improvement of image details are achieved.
Patent Information
- Application Number
- CN202510085926.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-01-20
AI Technical Summary
CT images show insufficient details on the sagittal and coronal planes, resulting in poor visibility of lesions or anatomical structures, and it is difficult to learn high-frequency information, which is prone to fit deviations.
The frequency domain spatial domain collaborative learning module is adopted to convert spatial domain features into frequency domain features through frequency domain branches, so that high-frequency features are characterized in the frequency domain; at the same time, the spatial domain branch captures local and global features at different scales through the SwinTrans layer, realizing the complementarity of features between the frequency domain and the spatial domain.
The global feature representation of CT images on different scales is improved, the learning ability of high-frequency information is enhanced, and the detailed display of images and the visibility of the lesion is improved.
Smart Images

Figure CN119991441A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image analysis, and in particular to a super-resolution interpolation reconstruction method, device, equipment and medium that combines spatial domain and frequency domain information for joint learning. Background Art
[0002] Computed Tomography (CT) is one of the important tools for doctors to diagnose diseases and preoperative planning. It has the advantages of fast scanning speed and clear imaging of bone structures. During the imaging process, in order to reduce radiation to patients and ensure acquisition efficiency, CT volume data is usually reconstructed from thicker slices. This causes most CT images to be anisotropic, which is manifested in that the slices in the horizontal plane are sparser than those in the sagittal and coronal planes. This difference in resolution will result in the inability to display sufficient details in the sagittal and coronal planes, thereby affecting the visibility of lesions or anatomical structures (especially small lesions or fine structures), and at the same time, it brings difficulties to the accuracy of quantitative evaluation when facing tasks such as volume measurement and shape analysis of tumors or organs. Therefore, increasing the inter-layer resolution is necessary for clinical practice.
[0003] In the process of super-resolution reconstruction of medical images, information such as organs, blood vessels and bone structures often contains more anatomical information. In the spatial domain, this information is often contained in high-dimensional features. However, during the training process, the convolution-based network structure tends to extract low-frequency information first and then further fit high-frequency information, which makes it more difficult to learn high-frequency information and is prone to fitting deviation. Summary of the invention
[0004] In view of the above problems, the present invention provides a super-resolution interpolation reconstruction method, apparatus, device and medium for overcoming the above problems or at least partially solving the above problems.
[0005] The present invention provides the following scheme:
[0006] A super-resolution interpolation reconstruction method, comprising:
[0007] Acquire a horizontal CT image to be interpolated, a coronal CT image, and a sagittal CT image generated by slicing from an axial angle;
[0008] Taking the horizontal position as the main feature extraction direction, feature extraction is performed on the to-be-interpolated horizontal CT image, the coronal CT image, and the sagittal CT image to obtain a first feature set, wherein the first feature set includes horizontal view features, coronal view features extracted along the channel and height dimensions, and sagittal view features extracted along the channel and height dimensions;
[0009] Using a first multi-view feature fusion module to fuse the first feature set to obtain a second feature set;
[0010] Inputting the second feature set into the frequency-domain-spatial-domain collaborative learning module, and using the second multi-view feature fusion module to fuse the features output by the frequency-domain-spatial-domain collaborative learning module to obtain a third feature set;
[0011] Using a feature fusion module to fuse the first feature set, the second feature set and the third feature set to obtain a dense interpolation result corresponding to the horizontal CT image to be interpolated;
[0012] Among them, the frequency domain and space domain collaborative learning module includes a frequency domain branch and a space domain branch. The frequency domain branch is used to convert the space domain features into frequency domain features so that the high-frequency features that are difficult to capture in the space domain can be represented in the frequency domain; the space domain branch is used to extract the spatial characteristic information of the image so as to capture the local features and global features at different scales through the SwinTrans layer.
[0013] Preferably, the frequency-domain-spatial-domain collaborative learning module includes a Transformer-based network structure.
[0014] Preferably: the frequency domain branch is used to convert the spatial domain features into frequency domain features by performing discrete cosine transform on the features.
[0015] Preferably, the frequency domain branch is used to convert the spatial domain features into frequency domain features by performing discrete cosine transform on the features, including:
[0016] Generate two discrete cosine transform coefficient matrices C along the height and width directions respectively h and C w , as follows:
[0017]
[0018] Where k1 and k2 are the frequency indices in the height and width directions, p represents the row index of the image, and q represents the column index of the image;
[0019] The spatial domain features Discrete cosine transform is performed in the height and width directions to obtain frequency domain features Right now:
[0020]
[0021] Where: f represents the feature set, and T represents the transpose of the matrix.
[0022] Preferably, the spatial domain branch is used to extract the spatial characteristic information of the image, including:
[0023] The spatial features of the input image are extracted through the convolution layer to obtain the features D, H, and W are the resolutions of the horizontal, coronal, and sagittal planes, respectively;
[0024] After linear embedding and dimension transformation, the features become C represents the number of channels;
[0025] Said Entering the SwinTrans layer, the local and global dependencies in the image are captured through the sliding window self-attention mechanism and the multi-head self-attention mechanism, realizing the global feature representation of CT images at different scales, as shown in the following formula:
[0026]
[0027] In the formula, is the i-th SwinTrans layer.
[0028] Preferably: Before entering the SwinTrans layer, the frequency domain features are layer normalized;
[0029] The sliding window of the SwinTrans layer divides the entire frequency band into non-overlapping M×M windows and performs self-attention operation;
[0030] The sliding window method is used to allow the features of different frequency bands to interact, realizing feature learning across frequency bands and channels;
[0031] The frequency domain features are transformed into spatial domain features through inverse discrete cosine transform
[0032] Preferably, the first multi-view feature fusion module is represented by the following formula:
[0033]
[0034] Where: F axial ∈R D×H×W Represents the given horizontal view feature, F sag ∈R W×D×H represents the sagittal view features extracted along the channel and height dimensions, F cor ∈R H×D×W Represents the coronal view features extracted along the channel and width dimensions.
[0035] A super-resolution interpolation reconstruction device, used to perform the above-mentioned super-resolution interpolation reconstruction method, the device comprising:
[0036] An image acquisition unit, used for acquiring a horizontal CT image to be interpolated, a coronal CT image, and a sagittal CT image generated by performing slice processing from an axial angle;
[0037] A first feature set acquisition unit is used to perform feature extraction on the horizontal CT image to be interpolated, the coronal CT image, and the sagittal CT image respectively with the horizontal position as the main feature extraction direction to obtain a first feature set, wherein the first feature set includes horizontal view features, coronal view features extracted along channel and height dimensions, and sagittal view features extracted along channel and height dimensions;
[0038] A second feature set acquisition unit, configured to fuse the first feature set using a first multi-view feature fusion module to obtain a second feature set;
[0039] A frequency-domain-spatial-domain collaborative learning unit, configured to input the second feature set into the frequency-domain-spatial-domain collaborative learning module, and fuse the features output by the frequency-domain-spatial-domain collaborative learning module using a second multi-view feature fusion module to obtain a third feature set;
[0040] A feature fusion unit, configured to use a feature fusion module to fuse the first feature set, the second feature set and the third feature set to obtain a dense interpolation result corresponding to the horizontal CT image to be interpolated;
[0041] Among them, the frequency domain and space domain collaborative learning module includes a frequency domain branch and a space domain branch. The frequency domain branch is used to convert the space domain features into frequency domain features so that the high-frequency features that are difficult to capture in the space domain can be represented in the frequency domain; the space domain branch is used to extract the spatial characteristic information of the image so as to capture the local features and global features at different scales through the SwinTrans layer.
[0042] A super-resolution interpolation reconstruction device, the device comprising a processor and a memory:
[0043] The memory is used to store program code and transmit the program code to the processor;
[0044] The processor is used to execute the above-mentioned super-resolution interpolation reconstruction method according to the instructions in the program code.
[0045] A computer-readable storage medium is used to store program codes, and the program codes are used to execute the above-mentioned super-resolution interpolation reconstruction method.
[0046] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:
[0047] The embodiments of the present application provide a super-resolution interpolation reconstruction method, device, equipment and medium, which can generate dense horizontal interpolation results from sparsely sampled images to increase inter-layer resolution. The method uses a frequency domain and spatial domain collaborative learning module to achieve feature complementarity in the frequency domain and spatial domain. High-frequency features can be learned fairly with low-frequency features, so as to better learn high-frequency information to make up for the shortcomings of the spatial domain. The horizontal position is used as the main feature extraction direction, and the sagittal and coronal information is fused to further improve the interpolation reconstruction effect. A multi-view feature fusion module is used to fuse information from different views, and complementary information is extracted from the three views by making features from different views interact in real time without increasing the amount of calculation too much.
[0048] Of course, any product implementing the present invention does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention, and for ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0050] Figure 1 is a flow chart of a super-resolution interpolation reconstruction method provided by an embodiment of the present invention;
[0051] Figure 2 is a network structure diagram of a super-resolution interpolation reconstruction method provided by an embodiment of the present invention;
[0052] Figure 3 It is a frequency domain and space domain collaborative learning structure diagram provided by an embodiment of the present invention;
[0053] Figure 4 is a multi-view feature fusion structure diagram provided by an embodiment of the present invention;
[0054] Figure 5 It is a quantitative comparison diagram of the method provided by the embodiment of the present invention and other advanced methods;
[0055] Figure 6 is a schematic diagram for visual comparison between the method provided by an embodiment of the present invention and other advanced methods;
[0056] Figure 7 is a schematic diagram of a super-resolution interpolation reconstruction device provided by an embodiment of the present invention;
[0057] Figure 8 Schematic diagram of a super-resolution interpolation reconstruction device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0058] The technical scheme in the embodiment of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiment of the present invention. Obviously, the described embodiment is only a part of the embodiment of the present invention, not all of the embodiments. Based on the embodiment of the present invention, all other embodiments obtained by ordinary technicians in this field belong to the scope of protection of the present invention.
[0059] See also Figure 1 , is a super-resolution interpolation reconstruction method provided by an embodiment of the present invention, such as Figure 1 As shown, the method may include:
[0060] S101: Acquire a horizontal CT image to be interpolated, a coronal CT image, and a sagittal CT image generated by performing slice processing from an axial angle;
[0061] S102: Taking the horizontal position as the main feature extraction direction, feature extraction is performed on the horizontal CT image to be interpolated, the coronal CT image, and the sagittal CT image to obtain a first feature set, wherein the first feature set includes horizontal view features, coronal view features extracted along the channel and height dimensions, and sagittal view features extracted along the channel and height dimensions;
[0062] S103: using a first multi-view feature fusion module to fuse the first feature set to obtain a second feature set;
[0063] S104: inputting the second feature set into the frequency-domain-spatial-domain collaborative learning module, and using a second multi-view feature fusion module to fuse the features output by the frequency-domain-spatial-domain collaborative learning module to obtain a third feature set;
[0064] S105: using a feature fusion module to fuse the first feature set, the second feature set and the third feature set to obtain a dense interpolation result corresponding to the horizontal CT image to be interpolated;
[0065] Among them, the frequency domain and space domain collaborative learning module includes a frequency domain branch and a space domain branch. The frequency domain branch is used to convert the space domain features into frequency domain features so that the high-frequency features that are difficult to capture in the space domain can be represented in the frequency domain; the space domain branch is used to extract the spatial characteristic information of the image so as to capture the local features and global features at different scales through the SwinTrans layer.
[0066] In specific implementation, the embodiment of the present application can provide that the frequency domain and space domain collaborative learning module includes a Transformer-based network structure.
[0067] The frequency domain branch is used to convert the spatial domain features into frequency domain features by performing discrete cosine transform on the features.
[0068] Generate two discrete cosine transform coefficient matrices C along the height and width directions respectively h and C w , as follows:
[0069]
[0070] Where k1 and k2 are the frequency indices in the height and width directions, p represents the row index of the image, and q represents the column index of the image;
[0071] The spatial domain features Discrete cosine transform is performed in the height and width directions to obtain frequency domain features Right now:
[0072]
[0073] Where: f represents the feature set, and T represents the transpose of the matrix.
[0074] The spatial domain branch is used to extract the spatial characteristic information of the image, including:
[0075] The spatial features of the input image are extracted through the convolution layer to obtain the features D, H, and W are the resolutions of the horizontal, coronal, and sagittal planes, respectively;
[0076] After linear embedding and dimension transformation, the features become C represents the number of channels;
[0077] Said Entering the SwinTrans layer, the local and global dependencies in the image are captured through the sliding window self-attention mechanism and the multi-head self-attention mechanism, realizing the global feature representation of CT images at different scales, as shown in the following formula:
[0078]
[0079] In the formula, is the i-th SwinTrans layer.
[0080] Said Before entering the SwinTrans layer, the frequency domain features are layer normalized;
[0081] The sliding window of the SwinTrans layer divides the entire frequency band into non-overlapping M×M windows and performs self-attention operation;
[0082] The sliding window method is used to allow the features of different frequency bands to interact, realizing feature learning across frequency bands and channels;
[0083] The frequency domain features are transformed into spatial domain features through inverse discrete cosine transform
[0084] The first multi-view feature fusion module is expressed by the following formula:
[0085]
[0086] Where: Faxial∈R D×H×W Represents the given horizontal bit view feature, Fsag∈R W×D×H represents the sagittal view features extracted along the channel and height dimensions, Fcor∈R H×D×W Represents the coronal view features extracted along the channel and width dimensions.
[0087] The super-resolution interpolation reconstruction method provided in the embodiment of the present application can generate dense horizontal interpolation results from sparsely sampled images and increase the inter-layer resolution. The method adopts a frequency domain and spatial domain collaborative learning module to realize the feature complementarity of the frequency domain and the spatial domain. High-frequency features can be learned fairly with low-frequency features, so as to better learn high-frequency information to make up for the shortcomings of the spatial domain. The horizontal position is used as the main feature extraction direction, and the sagittal and coronal information is fused to further improve the interpolation reconstruction effect. A multi-view feature fusion module is used to fuse information from different views, and complementary information is extracted from the three views by making the features from different views interact in real time without increasing the amount of calculation too much.
[0088] The super-resolution interpolation reconstruction method provided by this application is introduced in detail below.
[0089] The embodiment of the present application implements a method for interpolating a CT image in the horizontal position by using an interpolation network that combines spatial domain and frequency domain information for joint learning. The network structure is as follows: Figure 2 In the specific implementation, this method takes the horizontal position as the main feature extraction direction, and integrates the information of the sagittal and coronal positions to further improve the interpolation reconstruction effect.
[0090] The anisotropic CT data input to the network is V∈R D×H×W , where D, H and W are the resolutions of the horizontal, coronal and sagittal planes respectively. <H,D<W,H=W。
[0091] The 2D image sets of the horizontal, coronal and sagittal views are denoted as V axial ∈R D×W ,V cor ∈R D ×W and Vsag ∈R D×H .
[0092] The goal of this method is to generate a dense interpolation result from a sparsely sampled image V in S is the upsampling factor in the horizontal direction, and S-1 slices are interpolated between every two adjacent slices in the horizontal direction. The method mainly includes two core parts, namely the frequency domain and spatial domain collaborative learning module and the multi-view feature fusion module.
[0093] The frequency-space domain collaborative learning module is a Transformer-based network structure, which includes a frequency-domain branch and a space-domain branch.
[0094] The frequency domain branch converts the spatial domain to the frequency domain by performing discrete cosine transform on the features, so that high-frequency features that are difficult to capture in the spatial domain can be better represented in the frequency domain.
[0095] The spatial domain branch is used to extract the spatial characteristic information of the image. Through the proposed SwinTrans layer, local and global features at different scales can be efficiently captured. Finally, the features of the frequency domain and the spatial domain are complementary. SwinTrans is a hierarchical network with sliding window operations. It limits the attention mechanism to a window, which not only takes into account local information, but also saves computation.
[0096] In the spatial domain, it includes several steps including convolution, linear embedding, dimension transformation and global feature extraction, such as Figure 3 First, the spatial features of the input image are extracted through the convolution layer to obtain the features After linear embedding and dimension transformation, the features become Where C is the number of channels. Entering the SwinTrans layer, the local and global dependencies in the image are captured through the sliding window self-attention mechanism and the multi-head self-attention mechanism, which improves the global feature representation of CT images at different scales. As shown in the following formula:
[0097]
[0098] in, is the i-th SwinTrans layer. The final output is
[0099] In the frequency domain, since discrete cosine transform is more computationally efficient than Fourier transform in processing image data and can reduce the problem of discontinuity at the boundary, this method uses discrete cosine transform (DCT) to convert spatial domain features into frequency domain features. Transformed into frequency domain features through DCT Specifically, first generate two DCT coefficient matrices C along the height and width directions respectively h and C w , as follows:
[0100]
[0101] Among them, k1 and k2 are the frequency indexes in two directions. DCT transform is performed in two directions, namely:
[0102]
[0103] In order to reduce internal covariate shift and training stability, the frequency domain features are layer normalized before entering the SwinTrans layer. In order to allow features of different frequency bands to be learned equally, the sliding window of the SwinTrans layer first divides the entire frequency band into non-overlapping M×M windows and performs self-attention operations to ensure that the frequency band of each window can be kept within a small difference range. Then, a sliding window method is used to allow features of different frequency bands to interact, breaking the boundaries of window divisions to achieve cross-frequency band and cross-channel feature learning. Finally, the frequency domain features are transformed into spatial domain features through inverse discrete cosine transform (IDCT).
[0104] Since the axial view contains more information than other views, our interpolation network chooses to slice from the axial perspective. However, sometimes even if the information of truly complete organ tissues can be obtained from the axial view, the organ morphology results obtained in the coronal or parasagittal view may deviate from the correct direction. Therefore, it is necessary to combine the information of the sagittal and coronal views and supplement the features learned from the horizontal views. Existing methods process the three views separately and then fuse them in the final stage, which results in no real-time interaction between these features.
[0105] To this end, this application provides two multi-view feature fusion modules that fuse information from different views, such as Figure 4 As shown in the figure, by making the features from different views interact in real time, complementary information can be extracted from the three views without increasing the amount of computation. Given the horizontal view feature Faxial∈R D×H×W ,Sagittal view feature F sag ∈R W×D×H Extract along the channel and height dimensions, coronal view features F cor ∈R H×D×W Extract along the channel and width dimensions. The final output features It can be expressed as:
[0106]
[0107] After the horizontal CT image to be interpolated, the coronal CT image and the sagittal CT image are encoded by the encoder, the first feature set is first obtained. The first feature set First, it is input into the first multi-view feature fusion module for fusion to obtain the second feature set Second feature set The input frequency domain and space domain collaborative learning module performs frequency domain and space domain collaborative learning, and then the obtained features are fused through the second multi-view feature fusion module to obtain the third feature set. Then the first feature set is transformed into Second feature set And the third feature set Channel fusion is performed, and finally the decoder decodes the fused features to obtain a dense interpolation result.
[0108] During the network training phase, the network provided in the embodiment of the present application may use the L1 norm as a loss function.
[0109] In order to verify the effectiveness of the proposed method, experiments were conducted on 120 CT data provided by Wenzhou Eye and Optometry Hospital. The resolution of CT is 512×512×L, where L∈[102,179]. Figure 5 Quantitative and qualitative comparisons of this method with current state-of-the-art methods are presented. Figure 6 The visual comparison between the method provided by this application and other advanced methods is shown. The first row is the result of horizontal bit slices, and the second row is the error map. The brighter the image, the larger the error.
[0110] Specifically, the method provided in this application achieved the best performance at upsampling factors of x2, x4 and x6, achieving PSNR values of 40.24dB, 34.37dB and 31.39dB respectively; at the same time, the best results were achieved on SSIM.
[0111] It can be seen that the method provided in the embodiment of the present application can achieve good performance under different upsampling factors, and as the upsampling factor increases, the performance degradation is small.
[0112] See also Figure 7 , the embodiment of the present application can also provide a super-resolution interpolation reconstruction device, such as Figure 7 As shown, the apparatus for performing the above-mentioned super-resolution interpolation reconstruction method may include:
[0113] An image acquisition unit 701 is used to acquire a horizontal CT image to be interpolated, a coronal CT image, and a sagittal CT image generated by performing slice processing from an axial angle;
[0114] A first feature set acquisition unit 702 is configured to extract features of the to-be-interpolated horizontal CT image, the coronal CT image, and the sagittal CT image, respectively, with the horizontal position being the main feature extraction direction, to obtain a first feature set, wherein the first feature set includes horizontal view features, coronal view features extracted along channel and height dimensions, and sagittal view features extracted along channel and height dimensions;
[0115] A second feature set acquisition unit 703, configured to use a first multi-view feature fusion module to fuse the first feature set to obtain a second feature set;
[0116] The frequency-domain-spatial-domain collaborative learning unit 704 is used to input the second feature set into the frequency-domain-spatial-domain collaborative learning module, and use the second multi-view feature fusion module to fuse the features output by the frequency-domain-spatial-domain collaborative learning module to obtain a third feature set;
[0117] A feature fusion unit 705 is used to fuse the first feature set, the second feature set and the third feature set using a feature fusion module to obtain a dense interpolation result corresponding to the horizontal CT image to be interpolated;
[0118] Among them, the frequency domain and space domain collaborative learning module includes a frequency domain branch and a space domain branch. The frequency domain branch is used to convert the space domain features into frequency domain features so that the high-frequency features that are difficult to capture in the space domain can be represented in the frequency domain; the space domain branch is used to extract the spatial characteristic information of the image so as to capture the local features and global features at different scales through the SwinTrans layer.
[0119] The embodiment of the present application may also provide a super-resolution interpolation reconstruction device, the device comprising a processor and a memory:
[0120] The memory is used to store program code and transmit the program code to the processor;
[0121] The processor is used to execute the steps of the above-mentioned super-resolution interpolation reconstruction method according to the instructions in the program code.
[0122] like Figure 8 As shown, a super-resolution interpolation reconstruction device provided in an embodiment of the present application may include: a processor 10, a memory 11, a communication interface 12 and a communication bus 13. The processor 10, the memory 11 and the communication interface 12 communicate with each other through the communication bus 13.
[0123] In the embodiment of the present application, the processor 10 may be a central processing unit (CPU), an application specific integrated circuit, a digital signal processor, a field programmable gate array or other programmable logic devices, etc.
[0124] The processor 10 may call a program stored in the memory 11. Specifically, the processor 10 may execute operations in an embodiment of the super-resolution interpolation reconstruction method.
[0125] The memory 11 is used to store one or more programs, which may include program codes, and the program codes include computer operation instructions. In the embodiment of the present application, the memory 11 at least stores programs for implementing the following functions:
[0126] Acquire a horizontal CT image to be interpolated, a coronal CT image, and a sagittal CT image generated by slicing from an axial angle;
[0127] Taking the horizontal position as the main feature extraction direction, feature extraction is performed on the to-be-interpolated horizontal CT image, the coronal CT image, and the sagittal CT image to obtain a first feature set, wherein the first feature set includes horizontal view features, coronal view features extracted along the channel and height dimensions, and sagittal view features extracted along the channel and height dimensions;
[0128] Using a first multi-view feature fusion module to fuse the first feature set to obtain a second feature set;
[0129] Inputting the second feature set into the frequency-domain-spatial-domain collaborative learning module, and using the second multi-view feature fusion module to fuse the features output by the frequency-domain-spatial-domain collaborative learning module to obtain a third feature set;
[0130] Using a feature fusion module to fuse the first feature set, the second feature set and the third feature set to obtain a dense interpolation result corresponding to the horizontal CT image to be interpolated;
[0131] Among them, the frequency domain and space domain collaborative learning module includes a frequency domain branch and a space domain branch. The frequency domain branch is used to convert the space domain features into frequency domain features so that the high-frequency features that are difficult to capture in the space domain can be represented in the frequency domain; the space domain branch is used to extract the spatial characteristic information of the image so as to capture the local features and global features at different scales through the SwinTrans layer.
[0132] In one possible implementation, the memory 11 may include a program storage area and a data storage area, wherein the program storage area can store an operating system and application programs required for at least one function (such as a file creation function, a data reading and writing function), etc.; the data storage area can store data created during use, such as initialization data, etc.
[0133] In addition, the memory 11 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device or other volatile solid-state storage device.
[0134] The communication interface 12 may be an interface of a communication module, and is used to connect to other devices or systems.
[0135] Of course, it should be noted that Figure 8 The structure shown does not constitute a limitation on the super-resolution interpolation reconstruction device in the embodiment of the present application. In practical applications, the super-resolution interpolation reconstruction device may include Figure 8 More or fewer components than shown, or combinations of certain components.
[0136] The embodiment of the present application may also provide a computer-readable storage medium, wherein the computer-readable storage medium is used to store program code, and the program code is used to execute the steps of the above-mentioned super-resolution interpolation reconstruction method.
[0137] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0138] It can be known from the description of the above implementation methods that those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present application or certain parts of the embodiments.
[0139] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can refer to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment. The system and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without creative work.
[0140] The above description is only a preferred embodiment of the present invention and is not intended to limit the protection scope of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention are included in the protection scope of the present invention.
Claims
1. A super-resolution interpolation reconstruction method, characterized in that: include: Acquire a horizontal CT image to be interpolated, a coronal CT image, and a sagittal CT image generated by slicing from an axial angle; Taking the horizontal position as the main feature extraction direction, feature extraction is performed on the to-be-interpolated horizontal CT image, the coronal CT image, and the sagittal CT image to obtain a first feature set, wherein the first feature set includes horizontal view features, coronal view features extracted along the channel and height dimensions, and sagittal view features extracted along the channel and height dimensions; Using a first multi-view feature fusion module to fuse the first feature set to obtain a second feature set; Inputting the second feature set into the frequency-domain-spatial-domain collaborative learning module, and using the second multi-view feature fusion module to fuse the features output by the frequency-domain-spatial-domain collaborative learning module to obtain a third feature set; Using a feature fusion module to fuse the first feature set, the second feature set and the third feature set to obtain a dense interpolation result corresponding to the horizontal CT image to be interpolated; Among them, the frequency domain and space domain collaborative learning module includes a frequency domain branch and a space domain branch. The frequency domain branch is used to convert the space domain features into frequency domain features so that the high-frequency features that are difficult to capture in the space domain can be represented in the frequency domain; the space domain branch is used to extract the spatial characteristic information of the image so as to capture the local features and global features at different scales through the SwinTrans layer.
2. The super-resolution interpolation reconstruction method according to claim 1, characterized in that: The frequency-domain and space-domain collaborative learning module includes a Transformer-based network structure.
3. The super-resolution interpolation reconstruction method according to claim 2, characterized in that: The frequency domain branch is used to convert the spatial domain features into frequency domain features by performing discrete cosine transform on the features.
4. The super-resolution interpolation reconstruction method according to claim 3, characterized in that: The frequency domain branch is used to convert the spatial domain features into frequency domain features by performing discrete cosine transform on the features, including: Generate two discrete cosine transform coefficient matrices C along the height and width directions respectively h and C w , as follows: Where k1 and k2 are the frequency indices in the height and width directions, p represents the row index of the image, and q represents the column index of the image; The spatial domain features Discrete cosine transform is performed in the height and width directions to obtain frequency domain features Right now: Where: f represents the feature set, and T represents the transpose of the matrix.
5. The super-resolution interpolation reconstruction method according to claim 2, characterized in that: The spatial domain branch is used to extract the spatial characteristic information of the image, including: The spatial features of the input image are extracted through the convolution layer to obtain the features D, H, and W are the resolutions of the horizontal, coronal, and sagittal planes, respectively; After linear embedding and dimension transformation, the features become C represents the number of channels; Said Entering the SwinTrans layer, the local and global dependencies in the image are captured through the sliding window self-attention mechanism and the multi-head self-attention mechanism, realizing the global feature representation of CT images at different scales, as shown in the following formula: In the formula, is the i-th SwinTrans layer.
6. The super-resolution interpolation reconstruction method according to claim 5, characterized in that: Said Before entering the SwinTrans layer, the frequency domain features are layer normalized; The sliding window of the SwinTrans layer divides the entire frequency band into non-overlapping M×M windows and performs self-attention operation; The sliding window method is used to allow the features of different frequency bands to interact, realizing feature learning across frequency bands and channels; The frequency domain features are transformed into spatial domain features through inverse discrete cosine transform 7. The super-resolution interpolation reconstruction method according to claim 1, characterized in that: The first multi-view feature fusion module is expressed by the following formula: Where: F axial ∈R D×H×W Represents the given horizontal view feature, F sag ∈R W×D×H represents the sagittal view features extracted along the channel and height dimensions, F cor ∈R H×D×W Represents the coronal view features extracted along the channel and width dimensions.
8. A super-resolution interpolation reconstruction device, characterized in that: Used to perform the super-resolution interpolation reconstruction method according to any one of claims 1 to 7, the device comprising: An image acquisition unit, used for acquiring a horizontal CT image to be interpolated, a coronal CT image, and a sagittal CT image generated by performing slice processing from an axial angle; A first feature set acquisition unit is used to perform feature extraction on the horizontal CT image to be interpolated, the coronal CT image, and the sagittal CT image respectively with the horizontal position as the main feature extraction direction to obtain a first feature set, wherein the first feature set includes horizontal view features, coronal view features extracted along channel and height dimensions, and sagittal view features extracted along channel and height dimensions; A second feature set acquisition unit, configured to fuse the first feature set using a first multi-view feature fusion module to obtain a second feature set; A frequency-domain-spatial-domain collaborative learning unit, configured to input the second feature set into the frequency-domain-spatial-domain collaborative learning module, and fuse the features output by the frequency-domain-spatial-domain collaborative learning module using a second multi-view feature fusion module to obtain a third feature set; A feature fusion unit, configured to use a feature fusion module to fuse the first feature set, the second feature set and the third feature set to obtain a dense interpolation result corresponding to the horizontal CT image to be interpolated; Among them, the frequency domain and space domain collaborative learning module includes a frequency domain branch and a space domain branch. The frequency domain branch is used to convert the space domain features into frequency domain features so that the high-frequency features that are difficult to capture in the space domain can be represented in the frequency domain; the space domain branch is used to extract the spatial characteristic information of the image so as to capture the local features and global features at different scales through the SwinTrans layer.
9. A super-resolution interpolation reconstruction device, characterized in that: The device comprises a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is used to execute the super-resolution interpolation reconstruction method according to any one of claims 1 to 7 according to the instructions in the program code.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium is used to store program codes, and the program codes are used to execute the super-resolution interpolation reconstruction method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Low-resolution face image super-resolution reconstruction method
CN111612695A
Image super-resolution reconstruction method based on lightweight double-branch network
CN116664398A
Binocular image super-resolution reconstruction method and device and storage medium
CN118154432A
Lightweight image super-resolution reconstruction method based on frequency domain-spatial domain assisted Mama
CN119251051A
Systems and method for reconstructing 3D radio frequency tomographic images
US20180218519A1