Image reconstruction method and device, electronic equipment and storage medium

By dynamically multiplexing the shallow and deep feature extraction modules of the attention network, the problem of high computational complexity in large-size image processing is solved, and efficient image reconstruction is achieved.

CN120339067APending Publication Date: 2025-07-18SUZHOU GAIDE PHOTOELECTRIC TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510407153.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing image super-resolution method has high computational complexity when processing large-size images, and it is difficult to reduce computational complexity while ensuring image reconstruction quality.

Method used

A dynamic multiplexing attention network is adopted, including shallow feature extraction module, deep feature extraction module and image reconstruction module, and a high-similar attention matrix is selected to share a high-similar attention matrix between different layers through the dynamic multiplexing attention mechanism to reduce redundant calculations.

Benefits of technology

While ensuring the quality of image reconstruction, the calculation complexity is significantly reduced and the efficiency and effect of image reconstruction is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339067A_ABST
    Figure CN120339067A_ABST
Patent Text Reader

Abstract

The invention discloses an image reconstruction method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring a first image, and inputting the first image into a pre-constructed dynamic multiplexing attention network; wherein the dynamic multiplexing attention network comprises a shallow feature extraction module, a deep feature extraction module and an image reconstruction module; performing feature extraction on the first image through a shallow feature extraction module to obtain shallow features of the first image; performing feature extraction on the shallow features based on a dynamic multiplexing attention mechanism through a deep feature extraction module to obtain deep features of the first image; fusing the shallow-layer features and the deep-layer features through an image reconstruction module to generate first fusion features, and performing image reconstruction on the first fusion features to obtain a second image; wherein the resolution of the second image is higher than that of the first image. According to the scheme, the image reconstruction quality can be ensured, and meanwhile, redundant calculation is reduced so as to reduce the calculation complexity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly relates to an image reconstruction method, apparatus, electronic device, and storage medium. Background Art

[0002] Image Super-Resolution (SR) is a computer vision task aimed at generating high-resolution (HR) images from low-resolution (LR) images, and is widely applied in fields such as medical imaging, satellite image processing, and video enhancement. Traditional SR methods mainly rely on interpolation techniques such as bilinear interpolation and bicubic interpolation, but these methods often fail to recover lost high-frequency details. With the development of deep learning, SR methods based on convolutional neural networks have made significant progress. However, SR methods based on convolutional neural networks still have limitations in capturing long-range dependencies. To overcome this limitation, an attention mechanism has been introduced in the image super-resolution task to capture local and global context information. Nevertheless, the image super-resolution task still faces the problem of high computational complexity, especially when processing large-sized images. Summary of the Invention

[0003] The present invention provides an image reconstruction method, apparatus, electronic device, and storage medium, which can reduce redundant calculations to lower the computational complexity while ensuring the quality of image reconstruction.

[0004] According to one aspect of the present invention, an image reconstruction method is provided. The method includes:

[0005] Obtain a first image, and input the first image into a pre-constructed dynamic multiplexing attention network; wherein, the dynamic multiplexing attention network includes a shallow feature extraction module, a deep feature extraction module, and an image reconstruction module;

[0006] Extract features of the first image through the shallow feature extraction module to obtain shallow features of the first image;

[0007] Extract features of the shallow features through the deep feature extraction module based on the dynamic multiplexing attention mechanism to obtain deep features of the first image;

[0008] Fuse the shallow features and the deep features through the image reconstruction module to generate a first fused feature, and perform image reconstruction on the first fused feature to obtain a second image; wherein, the resolution of the second image is higher than that of the first image.

[0009] According to another aspect of the present invention, there is provided an image reconstruction device, the device comprising:

[0010] A first image acquisition module, configured to acquire a first image and input the first image into a pre-constructed dynamic multiplexing attention network; wherein, the dynamic multiplexing attention network includes a shallow feature extraction module, a deep feature extraction module, and an image reconstruction module;

[0011] A shallow feature determination module, configured to perform feature extraction on the first image through the shallow feature extraction module to obtain shallow features of the first image;

[0012] A deep feature determination module, configured to perform feature extraction on the shallow features based on a dynamic multiplexing attention mechanism through the deep feature extraction module to obtain deep features of the first image;

[0013] A second image determination module, configured to fuse the shallow features and the deep features through the image reconstruction module to generate a first fusion feature, and perform image reconstruction on the first fusion feature to obtain a second image; wherein, the resolution of the second image is higher than the resolution of the first image.

[0014] According to another aspect of the present invention, there is provided an electronic device, the electronic device comprising:

[0015] At least one processor; and

[0016] A memory communicatively connected to the at least one processor; wherein,

[0017] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the image reconstruction method according to any embodiment of the present invention.

[0018] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the image reconstruction method according to any embodiment of the present invention when executed.

[0019] The technical solution of the embodiment of the present invention is to obtain a first image and input the first image into a pre-constructed dynamic multiplexing attention network; wherein, the dynamic multiplexing attention network includes a shallow feature extraction module, a deep feature extraction module, and an image reconstruction module; the shallow feature extraction module extracts features from the first image to obtain shallow features of the first image; the deep feature extraction module extracts features from the shallow features based on the dynamic multiplexing attention mechanism to obtain deep features of the first image; the image reconstruction module fuses the shallow features and the deep features to generate a first fused feature, and performs image reconstruction on the first fused feature to obtain a second image; wherein, the resolution of the second image is higher than that of the first image. The technical solution of the embodiment of the present invention extracts features from shallow features based on the dynamic multiplexing attention mechanism, and selectively shares attention matrices with higher similarity between different layers, which can reduce redundant calculations to reduce computational complexity while ensuring the quality of image reconstruction.

[0020] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0022] Figure 1 is a flowchart of an image reconstruction method provided in Embodiment 1 of the present invention;

[0023] Figure 2 is a schematic structural diagram of a dynamic multiplexing attention network provided in Embodiment 1 of the present invention;

[0024] Figure 3 is a flowchart of an image reconstruction method provided in Embodiment 2 of the present invention;

[0025] Figure 4 is a schematic structural diagram of an attention calculation unit provided in Embodiment 2 of the present invention;

[0026] Figure 5 is a schematic structural diagram of an attention gating unit provided in Embodiment 2 of the present invention;

[0027] Figure 6It is a schematic structural diagram of an attention multiplexing unit provided in Embodiment 2 of the present invention;

[0028] Figure 7 It is a schematic structural diagram of a channel attention mechanism provided in Embodiment 2 of the present invention;

[0029] Figure 8 It is a schematic structural diagram of a spatial attention mechanism provided in Embodiment 2 of the present invention;

[0030] Figure 9 It is a schematic structural diagram of an image reconstruction device provided in Embodiment 3 of the present invention;

[0031] Figure 10 It is a schematic structural diagram of an electronic device for implementing the image reconstruction method of the embodiments of the present invention. Detailed implementation manners

[0032] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0033] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily need to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0034] Embodiment 1

[0035] Figure 1 This is a flowchart of an image reconstruction method provided in Embodiment 1 of the present invention. This embodiment is applicable to the case of reconstructing a low-resolution image into a high-resolution image. This method can be executed by an image reconstruction device, which can be implemented in the form of hardware and / or software, and the image reconstruction device can be configured in an electronic device. As Figure 1 shown, the method includes:

[0036] S110. Obtain a first image and input the first image into a pre-constructed dynamic multiplexing attention network; wherein, the dynamic multiplexing attention network includes a shallow feature extraction module, a deep feature extraction module, and an image reconstruction module.

[0037] The first image may be a low-resolution image to be reconstructed. Specifically, the first image may be an image with a resolution lower than a preset resolution threshold, and the preset resolution threshold may be set by those skilled in the art according to actual situations.

[0038] The dynamic multiplexing attention network includes a shallow feature extraction module, a deep feature extraction module, and an image reconstruction module. The shallow feature extraction module includes a 3×3 convolutional layer. The deep feature extraction module includes an attention calculation unit, an attention gating unit, and at least two attention multiplexing units. The image reconstruction module includes a 3×3 convolutional layer and a sub-pixel convolutional layer. Exemplarily, Figure 2 shows a schematic structural diagram of a dynamic multiplexing attention network.

[0039] In an embodiment of the present invention, a first image to be reconstructed can be obtained and input into a pre-constructed dynamic multiplexing attention network, so as to process the first image through the dynamic multiplexing attention network to obtain a first image with a higher resolution than the original image, realizing super-resolution reconstruction of the first image.

[0040] Optionally, the construction process of the dynamic multiplexing attention network includes: inputting a sample image into an initial network to obtain an output result, and calculating a loss value corresponding to the sample image and the output result according to a loss function; wherein, the loss function includes an absolute value loss function and a pruning loss function; performing backpropagation through the loss value to iteratively update the parameter weights in the initial network to obtain a dynamic multiplexing attention network.

[0041] The absolute value loss function is used to measure the pixel difference between the output result of the dynamic multiplexing attention network and the sample image, thereby improving the reconstruction accuracy of the dynamic multiplexing attention network. The pruning loss function is used to optimize the selection of the attention matrix when the dynamic multiplexing attention network multiplexes attention, so as to ensure the balance between image reconstruction and computational efficiency of the dynamic multiplexing attention network.

[0042] In an embodiment of the present invention, the sample image can be first input into the initial network to obtain an output result. According to the loss function composed of the absolute value loss function and the pruning loss function, the loss value corresponding to the sample image and the output result is calculated, and then backpropagation is performed through the loss value to iteratively update the parameter weights in the initial network until a preset condition is reached, and a dynamic reuse attention network is obtained. The preset condition can be that the loss value is less than a preset threshold, or the number of iterations reaches the preset upper limit of the number of iterations. By introducing a multi-task joint optimization strategy and combining the absolute value loss function and the pruning loss function, the initial network is iteratively optimized, which can improve the image reconstruction quality of the dynamic reuse attention network while adaptively selecting a suitable attention matrix, thereby significantly reducing the consumption of training time and computing resources, and enabling the dynamic reuse attention network to be efficiently trained and deployed in resource-constrained environments.

[0043] S120. Extract features of the first image through the shallow feature extraction module to obtain the shallow features of the first image.

[0044] Among them, the shallow features of the first image include the basic texture and edge information of the first image.

[0045] In an embodiment of the present invention, after the first image is input into the pre-constructed dynamic reuse attention network, the shallow features of the first image can be first extracted through the shallow feature extraction module in the network. Specifically, the shallow features of the first image can be extracted through the 3*3 convolutional layer in the shallow feature extraction module as the input of the deep feature extraction module. By extracting the features of the first image through the shallow feature extraction module, the basic texture and edge information of the first image can be effectively captured, laying a foundation for subsequent deep feature extraction. Optionally, extracting the features of the first image through the shallow feature extraction module can be expressed as:

[0046] f0 = F shallow (I LR );

[0047] Among them, f0 is the shallow feature of the first image, and F shallow represents the 3*3 convolutional layer, and I LR is the first image.

[0048] S130. Through the deep feature extraction module, based on the dynamic reuse attention mechanism, extract features from the shallow features to obtain the deep features of the first image.

[0049] Among them, the dynamic reuse attention mechanism can selectively reuse the attention matrix between adjacent layers by calculating the similarity of the attention matrices of different layers, thereby reducing the computational amount and improving the parallelization ability, enabling the dynamic reuse attention network to train and infer more efficiently when processing long sequences. The deep features of the first image include the texture details and structural information of the first image.

[0050] In the embodiment of the present invention, after the shallow feature extraction module extracts the features of the first image to obtain the shallow features of the first image, the deep feature extraction module in the dynamic reuse attention network can extract the features of the shallow features based on the dynamic reuse attention mechanism to obtain the deep features of the first image. Specifically, the attention calculation unit, attention gating unit, and at least two attention reuse units in the deep feature extraction module can cooperate with each other to jointly realize the feature extraction of the shallow features based on the dynamic reuse attention mechanism to obtain the deep features of the first image. By extracting the features of the shallow features based on the dynamic reuse attention mechanism, redundant calculations can be reduced while maintaining the image reconstruction quality. Compared with the traditional shared attention mechanism, this method can more flexibly utilize the differences in the attention matrices between different layers and avoid feature aggregation errors caused by simple sharing.

[0051] S140. The shallow features and deep features are fused through an image reconstruction module to generate a first fused feature, and the first fused feature is used for image reconstruction to obtain a second image; wherein, the resolution of the second image is higher than that of the first image.

[0052] In the embodiment of the present invention, after the deep feature extraction module extracts the features of the shallow features based on the dynamic reuse attention mechanism to obtain the deep features of the first image, the image reconstruction module in the dynamic reuse attention network can first fuse the shallow features and deep features to generate a first fused feature, and then perform image reconstruction on the first fused feature to obtain a second image. Specifically, the shallow features and deep features can be added through the image reconstruction module to obtain a first fused feature. Then, the first fused feature can be input into the 3×3 convolutional layer in the image reconstruction module to map the first fused feature to 3×S 2 channels, and then the first fused feature is subjected to a channel rearrangement operation through the sub-pixel convolutional layer in the image reconstruction module to generate a second image with a resolution higher than that of the first image, realizing the super-resolution reconstruction of the first image. By generating a high-resolution image output through the sub-pixel convolutional layer, the high-frequency information of the image can be effectively restored while maintaining the image details.

[0053] Optionally, the shallow features and the deep features are fused by an image reconstruction module to generate first fused features, and the first fused features are used for image reconstruction to obtain a second image, which can be expressed as:

[0054] I SR = F UM (f0 + f n );

[0055] where, I SR is the second image, f0 is the shallow feature of the first image, and f n is the deep feature of the first image, and F UM represents a 3×3 convolutional layer and a sub-pixel convolutional layer.

[0056] In the technical solution of the embodiment of the present invention, a first image is obtained and input into a pre-constructed dynamic multiplexing attention network; wherein, the dynamic multiplexing attention network includes a shallow feature extraction module, a deep feature extraction module, and an image reconstruction module; the shallow feature extraction module is used to extract features from the first image to obtain the shallow features of the first image; the deep feature extraction module is used to extract features from the shallow features based on the dynamic multiplexing attention mechanism to obtain the deep features of the first image; the image reconstruction module is used to fuse the shallow features and the deep features to generate first fused features, and perform image reconstruction on the first fused features to obtain a second image; wherein, the resolution of the second image is higher than that of the first image. In the technical solution of the embodiment of the present invention, by extracting features from the shallow features based on the dynamic multiplexing attention mechanism and selectively sharing the attention matrices with higher similarity between different layers, redundant calculations can be reduced to lower the computational complexity while ensuring the image reconstruction quality.

[0057] Embodiment 2

[0058] Figure 3 FIG. is a flowchart of an image reconstruction method provided by Embodiment 2 of the present invention. The embodiment of the present invention is optimized based on the above embodiment. For the solutions not described in detail in the embodiment of the present invention, refer to the above embodiment. As Figure 3 shown, the method includes:

[0059] S210. Obtain a first image and input the first image into a pre-constructed dynamic multiplexing attention network; wherein, the dynamic multiplexing attention network includes a shallow feature extraction module, a deep feature extraction module, and an image reconstruction module.

[0060] S220. Extract features from the first image by the shallow feature extraction module to obtain the shallow features of the first image.

[0061] S230. Calculate the shallow attention matrix corresponding to the shallow features through the attention calculation unit, and extract the first feature of the shallow features.

[0062] Among them, the attention calculation unit includes two normalization layers, a multi-head attention layer, and a feed-forward network. Exemplarily, Figure 4 shows a schematic structural diagram of an attention calculation unit.

[0063] In the embodiment of the present invention, the attention calculation unit is positioned as the first layer of the deep feature extraction module. After the shallow features are input into the deep feature extraction network, the shallow attention matrix corresponding to the shallow features can be calculated through the attention calculation unit, and the first feature of the shallow features can be extracted. Specifically, the attention calculation unit can perform a complete attention calculation on the shallow features to generate the shallow attention matrix, and the generated shallow attention matrix is then used by the attention gating unit to determine the first reusable attention matrix corresponding to the first feature. At the same time, the first feature can also be obtained by extracting features from the shallow features through the attention calculation unit.

[0064] Optionally, the calculating the shallow attention matrix corresponding to the shallow features through the attention calculation unit includes: performing layer normalization on the shallow features through the attention calculation unit to obtain the query vector and key vector of the shallow features; calculating the dot product of the query vector and the key vector, and converting the dot product into the shallow attention matrix of the shallow features through a preset function.

[0065] In the embodiment of the present invention, the shallow attention matrix corresponding to the shallow features can be calculated through the attention calculation unit. Specifically, when the shallow features pass through the normalization layer of the attention calculation unit, layer normalization is performed on the shallow features to obtain the query vector Q, key vector K, and value vector V of the shallow features. When calculating the attention matrix, only the query vector Q and the key vector K are needed. Specifically, the dot product of the query vector Q and the key vector K can be calculated first, and then the dot product of the query vector Q and the key vector K is converted into the shallow attention matrix of the shallow features through a preset function. Optionally, the preset function can be the Softmax function, and the embodiment of the present invention does not limit this.

[0066] S240. Determine the first reusable attention matrix corresponding to the first feature from the shallow attention matrix through the attention gating unit according to a preset strategy.

[0067] Among them, the attention gating unit is used to evaluate the similarity of the attention matrices of adjacent layers, and dynamically select and reuse the attention matrix based on the evaluated similarity result. In the training phase, the attention gating unit adopts the Gumbel-Softmax reparameterization technique to enable the network to learn the optimal reuse strategy; in the usage phase, the attention gating unit adjusts the reuse decision adaptively according to the input features through a dynamic threshold mechanism. Exemplarily, Figure 5 shows a schematic structural diagram of an attention gating unit.

[0068] In the embodiment of the present invention, since at this time in the attention gating unit, there is only the attention matrix corresponding to the first layer, that is, the shallow attention matrix generated after the attention calculation unit performs a complete attention calculation on the shallow features, therefore, it is necessary to determine the first reusable attention matrix corresponding to the first feature from the shallow attention matrix according to a preset strategy. Among them, the preset strategy can be set by those skilled in the art according to the actual situation, and the embodiment of the present invention does not limit this.

[0069] S250. Extract the second feature of the first feature according to the first reusable attention matrix through the first attention reuse unit, and determine the second reusable attention matrix corresponding to the second feature through the attention gating unit.

[0070] Among them, the attention reuse unit includes a convolutional branch and a Transformer branch. The convolutional branch can enhance the representation ability of local features, the Transformer branch can capture global dependencies, and the information interaction between the two branches can be realized through a channel attention mechanism and a spatial attention mechanism, so as to improve the detailed performance and overall performance of the features. Exemplarily, Figure 6 shows a schematic structural diagram of an attention reuse unit.

[0071] In the embodiment of the present invention, the second feature of the first feature can be extracted according to the first reusable attention matrix through the first attention reuse unit, and the second reusable attention matrix corresponding to the second feature can be determined through the attention gating unit.

[0072] Optionally, the first attention multiplexing unit extracts the second feature of the first feature according to the first reusable attention matrix, and the attention gating unit determines the second reusable attention matrix corresponding to the second feature, including: the first attention multiplexing unit extracts the second feature of the first feature according to the first reusable attention matrix, and calculates the first attention matrix corresponding to the first feature, and splices the first reusable attention matrix and the first attention matrix into a first spliced attention matrix; the attention gating unit determines the second reusable attention matrix corresponding to the second feature from the first reusable attention matrix and the first spliced attention matrix.

[0073] In an embodiment of the present invention, the first attention multiplexing unit can extract the second feature of the first feature according to the first reusable attention matrix, and during the feature extraction process, calculate the first attention matrix corresponding to the first feature, and its calculation method is the same as that of the shallow attention matrix in step S230. After obtaining the first attention matrix, splice it with the first reusable attention matrix into a first spliced attention matrix. Then, the attention gating unit can evaluate the similarity of the first reusable attention matrix and the first spliced attention matrix, and determine the second reusable attention matrix corresponding to the second feature according to the similarity evaluation result.

[0074] Optionally, the first attention multiplexing unit extracts the second feature of the first feature according to the first reusable attention matrix, including: performing depthwise separable convolution processing on the first feature to obtain the convolution feature of the first feature, and performing layer normalization on the first feature to obtain the query vector, key vector, and value vector of the first feature; performing weighted processing on the value vector of the first feature according to the convolution feature through a channel attention mechanism, and calculating the spatial attention feature of the first feature according to the query vector, key vector, weighted value vector of the first feature, and the first reusable attention matrix; fusing the convolution feature and the spatial attention feature through a spatial attention mechanism to generate a second fused feature, and updating the features of the second fused feature, the first feature, and the spatial attention feature through layer normalization to obtain the second feature of the first feature.

[0075] Among them, the channel attention mechanism includes a pooling layer, two convolutional layers, a normalization layer, and two activation layers. Exemplarily, Figure 7 shows a schematic structural diagram of a channel attention mechanism. The spatial attention mechanism includes a normalization layer, two convolutional layers, and two activation layers. Exemplarily, Figure 8 shows a schematic structural diagram of a spatial attention mechanism.

[0076] In the embodiments of the present invention, the depthwise separable convolution processing can be performed on the first feature through the convolutional branch of the attention multiplexing unit to obtain the convolutional feature of the first feature, and the layer normalization can be performed on the first feature through the Transformer branch of the attention multiplexing unit to obtain the query vector, key vector, and value vector of the first feature. The process can be expressed as follows:

[0077] f conv.1 = DWconv(f1);

[0078] [Q1, K1, V1] = LN(f1);

[0079] After that, in order to enhance the channel dimension information of the feature, the value vector of the first feature can be weighted by the channel attention mechanism according to the convolutional feature, and the spatial attention feature of the first feature can be calculated according to the query vector, key vector, weighted value vector of the first feature, and the first reusable attention matrix. The process can be expressed as follows:

[0080] V1' = V1 ⊙ CA(f conv.1 );

[0081] f sa.1 = AMRM(V1', K1, Q1, rattn1);

[0082] Finally, in order to further fuse the information of the convolutional branch and the Transformer branch, the convolutional feature and the spatial attention feature can be fused through the spatial attention mechanism to generate the second fused feature, and the second fused feature, the first feature, and the spatial attention feature can be updated through layer normalization to obtain the second feature of the first feature. The process can be expressed as follows:

[0083] f oc.1 = f conv.1 ⊙ SA(f sa.1 );

[0084] f2 = LN(FFN(f1 + f oc.1 + f sa.1 ) + f1 + f oc.1 + f sa.1 );

[0085] Among them, f1 is the first feature, f conv.1 is the convolutional feature of the first feature, DWconv represents depthwise separable convolution, Q1 is the query vector of the first feature, K1 is the key vector of the first feature, V1 is the value vector of the first feature, LN represents layer normalization, V1' is the weighted value vector, ⊙ represents element-wise multiplication, CA represents the channel attention mechanism, f sa.1The spatial attention feature of the first feature, AMRM represents the multiplexed multi-head attention mechanism, rattn1 is the first reusable attention matrix, SA represents the spatial attention mechanism, f2 is the second feature, and FFN represents the feed-forward network.

[0086] It can be understood that this design not only reduces redundant calculations through partial reuse of the attention mechanism, but also makes up for the deficiency of the Transformer in local feature modeling by introducing a convolutional branch. The introduction of the channel attention and spatial attention mechanisms further enhances the ability to represent feature details, enabling the network to achieve efficient feature extraction and high-quality image reconstruction with low computational overhead.

[0087] Optionally, determining the second reusable attention matrix corresponding to the second feature from the first reusable attention matrix and the first spliced attention matrix through the attention gating unit includes: calculating the total variation distance between the first reusable attention matrix and the first spliced attention matrix through the attention gating unit; generating a binary mask map according to the total variation distance, and determining the second reusable attention matrix corresponding to the second feature from the first reusable attention matrix and the first spliced attention matrix according to the binary mask map.

[0088] Among them, the total variation distance is a measurement method for measuring the difference between adjacent pixels (or elements) in an image or matrix. In the embodiments of the present invention, the similarity between the first reusable attention matrix and the first spliced attention matrix is evaluated by calculating the total variation distance therebetween. The binary mask map is used to identify the regions where the first reusable attention matrix has significant changes or dissimilarities compared to the first spliced attention matrix.

[0089] In the embodiments of the present invention, the total variation distance between the first reusable attention matrix and the first spliced attention matrix can be calculated first through the attention gating unit, then a binary mask map is generated according to the total variation distance, and the attention matrix with a higher similarity in the first reusable attention matrix and the first spliced attention matrix is indicated by the binary mask map and determined as the second reusable attention matrix corresponding to the second feature.

[0090] S260. The first attention multiplexing unit sends the second feature to the next attention multiplexing unit, and the attention gating unit sends the second reusable attention matrix to the next attention multiplexing unit until the target feature output by the last attention multiplexing unit is obtained, and the target feature is used as the deep feature of the first image.

[0091] In an embodiment of the present invention, the first attention reuse unit sends the second feature to the next attention reuse unit, and the attention gating unit sends the second reusable attention matrix to the next attention reuse unit for the next attention reuse unit to repeat the operations in step S250 until the target feature output by the last attention reuse unit is obtained, and the target feature is used as the deep feature of the first image.

[0092] S270. The shallow feature and the deep feature are fused by an image reconstruction module to generate a first fused feature, and the first fused feature is reconstructed to obtain a second image.

[0093] The technical solution of the embodiment of the present invention obtains a first image and inputs the first image into a pre-constructed dynamic reuse attention network. The dynamic reuse attention network includes a shallow feature extraction module, a deep feature extraction module, and an image reconstruction module. The shallow feature extraction module extracts the shallow feature of the first image, the attention calculation unit calculates the shallow attention matrix corresponding to the shallow feature and extracts the first feature of the shallow feature, the attention gating unit determines the first reusable attention matrix corresponding to the first feature from the shallow attention matrix according to a preset strategy, the first attention reuse unit extracts the second feature of the first feature according to the first reusable attention matrix, and the attention gating unit determines the second reusable attention matrix corresponding to the second feature. The first attention reuse unit sends the second feature to the next attention reuse unit, and the attention gating unit sends the second reusable attention matrix to the next attention reuse unit until the target feature output by the last attention reuse unit is obtained, and the target feature is used as the deep feature of the first image. The shallow feature and the deep feature are fused by the image reconstruction module to generate a first fused feature, and the first fused feature is reconstructed to obtain a second image. The technical solution of the embodiment of the present invention extracts features from the shallow feature based on the dynamic reuse attention mechanism, selectively shares the attention matrix with higher similarity between different layers, can reduce redundant calculations to reduce the computational complexity while ensuring the image reconstruction quality.

[0094] Embodiment III

[0095] Figure 9 It is a schematic structural diagram of an image reconstruction device provided in Embodiment III of the present invention. As Figure 9 shown, the device includes:

[0096] A first image acquisition module 310, configured to acquire a first image and input the first image into a pre-constructed dynamic reuse attention network. The dynamic reuse attention network includes a shallow feature extraction module, a deep feature extraction module, and an image reconstruction module.

[0097] The shallow feature determination module 320 is configured to extract features from the first image through the shallow feature extraction module to obtain the shallow features of the first image;

[0098] The deep feature determination module 330 is configured to extract features from the shallow features through the deep feature extraction module based on the dynamic reuse attention mechanism to obtain the deep features of the first image;

[0099] The second image determination module 340 is configured to fuse the shallow features and the deep features through the image reconstruction module to generate a first fused feature, and perform image reconstruction on the first fused feature to obtain a second image; wherein, the resolution of the second image is higher than that of the first image.

[0100] Optionally, the deep feature extraction module includes an attention calculation unit, an attention gating unit, and at least two attention reuse units;

[0101] The deep feature determination module 330 includes:

[0102] The first feature determination unit is configured to calculate a shallow attention matrix corresponding to the shallow features through the attention calculation unit and extract the first features of the shallow features;

[0103] The first reusable attention matrix determination unit is configured to determine a first reusable attention matrix corresponding to the first features from the shallow attention matrix through the attention gating unit according to a preset strategy;

[0104] The second feature determination unit is configured to extract the second features of the first features through the first attention reuse unit according to the first reusable attention matrix, and determine a second reusable attention matrix corresponding to the second features through the attention gating unit;

[0105] The deep feature determination unit is configured to send the second features by the first attention reuse unit to the next attention reuse unit, and send the second reusable attention matrix by the attention gating unit to the next attention reuse unit until the target features output by the last attention reuse unit are obtained, and use the target features as the deep features of the first image.

[0106] Optionally, the first feature determination unit is specifically configured to:

[0107] Perform layer normalization on the shallow features through the attention calculation unit to obtain a query vector and a key vector of the shallow features;

[0108] Calculate the dot product of the query vector and the key vector, and convert the dot product into a shallow attention matrix of the shallow features through a preset function.

[0109] Optionally, the second feature determination unit includes:

[0110] A second feature determination subunit, configured to extract a second feature of the first feature according to the first reusable attention matrix by a first attention multiplexing unit, calculate a first attention matrix corresponding to the first feature, and splice the first reusable attention matrix and the first attention matrix into a first spliced attention matrix;

[0111] A second reusable attention matrix determination subunit, configured to determine a second reusable attention matrix corresponding to the second feature from the first reusable attention matrix and the first spliced attention matrix through the attention gating unit.

[0112] Optionally, the second feature determination subunit is specifically configured to:

[0113] Perform depthwise separable convolution processing on the first feature to obtain a convolutional feature of the first feature, and perform layer normalization on the first feature to obtain a query vector, a key vector, and a value vector of the first feature;

[0114] Perform weighted processing on the value vector of the first feature according to the convolutional feature through a channel attention mechanism, and calculate a spatial attention feature of the first feature according to the query vector, the key vector, the weighted value vector of the first feature, and the first reusable attention matrix;

[0115] Fuse the convolutional feature and the spatial attention feature through a spatial attention mechanism to generate a second fused feature, and update the features of the second fused feature, the first feature, and the spatial attention feature through layer normalization to obtain a second feature of the first feature.

[0116] Optionally, the second reusable attention matrix determination subunit is specifically configured to:

[0117] Calculate the total variation distance between the first reusable attention matrix and the first spliced attention matrix through the attention gating unit;

[0118] Generate a binary mask graph according to the total variation distance, and determine a second reusable attention matrix corresponding to the second feature from the first reusable attention matrix and the first spliced attention matrix according to the binary mask graph.

[0119] Optionally, the construction process of the dynamic multiplexing attention network includes:

[0120] Input the sample image into the initial network to obtain an output result, and calculate the loss value corresponding to the sample image and the output result according to the loss function; wherein, the loss function includes an absolute value loss function and a pruning loss function;

[0121] Perform backpropagation through the loss value to iteratively update the parameter weights in the initial network to obtain a dynamic multiplexing attention network.

[0122] The image reconstruction device provided by the embodiments of the present invention can execute the image reconstruction method provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method.

[0123] Embodiment Four

[0124] Figure 10 FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0125] As Figure 10 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., wherein the memory stores a computer program executable by the at least one processor, and the processor 11 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.

[0126] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0127] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the image reconstruction method.

[0128] In some embodiments, the image reconstruction method can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the image reconstruction method described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the image reconstruction method by any other suitable means (e.g., by means of firmware).

[0129] The various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0130] A computer program for implementing the method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, a special purpose computer, or other programmable data processing device, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer program can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on a remote machine or server.

[0131] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0132] In order to provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0133] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0134] A computing system can include a client and a server. The client and the server are generally remote from each other and typically interact via a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0135] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.

[0136] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. An image reconstruction method, characterized in that, The method includes: Obtain a first image and input the first image into a pre-constructed dynamic multiplexing attention network; wherein, the dynamic multiplexing attention network includes a shallow feature extraction module, a deep feature extraction module, and an image reconstruction module; Extract features of the first image through the shallow feature extraction module to obtain shallow features of the first image; Based on a dynamic multiplexing attention mechanism, extract features of the shallow features through the deep feature extraction module to obtain deep features of the first image; Fuse the shallow features and the deep features through the image reconstruction module to generate a first fused feature, and perform image reconstruction on the first fused feature to obtain a second image; wherein, the resolution of the second image is higher than that of the first image.

2. The method according to claim 1, wherein The deep feature extraction module includes an attention calculation unit, an attention gating unit, and at least two attention multiplexing units; The step of, based on a dynamic multiplexing attention mechanism, extracting features of the shallow features through the deep feature extraction module to obtain deep features of the first image includes: Calculate a shallow attention matrix corresponding to the shallow features through the attention calculation unit, and extract a first feature of the shallow features; Determine a first reusable attention matrix corresponding to the first feature from the shallow attention matrix through the attention gating unit according to a preset strategy; Extract a second feature of the first feature according to the first reusable attention matrix through a first attention multiplexing unit, and determine a second reusable attention matrix corresponding to the second feature through the attention gating unit; The first attention multiplexing unit sends the second feature to the next attention multiplexing unit, and the attention gating unit sends the second reusable attention matrix to the next attention multiplexing unit until a target feature output by the last attention multiplexing unit is obtained, and the target feature is used as the deep feature of the first image.

3. The method according to claim 2, characterized in that, The step of calculating a shallow attention matrix corresponding to the shallow features through the attention calculation unit includes: Perform layer normalization on the shallow features through the attention calculation unit to obtain a query vector and a key vector of the shallow features; Calculate the dot product of the query vector and the key vector, and convert the dot product into a shallow attention matrix of the shallow features through a preset function.

4. The method according to claim 2, wherein The step of extracting a second feature of the first feature according to the first reusable attention matrix through a first attention multiplexing unit and determining a second reusable attention matrix corresponding to the second feature through the attention gating unit includes: Extract a second feature of the first feature according to the first reusable attention matrix through a first attention multiplexing unit, and calculate a first attention matrix corresponding to the first feature, and splice the first reusable attention matrix and the first attention matrix into a first spliced attention matrix; Determine the second reusable attention matrix corresponding to the second feature from the first reusable attention matrix and the first concatenated attention matrix through the attention gating unit.

5. The method according to claim 4, characterized in that The extracting, by the first attention reuse unit, of the second feature of the first feature according to the first reusable attention matrix includes: Performing depthwise separable convolution processing on the first feature to obtain the convolutional feature of the first feature, and performing layer normalization on the first feature to obtain the query vector, key vector, and value vector of the first feature; Performing weighted processing on the value vector of the first feature according to the convolutional feature through a channel attention mechanism, and calculating the spatial attention feature of the first feature according to the query vector, key vector, weighted value vector of the first feature, and the first reusable attention matrix; Fusing the convolutional feature and the spatial attention feature through a spatial attention mechanism to generate a second fused feature, and updating the features of the second fused feature, the first feature, and the spatial attention feature through layer normalization to obtain the second feature of the first feature.

6. The method according to claim 4, wherein The determining, by the attention gating unit, of the second reusable attention matrix corresponding to the second feature from the first reusable attention matrix and the first concatenated attention matrix includes: Calculating the total variation distance between the first reusable attention matrix and the first concatenated attention matrix through the attention gating unit; Generating a binary mask graph according to the total variation distance, and determining the second reusable attention matrix corresponding to the second feature from the first reusable attention matrix and the first concatenated attention matrix according to the binary mask graph.

7. The method according to claim 1, wherein The construction process of the dynamic reuse attention network includes: Inputting a sample image into an initial network to obtain an output result, and calculating the loss value corresponding to the sample image and the output result according to a loss function; wherein, the loss function includes an absolute value loss function and a pruning loss function; Performing backpropagation through the loss value to iteratively update the parameter weights in the initial network to obtain a dynamic reuse attention network.

8. An image reconstruction device, characterized in that, The device includes: A first image acquisition module, configured to acquire a first image and input the first image into a pre-constructed dynamic reuse attention network; wherein, the dynamic reuse attention network includes a shallow feature extraction module, a deep feature extraction module, and an image reconstruction module; A shallow feature determination module, configured to extract the shallow feature of the first image through the shallow feature extraction module; A deep feature determination module, configured to extract the deep feature of the first image through the deep feature extraction module based on a dynamic reuse attention mechanism; A second image determination module, configured to fuse the shallow feature and the deep feature through the image reconstruction module to generate a first fused feature, and perform image reconstruction on the first fused feature to obtain a second image; wherein, the resolution of the second image is higher than that of the first image.

9. An electronic device, characterized in that, The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to execute the image reconstruction method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for implementing the image reconstruction method according to any one of claims 1-7 when executed by a processor.