Image reconstruction method, system and device, storage medium and program product

Through the image reconstruction method of cascade feature extraction and modulation units, the problems of blurred image and unclear texture in the imaging system are solved, and high-quality image reconstruction and resolution improvement are achieved, which is suitable for devices with resource limitations.

CN120543379APending Publication Date: 2025-08-26BEIJING INFORMATION SCI & TECH UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510648731.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-04-28
Filing Date
2025-05-20
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

Existing imaging systems have problems of blurred image and unclear texture details, especially in areas such as medical and security that require high image quality, and image reconstruction quality based on super-resolution methods is poor.

Method used

The cascading shallow feature extraction module, deep feature extraction module and reconstruction module are adopted to obtain the first feature through the shallow feature extraction module. The deep feature extraction module performs multi-scale division and adaptive aggregation, and combines the adaptive feature modulation unit and the convolutional modulation unit to generate high-quality reconstruction images.

Benefits of technology

The image reconstruction quality is improved, the image resolution and feature display range are enhanced, the local features and multi-scale information of the image are captured, and the calculation amount is reduced, making the image reconstruction model suitable for resource-constrained devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120543379A_ABST
    Figure CN120543379A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an image reconstruction method, system and device, a storage medium and a program product, and relates to the field of computer vision, and the image reconstruction method can comprise the steps: inputting a to-be-processed image into a shallow feature extraction module, and obtaining a first feature which represents the content information and space structure information of the to-be-processed image; the first feature is input into a deep feature extraction module to obtain a second feature, the adaptive feature modulation unit is used for performing multi-scale division on the first feature and inputting the first feature after adaptive aggregation into the convolution modulation unit, and the second feature represents an association relationship between each pixel in the first feature and pixels in an adjacent region; and inputting the first feature and the second feature into a reconstruction module to obtain a reconstructed image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer vision, and more specifically to an image reconstruction method, system, device, storage medium, and program product. Background Art

[0002] Currently, imaging systems often suffer from problems such as blurred images and unclear texture details, which restrict the application of various imaging systems in various fields, especially application fields such as medical and security that have high requirements for image quality.

[0003] Super-resolution methods can be used to reconstruct images output by the original imaging system in order to improve image quality, expand image resolution, or enhance the range of image feature display. However, the image quality of the reconstructed image obtained by super-resolution methods is poor. Summary of the Invention

[0004] In view of this, embodiments of the present disclosure provide an image reconstruction method, system, device, storage medium, and program product.

[0005] One aspect of an embodiment of the present disclosure provides an image reconstruction method, wherein the image reconstruction model may include a cascaded shallow feature extraction module, a deep feature extraction module, and a reconstruction module. The method may include: inputting the image to be processed into the shallow feature extraction module to obtain a first feature; wherein the first feature represents the content information and spatial structure information of the image to be processed. Inputting the first feature into the deep feature extraction module to obtain a second feature, wherein the adaptive feature modulation unit is used to perform multi-scale division on the first feature, and inputting the second feature into the convolution modulation unit after adaptive aggregation, and the second feature represents the correlation relationship between each pixel in the first feature and the pixels in the adjacent area. Inputting the first feature and the second feature into the reconstruction module to obtain a reconstructed image.

[0006] Another aspect of the disclosed embodiments provides an image reconstruction system, which may include: a shallow feature extraction module for inputting a to-be-processed image into a convolutional layer to extract a first feature; a deep feature extraction module for performing multi-scale segmentation on the first feature, and then adaptively aggregating the first feature before inputting it into a convolutional modulation unit to obtain a second feature; and a reconstruction module for performing image reconstruction based on the first and second features to obtain a reconstructed image.

[0007] Another aspect of the present disclosure provides an electronic device, including:

[0008] one or more processors;

[0009] a memory for storing one or more programs,

[0010] When one or more programs are executed by one or more processors, the one or more processors implement the above method.

[0011] Another aspect of the embodiments of the present disclosure provides a computer-readable storage medium storing computer-executable instructions, which are used to implement the above method when executed.

[0012] Another aspect of an embodiment of the present disclosure provides a computer program product, which includes computer-executable instructions. When the instructions are executed, they are used to implement the above method.

[0013] According to an embodiment of the present disclosure, since the second feature can be obtained by processing the processed first feature by the convolution modulation unit, the second feature can be used to characterize the correlation relationship between the pixels in the first feature and the pixels in the adjacent area of ​​the pixel, thereby achieving adaptive attention to different areas of the image to be processed and being able to capture the local features of the image to be processed. On this basis, since the reconstructed image can be obtained by the reconstruction module performing image reconstruction based on the first feature and the second feature, the image reconstruction quality is improved. In addition, the processed first feature can be obtained by the adaptive feature modulation unit performing multi-scale division and adaptive aggregation on the image to be processed, therefore, the processed first feature can reflect the multi-scale information of the image to be processed. On this basis, since the second feature can be obtained based on the processed first feature, the accuracy of the correlation relationship between the constructed pixel and the pixels in the adjacent area is improved, thereby further improving the image reconstruction quality. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The above and other objects, features and advantages of the present disclosure will become more apparent through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, in which:

[0015] Figure 1 Schematically illustrates an exemplary system architecture to which the image reconstruction method and system according to an embodiment of the present disclosure may be applied;

[0016] Figure 2 The following schematically shows a flow chart of an image reconstruction method according to an embodiment of the present disclosure;

[0017] Figure 3 The following schematically shows the principle of the image reconstruction method according to an embodiment of the present disclosure;

[0018] Figure 4 Schematically shows a principle diagram of an image reconstruction method according to another embodiment of the present disclosure;

[0019] Figure 5 Schematically shows a principle diagram of an adaptive feature modulation unit according to a disclosed embodiment;

[0020] Figure 6 The following schematically shows a principle diagram of a convolution modulation unit according to an embodiment of the present disclosure;

[0021] FIG7 schematically illustrates the effect of image reconstruction using an embodiment of the present disclosure;

[0022] Figure 8 A block diagram schematically illustrates an image reconstruction system according to an embodiment of the present disclosure; and

[0023] Figure 9 A block diagram of an electronic device suitable for implementing an image reconstruction method according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION

[0024] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the detailed description below, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.

[0025] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise," "include," etc. used herein indicate the presence of the features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0026] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0027] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).

[0028] In the embodiments of this disclosure, the collection, updating, analysis, processing, use, transmission, provision, disclosure, and storage of all data involved (including, but not limited to, user personal information) comply with relevant laws and regulations, are used for legitimate purposes, and do not violate public order and good morals. In particular, necessary measures have been taken to prevent unauthorized access to user personal information data and to safeguard the security of user personal information, network security, and national security.

[0029] In the embodiments of the present disclosure, the user's authorization or consent is obtained before obtaining or collecting the user's personal information.

[0030] Super-resolution methods can include those based on deep learning. These methods are computationally intensive. Therefore, lightweight modules such as small-scale convolutions can be used to reduce the complexity and computational complexity of the image reconstruction model, making it suitable for resource-constrained electronic devices, such as edge devices or end devices.

[0031] In the process of implementing the inventive concept of the present disclosure, it was found that the lightweight module had difficulty in obtaining relatively accurate local features, resulting in poor image reconstruction quality.

[0032] To this end, an embodiment of the present disclosure provides an image reconstruction method. For example, the image to be processed can be input into a shallow feature extraction module to obtain a first feature. The first feature is input into a deep feature extraction module to obtain a second feature. The adaptive feature modulation unit can be used to perform multi-scale division on the first feature, and after adaptive aggregation, input into a convolution modulation unit. The second feature can characterize the correlation between each pixel in the first feature and the pixels in the adjacent area. On this basis, the first feature and the second feature are input into a reconstruction module to obtain a reconstructed image.

[0033] The image reconstruction method provided by the embodiment of the present disclosure is described in detail below with reference to the accompanying drawings.

[0034] Figure 1 The following schematically illustrates an exemplary system architecture 100 to which the image reconstruction method and system according to an embodiment of the present disclosure may be applied. Figure 1 The examples shown are merely examples of system architectures to which the embodiments of the present disclosure may be applied, to help those skilled in the art understand the technical content of the present disclosure, but do not mean that the embodiments of the present disclosure may not be used in other devices, systems, environments or scenarios.

[0035] like Figure 1As shown, the system architecture 100 according to this embodiment may include a first terminal device 101, a second terminal device 102, a third terminal device 103, a network 104, and a server 105. The network 104 is used as a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0036] A user may use a first terminal device 101, a second terminal device 102, or a third terminal device 103 to interact with a server 105 via a network 104 to receive or send messages, etc. Various communication client applications may be installed on the first terminal device 101, the second terminal device 102, or the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software (for example only).

[0037] The first terminal device 101 , the second terminal device 102 , and the third terminal device 103 may be various electronic devices having display screens and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, and the like.

[0038] The server 105 may be a server that provides various services, such as a background management server (for example only) that supports websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103. The background management server may analyze and process received data such as user requests, and feed back processing results (e.g., web pages, information, or data obtained or generated based on user requests) to the terminal devices.

[0039] It should be noted that the image reconstruction method provided in the embodiment of the present disclosure can generally be executed by the server 105. Accordingly, the image reconstruction system provided in the embodiment of the present disclosure can generally be set in the server 105. The image reconstruction method provided in the embodiment of the present disclosure can also be executed by a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105. Accordingly, the image reconstruction system provided in the embodiment of the present disclosure can also be set in a server or server cluster that is different from the server 105 and can communicate with the first terminal device 101, the second terminal device 102, the third terminal device 103 and / or the server 105. Alternatively, the image reconstruction method provided in the embodiment of the present disclosure can also be executed by the first terminal device 101, the second terminal device 102 or the third terminal device 103, or by other terminal devices different from the first terminal device 101, the second terminal device 102 or the third terminal device 103. Accordingly, the image reconstruction system provided in the embodiment of the present disclosure may also be arranged in the first terminal device 101, the second terminal device 102 or the third terminal device 103, or in other terminal devices different from the first terminal device 101, the second terminal device 102 or the third terminal device 103.

[0040] For example, the image to be processed may be originally stored in any one of the first terminal device 101, the second terminal device 102, or the third terminal device 103 (for example, the first terminal device 101, but not limited thereto), or stored on an external storage device and imported into the first terminal device 101. The first terminal device 101 may then locally execute the image processing method provided in the embodiment of the present disclosure, or send the image to be processed to another terminal device, server, or server cluster, and the other terminal device, server, or server cluster that receives the image to be processed may execute the image reconstruction method provided in the embodiment of the present disclosure.

[0041] It should be understood that Figure 1 The number of the first terminal device, the second terminal device, the third terminal device, the network and the server is only illustrative. According to implementation requirements, there can be any number of the first terminal device, the second terminal device, the third terminal device, the network and the server.

[0042] Figure 2 The flowchart of the image reconstruction method based on the image reconstruction model according to an embodiment of the present disclosure is schematically shown. The image reconstruction model may include a shallow feature extraction module, a deep feature extraction module, and a reconstruction module. The deep feature extraction module may include an adaptive feature modulation unit and a convolution modulation unit.

[0043] like Figure 2As shown, the method may include operations S201 to S203.

[0044] In operation S201, the image to be processed is input into a shallow feature extraction module to obtain a first feature.

[0045] In operation S202 , the first feature is input into a deep feature extraction module to obtain a second feature.

[0046] In operation S203 , the first feature and the second feature are input into a reconstruction module to obtain a reconstructed image.

[0047] The adaptive feature modulation unit can be used to perform multi-scale division on the first feature and input it into the convolution modulation unit after adaptive aggregation. The convolution modulation unit can be used to characterize the correlation between each pixel in the first feature and the pixels in the adjacent area as the second feature.

[0048] Reference below Figure 3 ,right Figure 2 The method shown is further explained. It should be noted that Figure 3 This is only an example and is not intended to limit the present disclosure.

[0049] Figure 3 The schematic diagram schematically shows the principle of the image reconstruction method according to an embodiment of the present disclosure.

[0050] like Figure 3 As shown, the image to be processed 301 can be input into the shallow feature extraction module 302. The image to be processed 301 can be a low-resolution image or any image that needs to have its resolution increased.

[0051] The shallow feature extraction module 302 can be a single-layer 3×3 convolution or other model structure that can be used to extract shallow features. The shallow feature extraction module 302 can be used to process the image to be processed to obtain the first feature 303. The first feature 303 can represent the content information and spatial structure information of the image to be processed. The content information of the image can include at least one of the following: objects, scenes, colors, textures, etc. The content information of the image can serve as a basis for recognition and classification. The spatial structure information describes the relative positions and spatial layout relationships between objects in the image. There is no limitation on the method for obtaining the first feature 303.

[0052] In some embodiments, the adaptive feature modulation unit 304-1 in the deep feature extraction module 304 may perform multi-scale segmentation of the first feature 303. The multi-scale segmentation may be performed along a channel, followed by performing different feature extraction operations on the resulting components, and then aggregating the components at the same resolution. Since the first feature 303 is segmented along the channel and then downsampled to obtain features of different resolutions, then subjected to feature extraction, and finally aggregated at a unified resolution, spanning multiple scales, the result of the adaptive aggregation is a comprehensive feature that can represent multiple scales in the image to be reconstructed.

[0053] According to an embodiment of the present disclosure, the result obtained by the adaptive feature modulation unit 304-1 is input into the convolution modulation unit 304-2. The convolution modulation unit 304-2 can achieve the effect of the self-attention mechanism through the operation of deep convolution, so that the second feature 305 output by it can associate the feature values ​​of the pixels in the input feature with the adjacent areas, that is, capture the spatial context information of the pixels in the local area. The size of the adjacent area can match the size of the convolution kernel in the convolution modulation unit. The size of the adjacent area can be adjusted by the size of the convolution kernel in the convolution modulation unit.

[0054] According to an embodiment of the present disclosure, image reconstruction is performed based on the first feature 303 and the second feature 305 to obtain a reconstructed image 306. For example, the first feature 303 and the second feature 305 can be first feature spliced, and the spliced ​​result can be reconstructed to obtain the reconstructed image 306. For example, the spliced ​​result can be subjected to a deconvolution operation, an upsampling operation, or a dilated convolution operation to obtain the reconstructed image. The upsampling operation can include bilinear interpolation or nearest neighbor interpolation. In addition, the spliced ​​result can be processed by a decoder or a generator to obtain a reconstructed image. The embodiment of the present disclosure does not limit the image reconstruction method.

[0055] Since the second feature can be obtained by the convolution modulation unit processing the processed first feature, the second feature can be used to characterize the correlation between the pixels in the first feature and the pixels in the adjacent area of ​​the pixel, thereby achieving adaptive attention to different areas of the image to be processed and being able to capture the local features of the image to be processed. On this basis, since the reconstructed image can be obtained by the reconstruction module performing image reconstruction based on the first feature and the second feature, the image reconstruction quality is improved. In addition, the processed first feature can be obtained by the adaptive feature modulation unit performing multi-scale division and adaptive aggregation on the image to be processed. Therefore, the processed first feature can reflect the multi-scale information of the image to be processed. On this basis, since the second feature can be obtained based on the processed first feature, the accuracy of the correlation between the constructed pixel and the pixels in the adjacent area is improved, thereby further improving the image reconstruction quality.

[0056] According to an embodiment of the present disclosure, the deep feature extraction module may include T cascaded deep feature extraction submodules. The tth deep feature extraction submodule may include a cascaded tth adaptive feature modulation unit and a tth convolution modulation unit. T may be an integer greater than or equal to 1.

[0057] Inputting the first feature into a deep feature extraction module to obtain the second feature may include the following operations.

[0058] When 1<t≤T, the t-1th intermediate feature is input into the tth deep feature extraction submodule to obtain the tth intermediate feature. When t=T, the tth intermediate feature is used as the second feature.

[0059] The tth adaptive feature modulation unit can be used to perform multi-scale segmentation on the t-1th intermediate feature, and after adaptive aggregation, input it into the tth convolution modulation unit. The tth convolution modulation unit can be used to represent the relationship between each pixel in the t-1th intermediate feature and pixels in adjacent areas. The first intermediate feature can be obtained by inputting the first feature into the first deep feature extraction submodule.

[0060] Figure 4 The figure schematically shows the principle of an image reconstruction method according to another embodiment of the present disclosure.

[0061] like Figure 4 As shown, the image to be processed 301 is input into the shallow feature extraction module 302 to obtain the first feature 303.

[0062] The first feature 303 is input into the deep feature extraction module 401 to obtain the second feature 305. The deep feature extraction module 401 may include T deep feature extraction submodules. Figure 4 The first deep feature extraction submodule 401 - 1 , the t-th deep feature extraction submodule 401 - t and the T-th deep feature extraction submodule 401 -T are schematically shown. t∈{1, . . . , T}.

[0063] In the case of t=1, the first feature 303 can be input into the first adaptive feature modulation unit 401-11 to obtain a processed first feature, and then the processed first feature is input into the first convolution modulation unit 401-12 to obtain a first intermediate feature 401-13.

[0064] When 1<t<T, the t-1th intermediate feature can be input into the tth adaptive feature modulation unit 401-t1 to obtain the processed t-1th intermediate feature. The processed t-1th intermediate feature is then input into the tth convolution modulation unit 401-t2 to obtain the tth intermediate feature 401-t3.

[0065] When t=T, the T-1th intermediate feature can be input into the T-th adaptive feature modulation unit 401-T1 to obtain the processed T-1th intermediate feature. The processed T-1th intermediate feature is then input into the T-th convolution modulation unit 401-T2 to obtain the second feature 305.

[0066] The first feature 303 and the second feature 305 may be firstly spliced ​​together, and then image reconstruction may be performed on the spliced ​​result.

[0067] In some embodiments, T may be equal to 12.

[0068] Since the first feature 303 is cascaded to extract T features, the first feature 303 can be gradually abstracted and fused. In the shallow layer of the network, relatively local and low-level features are extracted, such as edges or corners. As the cascade progresses, low-level features are combined and abstracted into higher-level and more semantic features. For example, in a face recognition task, the shallow feature extraction module can detect local features such as eyes or noses, while the deep feature extraction module can fuse local features to form a recognition of the entire face, thereby utilizing the global information in the image.

[0069] According to an embodiment of the present disclosure, inputting the first feature into a deep feature extraction module to obtain the second feature may include the following operations.

[0070] Perform multi-channel feature extraction on the first feature to obtain a first component and multiple second components. Input the first component into a deep convolution layer to obtain a processed first component. A fused feature is obtained based on the multiple processed second components and the processed first component. The processed second component is obtained by downsampling the second component, performing a deep convolution operation, and performing an upsampling operation. The resolution of the processed second component is the same as the resolution of the image to be processed. Input the fused feature into a convolution modulation unit to obtain a second feature.

[0071] In some embodiments, obtaining a fused feature based on the plurality of processed second components and the processed first component may include: performing a cascade operation and an aggregation operation on the processed first component and the plurality of processed second components along a channel dimension to obtain a third feature, and obtaining the fused feature based on the third feature and the first feature.

[0072] Figure 5 The schematic diagram schematically shows the principle of the adaptive feature modulation unit according to the embodiment of the present disclosure.

[0073] like Figure 5As shown, in some embodiments, the adaptive feature modulation unit performs a normalization operation on the input first feature 303 using the normalization layer 501 to obtain a normalized first feature, and the subsequent operation steps can be expressed by the following formula.

[0074]

[0075]

[0076]

[0077]

[0078] in, It can represent the first feature after normalization. They are the four components generated after the normalized first feature passes through the channel division module 502. The four components can include the first component and three different second components. The channel division operation of the channel division module 502 can be represented. It can be a depthwise convolution 506. Can be the first component Features after processing by depthwise separable convolutional layers. It can represent the upsampling operation 507. The upsampling operation 507 can be to upsample the feature to the original resolution using the nearest interpolation. . It can be expressed as downsampling the input features to size. Figure 5 Schematic showing downsampling to and size. It can be said that the three second components are upsampled to the original resolution , namely the three processed second components.

[0079] because It can be extracted through adaptive maximum pooling, so the key features can be extracted. It can represent cascade operations along the channel dimension. The convolution layer 509 may be 1×1. It may be the third feature after the processed first component and the plurality of processed second components are cascaded by the cascade module 508 along the channel dimension.

[0080] Then, the activation function module 510 is used to process the third feature , and use the normalization layer 511 to process the third feature after the activation function , and then perform Hadamard product with the normalized first feature to obtain the eighth feature, and then perform residual connection between the eighth feature and the first feature to obtain the fusion feature 512.

[0081] For example, the third feature , the first feature after normalization Satisfy the following formula:

[0082]

[0083] in, It can represent the activation function, Indicates the eighth feature.

[0084] Because feature downsampling uses an adaptive max pooling layer, representative key features can be dynamically selected. For example, when processing complex texture images, the pooling area can be dynamically adjusted based on the texture distribution to extract more detailed texture features.

[0085] Moreover, since the normalized first feature is divided into channels and adaptively pooled in each channel, the pooled features are finally unified in resolution and fused, the multi-scale long-distance feature information in the first feature can be integrated.

[0086] According to an embodiment of the present disclosure, the fused feature is input into a convolution modulation unit to obtain a second feature, which includes the following operations.

[0087] The normalized fused features are fed into the first and second fully connected layers, respectively, to obtain the fourth and fifth features. A depthwise separable convolution operation is performed on the fourth feature to obtain a feature modulation matrix. The feature modulation matrix is ​​Hadamard-multiplied with the fifth feature to obtain the sixth feature. A residual concatenation of the sixth feature with the fused features is performed to obtain the second feature.

[0088] Figure 6 The figure schematically shows the principle of a convolution modulation unit according to an embodiment of the present disclosure.

[0089] like Figure 6As shown, in some embodiments, a normalization layer 601 is used to process the fused features 512. The normalized fused features are input into the first fully connected layer 602 to obtain the fourth feature. The normalized fused features are input into the second fully connected layer 603 to obtain the fifth feature. The fourth feature is input into the depthwise separable convolution layer 604 to obtain a feature modulation matrix. The feature modulation matrix is ​​Hadamard-multiplied with the fifth feature output by the second fully connected layer 603 to obtain the sixth feature. The sixth feature is then processed by the fully connected layer 605, and the processed sixth feature is residually connected with the fused features 512 to obtain the second feature 305. The convolution kernel size of the depthwise separable convolution 604 can be k×k. k>5. For example, k can be 11.

[0090] The above operation process can be expressed by the following formula.

[0091]

[0092]

[0093]

[0094] in, It can be the normalized fusion feature. can represent the Hadamard product. The weight matrix of the first fully connected layer 602 can be represented. It can represent the weight matrix of the second fully connected layer 603. It can represent the depth convolution with a kernel size of k×k. A can represent the feature modulation matrix, that is, The output after being multiplied by the weight matrix and processed by the depthwise separable convolution 604. It can represent the fifth characteristic. It can represent the sixth characteristic. It can be used to modulate the fifth feature using a feature modulation matrix.

[0095] The convolution modulation operation in the convolution modulation unit can be used to establish an association relationship between a pixel and the pixels in the k×k adjacent area of ​​the pixel. The convolution modulation unit can capture spatial context information in the local area. Compared with a single self-attention module, the use of convolution to establish a relationship can reduce the memory overhead caused by a large number of matrix multiplications in the self-attention calculation. For example, in the case of processing high-resolution images, due to the large amount of data of high-resolution images, the memory overhead caused by self-attention calculation is large, and the convolution modulation unit can reduce memory usage. When memory is limited, the image reconstruction model can process larger-sized images, thereby reducing memory consumption more significantly. In addition, due to the modulation operation, the normalized fusion features of the input can be extracted into feature modulation values ​​to modulate the fifth feature, so that it can better adapt to the input.

[0096] According to an embodiment of the present disclosure, the first feature and the second feature are input into a reconstruction module to obtain a reconstructed image, which includes the following operations.

[0097] The first and second features are residually connected to obtain the seventh feature. The seventh feature is upsampled using a 1×1 convolutional layer and a pixel rearrangement layer to generate a reconstructed image.

[0098] In some embodiments, the second feature output by the convolutional modulation unit is residually connected to the first feature. This allows shallow detail features to be spliced ​​and fused with deep abstract features, enabling the image reconstruction model to capture both subtle local information of the image to be processed and the overall semantic structure, thereby enriching the feature representation. Taking image recognition as an example, shallow features can include detailed information such as image edges and textures, while deep features can be related to semantic information such as the category and overall shape of the object. By splicing features, the image reconstruction model can comprehensively utilize information from different levels to improve the accuracy of its understanding of image content.

[0099] The seventh feature may be upsampled using a 1×1 convolution layer and a pixel rearrangement layer to generate a reconstructed image 306 .

[0100] The pixel rearrangement layer efficiently implements upsampling. By rearranging the pixels in the low-resolution feature map according to certain rules to form a high-resolution feature map, it reduces the information loss or blurring caused by upsampling operations (such as bilinear interpolation) and more accurately restores image details. Furthermore, compared to some convolution-based upsampling methods, the pixel rearrangement layer does not require learning additional convolution kernel parameters, thus reducing the number of parameters and computational complexity in the image reconstruction model. This not only reduces the training cost and overfitting risk of the image reconstruction model, but also enables the image reconstruction model to run faster during inference, improving its real-time performance and practicality.

[0101] The image reconstruction method provided by the embodiment of the present disclosure combines a lightweight channel segmentation method, which can finely extract image features in the frequency domain, and assign corresponding convolution kernels for processing according to features of different dimensions, thereby realizing multi-scale feature extraction. After the features are normalized, they are input into the next feature extraction module to enhance multi-scale information interaction and perform further feature extraction. In order to enhance the flexibility and expressiveness of the model, the convolution modulation unit enables the model to adaptively focus on different areas in the image while processing input data, and can effectively capture local structural information. The image reconstruction method of the embodiment of the present disclosure can improve the quality of image reconstruction with a smaller number of model parameters.

[0102] According to the embodiments of the present disclosure, the image reconstruction system of the present disclosure was trained using a dataset commonly used in the field of super-resolution. It was also tested using a publicly available dataset from current mainstream applications. The peak signal-to-noise ratio (PSNR) and structural similarity index (SSIM) were used to evaluate the image reconstruction quality of the image reconstruction method provided by the embodiments of the present disclosure.

[0103] During the training phase of the image reconstruction system, the sample image is split into 64×64 blocks with a batch size of 16. T=12, the number of channels of the adaptive feature modulation unit and the convolution modulation unit is 64, and the Adam optimizer is used with parameters β1 = 0.9, β2 = 0.99, and ε = 10. -8 The model's performance was evaluated using the L1 loss function and the frequency loss function. The initial learning rate was set to 0.001 and the final learning rate was set to 1e-5. Each training session took approximately 54 hours.

[0104] Table 1 shows the evaluation indicators of different networks on four datasets with a scaling ratio of 2 times downsampling. The largest value in each column represents the best performance.

[0105] From the above table, it can be found that the system of the embodiment of the present disclosure achieved the best performance in 7 out of 8 indicators when compared with 15 super-resolution domain networks.

[0106] Table 1

[0107]

[0108] In Table 1, scale is the scaling ratio, Params is the number of parameters, FLOPS is the number of floating-point operations, and SET5, SET14, B100, and URBAN100 represent different datasets.

[0109] According to an embodiment of the present disclosure, a set of images 7a and 7b in the Google public dataset are used. Figure 7b To demonstrate the effect of image reconstruction according to the embodiment of the present disclosure, as shown in FIG7 , Figure 7a is the low-resolution image to be reconstructed, Figure 7b For Figure 7a The corresponding high-definition original image;

[0110] Will Figure 7a After the image in the training image reconstruction system is input, the output will be as follows Figure 7c As shown in the reconstructed image, it can be clearly seen from the figure that the texture details of the image reconstructed by the image reconstruction system in the embodiment of the present disclosure are significantly clearer than those of the unreconstructed image 7a, and the resolution is higher and closer to the high-definition original. Figure 7b .

[0111] Figure 8 The block diagram of the image reconstruction system according to the embodiment of the present disclosure is schematically shown.

[0112] like Figure 8 As shown, the image reconstruction system 800 may include a shallow feature extraction module 810 , a deep feature extraction module 820 and a reconstruction module 830 .

[0113] The shallow feature extraction module 810 is used to extract the first feature in the image to be processed, where the first feature represents the content information and spatial structure information of the image to be processed.

[0114] The deep feature extraction module 820 is used to perform multi-scale division on the first feature, and input the adaptively aggregated first feature into the convolution modulation unit to obtain the second feature.

[0115] The reconstruction module 830 is configured to perform image reconstruction based on the first feature and the second feature to obtain a reconstructed image.

[0116] According to an embodiment of the present disclosure, the deep feature extraction module 820 may include an adaptive feature modulation unit and a convolution modulation unit.

[0117] The adaptive feature modulation unit divides the first feature into multiple scales and inputs the first feature into the convolution modulation unit after adaptive aggregation.

[0118] The convolution modulation unit is used to input the fusion feature into the convolution modulation unit to obtain the second feature.

[0119] The adaptive feature modulation unit may include a first extraction subunit, a second extraction subunit, a third extraction subunit, a fourth extraction subunit, and a fifth extraction subunit.

[0120] A first extraction subunit is configured to perform multi-channel feature extraction on the first feature to obtain a first component and a plurality of second components;

[0121] The second extraction subunit inputs the first component into a depthwise convolution layer to obtain a processed first component.

[0122] The third extraction subunit is used to obtain a fusion feature based on multiple processed second components and the processed first components, wherein the processed second components are obtained by performing a downsampling operation, a depth convolution operation, and an upsampling operation on the second components, and the resolution of the processed second components is the same as the resolution of the image to be processed.

[0123] The fourth extraction subunit is configured to perform a cascade operation and an aggregation operation on the processed first component and the multiple processed second components along the channel dimension to obtain a third feature.

[0124] The fifth extraction subunit is used to obtain a fusion feature based on the third feature and the first feature.

[0125] The convolution modulation unit may include a sixth extraction subunit, a seventh extraction subunit, and an eighth extraction subunit.

[0126] The sixth extraction subunit is used to input the normalized fusion features into the first fully connected layer and the second fully connected layer respectively to obtain the fourth feature and the fifth feature.

[0127] The seventh extraction subunit is configured to perform a depthwise separable convolution operation on the fourth feature to obtain a feature modulation matrix.

[0128] The eighth extraction unit is used to perform a residual connection between the sixth feature and the fusion feature to obtain the second feature.

[0129] The reconstruction module may include a first reconstruction submodule and a second reconstruction submodule.

[0130] The first reconstruction submodule is used to perform a residual connection on the first feature and the second feature to obtain the seventh feature.

[0131] The second reconstruction submodule is used to upsample the seventh feature using a 1×1 convolution layer and a pixel rearrangement layer to generate a reconstructed image.

[0132] According to the embodiments of the present invention, any number of modules, sub-modules, units, and sub-units, or at least part of the functions of any number of them, can be implemented in one module. According to the embodiments of the present invention, any one or more of the modules, sub-modules, units, and sub-units can be split into multiple modules for implementation. According to the embodiments of the present invention, any one or more of the modules, sub-modules, units, and sub-units can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or can be implemented by hardware or firmware in any other reasonable way of integrating or packaging the circuit, or can be implemented in any one of the three implementation methods of software, hardware, and firmware, or in any appropriate combination of any of them. Alternatively, according to the embodiments of the present invention, one or more of the modules, sub-modules, units, and sub-units can be at least partially implemented as a computer program module, which can perform the corresponding functions when the computer program module is executed.

[0133] For example, any number of the shallow feature extraction module 810, the deep feature extraction module 820, and the reconstruction module 830 can be combined into a single module / unit / sub-unit, or any one of these modules / units / sub-units can be split into multiple modules / units / sub-units. Alternatively, at least part of the functionality of one or more of these modules / units / sub-units can be combined with at least part of the functionality of other modules / units / sub-units and implemented in a single module / unit / sub-unit. According to an embodiment of the present disclosure, at least one of the shallow feature extraction module 810, the deep feature extraction module 820, and the reconstruction module 830 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or can be implemented in hardware or firmware by any other reasonable means of integrating or packaging circuits, or can be implemented in any one of the three implementation methods of software, hardware, and firmware, or in any appropriate combination of any of them. Alternatively, at least one of the shallow feature extraction module 810 , the deep feature extraction module 820 , and the reconstruction module 830 may be at least partially implemented as a computer program module, which may perform corresponding functions when executed.

[0134] It should be noted that the image reconstruction system in the embodiment of the present disclosure corresponds to the image reconstruction method in the embodiment of the present disclosure. The description of the image reconstruction model part specifically refers to the image reconstruction method part and will not be repeated here.

[0135] Figure 9The block diagram schematically shows an electronic device suitable for implementing the image reconstruction method described above according to an embodiment of the present disclosure. Figure 8 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0136] like Figure 9 As shown, the electronic device 900 according to an embodiment of the present disclosure includes a processor 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage portion 908 into a random access memory (RAM) 903. The processor 901 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 901 may also include onboard memory for caching purposes. The processor 901 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.

[0137] Various programs and data required for the operation of the electronic device 900 are stored in the RAM 903. The processor 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. The processor 901 performs various operations of the method flow according to the embodiment of the present disclosure by executing the programs in the ROM 902 and / or the RAM 903. It should be noted that the programs may also be stored in one or more memories other than the ROM 902 and the RAM 903. The processor 901 may also perform various operations of the method flow according to the embodiment of the present disclosure by executing the programs stored in the one or more memories.

[0138] According to an embodiment of the present disclosure, electronic device 900 may further include an input / output (I / O) interface 905, which is also connected to bus 904. Electronic device 900 may also include one or more of the following components connected to I / O interface 905: an input section 906 including a keyboard, mouse, etc.; an output section 907 including devices such as a cathode ray tube (CRT), liquid crystal display (LCD), and speakers; a storage section 908 including a hard disk; and a communication section 909 including a network interface card such as a LAN card or modem. Communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to I / O interface 905 as needed. Removable media 911, such as a magnetic disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed in drive 910 as needed, so that computer programs read from the removable media can be installed into storage section 908 as needed.

[0139] According to an embodiment of the present disclosure, the method flow according to an embodiment of the present disclosure can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 909, and / or installed from the removable medium 911. When the computer program is executed by the processor 901, the above-mentioned functions defined in the system of the embodiment of the present disclosure are executed. According to an embodiment of the present disclosure, the system, equipment, device, module, unit, etc. described above can be implemented by a computer program module.

[0140] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not be incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, and when executed, implements the method according to the embodiments of the present disclosure.

[0141] According to embodiments of the present disclosure, a computer-readable storage medium may be a non-volatile computer-readable storage medium. Examples include, but are not limited to, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0142] For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include the ROM 902 and / or the RAM 903 described above and / or one or more memories other than the ROM 902 and the RAM 903 .

[0143] An embodiment of the present disclosure also includes a computer program product, which includes a computer program, which contains program code for executing the method provided by the embodiment of the present disclosure. When the computer program product is run on an electronic device, the program code is used to enable the electronic device to implement the image reconstruction method provided by the embodiment of the present disclosure.

[0144] When the computer program is executed by the processor 901, the above functions defined in the system / device of the embodiment of the present disclosure are performed. According to the embodiment of the present disclosure, the system, device, module, unit, etc. described above can be implemented by a computer program module.

[0145] In one embodiment, the computer program may be stored on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal on a network medium, downloaded and installed via the communication portion 909, and / or installed from a removable medium 911. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to wireless, wired, or any suitable combination thereof.

[0146] According to an embodiment of the present disclosure, the program code for executing the computer program provided by the embodiment of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).

[0147] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of the systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or part of the code, which contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the boxes can also occur in an order different from that marked in the accompanying drawings. For example, two boxes shown in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, as well as the combination of boxes in the block diagram or flowchart, can be implemented using a dedicated hardware-based system that performs the specified functions or operations, or can be implemented using a combination of dedicated hardware and computer instructions. It will be understood by those skilled in the art that the features described in the various embodiments of the present disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features described in the various embodiments of the present disclosure may be combined and / or coupled in various ways, and all of these combinations and / or couplings fall within the scope of the present disclosure.

[0148] The above describes the embodiments of the present disclosure. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be advantageously used in combination. Without departing from the scope of the present disclosure, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present disclosure.

Claims

1. An image reconstruction method, characterized in that: The image reconstruction model includes a shallow feature extraction module, a deep feature extraction module and a reconstruction module. The deep feature extraction module includes an adaptive feature modulation unit and a convolution modulation unit. The method includes: Inputting the image to be processed into the shallow feature extraction module to obtain a first feature, wherein the first feature represents content information and spatial structure information of the image to be processed; Inputting the first feature into the deep feature extraction module to obtain a second feature, wherein the adaptive feature modulation unit is used to perform multi-scale division on the first feature and input the second feature into the convolution modulation unit after adaptive aggregation, and the second feature represents the correlation relationship between each pixel in the first feature and pixels in an adjacent area; and The first feature and the second feature are input into a reconstruction module to obtain a reconstructed image.

2. The method according to claim 1, characterized in that The deep feature extraction module includes T cascaded deep feature extraction submodules, the t-th deep feature extraction submodule includes a cascaded t-th adaptive feature modulation unit and a t-th convolution modulation unit, where T is an integer greater than 1; The step of inputting the first feature into the deep feature extraction module to obtain the second feature includes: In the case of 1<t≤T, inputting the t-1th intermediate feature into the t-th deep feature extraction submodule to obtain the t-th intermediate feature, wherein the t-th adaptive feature modulation unit is used to perform multi-scale division on the t-1th intermediate feature, and input the adaptively aggregated intermediate feature into the t-th convolution modulation unit, the t-th convolution modulation unit is used to use the t-th intermediate feature to characterize the correlation relationship between each pixel in the t-1th intermediate feature and pixels in adjacent areas, and the first intermediate feature is obtained by inputting the first feature into the first deep feature extraction submodule; and When t=T, the t-th intermediate feature is used as the second feature.

3. The method according to claim 1, characterized in that Inputting the first feature into the deep feature extraction module to obtain the second feature includes: Performing multi-channel feature extraction on the first feature to obtain a first component and multiple second components; Inputting the first component into a depthwise convolutional layer to obtain a processed first component; Obtaining a fused feature based on a plurality of processed second components and the processed first component, wherein the processed second components are obtained by performing a downsampling operation, a depthwise convolution operation, and an upsampling operation on the second components, and a resolution of the processed second components is the same as a resolution of the image to be processed; and The fused feature is input into the convolution modulation unit to obtain the second feature.

4. The method according to claim 3, characterized in that The obtaining of a fusion feature according to the plurality of processed second components and the processed first component includes: Performing a concatenation operation and an aggregation operation on the processed first component and the plurality of processed second components along a channel dimension to obtain a third feature; and The fusion feature is obtained according to the third feature and the first feature.

5. The method according to claim 3 or 4, characterized in that Inputting the fused feature into the convolution modulation unit to obtain the second feature includes: The normalized fusion features are input into the first fully connected layer and the second fully connected layer respectively to obtain the fourth feature and the fifth feature; Performing a depthwise separable convolution operation on the fourth feature to obtain a feature modulation matrix; Performing a Hadamard product on the characteristic modulation matrix and the fifth characteristic to obtain a sixth characteristic; and Perform a residual connection on the sixth feature and the fusion feature to obtain the second feature.

6. The method according to any one of claims 1 to 4, characterized in that Inputting the first feature and the second feature into a reconstruction module to obtain a reconstructed image includes: Performing a residual connection between the first feature and the second feature to obtain a seventh feature; and An upsampling operation is performed on the seventh feature using a 1×1 convolution layer and a pixel rearrangement layer to generate the reconstructed image.

7. An image reconstruction system, characterized in that: include: A shallow feature extraction module is used to extract a first feature from the image to be processed, where the first feature represents content information and spatial structure information of the image to be processed; A deep feature extraction module, configured to perform multi-scale division on the first feature, and input the resultant features into a convolutional modulation unit after adaptive aggregation to obtain a second feature; and A reconstruction module is used to reconstruct an image based on the first feature and the second feature to obtain a reconstructed image.

8. An electronic device, characterized in that: include: one or more processors; A memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors are enabled to implement the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that Executable instructions are stored thereon, which, when executed by a processor, enable the processor to implement the method according to any one of claims 1 to 6.

10. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 6.