Image super-resolution reconstruction method, device, equipment, medium and program product
The image super-resolution reconstruction method enhances video monitoring clarity and resolution using a residual network model with context attention and wavelet fusion, addressing issues of blurry and low-resolution images in bank and hospital settings.
Patent Information
- Application Number
- CN202510492008.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-15
AI Technical Summary
The surveillance video image is blurred due to equipment limitations, monitoring angles, light and dark light, etc., and the resolution is low, so it cannot objectively reflect the real situation.
The image super-resolution reconstruction method is adopted, and the context attention module and wavelet fusion module in the residual network model are used to extract the global features and local detail texture information of the image, and the remote context dependency is captured through multiple context attention modules, and the global contour information and local detail texture information are fused to reconstruct the high-resolution image.
The resolution of the surveillance video images is improved, and clear and high-quality images are reconstructed to meet the needs of real-time video monitoring and analysis in bank outlets, hospitals, stations and other scenarios.
Smart Images

Figure CN120318072A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and in particular, to an image super-resolution reconstruction method, apparatus, device, medium, and program product. Background Art
[0002] With the development of technology, video surveillance systems have been installed in many places, especially bank branches, hospitals, stations, etc., in order to monitor the surrounding situation in real time.
[0003] However, due to the limitations of the monitoring equipment itself, monitoring angles, monitoring distances, light brightness, etc., some video images collected by the video surveillance system are blurred and have low resolution, resulting in the images being unable to objectively reflect the real situation.
[0004] Therefore, how to improve the resolution of video images is an urgent problem to be solved at present. Summary of the Invention
[0005] The present application provides an image super-resolution reconstruction method, apparatus, device, medium, and program product, which can solve the problems of blurred and unclear monitoring video images and low resolution. In particular, it meets the requirements for real-time monitoring and analysis of videos in scenarios such as bank branches, hospitals, and stations.
[0006] In a first aspect, an embodiment of the present application provides an image super-resolution reconstruction method, the method comprising:
[0007] Obtain an image to be reconstructed;
[0008] Input the image to be reconstructed into an image super-resolution model to obtain a reconstructed image output by the image super-resolution model; the resolution of the reconstructed image is higher than that of the image to be reconstructed;
[0009] Wherein, the image super-resolution model is based on a residual network model, which includes a plurality of context attention modules (CAM) and a plurality of wavelet fusion modules (WFM); the plurality of context attention modules are used to extract global features corresponding to the image to be reconstructed; each wavelet fusion module is used to fuse the global contour information and local detail texture information corresponding to a plurality of feature maps associated with the image to be reconstructed;
[0010] The reconstructed image is obtained by the image super-resolution model based on the fusion of the image to be reconstructed, global features, each global contour information, and each local detail texture information.
[0011] In a second aspect, an embodiment of the present application further provides an image super-resolution reconstruction apparatus, the apparatus comprising:
[0012] An acquisition module, configured to acquire an image to be reconstructed;
[0013] An image reconstruction module, configured to input the image to be reconstructed into an image super-resolution model to obtain a reconstructed image output by the image super-resolution model; the resolution of the reconstructed image is higher than that of the image to be reconstructed;
[0014] Wherein, the image super-resolution model is based on a residual network model, which includes a plurality of context attention modules and a plurality of wavelet fusion modules; the plurality of context attention modules are configured to extract global features corresponding to the image to be reconstructed; each wavelet fusion module is configured to fuse the global contour information and local detail texture information corresponding to a plurality of feature maps associated with the image to be reconstructed;
[0015] The reconstructed image is obtained after the image super-resolution model fuses the image to be reconstructed, the global features, the global contour information, and the local detail texture information.
[0016] In a third aspect, an embodiment of the present application provides an electronic device, including:
[0017] One or more processors;
[0018] A memory, configured to store one or more programs,
[0019] When the one or more programs are executed by the one or more processors, the one or more processors implement the image super-resolution reconstruction method according to any embodiment of the present application.
[0020] In a fourth aspect, an embodiment of the present application provides a storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the image super-resolution reconstruction method according to any embodiment of the present application.
[0021] In a fifth aspect, an embodiment of the present application further provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the image super-resolution reconstruction method according to any embodiment of the present application.
[0022] An embodiment of the present application provides an image super-resolution reconstruction method, apparatus, device, medium, and program product, including: obtaining an image to be reconstructed; inputting the image to be reconstructed into an image super-resolution model to obtain a reconstructed image output by the image super-resolution model; the resolution of the reconstructed image is higher than that of the image to be reconstructed; wherein, the image super-resolution model is based on a residual network model, which includes a plurality of context attention modules and a plurality of wavelet fusion modules; the plurality of context attention modules are used to extract global features corresponding to the image to be reconstructed; each wavelet fusion module is used to fuse the global contour information and local detail texture information corresponding to a plurality of feature maps associated with the image to be reconstructed; the reconstructed image is obtained after the image super-resolution model fuses the image to be reconstructed, global features, each global contour information, and each local detail texture information. That is to say, in the technical solution of the present application, by inputting the image to be reconstructed into the image super-resolution model, the image super-resolution model captures pixel-level long-range context dependencies on the basis of the residual network model framework, adaptively focuses on global information, and thus extracts global features corresponding to the image to be reconstructed; each wavelet fusion module is used to fuse the global contour information and local detail texture information corresponding to a plurality of feature maps associated with the image to be reconstructed; finally, the image super-resolution model fuses the image to be reconstructed, global features, each global contour information, and each local detail texture information, and can reconstruct a clear, high-quality, and high-resolution reconstructed image, thereby improving the resolution of the image to be reconstructed and solving the problem of blurred and unclear and low-resolution monitoring video images. In particular, it meets the requirements of real-time video monitoring and analysis in scenarios such as bank branches, hospitals, and stations. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 It is a schematic flowchart of an image super-resolution reconstruction method provided by an embodiment of the present application;
[0024] Figure 2 It is a schematic structural diagram of an image super-resolution model provided by an embodiment of the present application;
[0025] Figure 3 It is a schematic structural diagram of a context attention module provided by an embodiment of the present application;
[0026] Figure 4 It is a schematic structural diagram of an image super-resolution reconstruction apparatus provided by an embodiment of the present application;
[0027] Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0028] The present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present application, rather than limiting the present application. In addition, it should be noted that for the sake of description, only the parts related to the present application rather than all the structures are shown in the accompanying drawings.
[0029] To facilitate a clearer understanding of the embodiments of the present application, some related knowledge will be introduced as follows.
[0030] Image super-resolution is a pixel-level computer vision task, aiming to reconstruct a clear high-resolution image from a given low-resolution image to improve the image quality. In related technologies, the image super-resolution methods applied to the monitoring scenario mainly include interpolation-based methods, reconstruction-based methods, and deep learning-based methods.
[0031] The interpolation-based image super-resolution methods include nearest neighbor interpolation, bilinear interpolation, bicubic interpolation, etc. The algorithms are simple and the processing speed is fast. However, the processing effect at pixel mutation points such as edges and textures is poor. The reconstruction-based methods include frequency domain methods and spatial domain methods, but they cannot well simulate the real scenario. The deep learning-based methods have achieved good performance in the field of image super-resolution, but they often ignore the attention to the global information and frequency domain information of the image, resulting in a blurring effect due to irrelevant textures.
[0032] To improve the quality of monitoring images, the embodiments of the present application use image super-resolution technology to reconstruct clear high-resolution images, meeting the monitoring and analysis requirements of places such as bank branches, hospitals, and stations.
[0033] Figure 1 FIG. is a schematic flowchart of an image super-resolution reconstruction method provided by an embodiment of the present application. This method can be executed by an image super-resolution reconstruction device or an electronic device. The device or the electronic device can be implemented in a software and / or hardware manner, and the device or the electronic device can be integrated in any intelligent device with network communication functions. Such as Figure 1 , the image super-resolution reconstruction method may include the following steps:
[0034] Step 101, obtain the image to be reconstructed.
[0035] In the embodiments of the present application, the image to be reconstructed can be a low-resolution and blurred image intercepted from a video monitoring system; it can also be any image that needs to be reconstructed to improve the image resolution. The present application does not limit the applicable scenarios of the image to be reconstructed.
[0036] Step 102: Input the image to be reconstructed into the image super-resolution model to obtain the reconstructed image output by the image super-resolution model; the resolution of the reconstructed image is higher than that of the image to be reconstructed; wherein, the image super-resolution model is based on a residual network model, which includes multiple context attention modules and multiple wavelet fusion modules; the multiple context attention modules are used to extract the global features corresponding to the image to be reconstructed; each wavelet fusion module is used to fuse the global contour information and local detail texture information corresponding to multiple feature maps associated with the image to be reconstructed; the reconstructed image is obtained after the image super-resolution model fuses the image to be reconstructed, the global features, the global contour information, and the local detail texture information.
[0037] In the embodiment of the present application, the main architecture of the image super-resolution model is a residual network model.
[0038] In some embodiments, in the residual network model, by deploying multiple context attention modules between residual modules, the global features corresponding to the image to be reconstructed can be extracted; by deploying multiple wavelet fusion modules, the global contour information and local detail texture information corresponding to multiple feature maps associated with the image to be reconstructed can be fused.
[0039] It should be noted that the global contour information and local detail texture information are obtained by performing Haar wavelet transform and inverse Haar wavelet transform on the feature maps associated with the image to be reconstructed for frequency decomposition and recombination.
[0040] The image super-resolution reconstruction method provided by the embodiment of the present application inputs the image to be reconstructed into the image super-resolution model, enabling the image super-resolution model to capture pixel-level long-range context dependencies on the basis of the residual network model framework, adaptively focus on global information, and thus extract the global features corresponding to the image to be reconstructed; uses each wavelet fusion module to fuse the global contour information and local detail texture information corresponding to multiple feature maps associated with the image to be reconstructed; finally, the image super-resolution model fuses the image to be reconstructed, the global features, the global contour information, and the local detail texture information, and can reconstruct a clear, high-quality, and high-resolution reconstructed image, thereby improving the resolution of the image to be reconstructed and solving the problems of blurred and unclear, low-resolution monitoring video images. In particular, it meets the requirements of real-time video monitoring and analysis in scenarios such as bank branches, hospitals, and stations.
[0041] Optionally, Figure 2 For the structural schematic diagram of the image super-resolution model provided by an embodiment of the present application, see Figure 2As shown, the image super-resolution model sequentially includes: a first convolutional layer, a first residual module, a first downsampling module, a second residual module, a second downsampling module, a third residual module, a first context attention module (i.e., the first CAM module), a third downsampling module, a fourth residual module, a second context attention module (i.e., the second CAM module), a third context attention module (i.e., the third CAM module), a second convolutional layer, a first wavelet fusion module (i.e., the first WFM module), a third convolutional layer, a second wavelet fusion module (i.e., the second WFM module), a fourth convolutional layer, a third wavelet fusion module (i.e., the third WFM module), a fifth convolutional layer, and a sixth convolutional layer;
[0042] The residual network model may further include: a first wavelet transform module (i.e., the first WT module), a seventh convolutional layer, a second wavelet transform module (i.e., the second WT module), an eighth convolutional layer, a third wavelet transform module (i.e., the third WT module), and a ninth convolutional layer;
[0043] Among them, the input end of the first wavelet transform module is connected to the output end of the first downsampling module, the output end of the first wavelet transform module is connected to the input end of the seventh convolutional layer, and the output end of the seventh convolutional layer is connected to the input end of the third wavelet fusion module;
[0044] The input end of the second wavelet transform module is connected to the output end of the second downsampling module, the output end of the second wavelet transform module is connected to the input end of the eighth convolutional layer, and the output end of the eighth convolutional layer is connected to the input end of the second wavelet fusion module;
[0045] The input end of the third wavelet transform module is connected to the output end of the third downsampling module, the output end of the third wavelet transform module is connected to the input end of the ninth convolutional layer, and the output end of the ninth convolutional layer is connected to the input end of the first wavelet fusion module.
[0046] In the embodiments of the present application, the first convolutional layer to the ninth convolutional layer are all Conv3*3 convolutional layers, and the parameters of each convolutional layer are different and can be specifically set according to actual needs.
[0047] Each residual module sequentially has a Conv1*1 convolutional layer, a Conv3*3 convolutional layer, and a Conv1*1 convolutional layer from the input end to the output end.
[0048] Each wavelet fusion module includes a channel separation (Split) sub-module, a wavelet inverse transformation (Wavelet Inverse Transformation, WIT) sub-module, an upsampling (Upsample) sub-module, a feature concatenation (Concatenation) sub-module, and a Conv1*1 convolutional layer.
[0049] Among them, the channel separation sub-module is used for channel separation. For example, if the original number of channels C of a feature map is 512, the channel separation operation (i.e., the Split operation) will split the feature map into 2 feature maps with the same number of channels, and the number of channels is 256 and 256 respectively.
[0050] The feature concatenation sub-module is used for channel merging and addition (also known as fusion). For example, if the original number of channels C of a feature map is 256, and the number of channels C of another feature map is also 256, when the two are added together for channel merging, the number of channels becomes 512.
[0051] It should be noted that the inverse wavelet transform sub-module in each wavelet fusion module provided by the embodiments of the present application will perform the inverse Haar wavelet transform on the input feature map for frequency domain decomposition, and the wavelet transform module connected to the wavelet fusion module will perform the Haar wavelet transform on the input feature map for feature recombination; the wavelet fusion module will fuse the results of the Haar wavelet transform and the inverse Haar wavelet transform to achieve feature fusion between the shallow layer and the deep layer. See the following formula (3):
[0052]
[0053] Among them, L and H represent the low-pass filter and the high-pass filter respectively. The low-pass filter can capture global contour information, and the high-pass filter can extract local detail texture information. Through the combination of the two filters, four Haar wavelet kernels can be realized, including:
[0054]
[0055] The above four Haar wavelet kernels can be used to decompose the feature map into frequency domain components LL, LH, HL, and HH. That is to say, in the embodiments of the present application, the first wavelet transform module, the second wavelet transform module, and the third wavelet transform module are used to perform the Haar wavelet transform on the feature map input to the module, so as to decompose the feature map into frequency domain components LL, LH, HL, and HH.
[0056] Optionally, the super-resolution model obtains the reconstructed image in the following manner, which specifically includes the following steps:
[0057] Step 1): For the image to be reconstructed, use the first convolutional layer, the first residual module, and the first downsampling module to perform convolutional operations and downsampling operations in sequence to obtain a first feature map associated with the image to be reconstructed.
[0058] In an embodiment of the present application, it is assumed that the size of the image to be reconstructed is: height H, width W, and the number of channels is 3, that is: 3×H×W. Then, the image to be reconstructed is input into the first convolutional layer of Conv3*3 for convolution operation, and the output result is input into the first residual module. The output result is subjected to downsampling operation to obtain a first feature map with a size of C×H×W.
[0059] Step 2): For the first feature map, convolution operation and downsampling operation are sequentially performed using the second residual module and the second downsampling module to obtain a second feature map associated with the image to be reconstructed.
[0060] In an embodiment of the present application, the first feature map is input into the second residual module, and the output result is subjected to downsampling operation to obtain a second feature map with a size of 2C×H / 2×W / 2.
[0061] Step 3): For the second feature map, convolution operation, feature extraction operation, and downsampling operation are sequentially performed using the third residual module, the first context attention module, and the third downsampling module to obtain a third feature map associated with the image to be reconstructed.
[0062] In an embodiment of the present application, the second feature map is input into the third residual module, the output result is input into the first context attention module for feature extraction operation, and the output result is input into the third downsampling module for downsampling operation to obtain a third feature map with a size of 4C×H / 4×W / 4.
[0063] Step 4): For the third feature map, convolution operation and feature extraction operation are sequentially performed using the fourth residual module and the second context attention module to obtain a fourth feature map associated with the image to be reconstructed.
[0064] In an embodiment of the present application, the third feature map is input into the fourth residual module, and the output result is input into the second context attention module for feature extraction operation to obtain a fourth feature map with a size of 8C×H / 8×W / 8.
[0065] Step 5): For the fourth feature map, feature extraction operation and convolution operation are sequentially performed using the third context attention module and the second convolutional layer to obtain a fifth feature map associated with the image to be reconstructed. The fifth feature map is used to represent the global feature corresponding to the image to be reconstructed.
[0066] In an embodiment of the present application, the fourth feature map is input into the third context attention module for feature extraction operation, and the output result is input into the second convolutional layer of Conv3*3 for convolution operation to obtain a fifth feature map with a size of 16C×H / 8×W / 8. It should be noted that the fifth feature map is the global feature map corresponding to the image to be reconstructed.
[0067] Through the above steps 3) - 5), the global feature maps corresponding to the reconstructed images are gradually extracted using the first context attention module, the second context attention module, and the third context attention module, providing a reliable data basis for subsequent feature map fusion.
[0068] Step 6): Use the first wavelet fusion module to perform an inverse wavelet transform operation on the fifth feature map to obtain a first transformed feature. Perform a feature concatenation operation, an upsampling operation, and a convolution operation on the first transformed feature and the second transformed feature to obtain a first intermediate feature map. The first transformed feature represents the local detailed texture information corresponding to the fifth feature map. Use the third convolutional layer to perform a convolution operation on the first intermediate feature map to obtain a sixth feature map. The second transformed feature is obtained after the third wavelet transform module and the ninth convolutional layer perform a wavelet transform operation and a convolution operation on the third feature map. The second transformed feature represents the global contour information corresponding to the third feature map.
[0069] As Figure 2 shown, the sixth feature map is also called the first upsampled feature, and its size is 8C × H / 4 × W / 4. In the embodiments of the present application, the first wavelet fusion module needs to receive data from two aspects: On the first aspect, it receives the second transformed feature output from the ninth convolutional layer; on the second aspect, it receives the fifth feature map output from the second convolutional layer in step 5).
[0070] Regarding the first aspect:
[0071] The second transformed feature is obtained after the third wavelet transform module performs a wavelet transform operation on the third feature map and then inputs the output result into the ninth convolutional layer for a convolution operation. The second transformed feature represents the global contour information corresponding to the third feature map.
[0072] In practical applications, the third wavelet transform module performs a wavelet transform operation on the third feature map F, which can be expressed as:
[0073]
[0074] where respectively represent group convolutions with a stride of 2 and weights LL T LH T HL T HH T .
[0075] Regarding the second aspect:
[0076] After receiving the fifth feature map, the first wavelet fusion module needs to perform an inverse wavelet transform operation on the fifth feature map to obtain the first transformed feature F'. That is, the first wavelet fusion module aggregates the general features (LL, LH, HL) and high-frequency detail feature HH of the fifth feature map through the inverse Haar wavelet transform to reconstruct the image and obtain the first transformed feature F'. Specifically, it can be represented by the following formula (4):
[0077]
[0078] Among them, respectively represent four independent transposed convolutions, and the weights are LL T , LH T , HL T , HH T .
[0079] It should be noted that the concatenation between the first transformed feature and the second transformed feature is equivalent to the fusion between the local detail texture information of the global feature and the global contour information corresponding to the third feature map.
[0080] Step 7): Use the second wavelet fusion module to perform an inverse wavelet transform operation on the sixth feature map to obtain the third transformed feature, perform a feature splicing operation, an upsampling operation, and a convolution operation on the third transformed feature and the fourth transformed feature to obtain the second intermediate feature map. The third transformed feature represents the local detail texture information corresponding to the sixth feature map; use the fourth convolutional layer to perform a convolution operation on the second intermediate feature map to obtain the seventh feature map; the fourth transformed feature is obtained after the second wavelet transform module and the eighth convolutional layer perform a wavelet transform operation and a convolution operation on the second feature map, and the fourth transformed feature represents the global contour information corresponding to the second feature map.
[0081] As Figure 2 shown, the seventh feature map is also called the second upsampled feature, and its size is 4C×H / 2×W / 2. In the embodiment of the present application, the second wavelet fusion module needs to receive two aspects of data: on the first hand, it receives the fourth transformed feature output from the eighth convolutional layer; on the second hand, it receives the sixth feature map output from the third convolutional layer in step 6).
[0082] Regarding the first aspect:
[0083] The fourth transformed feature is obtained after the second wavelet transform module performs a wavelet transform operation on the second feature map and then inputs the output result into the eighth convolutional layer for convolution operation. Among them, the fourth transformed feature represents the global contour information corresponding to the second feature map.
[0084] The process of the second wavelet transform module performing wavelet transform operation on the second feature map is similar to the process of the first wavelet transform module performing wavelet transform operation on the third feature map mentioned above. To avoid repetition, the corresponding parts in this and subsequent embodiments will not be described again.
[0085] Regarding the second aspect:
[0086] After receiving the sixth feature map, the second wavelet fusion module needs to perform an inverse wavelet transform operation on the sixth feature map to obtain the third transformed feature.
[0087] The process of the second wavelet fusion module performing an inverse wavelet transform operation on the sixth feature map is similar to the process of the first wavelet fusion module performing an inverse wavelet transform operation on the fifth feature map mentioned above. To avoid repetition, the corresponding parts in this and subsequent embodiments will not be described again.
[0088] Step 8): Use the third wavelet fusion module to perform an inverse wavelet transform operation on the seventh feature map to obtain the fifth transformed feature, perform a feature splicing operation, an upsampling operation, and a convolution operation on the fifth transformed feature and the sixth transformed feature to obtain the third intermediate feature map. The fifth transformed feature represents the local detailed texture information corresponding to the seventh feature map; use the fifth convolutional layer to perform a convolution operation on the third intermediate feature map to obtain the eighth feature map; the sixth transformed feature is obtained after the first wavelet transform module and the seventh convolutional layer perform wavelet transform operation and convolution operation on the first feature map. The sixth transformed feature represents the global contour information corresponding to the first feature map.
[0089] As Figure 2 shown, the eighth feature map is also called the third upsampled feature, and its size is 2C×H×W. In the embodiment of the present application, the third wavelet fusion module needs to receive data from two aspects: on the first aspect, it receives the sixth transformed feature output from the seventh convolutional layer; on the second aspect, it receives the seventh feature map output from the fourth convolutional layer in step 7).
[0090] Regarding the first aspect:
[0091] The sixth transformed feature is obtained after the first wavelet transform module performs a wavelet transform operation on the first feature map and then inputs the output result into the seventh convolutional layer for convolution operation. Among them, the sixth transformed feature represents the global contour information corresponding to the first feature map.
[0092] Regarding the second aspect:
[0093] After receiving the seventh feature map, the third wavelet fusion module needs to perform an inverse wavelet transform operation on the seventh feature map to obtain the fifth transformed feature.
[0094] Step 9): Use the sixth convolutional layer to perform a convolution operation on the eighth feature map to obtain the first-stage image.
[0095] Step 10): Add the image feature maps of the image to be reconstructed and the first-stage image to obtain the reconstructed image.
[0096] In the embodiment of the present application, after performing a convolution operation on the eighth feature map with a size of 2C×H×W in the second convolutional layer of Conv3*3 to obtain the first-stage image, it is necessary to fuse the image to be reconstructed with a size of 3×H×W and the first-stage image (i.e., the above-mentioned image feature map addition process) to obtain a high-resolution reconstructed image.
[0097] Optionally, for any context attention module, it includes a tenth convolutional layer, an eleventh convolutional layer, and a twelfth convolutional layer; the parameters between the tenth convolutional layer, the eleventh convolutional layer, and the twelfth convolutional layer are different from each other.
[0098] In the embodiment of the present application, the tenth convolutional layer, the eleventh convolutional layer, and the twelfth convolutional layer are all Conv1*1, but the parameters are different from each other.
[0099] Figure 3 For the structural schematic diagram of the context attention module provided in an embodiment of the present application, see Figure 3 As shown, the feature extraction operation is implemented in the following manner, specifically including the following steps a)-step e):
[0100] Step a): Use the tenth convolutional layer and the eleventh convolutional layer to perform convolution operations on the input feature map respectively to obtain the fourth intermediate feature map M output by the tenth convolutional layer and the fifth intermediate feature map N output by the eleventh convolutional layer.
[0101] In the embodiment of the present application, the input feature map is input to the tenth convolutional layer so that the tenth convolutional layer performs a reshaping operation (i.e., Reshape operation) or a transpose operation (i.e., Transpose operation) to obtain the fourth intermediate feature map M; the input feature map is input to the eleventh convolutional layer so that the eleventh convolutional layer performs a reshaping operation (i.e., Reshape operation) to obtain the fifth intermediate feature map N.
[0102] For example, given an input feature map F∈R C×H×W , use the tenth convolutional layer W m and W n to process respectively to obtain the fourth intermediate feature map M and the fifth intermediate feature map N; specifically, it can be expressed as: and where {M,N}∈R C×H×W .
[0103] Step b): Based on the fourth intermediate feature map and the fifth intermediate feature map, determine the correlation matrix; the correlation matrix is used to characterize the correlation degree between each pixel in the fourth intermediate feature map and each pixel in the fifth intermediate feature map.
[0104] In the embodiment of the present application, first, the real number fields of the fourth intermediate feature map M and the fifth intermediate feature map N need to be transformed into R C×K , where K = M × N. Then, the transpose M T of the fourth intermediate feature map M and the fifth intermediate feature map N are subjected to matrix multiplication, and then an activation function (such as the Softmax function) is applied to calculate and obtain the correlation matrix.
[0105] Optionally, based on the fourth intermediate feature map and the fifth intermediate feature map, to determine the correlation matrix, it can be calculated through the following formula (1):
[0106]
[0107] where p ji represents the correlation degree between the i-th pixel M i in the fourth intermediate feature map and the j-th pixel N j in the fifth intermediate feature map; K = M × N.
[0108] Step c): Use the twelfth convolutional layer to perform a convolutional operation on the input feature map to obtain the sixth intermediate feature map output by the twelfth convolutional layer.
[0109] In the embodiment of the present application, use the twelfth convolutional layer of Conv1*1 to perform a Reshape operation on the input feature map F and transform it into the sixth intermediate feature map B; specifically, it can be expressed as: where B ∈ R C×K .
[0110] Step d): Multiply the correlation matrix and the sixth intermediate feature map matrix to obtain the seventh intermediate feature map.
[0111] In the embodiment of the present application, perform a matrix multiplication operation on the correlation matrix P and the sixth intermediate feature map B, and transform its result into R C×H×W .
[0112] Step e): Based on the input feature map and the seventh intermediate feature map, generate the output feature map after feature extraction of the input feature map.
[0113] In the embodiment of the present application, the input feature map F ∈ R C×H×W passes through a skip connection and performs an element-wise sum with the seventh intermediate feature map, and finally obtains the output feature map F' ∈ R C×H×W .
[0114] Among them, the input feature map is any one of the output feature maps of the third residual module, the output feature map of the fourth residual module, and the fourth feature map; the output feature map is any one of the output feature map of the first context attention module, the fourth feature map, and the output feature map of the third context attention module; the output feature map of the first context attention module is the feature output map after feature extraction of the output feature map of the third residual module, the fourth feature map is the feature output map after feature extraction of the output feature map of the fourth residual module, and the output feature map of the third context attention module is the feature output map after feature extraction of the fourth feature map.
[0115] Optionally, the image super-resolution model is trained in the following way:
[0116] Based on the training data set, the image super-resolution model is iteratively trained until the target loss function reaches a preset threshold, and the trained image super-resolution model is obtained;
[0117] The target loss function is represented by the following formula (2):
[0118]
[0119] Among them, L1(y′,y) represents the target loss function, y (i) represents the predicted pixel value of the i-th pixel point of the image sample in the training data set output by the image super-resolution model, and y′ (i) represents the true pixel value of the i-th pixel point of the image sample, and m represents the number of pixel points in the image sample.
[0120] In practical applications, the initial image super-resolution model is iteratively trained based on multiple image samples, which can be specifically implemented through the following steps:
[0121] (1). Prepare the training data set. For example, collect RGB color images in the network point monitoring scenario of the bank, complete the annotation of positive and negative samples of the data set, and obtain multiple image samples.
[0122] (2). Adjust the size of each image sample in the training data set. For example, by the bilinear interpolation method, the size of each image sample in the training data set is transformed into a height H = 512 and a width W = 512.
[0123] (3). Uniformly represent all trainable parameters in the above convolutional neural network deep learning model as θ R (The parameters are randomly initialized at the beginning and are continuously iteratively updated during the training process).
[0124] (4). Define the target loss function \(L1(y′, y)\) of the image super-resolution model, and calculate the sum of the absolute differences between the true pixel values and the predicted pixel values, which is achieved through the above formula (2).
[0125] (5). Randomly select \(n\) image samples from the training set. For example, in this embodiment, the training batch size Batch_size of the neural network is set to 8.
[0126] (6). Input the image samples selected in step (5) into the initial image super-resolution model.
[0127] (7). Calculate the gradient of the target loss function \(L1(y′, y)\) with respect to the parameter \(\theta\) R :
[0128] (8). Adopt the stochastic gradient descent method to update and optimize the parameters in the network
[0129] where \(\epsilon\) is the learning rate of the network, usually set to 0.0001.
[0130] (9). Repeat steps (2) - (8), continuously update the parameter \(\theta\) in the network R , and one pass through all the image samples in the training set is recorded as one iteration. Stop until the preset number of iterations is reached, or the target loss function reaches the preset threshold, and obtain the trained image super-resolution model.
[0131] Figure 4 is the structural schematic diagram of an image super-resolution reconstruction device provided by an embodiment of the present application. As Figure 4 shown, the image super-resolution reconstruction device includes:
[0132] An acquisition module 401, configured to acquire the image to be reconstructed;
[0133] An image reconstruction module 402, configured to input the image to be reconstructed into the image super-resolution model to obtain the reconstructed image output by the image super-resolution model; the resolution of the reconstructed image is higher than that of the image to be reconstructed;
[0134] Among them, the image super-resolution model is based on a residual network model, which includes a plurality of context attention modules and a plurality of wavelet fusion modules; the plurality of context attention modules are used to extract the global features corresponding to the image to be reconstructed; each wavelet fusion module is used to fuse the global contour information and local detail texture information corresponding to a plurality of feature maps associated with the image to be reconstructed;
[0135] The reconstructed image is obtained by the image super-resolution model through fusing the image to be reconstructed, the global features, the global contour information, and the local detail texture information.
[0136] The image super-resolution reconstruction device provided by the embodiment of the present application inputs the image to be reconstructed into an image super-resolution model. Based on the residual network model framework, the image super-resolution model captures pixel-level long-range context dependencies through multiple context attention modules, adaptively focuses on global information, and thus extracts the global features corresponding to the image to be reconstructed. Each wavelet fusion module is used to fuse the global contour information and local detail texture information corresponding to multiple feature maps associated with the image to be reconstructed. Finally, the image super-resolution model fuses the image to be reconstructed, the global features, each global contour information, and each local detail texture information, and can reconstruct a clear, high-quality, and high-resolution reconstructed image, thereby improving the resolution of the image to be reconstructed and solving the problems of blurred and unclear monitoring video images with low resolution. In particular, it meets the requirements of real-time video monitoring and analysis in scenarios such as bank branches, hospitals, and stations.
[0137] The embodiment of the present invention also provides a computer program product.
[0138] The various embodiments of the systems and technologies described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer program products, the one or more computer program products can include one or more computer programs, the one or more computer programs can be executed and / or interpreted on a programmable system including at least one programmable processor, the programmable processor can be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0139] Figure 5 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Refer to Figure 5 , Figure 5 The displayed electronic device 12 is only an example and should not bring any limitations to the functions and usage scopes of the embodiments of the present application. As Figure 5 shown, the electronic device 12 is presented in the form of a general computing device. The components of the electronic device 12 can include but are not limited to: one or more processors or processing units 16, a system memory 28, and a bus 18 connecting different system components (including the system memory 28 and the processing unit 16).
[0140] Bus 18 represents one or more of several types of bus architectures, including a memory bus or memory controller, a peripheral bus, an Accelerated Graphics Port, a processor bus, or a local bus using any of a variety of bus architectures. By way of example, and without limitation, these architectures include the Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnect (PCI) bus.
[0141] Electronic device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by electronic device 12, including both volatile and nonvolatile media, removable and non-removable media.
[0142] System memory 28 can include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. Electronic device 12 may further include other removable / non-removable, volatile / nonvolatile computer system storage media. By way of example only, storage system 34 can be used for reading and writing non-removable, nonvolatile magnetic media ( Figure 5 not shown and typically called a "hard disk drive"). Although Figure 5 not shown in the figures, a disk drive for reading and writing removable nonvolatile disks (such as a "floppy disk"), and an optical disk drive for reading and writing removable nonvolatile optical disks (such as a CD-ROM, DVD-ROM or other optical media) can be provided. In such cases, each drive can be connected to bus 18 by one or more data media interfaces. Memory 28 may include at least one program product having a set (e.g., at least one) of program modules that are configured to carry out the functions of the embodiments of the present application.
[0143] A program / utility 40 having a set (at least one) of program modules 46 can be stored, for example, in memory 28, such program modules 46 including, but not limited to, an operating system, one or more application programs, other program modules, and program data, each of which examples or some combination thereof may include an implementation of a network environment. Program modules 46 typically carry out the functions and / or methods of the embodiments described herein.
[0144] The electronic device 12 can also communicate with one or more external devices 14 (such as a keyboard, a pointing device, a display 24, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 12, and / or communicate with any device that enables the electronic device 12 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 22. Moreover, the electronic device 12 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 20. As shown in the figure, the network adapter 20 communicates with other modules of the electronic device 12 through the bus 18. It should be understood that although Figure 5 not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0145] The processing unit 16 executes various functional applications and data processing by running programs stored in the system memory 28. For example, it implements an image super-resolution reconstruction method provided by an embodiment of the present invention: obtaining an image to be reconstructed; inputting the image to be reconstructed into an image super-resolution model to obtain a reconstructed image output by the image super-resolution model; the resolution of the reconstructed image is higher than that of the image to be reconstructed; wherein, the image super-resolution model is based on a residual network model, which includes a plurality of context attention modules and a plurality of wavelet fusion modules; the plurality of context attention modules are used to extract global features corresponding to the image to be reconstructed; each wavelet fusion module is used to fuse the global contour information and local detail texture information corresponding to a plurality of feature maps associated with the image to be reconstructed; the reconstructed image is obtained after the image super-resolution model fuses the image to be reconstructed, global features, each global contour information, and each local detail texture information.
[0146] An embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements an image super-resolution reconstruction method provided by all embodiments of the present invention: obtaining an image to be reconstructed; inputting the image to be reconstructed into an image super-resolution model to obtain a reconstructed image output by the image super-resolution model; the resolution of the reconstructed image is higher than that of the image to be reconstructed; wherein, the image super-resolution model is based on a residual network model, which includes a plurality of context attention modules and a plurality of wavelet fusion modules; the plurality of context attention modules are used to extract global features corresponding to the image to be reconstructed; each wavelet fusion module is used to fuse the global contour information and local detail texture information corresponding to a plurality of feature maps associated with the image to be reconstructed; the reconstructed image is obtained after the image super-resolution model fuses the image to be reconstructed, global features, each global contour information and each local detail texture information. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, electronic devices, devices or components of electricity, magnetism, light, electromagnetic, infrared, or semiconductor, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction-executing electronic device, device or component.
[0147] The computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate or transmit a program for use by or combined with an instruction-executing electronic device, device or component.
[0148] The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0149] Computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or, alternatively, may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0150] Note that the above is only the preferred embodiment of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments herein, and various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, more other equivalent embodiments may be included, and the scope of the present invention is determined by the scope of the appended claims.
Claims
1. An image super-resolution reconstruction method, characterized in that, The method includes: Obtaining an image to be reconstructed; Inputting the image to be reconstructed into an image super-resolution model to obtain a reconstructed image output by the image super-resolution model; the resolution of the reconstructed image is higher than that of the image to be reconstructed; Wherein, the image super-resolution model is based on a residual network model, which includes a plurality of context attention modules and a plurality of wavelet fusion modules; the plurality of context attention modules are used to extract global features corresponding to the image to be reconstructed; each wavelet fusion module is used to fuse global contour information and local detail texture information corresponding to a plurality of feature maps associated with the image to be reconstructed; The reconstructed image is obtained after the image super-resolution model fuses the image to be reconstructed, the global features, the global contour information, and the local detail texture information.
2. The method according to claim 1, wherein The image super-resolution model sequentially includes: a first convolutional layer, a first residual module, a first downsampling module, a second residual module, a second downsampling module, a third residual module, a first context attention module, a third downsampling module, a fourth residual module, a second context attention module, a third context attention module, a second convolutional layer, a first wavelet fusion module, a third convolutional layer, a second wavelet fusion module, a fourth convolutional layer, a third wavelet fusion module, a fifth convolutional layer, and a sixth convolutional layer; The residual network model further includes: a first wavelet transform module, a seventh convolutional layer, a second wavelet transform module, an eighth convolutional layer, a third wavelet transform module, and a ninth convolutional layer; Wherein, the input end of the first wavelet transform module is connected to the output end of the first downsampling module, the output end of the first wavelet transform module is connected to the input end of the seventh convolutional layer, and the output end of the seventh convolutional layer is connected to the input end of the third wavelet fusion module; The input end of the second wavelet transform module is connected to the output end of the second downsampling module, the output end of the second wavelet transform module is connected to the input end of the eighth convolutional layer, and the output end of the eighth convolutional layer is connected to the input end of the second wavelet fusion module; The input end of the third wavelet transform module is connected to the output end of the third downsampling module, the output end of the third wavelet transform module is connected to the input end of the ninth convolutional layer, and the output end of the ninth convolutional layer is connected to the input end of the first wavelet fusion module.
3. The method according to claim 2, wherein The super-resolution model obtains the reconstructed image in the following manner: For the image to be reconstructed, a first feature map associated with the image to be reconstructed is obtained by sequentially performing convolution operations and downsampling operations using the first convolutional layer, the first residual module, and the first downsampling module; For the first feature map, a second feature map associated with the image to be reconstructed is obtained by sequentially performing convolution operations and downsampling operations using the second residual module and the second downsampling module; For the second feature map, a third feature map associated with the image to be reconstructed is obtained by sequentially performing convolution operations, feature extraction operations, and downsampling operations using the third residual module, the first context attention module, and the third downsampling module; For the third feature map, perform convolution operations and feature extraction operations in sequence using the fourth residual module and the second context attention module to obtain a fourth feature map associated with the image to be reconstructed; For the fourth feature map, perform feature extraction operations and convolution operations in sequence using the third context attention module and the second convolutional layer to obtain a fifth feature map associated with the image to be reconstructed, where the fifth feature map is used to represent the global features corresponding to the image to be reconstructed; Perform an inverse wavelet transform operation on the fifth feature map using the first wavelet fusion module to obtain a first transformed feature. Perform a feature concatenation operation, an upsampling operation, and a convolution operation on the first transformed feature and a second transformed feature to obtain a first intermediate feature map. The first transformed feature represents the local detailed texture information corresponding to the fifth feature map; perform a convolution operation on the first intermediate feature map using the third convolutional layer to obtain a sixth feature map; the second transformed feature is obtained by performing a wavelet transform operation and a convolution operation on the third feature map using the third wavelet transform module and the ninth convolutional layer, and the second transformed feature represents the global contour information corresponding to the third feature map; Perform an inverse wavelet transform operation on the sixth feature map using the second wavelet fusion module to obtain a third transformed feature. Perform a feature concatenation operation, an upsampling operation, and a convolution operation on the third transformed feature and a fourth transformed feature to obtain a second intermediate feature map. The third transformed feature represents the local detailed texture information corresponding to the sixth feature map; perform a convolution operation on the second intermediate feature map using the fourth convolutional layer to obtain a seventh feature map; the fourth transformed feature is obtained by performing a wavelet transform operation and a convolution operation on the second feature map using the second wavelet transform module and the eighth convolutional layer, and the fourth transformed feature represents the global contour information corresponding to the second feature map; Perform an inverse wavelet transform operation on the seventh feature map using the third wavelet fusion module to obtain a fifth transformed feature. Perform a feature concatenation operation, an upsampling operation, and a convolution operation on the fifth transformed feature and a sixth transformed feature to obtain a third intermediate feature map. The fifth transformed feature represents the local detailed texture information corresponding to the seventh feature map; perform a convolution operation on the third intermediate feature map using the fifth convolutional layer to obtain an eighth feature map; the sixth transformed feature is obtained by performing a wavelet transform operation and a convolution operation on the first feature map using the first wavelet transform module and the seventh convolutional layer, and the sixth transformed feature represents the global contour information corresponding to the first feature map; Perform a convolution operation on the eighth feature map using the sixth convolutional layer to obtain a first-stage image; Perform an addition process on the feature maps of the image to be reconstructed and the first-stage image to obtain the reconstructed image.
4. The method according to claim 3, characterized in that, For any context attention module, it includes a tenth convolutional layer, an eleventh convolutional layer, and a twelfth convolutional layer; the parameters between the tenth convolutional layer, the eleventh convolutional layer, and the twelfth convolutional layer are different from each other; The feature extraction operation is implemented in the following manner: Use the tenth convolutional layer and the eleventh convolutional layer to perform convolutional operations on the input feature map respectively, to obtain a fourth intermediate feature map output by the tenth convolutional layer and a fifth intermediate feature map output by the eleventh convolutional layer; Based on the fourth intermediate feature map and the fifth intermediate feature map, determine a correlation matrix; the correlation matrix is used to represent the correlation degree between each pixel in the fourth intermediate feature map and each pixel in the fifth intermediate feature map; Use the twelfth convolutional layer to perform a convolutional operation on the input feature map, to obtain a sixth intermediate feature map output by the twelfth convolutional layer; Multiply the correlation matrix and the sixth intermediate feature map to obtain a seventh intermediate feature map; Based on the input feature map and the seventh intermediate feature map, generate an output feature map after feature extraction of the input feature map; Wherein, the input feature map is any one of the output feature map of the third residual module, the output feature map of the fourth residual module, and the fourth feature map; the output feature map is any one of the output feature map of the first context attention module, the fourth feature map, and the output feature map of the third context attention module; the output feature map of the first context attention module is a feature output map after feature extraction of the output feature map of the third residual module, the fourth feature map is a feature output map after feature extraction of the output feature map of the fourth residual module, and the output feature map of the third context attention module is a feature output map after feature extraction of the fourth feature map.
5. The method according to claim 4, characterized in that, The determining the correlation matrix based on the fourth intermediate feature map and the fifth intermediate feature map is calculated through the following formula (1): where p ji represents the correlation degree between the i-th pixel M in the fourth intermediate feature map i and the j-th pixel N in the fifth intermediate feature map j ; K = M × N.
6. The method according to any one of claims 1 to 5, characterized in that, The image super-resolution model is trained in the following manner: Based on the training data set, perform iterative training on the image super-resolution model until the target loss function reaches a preset threshold, to obtain the trained image super-resolution model; The target loss function is represented by the following formula (2): Among them, L1(y′,y) represents the target loss function, where y (i) represents the predicted pixel value of the i-th pixel point of the image sample in the training dataset output by the image super-resolution model, and y′ (i) represents the true pixel value of the i-th pixel point of the image sample, and m represents the number of pixel points in the image sample.
7. An image super-resolution reconstruction device, characterized in that, The device includes: An acquisition module, configured to acquire an image to be reconstructed; An image reconstruction module, configured to input the image to be reconstructed into the image super-resolution model, to obtain a reconstructed image output by the image super-resolution model; the resolution of the reconstructed image is higher than the resolution of the image to be reconstructed; Wherein, the image super-resolution model is based on a residual network model, which includes a plurality of context attention modules and a plurality of wavelet fusion modules; the plurality of context attention modules are used to extract global features corresponding to the image to be reconstructed; each wavelet fusion module is used to fuse global contour information and local detail texture information corresponding to a plurality of feature maps associated with the image to be reconstructed. The reconstructed image is obtained after the image super-resolution model fuses the image to be reconstructed, the global features, the global contour information, and the local detail texture information.
8. An electronic device, characterized in that, Comprising: One or more processors; A memory for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the image super-resolution reconstruction method according to any one of claims 1 to 6.
9. A storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the image super-resolution reconstruction method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the image super-resolution reconstruction method according to any one of claims 1 to 6.
Citation Information
Cited By
Data processing method and device, equipment, storage medium and computer program product
CN120852165A