Image restoration method, device and equipment and computer readable storage medium

The position information processing is enhanced through the coordinate perception attention module and the DenseHourglass structure, which solves the problem of Transformer ignoring details and missing features in image recovery, achieving better image recovery effect.

CN120543398APending Publication Date: 2025-08-26GUIZHOU BUSINESS SCHOOL +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510479308.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

The position information of the standard Transformer is embedded in image recovery ignores details such as fog density, raindrop stacking and noise distribution, resulting in poor performance when restoring multiple degraded types of images, and often lose edge features and local details during downsampling.

Method used

The coordinate perception attention module is used to enhance position information, and multi-scale feature processing is performed by constructing a dense jump connection DenseHourglass structure and dynamic prompt mechanism to make up for the lost local detailed features during multiple downsampling.

Benefits of technology

The effect of image recovery is significantly improved, and the potential ability of location information encoding is fully utilized to restore detailed features in various degraded images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120543398A_ABST
    Figure CN120543398A_ABST
Patent Text Reader

Abstract

The invention discloses an image restoration method, device and equipment and a computer readable storage medium, and the method comprises the steps: firstly, enhancing position information through a coordinate perception attention module, and capturing local detail features; and on the basis of a dense jump connection DenseHurglass structure constructed in the hierarchical feature recovery module and a dynamic prompt mechanism, local detail features lost in a multi-time down-sampling process are made up, and the problem that position information embedding of a standard Transform often neglects various details such as fog density, raindrop stacking and noise distribution in image recovery is solved. The technical problems that the performance is poor when multiple degradation types of images are recovered, the potential capability of position information coding cannot be fully utilized, and edge features and local details are often lost in the down-sampling process are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image restoration technology, and in particular to an image restoration method, apparatus, device, and computer-readable storage medium. Background Art

[0002] The goal of image restoration is to recover high-quality, clear images from degraded images caused by camera physical limitations, environmental factors, or other factors. These degradations can affect image quality and even severely impact performance in some applications, such as autonomous driving, surveillance, and medical image analysis. Traditional image restoration methods are typically optimized for specific types of degradation. However, with the increasing diversity of requirements and complexity of scenarios, a single approach often fails to meet the efficient restoration requirements in real-world applications.

[0003] With the development of deep learning methods, significant progress has been made in the field of image restoration. Given the differences between different types of degraded images, methods based on deep neural networks have adopted a variety of strategies to address the challenges of image restoration. Some methods incorporate explicit prior knowledge into the network to address specific restoration tasks such as denoising, deraining, deblurring, and dehazing. These methods are developed based on their specific prior knowledge, which limits their generalization capabilities across different degradation types. At the same time, other studies focus on developing robust neural networks that can implicitly learn prior knowledge and features. These methods either lack generalization capabilities across multiple degradation scenarios or require separate training of independent network instances for datasets with different extreme conditions. Such methods are computationally intensive and cumbersome, posing significant challenges to their deployment on resource-limited platforms such as mobile devices and edge devices.

[0004] In recent years, deep learning-based technologies have become the mainstream in the field of image restoration, and integrated image restoration methods have emerged, which aim to handle all types of image degradation problems through a unified model. AirNet solves the integrated restoration task through contrastive learning paradigm learning; IDR proposes a two-stage framework, first learning potential prior knowledge based on physical properties related to the degradation type, and then performing image restoration; PromptIR uses the learned prompt embedding for unified blind image restoration; ProRes adopts a fast learning strategy for integrated image restoration. The position information embedding of the standard Transformer often ignores various details in image restoration such as fog density, raindrop stacking and noise distribution, resulting in poor performance in restoring images of various degraded types. Although these methods have shown remarkable results in image restoration tasks, they fail to fully utilize the potential of position information encoding, and often lose edge features and local details during the downsampling process. Summary of the Invention

[0005] The present application provides an image restoration method, apparatus, device and computer-readable storage medium, which solves the technical problems that the position information embedding of the standard Transformer often ignores various details in image restoration, such as fog density, raindrop stacking and noise distribution, resulting in poor performance in restoring various types of degraded images, failing to fully utilize the potential capabilities of position information encoding, and often losing edge features and local details during the downsampling process.

[0006] In view of this, the first aspect of the present application provides an image restoration method, the method comprising:

[0007] S1, obtain the original image and perform image preprocessing to obtain the initial feature image;

[0008] S2, based on the coordinate perception attention module, the initial feature image is enhanced in position information, and the attention mechanism is used to adjust the attention to the position information to obtain an enhanced feature image;

[0009] S3. Based on the hierarchical feature recovery module, multi-scale feature processing is performed on the enhanced feature image and the initial feature image by constructing a dense jump connection DenseHourglass structure and a dynamic prompt mechanism to obtain a restored image.

[0010] Optionally, step S1 specifically includes:

[0011] S11, obtaining the original image;

[0012] S12, performing pre-processing of resizing and normalization on the original image to obtain an original image that meets preset requirements;

[0013] S13. Performing initial feature extraction on the original image that meets the preset requirements through a preset feature extractor to generate an initial feature map.

[0014] Optionally, step S2 specifically includes:

[0015] S21, perform layer normalization on the initial feature map to obtain a normalized tensor X∈RHxWxC;

[0016] S22, perform feature aggregation on the normalized tensor X through 1×1 convolution to obtain an aggregated feature map;

[0017] S23. Use 3×3 convolution to encode the aggregated feature map to generate query Q, key K and value V. Then project query Q and key K and calculate the dot product of query Q and key K to obtain the transposed attention map A∈RC×C.

[0018] S25. Apply the Softmax function to the attention map A so that the value is between 0 and 1;

[0019] S26. Apply a 3×3 convolution to the attention map A based on the potential coordinate encoding block to obtain an extended feature map. After obtaining the vertical position information set and the horizontal position information set of each pixel in the extended feature map, perform global adjustment and local adjustment on the vertical position information set and the horizontal position information set, and obtain an offset based on the residual operation between the global adjustment and the local adjustment to obtain a feature map with enhanced position information features.

[0020] S27. After element-wise summing the output of the Softmax function and the enhanced position information feature, the output is interacted with the value V component through a dot product operation. After applying a deep convolution module to the value V component, the output is element-wise summed with the feature map of the enhanced position information feature to obtain a feature map of the secondary enhanced position information feature.

[0021] S28. Filter the inefficient information in the feature map of the secondary enhanced position information feature through the GDFN module to obtain an enhanced feature image.

[0022] Optionally, step S3 specifically includes:

[0023] S31, performing feature fusion on the acquired enhanced feature image and the initial feature image to generate a fused feature image;

[0024] S32. Construct a DenseHourglass structure, which includes feature downsampling, dense skip connections, and feature upsampling. In each layer of the DenseHourglass structure, an automatic attention mechanism is applied to adjust and identify important features of the fused feature image.

[0025] S33, restoring the lost local features and edge features in the fused feature image based on the dynamic prompt mechanism, and generating an enhanced fused feature image;

[0026] S34, converting the enhanced fusion feature image back to the image domain to generate a restored image.

[0027] A second aspect of the present application provides an image restoration device, the device comprising:

[0028] An acquisition unit, used for acquiring an original image and performing image preprocessing to obtain an initial feature image;

[0029] An enhancement unit, configured to enhance the position information of the initial feature image based on the coordinate perception attention module, and adjust the attention to the position information using the attention mechanism to obtain an enhanced feature image;

[0030] The restoration unit is used to perform multi-scale feature processing on the enhanced feature image and the initial feature image based on the hierarchical feature recovery module by constructing a dense jump connection DenseHourglass structure and a dynamic prompt mechanism to obtain a restored image.

[0031] Optionally, the acquiring unit is specifically configured to:

[0032] Get the original image;

[0033] Performing resizing and normalization preprocessing on the original image to obtain the original image that meets the preset requirements;

[0034] The preset feature extractor is used to extract initial features from the original image that meets the preset requirements to generate an initial feature map.

[0035] Optionally, the enhancement unit is specifically configured to:

[0036] Perform layer normalization on the initial feature map to obtain the normalized tensor X∈RHxWxC;

[0037] Perform feature aggregation on the normalized tensor X through 1×1 convolution to obtain the aggregated feature map;

[0038] Use 3×3 convolution to encode the aggregated feature map to generate query Q, key K and value V, then project the query Q and key K, and calculate the dot product of the query Q and key K to obtain the transposed attention map A∈RC×C;

[0039] Apply the Softmax function to the attention map A so that the value is between 0 and 1;

[0040] Apply 3×3 convolution to the attention map A based on the potential coordinate encoding block to obtain an extended feature map. After obtaining the vertical position information set and horizontal position information set of each pixel in the extended feature map, perform global and local adjustments on the vertical and horizontal position information sets, and obtain the offset based on the residual operation between the global and local adjustments to obtain a feature map with enhanced position information features.

[0041] After element-wise summing of the output of the Softmax function and the enhanced position information feature, the output is interacted with the value V component through a dot product operation. After applying a deep convolution module to the value V component, the output is element-wise summed with the feature map of the enhanced position information feature to obtain a feature map of the secondary enhanced position information feature.

[0042] The GDFN module is used to filter out the inefficient information in the feature map of the secondary enhanced position information feature to obtain an enhanced feature image.

[0043] Optionally, the recovery unit is specifically configured to:

[0044] Perform feature fusion on the acquired enhanced feature image and the initial feature image to generate a fused feature image;

[0045] Construct a DenseHourglass structure, which includes feature downsampling, dense skip connections, and feature upsampling. In each layer of the DenseHourglass structure, an automatic attention mechanism is applied to adjust and identify important features of the fused feature image.

[0046] Based on the dynamic hint mechanism, the lost local features and edge features in the fused feature image are restored to generate an enhanced fused feature image;

[0047] The enhanced fused feature image is converted back to the image domain to generate the restored image.

[0048] A third aspect of the present application provides an image restoration device, the device comprising a processor and a memory:

[0049] The memory is used to store program code and transmit the program code to the processor;

[0050] The processor is configured to execute the steps of the image restoration method described in the first aspect according to the instructions in the program code.

[0051] A fourth aspect of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium is used to store program code, and the program code is used to execute the method described in the first aspect above.

[0052] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:

[0053] In the present application, an image restoration method, apparatus, device and computer-readable storage medium are provided. First, the coordinate-aware attention module is used to enhance the position information and capture the local detail features. Then, the dense jump connection DenseHourglass structure and dynamic prompt mechanism constructed in the hierarchical feature recovery module compensate for the local detail features lost during multiple downsampling processes. This solves the technical problems that the position information embedding of the standard Transformer often ignores various details in image restoration such as fog density, raindrop stacking and noise distribution, resulting in poor performance in restoring various degraded images, failing to fully utilize the potential capabilities of position information encoding, and often losing edge features and local details during the downsampling process. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 This is a flow chart of an image restoration method in an embodiment of the present application;

[0055] Figure 2 This is a schematic diagram of the structure of an image restoration network in an embodiment of the present application;

[0056] Figure 3 This is a schematic diagram of the structure of the coordinate perception attention module in an embodiment of the present application;

[0057] Figure 4 This is a structural diagram of a hierarchical feature recovery module in an embodiment of the present application;

[0058] Figure 5 This is a schematic structural diagram of an image restoration device in an embodiment of the present application;

[0059] Figure 6 Schematic diagram of the structure of the image restoration device in an embodiment of the present application. DETAILED DESCRIPTION

[0060] In order to help those skilled in the art better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.

[0061] This application designs an image restoration method, apparatus, device and computer-readable storage medium to solve the technical problems that the standard Transformer's position information embedding often ignores various details in image restoration, such as fog density, raindrop stacking and noise distribution, resulting in poor performance in restoring various degraded images, failing to fully utilize the potential capabilities of position information encoding, and often losing edge features and local details during the downsampling process.

[0062] For easier understanding, see Figures 1 to 4 ,in, Figure 1 This is a flow chart of the image restoration method in the embodiment of the present application. Figures 1 to 4 As shown, specifically:

[0063] S1, obtain the original image and perform image preprocessing to obtain the initial feature image;

[0064] Furthermore, step S1 specifically includes:

[0065] S11, obtaining the original image;

[0066] S12, performing pre-processing of resizing and normalization on the original image to obtain an original image that meets preset requirements;

[0067] S13. Performing initial feature extraction on the original image that meets the preset requirements through a preset feature extractor to generate an initial feature map.

[0068] It should be noted that the original degraded images are obtained from cameras, surveillance equipment or other image acquisition devices. These images may be affected by many factors, such as weather conditions (fog, rain), insufficient light or motion blur.

[0069] The original image is resized to meet the size requirements of the model input; the image is normalized to adjust the pixel values ​​to a specific range (such as 0 to 1).

[0070] For example, the original image is resized to 256×256 pixels and the pixel values ​​are normalized to the range of 0 to 1.

[0071] Use a preset feature extractor (such as a convolutional neural network) to extract features from the preprocessed image and generate an initial feature map.

[0072] For example, a pre-trained convolutional neural network is used to extract features from the resized and normalized image to generate a feature map with 64 channels.

[0073] S2, based on the coordinate perception attention module, the initial feature image is enhanced in position information, and the attention mechanism is used to adjust the attention to the position information to obtain an enhanced feature image;

[0074] Furthermore, step S2 specifically includes:

[0075] S21, perform layer normalization on the initial feature map to obtain a normalized tensor X∈RHxWxC;

[0076] It should be noted that the initial feature map is layer normalized to stabilize the training process and accelerate convergence.

[0077] For example: normalize each channel of the initial feature map so that the mean of each channel is 0 and the variance is 1.

[0078] S22, perform feature aggregation on the normalized tensor X through 1×1 convolution to obtain an aggregated feature map;

[0079] It should be noted that 1×1 convolution is used to compress the number of channels of the feature map and aggregate context information.

[0080] For example, a feature map with 64 channels is compressed to 32 channels through 1×1 convolution.

[0081] S23. Use 3×3 convolution to encode the aggregated feature map to generate query Q, key K and value V. Then project query Q and key K and calculate the dot product of query Q and key K to obtain the transposed attention map A∈RC×C.

[0082] It should be noted that 3×3 convolution is used to generate query, key, and value matrices, and then their dot products are calculated to generate the attention map.

[0083] For example, a 3×3 convolution is applied to the aggregated feature map to generate query, key, and value matrices. Assuming the number of channels is 32, the dimensions of the generated query and key matrices are 32×32, and the dimension of the value matrix is ​​32×32.

[0084] S25. Apply the Softmax function to the attention map A so that the value is between 0 and 1;

[0085] It should be noted that the Softmax function is applied to normalize the value of the attention map to between 0 and 1, indicating the correlation of the features.

[0086] For example, we apply the Softmax function to the generated attention map so that the sum of the attention weights at each position is 1.

[0087] S26. Apply a 3×3 convolution to the attention map A based on the potential coordinate encoding block to obtain an extended feature map. After obtaining the vertical position information set and the horizontal position information set of each pixel in the extended feature map, perform global adjustment and local adjustment on the vertical position information set and the horizontal position information set, and obtain an offset based on the residual operation between the global adjustment and the local adjustment to obtain a feature map with enhanced position information features.

[0088] It should be noted that the position information is enhanced by the latent coordinate encoding block to capture global and local position features.

[0089] Global adjustments capture a wider range of positional information, while local adjustments provide finer adjustments, focusing on local details and effectively capturing detailed positional information. The offset is obtained through the residual operation between global and local adjustments, providing the model with the potential features of different degraded images for better decoupling capabilities.

[0090] For example, a 3×3 convolution is applied to the attention map to generate an expanded feature map, and then the vertical and horizontal position information of each pixel is globally and locally adjusted to obtain an offset to enhance the position information.

[0091] S27. After element-wise summing the output of the Softmax function and the enhanced position information feature, the output is interacted with the value V component through a dot product operation. After applying a deep convolution module to the value V component, the output is element-wise summed with the feature map of the enhanced position information feature to obtain a feature map of the secondary enhanced position information feature.

[0092] It should be noted that the enhanced position information is interacted with the value V component, and a deep convolution module is applied to further extract features.

[0093] S28. Filter the inefficient information in the feature map of the secondary enhanced position information feature through the GDFN module to obtain an enhanced feature image.

[0094] It should be noted that the GDFN module is used to filter out inefficient information and retain key features.

[0095] S3. Based on the hierarchical feature recovery module, multi-scale feature processing is performed on the enhanced feature image and the initial feature image by constructing a dense jump connection DenseHourglass structure and a dynamic prompt mechanism to obtain a restored image.

[0096] Furthermore, step S3 specifically includes:

[0097] S31, performing feature fusion on the acquired enhanced feature image and the initial feature image to generate a fused feature image;

[0098] It should be noted that the enhanced feature image and the initial feature image are fused, usually by feature concatenation or addition.

[0099] S32. Construct a DenseHourglass structure, which includes feature downsampling, dense skip connections, and feature upsampling. In each layer of the DenseHourglass structure, an automatic attention mechanism is applied to adjust and identify important features of the fused feature image.

[0100] It should be noted that the DenseHourglass structure is constructed to extract multi-scale features through downsampling and upsampling, and the attention mechanism is applied to identify important features.

[0101] For example, we construct a 4-layer DenseHourglass structure with dense skip connections between each layer, and apply the CBAM attention mechanism to identify important features.

[0102] S33, restoring the lost local features and edge features in the fused feature image based on the dynamic prompt mechanism, and generating an enhanced fused feature image;

[0103] S34, converting the enhanced fusion feature image back to the image domain to generate a restored image.

[0104] It should be noted that by supplementing local details and edge features, combining a multi-scale feature extraction structure with dense connections, a dynamic attention mechanism and degradation information, the feature representation capability is significantly enhanced.

[0105] See also Figure 5 , Figure 5 FIG. 1 is a schematic diagram of the structure of the image restoration device in an embodiment of the present application. Figure 5 As shown, specifically:

[0106] An acquisition unit 501 is used to acquire an original image and perform image preprocessing to obtain an initial feature image;

[0107] An enhancement unit 502 is configured to enhance the position information of the initial feature image based on the coordinate perception attention module and adjust the attention to the position information using the attention mechanism to obtain an enhanced feature image;

[0108] The restoration unit 503 is used to perform multi-scale feature processing on the enhanced feature image and the initial feature image based on the hierarchical feature restoration module by constructing a dense hourglass structure with dense skip connections and a dynamic prompting mechanism to obtain a restored image.

[0109] Furthermore, the acquiring unit 501 is specifically configured to:

[0110] Get the original image;

[0111] Performing resizing and normalization preprocessing on the original image to obtain the original image that meets the preset requirements;

[0112] The preset feature extractor is used to extract initial features from the original image that meets the preset requirements to generate an initial feature map.

[0113] Furthermore, the enhancement unit is specifically used to:

[0114] Perform layer normalization on the initial feature map to obtain the normalized tensor X∈RHxWxC;

[0115] Perform feature aggregation on the normalized tensor X through 1×1 convolution to obtain the aggregated feature map;

[0116] Use 3×3 convolution to encode the aggregated feature map to generate query Q, key K and value V, then project the query Q and key K, and calculate the dot product of the query Q and key K to obtain the transposed attention map A∈RC×C;

[0117] Apply the Softmax function to the attention map A so that the value is between 0 and 1;

[0118] Apply 3×3 convolution to the attention map A based on the potential coordinate encoding block to obtain an extended feature map. After obtaining the vertical position information set and horizontal position information set of each pixel in the extended feature map, perform global and local adjustments on the vertical and horizontal position information sets, and obtain the offset based on the residual operation between the global and local adjustments to obtain a feature map with enhanced position information features.

[0119] After element-wise summing of the output of the Softmax function and the enhanced position information feature, the output is interacted with the value V component through a dot product operation. After applying a deep convolution module to the value V component, the output is element-wise summed with the feature map of the enhanced position information feature to obtain a feature map of the secondary enhanced position information feature.

[0120] The GDFN module is used to filter out the inefficient information in the feature map of the secondary enhanced position information feature to obtain an enhanced feature image.

[0121] Furthermore, the recovery unit is specifically configured to:

[0122] Perform feature fusion on the acquired enhanced feature image and the initial feature image to generate a fused feature image;

[0123] Construct a DenseHourglass structure, which includes feature downsampling, dense skip connections, and feature upsampling. In each layer of the DenseHourglass structure, an automatic attention mechanism is applied to adjust and identify important features of the fused feature image.

[0124] Based on the dynamic hint mechanism, the lost local features and edge features in the fused feature image are restored to generate an enhanced fused feature image;

[0125] The enhanced fused feature image is converted back to the image domain to generate the restored image.

[0126] The present application also provides another stereo matching device based on geometric information, such as Figure 6 As shown, the device 10 includes:

[0127] One or more processors 110 and memory 120, Figure 6 In the description, a processor 110 is used as an example. The processor 110 and the memory 120 may be connected via a bus or other means. Figure 6 The bus connection is taken as an example.

[0128] The processor 110 is used to implement various control logics of the device 10. It can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a single-chip microcomputer, an ARM (Acorn RICS Machine) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination of these components. In addition, the processor 110 can also be any traditional processor, microprocessor or state machine. The processor 110 can also be implemented as a combination of computing devices, for example, a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP and / or any other such configuration.

[0129] Memory 120, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as program instructions corresponding to the method for constructing a multilingual phoneme representation model in the embodiments of the present invention. Processor 110 executes the non-volatile software programs, instructions, and modules stored in memory 120 to execute various functional applications and data processing of device 10, thereby implementing the method for constructing a multilingual phoneme representation model in the aforementioned method embodiment.

[0130] The memory 120 may include a program storage area and a data storage area. The program storage area may store an operating system and applications required for at least one function; the data storage area may store data created based on the use of the device 10. Furthermore, the memory 120 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state memory device. In some embodiments, the memory 120 may optionally include a memory remotely located relative to the processor 110, and such remote memory may be connected to the device 10 via a network. Examples of such networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0131] One or more units are stored in the memory 120, and when executed by one or more processors 110, implement the following steps:

[0132] S1, obtain the original image and perform image preprocessing to obtain the initial feature image;

[0133] S2, based on the coordinate perception attention module, the initial feature image is enhanced in position information, and the attention mechanism is used to adjust the attention to the position information to obtain an enhanced feature image;

[0134] S3. Based on the hierarchical feature recovery module, multi-scale feature processing is performed on the enhanced feature image and the initial feature image by constructing a dense jump connection DenseHourglass structure and a dynamic prompt mechanism to obtain a restored image.

[0135] An embodiment of the present application further provides a computer-readable storage medium for storing program code, wherein the program code is used to execute any one of the implementations of the image restoration method described in the aforementioned embodiments.

[0136] In an embodiment of the present application, an image restoration method, apparatus, device and computer-readable storage medium are provided. First, position information is enhanced and local detail features are captured through a coordinate-aware attention module. Then, based on the densely skipped hourglass structure and dynamic prompt mechanism constructed in the hierarchical feature recovery module, local detail features lost during multiple downsampling processes are compensated. This solves the technical problem that the position information embedding of the standard Transformer often ignores various details in image restoration such as fog density, raindrop stacking and noise distribution, resulting in poor performance in restoring various types of degraded images, failing to fully utilize the potential capabilities of position information encoding, and often losing edge features and local details during the downsampling process.

[0137] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0138] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can, for example, be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0139] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0140] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0141] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0142] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0143] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (English full name: Read-Only Memory, English abbreviation: ROM), random access memory (English full name: Random Access Memory, English abbreviation: RAM), disk or optical disk and other media that can store program code.

[0144] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. An image restoration method, characterized in that: include: S1, obtain the original image and perform image preprocessing to obtain the initial feature image; S2, based on the coordinate perception attention module, the initial feature image is enhanced in position information, and the attention mechanism is used to adjust the attention to the position information to obtain an enhanced feature image; S3. Based on the hierarchical feature recovery module, multi-scale feature processing is performed on the enhanced feature image and the initial feature image by constructing a dense jump connection DenseHourglass structure and a dynamic prompt mechanism to obtain a restored image.

2. The image restoration method according to claim 1, wherein: The step S1 specifically includes: S11, obtaining the original image; S12, performing pre-processing of resizing and normalization on the original image to obtain an original image that meets preset requirements; S13. Performing initial feature extraction on the original image that meets the preset requirements through a preset feature extractor to generate an initial feature map.

3. The image restoration method according to claim 1, wherein: The step S2 specifically includes: S21, perform layer normalization on the initial feature map to obtain a normalized tensor X∈RHxWxC; S22, perform feature aggregation on the normalized tensor X through 1×1 convolution to obtain an aggregated feature map; S23. Use 3×3 convolution to encode the aggregated feature map to generate query Q, key K and value V. Then project query Q and key K and calculate the dot product of query Q and key K to obtain the transposed attention map A∈RC×C. S25. Apply the Softmax function to the attention map A so that the value is between 0 and 1; S26. Apply a 3×3 convolution to the attention map A based on the potential coordinate encoding block to obtain an extended feature map. After obtaining the vertical position information set and the horizontal position information set of each pixel in the extended feature map, perform global adjustment and local adjustment on the vertical position information set and the horizontal position information set, and obtain an offset based on the residual operation between the global adjustment and the local adjustment to obtain a feature map with enhanced position information features. S27. After element-wise summing the output of the Softmax function and the enhanced position information feature, the output is interacted with the value V component through a dot product operation. After applying a deep convolution module to the value V component, the output is element-wise summed with the feature map of the enhanced position information feature to obtain a feature map of the secondary enhanced position information feature. S28. Filter the inefficient information in the feature map of the secondary enhanced position information feature through the GDFN module to obtain an enhanced feature image.

4. The image restoration method according to claim 1, wherein: The step S3 specifically includes: S31, performing feature fusion on the acquired enhanced feature image and the initial feature image to generate a fused feature image; S32. Construct a DenseHourglass structure, which includes feature downsampling, dense skip connections, and feature upsampling. In each layer of the DenseHourglass structure, an automatic attention mechanism is applied to adjust and identify important features of the fused feature image. S33, restoring the lost local features and edge features in the fused feature image based on the dynamic prompt mechanism, and generating an enhanced fused feature image; S34, converting the enhanced fusion feature image back to the image domain to generate a restored image.

5. An image restoration device, characterized in that: include: An acquisition unit, used for acquiring an original image and performing image preprocessing to obtain an initial feature image; An enhancement unit, configured to enhance the position information of the initial feature image based on the coordinate perception attention module, and adjust the attention to the position information using the attention mechanism to obtain an enhanced feature image; The restoration unit is used to perform multi-scale feature processing on the enhanced feature image and the initial feature image based on the hierarchical feature recovery module by constructing a dense jump connection DenseHourglass structure and a dynamic prompt mechanism to obtain a restored image.

6. The image restoration device according to claim 5, characterized in that The acquisition unit is specifically configured to: Get the original image; Performing resizing and normalization preprocessing on the original image to obtain the original image that meets the preset requirements; The preset feature extractor is used to extract initial features from the original image that meets the preset requirements to generate an initial feature map.

7. The image restoration device according to claim 5, characterized in that: The enhancement unit is specifically used for: Perform layer normalization on the initial feature map to obtain the normalized tensor X∈RHxWxC; Perform feature aggregation on the normalized tensor X through 1×1 convolution to obtain the aggregated feature map; Use 3×3 convolution to encode the aggregated feature map to generate query Q, key K and value V, then project the query Q and key K, and calculate the dot product of the query Q and key K to obtain the transposed attention map A∈RC×C; Apply the Softmax function to the attention map A so that the value is between 0 and 1; Apply 3×3 convolution to the attention map A based on the potential coordinate encoding block to obtain an extended feature map. After obtaining the vertical position information set and horizontal position information set of each pixel in the extended feature map, perform global and local adjustments on the vertical and horizontal position information sets, and obtain the offset based on the residual operation between the global and local adjustments to obtain a feature map with enhanced position information features. After element-wise summing of the output of the Softmax function and the enhanced position information feature, the output is interacted with the value V component through a dot product operation. After applying a deep convolution module to the value V component, the output is element-wise summed with the feature map of the enhanced position information feature to obtain a feature map of the secondary enhanced position information feature. The GDFN module is used to filter out the inefficient information in the feature map of the secondary enhanced position information feature to obtain an enhanced feature image.

8. The image restoration device according to claim 5, characterized in that: The recovery unit is specifically used for: Perform feature fusion on the acquired enhanced feature image and the initial feature image to generate a fused feature image; Construct a DenseHourglass structure, which includes feature downsampling, dense skip connections, and feature upsampling. In each layer of the DenseHourglass structure, an automatic attention mechanism is applied to adjust and identify important features of the fused feature image. Based on the dynamic hint mechanism, the lost local features and edge features in the fused feature image are restored to generate an enhanced fused feature image; The enhanced fused feature image is converted back to the image domain to generate the restored image.

9. An image restoration device, characterized in that: The device includes a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is configured to execute the image restoration method according to any one of claims 1 to 4 according to instructions in the program code.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium is used to store program code, and the program code is used to execute the image restoration method according to any one of claims 1 to 4.