An ultra-resolution image reconstruction method, device, equipment and storage medium
By combining feature encoding networks, feature separation and aggregation networks, and image decoding networks, the limitations of image super-resolution reconstruction in existing technologies are solved, and clearer image textures and edge effects are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-15
- Publication Date
- 2026-04-07
AI Technical Summary
Existing super-resolution algorithms based on the original single low-resolution image have limited super-resolution reconstruction effects and are prone to generating abnormal noise.
A combination of feature encoding network, feature separation and aggregation network, and image decoding network is adopted. The APS image to be processed and the event image are input into the feature encoding network for feature extraction, and the shared features are output. Feature separation and aggregation are performed in the feature separation and aggregation network, and finally feature decoding is performed in the image decoding network to generate a super-resolution APS image.
It effectively suppresses anomalous noise and improves the clarity of textures and edges in super-resolution image reconstruction results.
Smart Images

Figure CN115908131B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to a method, apparatus, device and storage medium for super-resolution image reconstruction. Background Technology
[0002] Image super-resolution reconstruction is a classic problem in computer vision, aiming to reconstruct high-resolution images from low-resolution images. Currently, deep learning-based super-resolution image reconstruction methods have been proposed, such as super-resolution algorithms based on convolutional neural networks, super-resolution algorithms based on accelerated convolutional neural networks, and super-resolution algorithms based on skip connection residual networks. However, these algorithms all rely on a single, original low-resolution image, resulting in a relatively simple image signal that cannot provide sufficient image details. This often leads to abnormal noise during super-resolution reconstruction, limiting the effectiveness of the reconstruction. Summary of the Invention
[0003] This application provides a super-resolution image reconstruction method, apparatus, device, and storage medium, which can at least solve the problem that the super-resolution reconstruction effect is relatively limited due to the use of super-resolution algorithms based on original single low-resolution images in related technologies.
[0004] The first aspect of this application provides a super-resolution image reconstruction method, applied to an image reconstruction model including a feature encoding network, a feature separation and aggregation network, and an image decoding network. The super-resolution image reconstruction method includes: inputting an APS image to be processed and a corresponding event image into the feature encoding network for feature extraction, outputting shared features; inputting the shared features into the feature separation module of the feature separation and aggregation network for feature separation, obtaining APS image shared features and event image shared features; inputting the APS image shared features and event image shared features into the feature aggregation module of the feature separation and aggregation network for aggregation, outputting aggregated features; and inputting the aggregated features into the image decoding network for feature decoding, outputting a super-resolution APS image corresponding to the APS image to be processed.
[0005] A second aspect of this application provides a super-resolution image reconstruction apparatus applied to an image reconstruction model including a feature encoding network, a feature separation and aggregation network, and an image decoding network. The super-resolution image reconstruction apparatus includes: an encoding module for inputting an APS image to be processed and a corresponding event image into the feature encoding network for feature extraction and outputting shared features; a separation module for inputting the shared features into the feature separation module of the feature separation and aggregation network for feature separation, obtaining APS image shared features and event image shared features; an aggregation module for inputting the APS image shared features and event image shared features into the feature aggregation module of the feature separation and aggregation network for aggregation, outputting aggregated features; and a decoding module for inputting the aggregated features into the image decoding network for feature decoding, outputting a super-resolution APS image corresponding to the APS image to be processed.
[0006] A third aspect of this application provides an electronic device, including a memory and a processor, wherein the processor is configured to execute a computer program stored in the memory, and when the processor executes the computer program, it implements the steps of the super-resolution image reconstruction method provided in the first aspect of this application.
[0007] The fourth aspect of this application provides a computer-readable storage medium storing a computer program thereon. When the computer program is executed by a processor, it implements the steps of the super-resolution image reconstruction method provided in the first aspect of this application.
[0008] As can be seen from the above, according to the super-resolution image reconstruction method, apparatus, device, and storage medium provided in this application, the APS image to be processed and the corresponding event image are input into a feature encoding network for feature extraction, and shared features are output. The shared features are then input into the feature separation module of a feature separation and aggregation network for feature separation, resulting in shared features of the APS image and shared features of the event image. These shared features are then input into the feature aggregation module of the feature separation and aggregation network for aggregation, and aggregated features are output. Finally, the aggregated features are input into an image decoding network for feature decoding, and a super-resolution APS image corresponding to the APS image to be processed is output. By implementing this application, combining event camera signal-guided super-resolution image reconstruction, and learning the shared features of the APS image and event image based on the feature separation and aggregation network, abnormal noise can be effectively suppressed, and the clarity of textures and edges in the super-resolution image reconstruction results can be improved. Attached Figure Description
[0009] Figure 1 A schematic diagram illustrating an application scenario provided in one embodiment of this application;
[0010] Figure 2This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;
[0011] Figure 3 A schematic diagram of the basic process of a super-resolution image reconstruction method provided in an embodiment of this application;
[0012] Figure 4 A network structure diagram of an encoder provided in one embodiment of this application;
[0013] Figure 5 This application provides a schematic diagram comparing images before and after reconstruction, as part of an embodiment of the present application.
[0014] Figure 6 A network structure diagram of an image decoding network provided in one embodiment of this application;
[0015] Figure 7 A network structure diagram of an image reconstruction model used in the training process is provided in one embodiment of this application;
[0016] Figure 8 A detailed flowchart illustrating a super-resolution image reconstruction method provided in an embodiment of this application;
[0017] Figure 9 This is a schematic diagram of the program modules of a super-resolution image reconstruction apparatus provided in an embodiment of this application. Detailed Implementation
[0018] To make the inventive objectives, features, and advantages of this application more apparent and understandable, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0019] In the description of the embodiments of this application, it should be understood that the terms "length", "width", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings. They are only for the convenience of describing the embodiments of this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limiting the present invention.
[0020] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0021] In the embodiments of this application, unless otherwise explicitly specified and limited, the terms "installation," "connection," "linking," "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components. For those skilled in the art, the specific meaning of the above terms in the embodiments of this application can be understood according to the specific circumstances.
[0022] The following will describe in detail, with reference to the accompanying drawings, a method, apparatus, device, and storage medium for super-resolution image reconstruction according to embodiments of this application.
[0023] To address the limitation in image super-resolution reconstruction results caused by super-resolution algorithms based on a single original low-resolution image in related technologies, this application provides an embodiment of a super-resolution image reconstruction method applicable to, for example,... Figure 1 The scenario shown may include an APS (Active-Pixel Sensor) camera 101, an EVS (Event-based Vision Sensor) camera 102, and an electronic device 103.
[0024] It is worth noting that active pixel sensors are a commonly used type of image sensor, in which each pixel sensor unit has a photodetector and at least one active transistor. In metal-oxide-semiconductor (MOS) active pixel sensors, MOS field-effect transistors (MOSFETs) are used as amplifiers. There are various types of APS, including the early NMOS type APS and the more common complementary MOS (CMOS) type APS. Event monitoring vision sensors are a new type of sensor that simulates the human retina and responds to pixel pulses caused by changes in brightness due to motion. Therefore, it can capture changes in scene brightness (i.e., changes in light intensity) at an extremely high frame rate, record events at specific times and specific locations in the image, forming an event stream instead of a frame stream. This can solve the problems of information redundancy, large data storage, and large real-time processing requirements of traditional cameras.
[0025] It should be noted that the APS image sensor and EVS image sensor in this embodiment can be discrete image sensors or integrated image sensors. For integrated image sensors, the overall photosensitive area of the integrated image sensor is divided into multiple sub-photosensitive areas, and the pixel arrays of the multiple sub-photosensitive areas correspond to the APS data mode and the EVS data mode, respectively. Compared with multiple discrete sensor modules, this effectively reduces the device size of the sensor module and is more conducive to the miniaturization of the overall hardware architecture.
[0026] In addition, electronic device 103 is a variety of terminal devices with data processing capabilities, including but not limited to smartphones, tablets, laptops, and desktop computers.
[0027] exist Figure 1 In the application scenario shown, an APS image to be processed (i.e., a low-resolution APS image) can be acquired by an APS image sensor 101, and a corresponding event image can be acquired simultaneously by an EVS image sensor 102. Both sensors then send their respective acquired images to an electronic device 103. The electronic device 103 performs the following super-resolution image reconstruction method on the received APS image to be processed and the event image: First, the APS image to be processed and the corresponding event image are input into a feature encoding network for feature extraction, outputting shared features. Then, the shared features are input into the feature separation module of a feature separation and aggregation network for feature separation, obtaining shared features of the APS image and shared features of the event image. These shared features are then input into the feature aggregation module of the feature separation and aggregation network for aggregation, outputting aggregated features. Finally, the aggregated features are input into an image decoding network for feature decoding, outputting a super-resolution APS image corresponding to the APS image to be processed.
[0028] like Figure 2 The diagram shown is a schematic representation of an electronic device according to an embodiment of this application. The electronic device mainly includes a memory 201 and a processor 202. The number of processors 202 can be one or more. The memory 201 stores a computer program 203 that can run on the processor 202. The memory 201 and the processor 202 are communicatively connected. When the processor 202 executes the computer program 203, it implements the aforementioned super-resolution image reconstruction method.
[0029] It should be noted that memory 201 can be high-speed random access memory (RAM) or non-volatile memory, such as disk storage. Memory 201 is used to store executable program code, and processor 202 is coupled to memory 201.
[0030] One embodiment of this application also provides a computer-readable storage medium, which may be disposed in the aforementioned electronic device. The computer-readable storage medium may be as described above. Figure 2 The memory in the illustrated embodiment.
[0031] The computer-readable storage medium stores a computer program that, when executed by a processor, implements the aforementioned super-resolution image reconstruction method. Furthermore, the computer-readable storage medium can also be a USB flash drive, external hard drive, read-only memory (ROM), RAM, magnetic disk, or optical disk, or any other medium capable of storing program code.
[0032] like Figure 3 This is a basic flowchart of a super-resolution image reconstruction method provided in an embodiment of this application. This super-resolution image reconstruction method is applied to an image reconstruction model including a feature encoding network, a feature separation and aggregation network, and an image decoding network, and can be derived from... Figure 1 or Figure 2 The electronic device in the process executes the following steps:
[0033] Step 301: Input the APS image to be processed and the corresponding event image into the feature encoding network for feature extraction and output the shared features.
[0034] In the image reconstruction process of this embodiment, the low-resolution APS image to be processed is I lr and the corresponding event image E as the input to the model, where I lr E is an H×W×1 matrix, and E is an H×W×M matrix.
[0035] In one optional embodiment of this example, the feature coding network includes a feature extraction module and a shared encoder. The shared encoder includes a frequency domain module, which includes a global link branch, a high-frequency feature extraction branch, and a low-frequency feature extraction branch.
[0036] First, in this embodiment, the APS image to be processed and the corresponding event image are respectively input into the corresponding feature extraction module in the feature coding network for feature extraction to obtain intermediate APS image features and intermediate event image features. Then, the stitched features obtained by stitching the intermediate APS image features and intermediate event image features are input into the shared encoder for feature extraction, and the extracted different frequency domain features are aggregated to output shared features.
[0037] Specifically, the feature extraction module in this embodiment includes two convolutional modules corresponding to the APS image and the event image, respectively, and the low-resolution image I to be processed... lrThe corresponding event signal E is passed through the corresponding convolution module C. i and C e After processing, the intermediate APS image features F are obtained. i,0 and intermediate event image features F e,0 Then, the two are concatenated (i.e., concat) to obtain the concatenated feature F. ie,0 Next, we will splice feature F. ie,0 The input is fed into the ShareEncoderε to obtain the shared features F between the APS image to be processed and the corresponding event image. ie, s.
[0038] like Figure 4 The diagram shows a network structure of an encoder provided in this embodiment. This network structure preferably includes multiple frequency domain modules (Freq Blocks, FBs), which are cascaded and converted to a higher channel number using Conv1 before output. Each FB has three branches: a global connection branch, a high-frequency feature extraction branch, and a low-frequency feature extraction branch. These three branches model different frequency domains of the input features and sum them before output to obtain the output features. The theoretical HF operator is used in the High Frequency Branch for high-frequency feature extraction, and the Down Sample operator is used in the Low Frequency Branch for low-frequency feature extraction. The HF operator is represented as follows:
[0039] F h =F-↑ nearest (↓ average (F))
[0040] Among them, ↑ nearest Indicates the nearest neighbor upsample, ↓ average This represents the downsampled value of Average (DownSample).
[0041] Step 302: Input the shared features into the feature separation module of the feature separation and aggregation network to perform feature separation, and obtain the APS image shared features and the event image shared features.
[0042] Step 303: Input the APS image shared features and event image shared features into the feature aggregation module of the feature separation and aggregation network for aggregation, and output the aggregated features.
[0043] Specifically, the feature separation and aggregation network in this embodiment includes a cascaded feature separation module and a feature aggregation module, which will share features Fie,s After splitting (i.e., splitting), the image shared feature F is obtained. i,s Event Shared Feature (F) e,s Next, let's talk about F. i,s F e,s After performing aggregation (i.e., Add) processing, the aggregated feature F is obtained. ie,d .
[0044] Step 304: Input the aggregated features into the image decoding network for feature decoding, and output the super-resolution APS image corresponding to the APS image to be processed.
[0045] Specifically, the purpose of image decoding is to use a decoder to decode the aggregated features obtained after feature separation and aggregation, thereby obtaining the super-resolution reconstructed image. For example... Figure 5 The diagram shows a comparison of images before and after reconstruction provided in this embodiment. (a) is a low-resolution APS image, (b) is a super-resolution APS image generated by the algorithm provided in this embodiment, and (c) is the ground truth of the high-resolution APS image. It can be seen that the super-resolution image reconstruction algorithm provided in this embodiment can effectively ensure the image enhancement effect, and the texture and edges of the finally generated reconstructed image can be fully guaranteed to be clear.
[0046] like Figure 6 The diagram shown is a network structure diagram of an image decoding network provided in this embodiment. In an optional implementation of this embodiment, the image decoding network includes a feature-to-vector transformation module (M2V, Feature Map to Vector) and a self-attention module (e.g., [example missing]). Figure 6 The system includes the Transformer module, the Vector-Feature Transformation module (V2M, Feature VectorToMap), and the Depth to Space+ module.
[0047] First, the aggregated features are input to the feature-vector transformation module of the image decoding network to convert the two-dimensional feature map into the corresponding feature vector. Then, the feature vector is input to the self-attention module for decoding and outputs the result vector. Next, the result vector is input to the vector-feature transformation module to convert it into the corresponding decoded features. Finally, the decoded features are input to the depth space feature transformation module for image reconstruction and output the super-resolution APS image corresponding to the APS image to be processed.
[0048] Please refer to it again. Figure 6 The self-attention module includes the first normalization module (i.e., Figure 6The first LayerNorm), attention module (Multi-Head Attention), and second normalization module (i.e. Figure 6 The latter is a LayerNorm and a multi-layer perceptron (MLP).
[0049] In the process of modeling global features through feature decoding in Transformer, firstly, the feature vector is input to the first normalization module of the self-attention module for normalization processing, and the first normalized feature vector is output. Then, the first normalized feature vector is input to the attention module to obtain the mask corresponding to the feature vector. Further, the feature vector and the mask are fused and input to the second normalization module for normalization processing, and the second normalized feature vector is output. Finally, the second normalized feature vector is input to the multilayer perception module for classification and regression, and the result vector is output.
[0050] Specifically, in this embodiment, the first normalization module calculates the mean and variance of the feature vector to normalize the feature map. According to the normalization principle, the weight of each pixel in the feature map can ensure that the original distribution of the data is centrally symmetric. Next, the normalized feature vector is input into the attention module for weight optimization to obtain the corresponding mask of the input feature vector, extracting more effective features. Then, the mask is fused with the input feature vector to retain the lower-level features. Further, the fused feature vector is normalized and input into the multilayer perception module for classification and regression, outputting the result vector.
[0051] like Figure 7 The diagram shown illustrates the network structure of an image reconstruction model used during training in this embodiment. Feature Encoding represents the feature encoding network, Event and Image Feature Disentanglement represents the feature separation and aggregation network, and SR Image Reconstruction represents the image decoding network. LR IntensityImage represents a low-resolution APS image sample I. lr Events represents the corresponding event image samples, and Image Encoder represents the APS private encoder used during the model training phase to extract APS image private features (F... i,p ImageSpecific Feature (ESI), Event Encoder represents the EVS proprietary encoder used during the model training phase to extract event image proprietary features (ESI). e,pIt should be understood that the APS private encoder, EVS private encoder and their respective extracted private features in the feature encoding network of this embodiment are only used during the model training stage, and are not required to be used in the image super-resolution reconstruction process after the model training is completed. The encoder required in the actual image super-resolution reconstruction process only includes the shared encoder. In addition, the network structure used by the APS private encoder and EVS private encoder in this embodiment can be implemented using the same network structure as the shared encoder mentioned above.
[0052] In one optional implementation of this embodiment, the overall loss function of the above image reconstruction model is expressed as:
[0053]
[0054] in, Represents the overall loss function. I represents the loss function of the image decoding network. sr I represents the super-resolution APS image output by the image reconstruction model during model training. hr This represents high-resolution APS image samples, that is... Figure 7 HRIntensity Image Let λ represent the loss function of the feature aggregation module. s F represents the loss weight of the feature aggregation module. i,s F represents the shared features of APS images corresponding to low-resolution APS image samples. e,s This indicates the shared features of the event images corresponding to the event image samples. Let λ represent the loss function of the feature separation module. d F represents the loss weight of the feature separation module. i,p F represents the APS image private features extracted by the APS private encoder from the intermediate APS image features during the model training phase. e,p This represents the private features of the event images extracted by the EVS private encoder during the intermediate event image features of the model training phase.
[0055] Furthermore, the loss function of the feature separation module is expressed as:
[0056]
[0057] in, Indicates transpose, ||·|| FThis represents the Frobenius norm. In this embodiment, this loss function is used to guide the network to learn the shared and private features of the APS image and the event image, which can effectively suppress noise and other information in the event that is detrimental to multimodal fusion.
[0058] The loss function of the feature aggregation module is expressed as:
[0059]
[0060] Where n and m represent the number of eigenvectors, i = 1, 2, ..., n and j = 1, 2, ..., m, φ(·) represents the mapping function, f i,s f represents the shared features of each channel in an APS image. e,s This represents the shared features of each channel in the event image, ||·|| H Represents the Hilbert space function of the regenerating kernel
[0061] It is worth noting that the feature separation and aggregation network is used to process the F-value of the feature encoding output. i,p F i,s F e,s F e,p Separation and aggregation are performed. Specifically, the feature separation loss function used in this embodiment... Used to increase F i,s F i,p and F e, s F e,p The distance between them, and the feature similarity loss function used. Used to reduce F i,s F e,s The distance between them.
[0062] Figure 8 The method described in this application is a refined super-resolution image reconstruction method provided in an embodiment of the present application. It is applied to an image reconstruction model including a feature encoding network, a feature separation and aggregation network, and an image decoding network. The feature encoding network includes a feature extraction module and a shared encoder. The shared encoder includes a frequency domain module, which includes a global link branch, a high-frequency feature extraction branch, and a low-frequency feature extraction branch. The feature separation and aggregation network includes a feature separation module and a feature aggregation module. The image decoding network includes a feature-vector transformation module, a self-attention module, a vector-feature transformation module, and a depth space feature transformation module.
[0063] Specifically, the implementation process of this super-resolution image reconstruction method includes the following steps:
[0064] Step 801: Input the APS image to be processed and the corresponding event image into the corresponding feature extraction module in the feature coding network to extract features, and obtain the intermediate APS image features and the intermediate event image features.
[0065] Step 802: The stitched features obtained by stitching the intermediate APS image features and the intermediate event image features are input into the frequency domain module of the shared encoder for feature extraction, and the extracted different frequency domain features are aggregated to output the shared features.
[0066] Step 803: Input the shared features into the feature separation module of the feature separation and aggregation network to perform feature separation, and obtain the APS image shared features and the event image shared features;
[0067] Step 804: Input the APS image shared features and event image shared features into the feature aggregation module of the feature separation and aggregation network for aggregation, and output the aggregated features;
[0068] Step 805: Input the aggregated features into the feature-vector conversion module of the image decoding network to convert them into corresponding feature vectors;
[0069] Step 806: Input the feature vector into the self-attention module for decoding and output the result vector;
[0070] Step 807: Input the result vector into the vector-feature conversion module to convert it into the corresponding decoded features;
[0071] Step 808: Input the decoded features into the depth space feature transformation module for image reconstruction, and output a super-resolution APS image corresponding to the APS image to be processed.
[0072] It should be understood that the sequence number of each step in this embodiment does not imply the order in which the steps are executed. The execution order of each step should be determined by its function and internal logic, and should not constitute a unique limitation on the implementation process of this application embodiment.
[0073] Figure 9 This application provides a super-resolution image reconstruction apparatus according to an embodiment. This super-resolution image reconstruction apparatus can be used to implement the super-resolution image reconstruction method in the foregoing embodiments. The super-resolution image reconstruction apparatus mainly includes:
[0074] Encoding module 901 is used to input the APS image to be processed and the corresponding event image into the feature encoding network for feature extraction and output shared features;
[0075] The separation module 902 is used to input the shared features into the feature separation module of the feature separation and aggregation network for feature separation to obtain APS image shared features and event image shared features.
[0076] The aggregation module 903 is used to input the shared features of APS images and shared features of event images into the feature aggregation module of the feature separation and aggregation network for aggregation, and output aggregated features;
[0077] The decoding module 904 is used to input the aggregated features into the image decoding network for feature decoding and output a super-resolution APS image corresponding to the APS image to be processed.
[0078] In some embodiments of this example, the feature coding network includes a feature extraction module and a shared encoder. The shared encoder includes a frequency domain module, which includes a global link branch, a high-frequency feature extraction branch, and a low-frequency feature extraction branch. Accordingly, the coding module is specifically used to: input the APS image to be processed and the corresponding event image into the corresponding feature extraction module in the feature coding network for feature extraction, obtaining intermediate APS image features and intermediate event image features; input the concatenated features obtained by concatenating the intermediate APS image features and intermediate event image features into the shared encoder for feature extraction, and aggregate the extracted different frequency domain features to output shared features.
[0079] In some embodiments of this example, the image decoding network includes a feature-vector conversion module, a self-attention module, a vector-feature conversion module, and a depth space feature transformation module. Specifically, the decoding module is used to: input aggregated features to the feature-vector conversion module of the image decoding network, converting them into corresponding feature vectors; input the feature vectors to the self-attention module for decoding, outputting a result vector; input the result vector to the vector-feature conversion module, converting it into corresponding decoded features; and input the decoded features to the depth space feature transformation module for image reconstruction, outputting a super-resolution APS image corresponding to the APS image to be processed.
[0080] Furthermore, in some embodiments of this example, the self-attention module includes a first normalization module, an attention module, a second normalization module, and a multilayer perception module. Correspondingly, when the decoding module performs the function of inputting the feature vector into the self-attention module for decoding and outputting a result vector, it is specifically used for: inputting the feature vector into the first normalization module of the self-attention module for normalization processing, and outputting a first normalized feature vector; inputting the first normalized feature vector into the attention module to obtain the mask corresponding to the feature vector; fusing the feature vector and the mask and inputting it into the second normalization module for normalization processing, and outputting a second normalized feature vector; inputting the second normalized feature vector into the multilayer perception module for classification and regression, and outputting a result vector.
[0081] It should be noted that the super-resolution image reconstruction methods in the foregoing embodiments can all be implemented based on the super-resolution image reconstruction device provided in this embodiment. Those skilled in the art can clearly understand that, for the sake of convenience and brevity, the specific working process of the super-resolution image reconstruction device described in this embodiment can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0082] Based on the technical solution of the above embodiments of this application, the APS image to be processed and the corresponding event image are input into a feature encoding network for feature extraction, and shared features are output. The shared features are input into the feature separation module of a feature separation and aggregation network for feature separation, obtaining shared features of the APS image and shared features of the event image. The shared features of the APS image and the shared features of the event image are input into the feature aggregation module of the feature separation and aggregation network for aggregation, and aggregated features are output. The aggregated features are input into an image decoding network for feature decoding, and a super-resolution APS image corresponding to the APS image to be processed is output. By implementing the solution of this application, combining event camera signal-guided super-resolution image reconstruction, and learning the shared features of the APS image and event image based on the feature separation and aggregation network, abnormal noise can be effectively suppressed, and the clarity of texture and edges in the super-resolution image reconstruction result can be improved.
[0083] It should be noted that the apparatuses and methods disclosed in the several embodiments provided in this application can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.
[0084] The modules described as separate components may or may not be physically separate. Similarly, the components shown as modules may or may not be physical modules; they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0085] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated modules described above can be implemented in hardware or as software functional modules.
[0086] If the integrated module is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned readable storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.
[0087] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0088] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0089] The above is a description of the super-resolution image reconstruction method, apparatus, device and storage medium provided in this application. For those skilled in the art, based on the ideas of the embodiments of this application, there will be changes in the specific implementation and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A super-resolution image reconstruction method, characterized in that, This method is applied to an image reconstruction model that includes a feature coding network, a feature separation and aggregation network, and an image decoding network. The feature coding network includes a feature extraction module and a shared encoder. The shared encoder includes a frequency domain module, which includes a global link branch, a high-frequency feature extraction branch, and a low-frequency feature extraction branch. The super-resolution image reconstruction method includes: The APS image to be processed and the corresponding event image are input into the feature encoding network for feature extraction, and shared features are output. The shared features are input into the feature separation module of the feature separation and aggregation network for feature separation to obtain APS image shared features and event image shared features. The APS image sharing features and event image sharing features are input into the feature aggregation module of the feature separation and aggregation network for aggregation, and the aggregated features are output. The aggregated features are input into the image decoding network for feature decoding, and a super-resolution APS image corresponding to the APS image to be processed is output. The step of inputting the APS image to be processed and the corresponding event image into the shared encoder of the feature coding network for feature extraction and outputting shared features includes: The APS image to be processed and the corresponding event image are respectively input into the corresponding feature extraction module in the feature coding network for feature extraction to obtain intermediate APS image features and intermediate event image features; wherein, the APS image is an image acquired by an active pixel sensor; The stitched features obtained by stitching the intermediate APS image features and intermediate event image features are input into the shared encoder for feature extraction, and the extracted different frequency domain features are aggregated to output shared features.
2. The super-resolution image reconstruction method according to claim 1, characterized in that, The image decoding network includes a feature-vector transformation module, a self-attention module, a vector-feature transformation module, and a depth space feature transformation module; The step of inputting the aggregated features into the image decoding network for feature decoding and outputting a super-resolution APS image corresponding to the APS image to be processed includes: The aggregated features are input into the feature-vector conversion module of the image decoding network and converted into corresponding feature vectors; The feature vector is input into the self-attention module for decoding, and a result vector is output. The result vector is input into the vector-feature conversion module and converted into corresponding decoded features; The decoded features are input into the depth space feature transformation module for image reconstruction, and a super-resolution APS image corresponding to the APS image to be processed is output.
3. The super-resolution image reconstruction method according to claim 2, characterized in that, The self-attention module includes a first normalization module, an attention module, a second normalization module, and a multilayer perception module; The step of inputting the feature vector into the self-attention module for decoding and outputting a result vector includes: The feature vector is input into the first normalization module of the self-attention module for normalization processing, and the first normalized feature vector is output. The first normalized feature vector is input into the attention module to obtain the mask corresponding to the feature vector; The feature vector is fused with the mask and then input into the second normalization module for normalization processing, and the second normalized feature vector is output. The second normalized feature vector is input into the multilayer perception module for classification and regression, and the result vector is output.
4. The super-resolution image reconstruction method according to claim 1, characterized in that, The overall loss function of the image reconstruction model is expressed as: in, This represents the overall loss function. This represents the loss function of the image decoding network. This refers to the super-resolution APS image output by the image reconstruction model during model training. Represents high-resolution APS image samples. This represents the loss function of the feature aggregation module. This represents the loss weight of the feature aggregation module. This represents the APS image shared features corresponding to low-resolution APS image samples. The shared features of the event images corresponding to the event image samples are represented. This represents the loss function of the feature separation module. This represents the loss weight of the feature separation module. This refers to the APS image private features extracted by the APS private encoder from the intermediate APS image features during the model training phase. This refers to the private features of the event images extracted by the EVS private encoder during the model training phase, representing the intermediate event image features.
5. The super-resolution image reconstruction method according to claim 4, characterized in that, The loss function of the feature separation module is expressed as: in, Indicates transpose. This represents the Frobenius norm.
6. The super-resolution image reconstruction method according to claim 4, characterized in that, The loss function of the feature aggregation module is expressed as: in, n and m Indicates the number of eigenvectors. i =1,2,…, n as well as j =1,2,…, m , Represents a mapping function. This represents the shared features of each channel in the APS image. This indicates that the event image shares features in each channel. This represents the regenerating kernel Hilbert space function.
7. A super-resolution image reconstruction apparatus, characterized in that, This method is applied to an image reconstruction model that includes a feature coding network, a feature separation and aggregation network, and an image decoding network. The feature coding network includes a feature extraction module and a shared encoder. The shared encoder includes a frequency domain module, which includes a global link branch, a high-frequency feature extraction branch, and a low-frequency feature extraction branch. The super-resolution image reconstruction apparatus includes: The encoding module is used to input the APS image to be processed and the corresponding event image into the feature encoding network for feature extraction and output shared features; The separation module is used to input the shared features into the feature separation module of the feature separation and aggregation network for feature separation to obtain APS image shared features and event image shared features. The aggregation module is used to input the APS image shared features and event image shared features into the feature aggregation module of the feature separation and aggregation network for aggregation, and output aggregated features; The decoding module is used to input the aggregated features into the image decoding network for feature decoding and output a super-resolution APS image corresponding to the APS image to be processed. The encoding module is specifically used for: inputting the APS image to be processed and the corresponding event image into the corresponding feature extraction module in the feature encoding network for feature extraction to obtain intermediate APS image features and intermediate event image features; wherein, the APS image is an image acquired by an active pixel sensor; inputting the stitched features obtained by stitching the intermediate APS image features and intermediate event image features into the shared encoder for feature extraction, and aggregating the extracted different frequency domain features to output shared features.
8. An electronic device, characterized in that, Includes memory and processor, of which: The processor is used to execute computer programs stored in the memory; When the processor executes the computer program, it implements the steps in the super-resolution image reconstruction method according to any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps in the super-resolution image reconstruction method according to any one of claims 1 to 6.
Citation Information
Patent Citations
High-quality and high-frame-rate image reconstruction method based on event camera
CN111667442A
Heart MRI image multi-task segmentation method based on multi-modal complementary information exploration
CN113129316A