A high dynamic range imaging system and imaging method based on the unity of image block level and pixel level

Through a high dynamic range imaging system with unified image block level and pixel level, the content completion network and deformable fusion network are used to solve the artifact problem in dynamic scenarios and achieve high-quality high dynamic range image reconstruction.

CN115760647BActive Publication Date: 2025-07-22XIDIAN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211583428.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-10
Publication Date
2025-07-22
Estimated Expiration
2042-12-10

AI Technical Summary

Technical Problem

Existing high dynamic range imaging technologies are prone to artifacts when processing dynamic scenes, especially in the motion and saturated areas recovery effects, and traditional convolutional neural networks cannot adaptively adjust the fusion weights of different areas.

Method used

The content completion network and the transformer-based deformable large-scale fusion network are adopted. The internal information of the motion and saturated areas is restored through the image block collection module and the attention mechanism module, and dynamic cross-fusion is realized through the gate module. Combined with the residual deformable transformer module, the weights of different exposure areas are adaptively adjusted.

Benefits of technology

Effectively restore high dynamic range images in dynamic scenes, remove artifacts and improve image quality, especially in the detailed recovery effect of moving and saturated areas, reducing the computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115760647B_ABST
    Figure CN115760647B_ABST
Patent Text Reader

Abstract

A high dynamic range imaging system and imaging method based on the unification of image patch level and pixel level. The system includes a content completion network and a deformable large-scale fusion network based on a transformer. The method is as follows: the content completion network uses a module based on an image patch set to collect features with similar textures in a large range to fill the artifact regions. At the same time, in order to reduce the risk of reconstruction deviation, a gating module is used to leverage the advantages of the module based on the image patch set and the attention mechanism module, and fuse the effective information inside and at the edges of the motion and saturation regions. The adaptive deformable fusion network based on the transformer fuses the features of large motions through a residual deformable transformer module, and can also dynamically adjust the weights of different regions to adaptively fuse the information of different exposure regions. Finally, a high-quality high dynamic range image is reconstructed, effectively removing artifacts and improving the quality of the reconstructed image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of high dynamic range imaging, and particularly relates to a high dynamic range imaging system and an imaging method based on the unification of image block level and pixel level. Background Art

[0002] The dynamic range refers to the span between the maximum brightness and the minimum brightness in a scene. The illumination dynamic range of natural scenes is very wide. Through millions of years of evolution, the human eye can adaptively adjust to a wide span of illumination dynamic ranges, so as to perceive the details of objects with different brightness levels in the scene. However, the sensors of ordinary cameras can only measure a limited illumination dynamic range. The low dynamic range pictures captured by ordinary cameras usually have overexposed or underexposed areas, and such areas will seriously affect the visual effect due to the lack of details. The high dynamic range imaging technology emerges as the times require to solve these defects, and it can make the areas with different brightness levels in the picture present rich detail information. Therefore, it has been widely applied in many fields, such as the medical field, autonomous driving, 3D reconstruction, etc. In order to obtain a high dynamic range picture, the most common method is to use a series of low dynamic range pictures with different exposure times for software synthesis. This method can only synthesize a high-quality high dynamic range picture when both the scene and the camera are stationary, but artifacts will be generated when the scene moves or the camera moves.

[0003] Some methods have been proposed to solve the artifact problem of high dynamic range imaging in dynamic scenes, and they are mainly: alignment-based methods and rejection-based methods. The alignment-based methods are mainly divided into: rigid alignment and non-rigid alignment. Rigid alignment uses homography transformation to correct the globally misaligned areas, but it cannot handle complex foreground movements. In order to detect foreground movements, non-rigid alignment-based methods (such as: optical flow method) adopt a more detailed alignment method for alignment. However, due to occlusion problems, it is difficult to accurately align the moving areas, so this method will generate artifacts due to incorrect motion matching. Therefore, rejection-based methods are proposed. They first detect the moving areas and then discard the misaligned moving areas. However, due to inaccurate motion detection methods, useful information will be lost, resulting in a poor high dynamic range reconstruction result in the end.

[0004] In recent years, with the rise of deep learning, many studies have used convolutional neural networks to directly learn the mapping from low dynamic range to high dynamic range end-to-end. These models generally follow the paradigm of alignment first and then fusion. First, optical flow or homography transformation is used for pre-alignment, and then a convolutional neural network is used to reconstruct high dynamic range images. However, artifacts are inevitably generated due to the estimation errors of optical flow or homography transformation. Subsequently, methods based on the attention mechanism have been proposed to solve these problems. They first automatically generate an attention map and then multiply the features of the non-reference image pixel by pixel, which enables the network to automatically remove moving and saturated regions according to the weights of the attention, while highlighting the information in more effective regions, thus achieving the purpose of artifact removal and desaturation.

[0005] For moving or saturated regions, due to the obvious changes between the reference frame and the non-reference frame, the attention map generated by the method based on the attention mechanism can naturally learn effective weights to remove the contaminated regions in a pixel-by-pixel multiplication calculation manner. Therefore, even if the input images are not aligned, the attention mechanism can handle the moving edge regions well. However, when both saturation and motion exist, the attention mechanism will produce unsatisfactory results. Specifically, due to the pixel-by-pixel multiplication calculation method, when the reference frame is saturated (unable to utilize information) and does not aggregate the information of similar regions in the non-reference frame, the attention mechanism forces the use of non-saturated information in the non-reference frame (which will produce artifacts if there is motion), so it cannot well recover the information inside the regions occluded by motion and saturated, and the remaining artifacts are relatively obvious. In addition, how to fuse the features of multiple frames of images is particularly important for artifact removal in high dynamic range images. Previous high dynamic range reconstruction networks are all based on convolutional neural networks. Since the weights of convolutional neural networks are static, the attention to overexposed, well-exposed, and underexposed regions is the same, and the network cannot adaptively adjust the fusion regions according to different input images. Moreover, the receptive field of convolutional neural networks is limited, and they can only fuse local features, which is not conducive to the recovery of large motion regions and large-scale saturated regions. Summary of the Invention

[0006] In order to overcome the above deficiencies of the prior art, the purpose of the present invention is to propose a high dynamic range imaging system and its imaging method unified at the image block level and pixel level. Through a content completion network, a feature aggregation module based on image blocks is used to aggregate features with similar textures in a large range to fill the artifact regions; a gating module is also used to enable the attention mechanism module and the feature aggregation module based on image blocks to complement each other, so that the inside of the moving and saturated regions can be better recovered, while making the edges of the moving and saturated regions sharp and clear; an adaptive deformable fusion network based on a transformer is used to fuse the features of large motions through a residual deformable transformer module, and can also dynamically adjust the weights of different regions, adaptively fuse the information of different exposure regions, and finally reconstruct a high-quality high dynamic range image.

[0007] To achieve the above object, the present invention adopts the following technical solutions:

[0008] A high dynamic range imaging system based on the unification of image patch level and pixel level, including a content completion network and a deformable large-range fusion network based on a transformer,

[0009] The content completion network includes an image patch set module, an attention mechanism module, and a gating module;

[0010] Among them, the image patch set module collects similar image patches around the area to be processed, so as to restore the texture information inside the motion and saturation areas;

[0011] The attention mechanism module is used to suppress non-aligned and high-exposure or low-exposure information point by point, so as to achieve the purpose of removing artifacts and desaturating, and making the edges of the motion and saturation areas sharp and clear;

[0012] The gating module is used to perform dynamic cross-pixel multiplication on the output features of the image patch set module and the attention mechanism module, so that the inside of the motion and saturation areas can be better restored, and at the same time, the edges of the motion and saturation areas are sharp and clear;

[0013] The deformable large-range fusion network based on the transformer mainly includes a residual deformable transformer module;

[0014] The residual deformable transformer module is used to dynamically adjust the weights of different exposure areas, model long-range dependencies, fuse the image features of large-range different exposure areas, and effectively restore the information of large-range motion and saturation areas.

[0015] A high dynamic range imaging method based on the unification of image patch level and pixel level specifically includes the following steps:

[0016] Step 1, Process by the content completion network

[0017] 1) For the attention mechanism module, use the convolutional layer to automatically generate the attention weights between the non-reference frame image patches and the reference frame image patches, and then multiply the attention weights with the non-reference frame image patches pixel by pixel, so that the non-reference frame image patches highlight the well-exposed areas in the image according to the size of the attention weights, and at the same time can suppress the motion and saturation areas in the non-reference frame image patches;

[0018] 2) For the image patch set module, using matrix multiplication, calculate the inner product of the feature matrix of the non-reference frame image patches and the transpose of the feature matrix of the reference frame patches, i.e., the similarity matrix. Subsequently, multiply the non-reference frame image patches with the similarity matrix so that the non-reference frame image patches selectively aggregate similar texture information in the non-reference frame image patches according to the similarity with the reference frame image patches, which can fill the information inside the motion and saturation regions;

[0019] a) For the motion region, based on the image patch set module, according to the similarity weight between the reference frame image patch and the non-reference frame image patch features, perform weighted summation on the non-aligned regions in the non-reference frame image patch features, aggregate the texture features similar to the reference frame image patch features in the non-reference frame image patch features, and fill the motion non-aligned regions of the non-reference frame image patches to achieve the purpose of aligning with the reference frame image patches;

[0020] b) For the saturation region, based on the image patch set module, perform weighted summation on the features around the saturation region in the non-reference frame image patches according to the similarity weight, aggregate the similar texture features, and make the pixel values closer to the illumination of the real scene;

[0021] 3) For the problem that the image features after the first step of step 1 cannot effectively fill the information in the motion and saturation regions, and for the problem that there is a risk of reconstruction offset due to other background information introduced by the image features after the second step of step 1, use a gating module to predict the dynamic weights of the output features of the attention mechanism module and the image patch set module respectively using convolutional layers, and then multiply these dynamic weights cross-pixel by pixel with the corresponding features. In this way, the cross-pixel multiplication of the dynamic weights can automatically interact and fuse the features of the attention mechanism module and the image patch set module.

[0022] Step 2. Process the deformable large-range fusion network based on the transformer

[0023] For the deformable large-range fusion network based on the transformer, input the features of the content completion network into the residual deformable transformer module. This module uses multiple moving window transformer layers as the basic layer and an attention mechanism layer based on the deformable window to fuse long-range information; specifically:

[0024] 1) The moving window transformer layer first divides the image into non-overlapping windows, then maps the image patches within the non-overlapping windows into query features, key features, and value features through three fully connected layers. Then, through the self-attention mechanism layer, calculate the inner product between the query features and the key features within and between the non-overlapping windows to obtain the similarity matrix. Finally, fuse the features of different exposure and motion regions by multiplying the similarity matrix and the value features through matrix multiplication;

[0025] 2) The attention mechanism layer based on deformable windows first divides the input features into non-overlapping windows, then maps the image patches within the non-overlapping windows into query features through a fully connected layer, predicts the offset of the query features through an offset network, adds the query features and the offset to obtain the offset query features, passes the offset query features through two fully connected layers to map them into key features and value features, and then calculates the inner product of the offset query features and the key features within and between the non-overlapping windows through the self-attention mechanism to obtain a similarity matrix. Finally, the similarity matrix and the value features are dynamically fused using matrix multiplication; to avoid the limitation of the modeling ability for long-range relationships and reduce the computational complexity at the same time.

[0026] Compared with the prior art, the present invention has the following advantages:

[0027] For the motion area, the image patch set module of the present invention can utilize the texture features of a relatively large range of similar areas to fill the information inside the motion area, so it can better restore the inside of the artifact area. For the saturation area, instead of suppressing the saturation area like the attention mechanism, the image patch set module performs weighted fusion on the areas around the saturation, making the pixel values closer to the illumination of the real scene, and at the same time can refer to the information of a larger receptive field, so it can more easily restore the saturation area. Moreover, the gating module invented by us can make the advantages of the attention mechanism module and the image patch set module complementary, enabling the inside of the motion and saturation areas to be better restored, while making the edges of the motion and saturation areas sharp and clear, which could not be achieved by previous technologies.

[0028] At the same time, the deformable large-range fusion network based on transformers of the present invention utilizes the information of dynamic weights and global receptive fields, and can adaptively fuse and restore large-range motion and saturation areas. However, previous high-dynamic-range fusion networks only used convolutional neural networks and could only fuse information in a small range with fixed weights, and their performance was poor in complex motion and exposure scenarios.

[0029] The attention mechanism based on deformable windows in the present invention has a larger receptive field because it can predict the offset of features. At the same time, since the offset of the features is automatically learned through the offset network, it can dynamically adjust the areas of attention between windows, which is more conducive to information fusion, and thus more conducive to synthesizing a high-dynamic-range image without artifacts and saturation. The attention mechanism based on deformable windows calculates the similarity of image patches within and between windows using windows. Compared with previous transformer networks, it greatly reduces the computational complexity while maintaining good performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 It is an overview diagram of the model of the present invention.

[0031] Figure 2 This is the schematic diagram of the attention mechanism module of the present invention.

[0032] Figure 3 This is the schematic diagram of the module based on the set of image patches of the present invention.

[0033] Figure 4 This is the schematic diagram of the gating module of the present invention.

[0034] Figure 5 This is the schematic diagram of the attention mechanism layer based on the deformable window of the present invention Detailed implementation manners

[0035] The present invention will be further described in detail below with reference to the accompanying drawings.

[0036] A high-dynamic-range imaging system based on the unification of image patch level and pixel level, including a content completion network and a deformable large-range fusion network based on a transformer,

[0037] The content completion network includes a module based on a set of image patches, an attention mechanism module, and a gating module;

[0038] Among them, the module based on the set of image patches aggregates similar image patches around the area to be processed, so as to restore the texture information inside the motion and saturation areas;

[0039] The attention mechanism module is used to suppress non-aligned and high-exposure or low-exposure information point by point, so as to achieve the purpose of removing artifacts and desaturating and making the edges of the motion and saturation areas sharp and clear;

[0040] The gating module is used to perform dynamic cross-pixel multiplication on the output features of the module based on the set of image patches and the attention mechanism module, so that the inside of the motion and saturation areas can be better restored, and at the same time, the edges of the motion and saturation areas are sharp and clear;

[0041] The deformable large-range fusion network based on a transformer mainly includes a residual deformable transformer module;

[0042] The residual deformable transformer module is used to dynamically adjust the weights of different exposure areas, model long-range dependencies, fuse the image features of different exposure areas in a large range, and effectively restore the information of large-range motion and saturation areas.

[0043] A high-dynamic-range imaging method based on the unification of image patch level and pixel level specifically includes the following steps:

[0044] Step 1, Process by the content completion network

[0045] 1) For the attention mechanism module, a convolutional layer is used to automatically generate the attention weights between the non-reference frame image patches and the reference frame image patches. Subsequently, the attention weights are multiplied pixel by pixel with the non-reference frame image patches, enabling the non-reference frame image patches to highlight the well-exposed regions in the image according to the magnitude of the attention weights, while suppressing the moving and saturated regions in the non-reference frame image patches;

[0046] 2) For the module based on the set of image patches, matrix multiplication is used to calculate the transpose of the non-reference frame image patch feature matrix and the reference frame patch image patch feature matrix, that is: the inner product, to obtain the similarity matrix. Subsequently, the non-reference frame image patches are multiplied with the similarity matrix through matrix multiplication, enabling the non-reference frame image patches to selectively aggregate the similar texture information in the non-reference frame image patches according to the similarity with the reference frame image patches, and being able to fill the information inside the moving and saturated regions;

[0047] 3) Regarding the problem that the image features after the first step of step 1 cannot effectively fill the information in the moving and saturated regions, and the problem that there is a risk of reconstruction offset for the other background information introduced by the image features after the second step of step 1, a gating module is used. The convolutional layer is respectively used to predict the dynamic weights of the output features of the attention mechanism module and the module based on the set of image patches, and then this dynamic weight is multiplied pixel by pixel with the corresponding features crosswise. In this way, the features of the attention mechanism module and the module based on the set of image patches can be automatically interactively fused by the way of multiplying the dynamic weights pixel by pixel crosswise;

[0048] Step 2: Process the deformable large-range fusion network based on the transformer

[0049] The deformable large-range fusion network based on the transformer inputs the features of the content completion network into the residual deformable transformer module. This module uses multiple moving window transformer layers as the basic layer and an attention mechanism layer based on the deformable window to fuse long-distance information; specifically:

[0050] 1) For the moving window transformer layer, first the image is divided into non-overlapping windows, and then the image patches within the non-overlapping windows are mapped into query features, key features, and value features through three fully connected layers. Furthermore, the inner product between the query features and the key features within and between the non-overlapping windows is calculated through the self-attention mechanism layer to obtain the similarity matrix. Finally, the similarity matrix and the value features are fused through matrix multiplication to fuse the features of different exposure and motion regions;

[0051] 2) The attention mechanism layer based on deformable windows first divides the input features into non-overlapping windows, then maps the image patches within the non-overlapping windows into query features through a fully connected layer, predicts the offset of the query features through an offset network for the query features, then adds the query features and the offset to obtain the offset query features, passes the offset query features through two fully connected layers to map them into key features and value features, and then calculates the inner product of the offset query features and the key features within and between the non-overlapping windows through the self-attention mechanism to obtain a similarity matrix. Finally, the similarity matrix and the value features are dynamically fused using matrix multiplication; to avoid the limitation of the modeling ability for long-range relationships and reduce the computational complexity at the same time.

[0052] 3. A method for high dynamic range imaging based on the unification of image patch level and pixel level according to claim 2, characterized in that: the specific method for filling the motion and saturation regions in step 1, step 2 is as follows:

[0053] a) For the motion region, based on the image patch set module, according to the similarity weight between the reference frame image patch and the non-reference frame image patch features, the non-aligned regions in the non-reference frame image patch features are weighted and summed, and the texture features similar to the reference frame image patch features in the non-reference frame image patch features are collected to fill the motion non-aligned regions of the non-reference frame image patches to achieve the purpose of alignment with the reference frame image patches;

[0054] b) For the saturation region, based on the image patch set module, the features around the saturation region in the non-reference frame image patches are weighted and summed according to the similarity weight, and the similar texture features are collected to make the pixel values closer to the illumination of the real scene.

[0055] Embodiment 1

[0056] See Figure 1 , a method for high dynamic range imaging based on the unification of image patch level and pixel level, specifically includes the following steps:

[0057] For the content completion network, given three input low dynamic range pictures, first extract shallow features through three convolutional encoders, then we pass the shallow features through the attention mechanism module and the image patch set module respectively, and finally interact and fuse the output features of the attention mechanism module and the image patch set module through the gating module to jointly exert their advantages.

[0058] Such as Figure 2As shown, the attention mechanism module automatically generates attention maps for the reference frame features and non-reference frame features through a convolutional layer, and then multiplies the non-reference frame features and the attention maps pixel by pixel to obtain the output features of the attention mechanism module. The attention mechanism module works in a pixel-by-pixel multiplication calculation method and can better handle moving and saturated edge regions. The patch-based aggregation module first divides the shallow features into non-overlapping small patches, then calculates the similarity between the non-reference frame features and the reference frame features through inner product, and finally selectively fuses multiple similar feature patches of the non-reference frame features to obtain the output features of the patch-based aggregation module, as Figure 3 shown. Since the patch-based aggregation module can dynamically fuse multiple small patch features in the neighborhood, it can effectively fill in the information in the moving and saturated regions, further improving the quality of high dynamic range image synthesis.

[0059] As Figure 4 shown, the gating module respectively predicts the dynamic weights of the output features of the attention mechanism module and the patch-based aggregation module, then cross-multiplies the corresponding features with this dynamic weight pixel by pixel, and finally obtains the output features of the content completion network. The gating module can make the advantages of the attention mechanism module and the patch-based aggregation module complementary, enabling better recovery inside the moving and saturated regions, and at the same time making the edges of the moving and saturated regions sharp and clear.

[0060] See Figure 1 , the deformable large-scale fusion network based on transformers includes multiple residual deformable transformer modules. This module uses multiple moving window transformer layers as the basic layer to fuse long-range information. The moving window transformer layer first divides the image into non-overlapping windows, then maps the image patches within the windows through a fully connected layer into different features, and then calculates the similarity within and between the windows through the self-attention mechanism layer, and finally fuses the features of different exposure and motion regions. However, this in-window attention mechanism still limits the ability to model long-range relationships. At the same time, to reduce the computational complexity, we further propose an attention mechanism layer based on deformable windows, and its core operation is the attention mechanism based on deformable windows, as Figure 5As shown. The deformable window-based attention mechanism first divides the input features into non-overlapping windows, then maps the image patches within the non-overlapping windows into query features through a fully connected layer, and then predicts the offset of the query features through an offset network. Subsequently, the query features are added to the offset to obtain the offset query features. Then, the offset query features pass through two fully connected layers, which are respectively mapped into key features and value features. Furthermore, the self-attention mechanism calculates the inner product of the offset query features and the key features within and between the non-overlapping windows to obtain a similarity matrix. Finally, the similarity matrix and the value features are fused using matrix multiplication. The deformable window-based attention mechanism has a larger receptive field, and at the same time, it can dynamically adjust the regions of interest between windows, which is more conducive to information fusion. Therefore, it is more conducive to synthesizing a high-dynamic-range image without artifacts and saturation.

Claims

1. A high-dynamic-range imaging method based on the unification of image block level and pixel level, characterized in that: Specifically, it includes the following steps: Step 1, Content Completion Network Processing 1) For the attention mechanism module, the convolutional layer is used to automatically generate the attention weights between the non-reference frame image patches and the reference frame image patches. Subsequently, the attention weights are multiplied pixel by pixel with the non-reference frame image patches, enabling the non-reference frame image patches to highlight the well-exposed areas in the image according to the magnitude of the attention weights, while suppressing the moving and saturated areas in the non-reference frame image patches; 2) For the module based on the set of image patches, matrix multiplication is used to calculate the transpose of the non-reference frame image patch feature matrix and the reference frame patch image patch feature matrix, that is: the inner product, to obtain the similarity matrix. Subsequently, the non-reference frame image patches are multiplied with the similarity matrix through matrix multiplication, enabling the non-reference frame image patches to selectively aggregate the similar texture information in the non-reference frame image patches according to the similarity with the reference frame image patches, and being able to fill the information inside the moving and saturated areas; 3) Regarding the problem that the image features after the first step of Step 1 cannot effectively fill the information in the moving and saturated areas, and the problem that there is a risk of reconstruction offset for the other background information introduced by the image features after the second step of Step 1, the gating module is used to predict the dynamic weights of the output features of the attention mechanism module and the module based on the set of image patches respectively using the convolutional layer, and then multiply the corresponding features cross-pixel by pixel with this dynamic weight. In this way, the features of the attention mechanism module and the module based on the set of image patches can be interactively fused automatically by multiplying the dynamic weights cross-pixel by pixel; Step 2, Processing for the Transformer-based Deformable Large-scale Fusion Network The Transformer-based Deformable Large-scale Fusion Network inputs the features of the content completion network into the residual deformable Transformer module. This module uses multiple moving window Transformer layers as the basic layer and an attention mechanism layer based on the deformable window to fuse long-distance information; specifically: 1) The moving window Transformer layer first divides the image into non-overlapping windows, and then maps the image patches within the non-overlapping windows into query features, key features, and value features through three fully connected layers. Furthermore, through the self-attention mechanism layer, the inner product between the query features and the key features is calculated within and between the windows to obtain the similarity matrix. Finally, the similarity matrix and the value features are fused through matrix multiplication for the features of different exposure and moving areas; 2) For the attention mechanism layer based on the deformable window, first divide the input features into non-overlapping windows, then map the image patches within the non-overlapping windows into query features through a fully connected layer, then predict the offset of the query features through the offset network for the query features, and then add the query features and the offset to obtain the offset query features. Then, map the offset query features into key features and value features through two fully connected layers. Furthermore, through the self-attention mechanism, calculate the inner product between the offset query features and the key features within and between the non-overlapping windows to obtain the similarity matrix. Finally, dynamically fuse the similarity matrix and the value features using matrix multiplication; to avoid the limitation of the modeling ability for long-distance relationships and at the same time reduce the computational complexity.

2. The high-dynamic-range imaging method based on the unification of image block level and pixel level according to claim 1, characterized in that: The specific method for the filling motion and the interior of the saturation region described in step 1, step 2) is as follows: a) For the motion region, based on the image block set module, according to the similarity weight between the reference frame image block and the non-reference frame image block features, the non-aligned regions in the non-reference frame image block features are weighted and summed, and the texture features similar to the reference frame image block features in the non-reference frame image block features are aggregated to fill the motion non-aligned regions of the non-reference frame image block, so as to achieve the purpose of aligning with the reference frame image block; b) For the saturation region, based on the image block set module, the features around the saturation region in the non-reference frame image block are weighted and summed according to the similarity weight, and the similar texture features are aggregated to make the pixel value closer to the illumination of the real scene.

3. An imaging system for the high dynamic range imaging method based on the unification of the image block level and the pixel level according to claim 1 or 2, comprising a content completion network and a deformable large-range fusion network based on a transformer; characterized in that: The content completion network includes an image block set module, an attention mechanism module and a gating module; Among them, the image block set module aggregates similar image blocks around the region to be processed, so as to achieve the purpose of restoring the texture information inside the motion and saturation regions; The attention mechanism module is used to suppress the non-aligned and high-exposure or low-exposure information point by point, so as to achieve the purpose of removing artifacts and desaturating and making the edges of the motion and saturation regions sharp and clear; The gating module is used to perform dynamic cross-pixel multiplication on the output features of the image block set module and the attention mechanism module, so that the interior of the motion and saturation regions can be better restored, and at the same time, the edges of the motion and saturation regions are sharp and clear; The deformable large-range fusion network based on the transformer mainly includes a residual deformable transformer module; The residual deformable transformer module is used to dynamically adjust the weights of different exposure regions, model long-range dependencies, fuse the image features of large-range different exposure regions, and effectively restore the information of large-range motion and saturation regions.

Citation Information

Patent Citations

  • High dynamic range imaging and ghosting removal method based on generative adversarial network

    CN115018733A

  • Method and apparatus for deep neural network based inter-frame prediction in video coding

    US20220210402A1