Image processing method and device
Through the image processing method based on the cross attention mechanism, the moving objects between the first image feature and the second image feature are captured and deformable alignment processing, which solves the problem of ghosting artifacts when the existing HDR imaging technology is used to process the movement of long-distance objects, and realizes high-quality HDR image synthesis.
Patent Information
- Application Number
- CN202510542447.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-28
AI Technical Summary
Existing HDR imaging technologies are difficult to effectively align when processing long-distance objects, resulting in ghosting artifacts in the synthetic HDR images, affecting image quality.
Using an image processing method based on the cross attention mechanism, the intermediate features are determined by capturing the moving object between the first image feature and the second image feature, and performing deformable alignment processing based on the intermediate features, the combined image features are generated to determine the target image.
It effectively alleviates the problems of ghosting artifacts and fusion distortion, improves the quality of HDR images, and improves the clarity, stability and consistency of the images.
Smart Images

Figure CN120070208A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technologies, and in particular, to an image processing method and device. Background Art
[0002] With the development of digital imaging technologies, the demand for image quality has been increasing in various fields. Since traditional low dynamic range (LDR) images can only represent limited luminance information, in scenes with extremely large luminance differences, overexposure often occurs in overbright regions, details are lost, and the scenes seen by the human eye cannot be truly reflected, affecting the visual experience. To solve this problem, high dynamic range (HDR) imaging technologies collect multiple images with different exposure levels and fuse them into an HDR image with a wide luminance range, improving the visual effect and user experience.
[0003] Existing HDR imaging technologies generally include capturing multiple LDR images with different exposure levels, aligning these images to eliminate displacement errors caused by object movement, and merging the aligned images into an HDR image.
[0004] However, due to possible long-distance object movement during the shooting process, existing alignment algorithms are difficult to effectively handle, resulting in problems such as ghost artifacts in the synthesized HDR image, which affect the image quality. Summary of the Invention
[0005] Embodiments of this application provide an image processing method and device to achieve the effect of improving image quality.
[0006] In a first aspect, embodiments of this application provide an image processing method, including:
[0007] Based on a cross-attention mechanism, capturing moving objects between a first image feature and a second image feature, and determining an intermediate feature between the first image feature and the second image feature, where the first image feature is an image feature in a first multi-channel image, the second image feature is an image feature in a second multi-channel image, and the first multi-channel image and the second multi-channel image have different exposure levels;
[0008] Determining the position offset of the moving object between the first image feature and the second image feature, and the position offset feature of the position offset according to the intermediate feature between the first image feature and the second image feature;
[0009] Performing deformable alignment processing on the second image feature towards the first image feature according to the position offset feature of the moving object between the first image feature and the second image feature to obtain an aligned feature;
[0010] Obtain the combined image features based on the alignment features and the first image features;
[0011] Determine the target image based on the combined image features.
[0012] In one possible implementation, based on the cross-attention mechanism, capture the moving objects between the first image features and the second image features, and determine the intermediate features between the first image features and the second image features, including:
[0013] Determine the first window features of the first image features and the second window features of the second image features;
[0014] Based on the cross-attention mechanism, capture the first moving objects between the first window features and the second window features to obtain the initial intermediate features;
[0015] Obtain the first adjustment features and the second adjustment features according to the first window features and the initial intermediate features;
[0016] Based on the cross-attention mechanism, capture the second moving objects between the first adjustment features and the second adjustment features to obtain the intermediate features, where the moving distance of the first moving objects is less than the moving distance of the second moving objects.
[0017] In one possible implementation, obtain the first adjustment features and the second adjustment features according to the initial intermediate features and the second image features, including:
[0018] Perform window merging on the first window features and the initial intermediate features respectively to obtain the first window merged features and the second window merged features after window merging;
[0019] Perform downsampling processing on the first window merged features and the second window merged features to obtain the first adjustment features and the second adjustment features.
[0020] In one possible implementation, the method further includes:
[0021] Perform non-linear transformation processing on the first low-dynamic range image to obtain the first transformed image;
[0022] Perform channel dimension splicing processing on the first transformed image and the first low-dynamic range image to obtain the first multi-channel image.
[0023] In one possible implementation, the method further includes:
[0024] Input the first multi-channel image into a convolutional block for convolutional processing to obtain the first initial features;
[0025] Input the first initial features into a residual block for residual processing to obtain the first image features.
[0026] In a possible implementation manner, after performing deformable alignment processing on the second image feature toward the first image feature according to the position offset feature of the moving object between the first image feature and the second image feature to obtain an aligned feature, the method further includes:
[0027] Taking the aligned feature as the second image feature, and repeatedly executing the step of capturing the moving object between the first image feature and the second image feature based on the cross-attention mechanism to determine the intermediate feature between the first image feature and the second image feature until the number of repeated executions meets the preset number requirement, so as to obtain the aligned feature that meets the number of repeated executions.
[0028] In a possible implementation manner, obtaining a merged image feature according to the aligned feature and the first image feature includes:
[0029] Performing splicing processing on the aligned feature and the first image feature to obtain a spliced image feature;
[0030] Performing feature merging processing on the spliced image feature and the first image feature to obtain a merged image feature.
[0031] In a possible implementation manner, performing feature merging processing on the spliced image feature and the first image feature to obtain a merged image feature includes:
[0032] Determining the spatial dynamic weight of the spliced image feature;
[0033] Obtaining an initial merged feature according to the spatial dynamic weight and the spliced image feature;
[0034] Obtaining a merged image feature according to the initial merged feature and the first image feature.
[0035] In a possible implementation manner, after obtaining the initial merged feature according to the spatial dynamic weight and the spliced image feature, the method further includes:
[0036] Taking the initial merged feature as the spliced image feature, and repeatedly executing the step of determining the spatial dynamic weight of the spliced image feature until the number of repeated executions meets the preset number, so as to obtain a merged image feature.
[0037] In a second aspect, an image processing apparatus provided in an embodiment of the present application includes:
[0038] A capturing module, configured to capture a moving object between a first image feature and a second image feature based on a cross-attention mechanism to determine an intermediate feature between the first image feature and the second image feature, where the first image feature is an image feature in a first multi-channel image, the second image feature is an image feature in a second multi-channel image, and the exposure degrees of the first multi-channel image and the second multi-channel image are different;
[0039] A determination module, configured to determine a position offset of a moving object between a first image feature and a second image feature, and a position offset feature of the position offset, according to an intermediate feature between the first image feature and the second image feature;
[0040] An alignment module, configured to perform deformable alignment processing on the second image feature toward the first image feature according to the position offset feature of the moving object between the first image feature and the second image feature, to obtain an aligned feature;
[0041] A merging module, configured to obtain a merged image feature according to the aligned feature and the first image feature;
[0042] A determination module, configured to determine a target image according to the merged image feature.
[0043] In a third aspect, an embodiment of the present application provides an electronic device, including: a memory, a processor;
[0044] The memory stores computer-executable instructions;
[0045] The processor executes the computer-executable instructions stored in the memory, so that the processor executes the first aspect and / or various possible implementation manners of the first aspect as above.
[0046] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed by a processor, they are used to implement the first aspect and / or various possible implementation manners of the first aspect as above.
[0047] In a fifth aspect, an embodiment of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the first aspect and / or various possible implementation manners of the first aspect as above.
[0048] An image processing method and device provided by an embodiment of the present application are based on a cross-attention mechanism. By determining the correlation between a first image feature in a first multi-channel image and a second image feature in a second multi-channel image with different exposure degrees, focusing on and capturing the moving object feature, suppressing irrelevant information, generating an intermediate feature, thereby determining the position offset of the moving object and the position offset feature, and then performing transformations such as translation and rotation on the second image feature according to the feature to align it with the first image feature, obtaining an aligned feature, and then merging the aligned feature with the first image feature, integrating the information of the two, generating a target image, so that the advantages of different exposure images can be effectively integrated, the details of the moving object can be highlighted, ghosting and blurring can be avoided, and the clarity, stability and consistency of the image are improved, thereby achieving the effect of improving the image quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0050] Figure 1 It is a schematic flowchart of the existing image processing method provided by the present application;
[0051] Figure 1a It is a schematic structural diagram of the spatial attention module of the existing image processing method provided by the present application;
[0052] Figure 1b It is a schematic structural diagram of the alignment module of the existing image processing method provided by the present application;
[0053] Figure 2 It is a schematic flowchart of the image processing method provided by the present application Figure 1 ;
[0054] Figure 3 It is a schematic flowchart of the image processing method provided by the present application Figure 2 ;
[0055] Figure 4 It is a schematic flowchart of the image processing method provided by the present application Figure 3 ;
[0056] Figure 5 It is a schematic flowchart of a model training method provided by an embodiment of the present application;
[0057] Figure 5a It is a schematic diagram of the use of the high-dynamic range imaging network model provided by an embodiment of the present application;
[0058] Figure 6 It is a schematic diagram of another image processing method provided by an embodiment of the present application;
[0059] Figure 7 It is a schematic structural diagram of the image processing device provided by the present application;
[0060] Figure 8 It is a schematic structural diagram of the electronic device provided by the present application.
[0061] Through the above-mentioned accompanying drawings, specific embodiments of the present application have been shown, and there will be more detailed descriptions hereinafter. These accompanying drawings and written descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Detailed implementation manners
[0062] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0063] First, the terms related to the present application are explained:
[0064] Exposure refers to the total amount of light received by an image sensor or film during the photography and imaging process. Exposure is jointly determined by three key parameters: shutter speed, aperture size, and sensitivity (ISO). Exposure can determine the brightness level of a photo or image. Among them, when the exposure is insufficient (underexposed), the photo or image will appear dim, and details in the shadow part will be lost; when the exposure is excessive (overexposed), the photo or image will be too bright, and details in the highlight part will be lost.
[0065] The cross-attention mechanism can refer to a technique in deep learning for processing multi-input sequences or feature dependencies. Its principle is to dynamically allocate attention weights by calculating the similarity between different inputs, breaking the equal processing method of traditional neural networks for input information. By converting the input into query, key, and value vectors, calculating the similarity between the query and key vectors and normalizing it by Softmax to obtain the attention weights, and then weighted summing the value vectors, dynamic selection and focusing of input information are achieved.
[0066] A multi-channel image can refer to an image containing multiple information channels obtained by processing a low dynamic range (LDR) image.
[0067] Image features can refer to specific attributes or information segments that can describe and characterize the content of an image. These features can be visual elements such as color, texture, shape, edges, and corners, or more complex statistical information or structural patterns extracted from image data.
[0068] A convolutional block can refer to the basic building block of a convolutional neural network (CNN) in deep learning, which can be composed of one or more convolutional layers and other functional layers paired with them.
[0069] A residual block can refer to the core component of a residual network (ResNet) in deep learning, which can solve problems such as vanishing gradients and degradation that occur as the depth of the neural network increases.
[0070] The Patch embedding layer can refer to a commonly used component when processing data such as images in a deep learning model. It divides the input data (such as an image) into multiple small patches, which are similar to the pieces of a jigsaw puzzle, and each patch contains local information of the input data. Subsequently, these patches are embedded into a low-dimensional feature space through linear projection to generate corresponding feature vectors for each patch.
[0071] Figure 1 It is a schematic flowchart of the existing image processing method provided in this application. As Figure 1 shown, the execution entity of the existing image processing method can be a server, and the server can be devices such as mobile phones, computers, and tablets. The existing image processing method can include:
[0072] Input three different-exposure LDR images of the same scene , and regard the low-dynamic-range image with medium exposure as the reference image, and regard the remaining images as non-reference images;
[0073] Perform gamma correction on the LDR image to convert it to the high-dynamic-range domain;
[0074] Use the spatial attention module to extract features from the input LDR image , and use the alignment module to perform rough alignment on small motions for the exposure-corrected image ;
[0075] For the LDR image , first use a 3×3 convolution layer in three branches respectively to extract the image to the feature layer to obtain shallow features , then use the spatial attention module on the branch of the non-reference image to extract features to obtain non-reference features and ;
[0076] Figure 1a It is a schematic structural diagram of the spatial attention module of the existing image processing method provided in this application. As Figure 1a shown, input the non-reference features and reference features. First, use concatenation, two 3×3 convolutions, and the Sigmoid function to estimate the attention map, and then perform pixel-level multiplication on the input non-reference features and the attention map to obtain intermediate non-reference features;
[0077] For the gamma-corrected image , first use a 3×3 convolution layer in three branches respectively to extract the image to the feature layer to obtain shallow features , and then on the branch of the non-reference image, an alignment module is used to extract coarsely aligned features to obtain non-reference features and ;
[0078] Figure 1b is a schematic structural diagram of the alignment module of the existing image processing method provided by this application. As Figure 1b shown, the non-reference features and reference features are input. First, an offset estimator based on CNN is used to estimate the offset. The offset estimator includes a concatenation, two 3×3 convolutions, and a ReLu function. Then, a 3×3 deformable convolution is used to perform a geometric transformation on the input non-reference features to output intermediate non-reference features;
[0079] The obtained features are concatenated along the channels, and a 3×3 convolution is used for channel compression to obtain preliminary fusion features ;
[0080] Further, the preliminary fusion features are subjected to feature fusion through 3 dense residual dilated modules to obtain high-level fusion features;
[0081] Specifically, the dense residual dilated module includes 4 3×3 dilated convolutions and a residual connection. The input features of each 3×3 dilated convolution include the output features of all previous dilated convolutions;
[0082] Finally, the obtained high-level fusion features are reduced in channels through a 3×3 convolution, the features are transformed into the image domain, and then a sigmoid function is used to convert the values to between (0, 1) to output the final HDR image.
[0083] Combined with the above scenarios, in the prior art, the offset estimator is constructed based on CNN, and its receptive field has locality, which makes it difficult to accurately perceive long-distance object motion. When aligning LDR images, this limitation is particularly obvious, resulting in ghost artifacts in the synthesized HDR image. In addition, there are generally underexposed areas in LDR images, and the prior art fails to effectively suppress these areas during the fusion process, thereby causing fusion distortion and resulting in performance degradation of the synthesized HDR image.
[0084] The image processing method provided by this application effectively aligns moving objects between LDR images through the cross-attention deformable alignment method, and then uses the spatial attention fusion method to adaptively suppress the content of underexposed areas and select beneficial content to merge multi-exposure features, so that this patent effectively alleviates the problems of ghost artifacts and fusion distortion and improves the quality of HDR images.
[0085] The following will use specific embodiments to elaborate in detail on the technical solutions of the present application and how the technical solutions of the present application solve the above technical problems. The following several specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0086] Figure 2 Flow schematic of the image processing method provided by the present application Figure 1 , such as Figure 2 shown, the method includes:
[0087] S201. Based on the cross-attention mechanism, capture the moving objects between the first image feature and the second image feature, and determine the intermediate feature between the first image feature and the second image feature. The first image feature is the image feature in the first multi-channel image, the second image feature is the image feature in the second multi-channel image, and the exposure degrees of the first multi-channel image and the second multi-channel image are different.
[0088] Among them, the first multi-channel image and the second multi-channel image can refer to images with different exposure degrees. For example, the first multi-channel image can refer to a picture with normal exposure, and there can be multiple second multi-channel images, and they can include images with overexposed exposure and underexposed exposure. For example, the first multi-channel image can refer to a picture with medium exposure. That is, in the embodiments of the present application, the first multi-channel image can be a reference image, there can be multiple second multi-channel images, and they can include images with overexposed exposure and underexposed exposure. That is, in the embodiments of the present application, the second multi-channel image can be a non-reference image.
[0089] The first image feature can refer to the image feature of the first multi-channel image, and the second image feature can refer to the image feature of the second multi-channel image.
[0090] The moving object between the first image feature and the second image feature can refer to an object in the image sequence whose position in the image changes significantly, and the features such as its edges, corners, and textures show differences in position, shape, etc. in different images. Through feature extraction and comparison, the object whose feature change exceeds a certain threshold.
[0091] Based on the cross-attention mechanism, capturing the moving object between the first image feature and the second image feature can refer to using the first image feature as the query, the second image feature as the key and value, or using the second image feature as the query, the first image feature as the key and value. By calculating the correlation scores between feature elements, dynamically allocate attention weights, highlight the features related to the moving object, thereby capturing the position changes, movement trajectories, and state information of the moving object, fusing key information to enhance its feature representation, and providing high-quality input for subsequent image analysis tasks.
[0092] For example, when performing the cross-attention mechanism, query Q can be generated based on the first image feature , and key K and value V can be generated based on the second image feature ; where:
[0093]
[0094]
[0095] Among them, is the corresponding weight, is the number of channel dimensions of the feature.
[0096] The intermediate feature between the first image feature and the second image feature can refer to the fusion result generated during the attention calculation process. By regarding the first image feature as the query and the second image feature as the key and value, the correlation score between them is determined, and the attention weight is dynamically allocated according to the correlation score, and the second image feature is weighted and summed, so that the intermediate feature can be generated. Thus, the part of the second image feature related to the moving object in the first image feature is enhanced through the intermediate feature, and at the same time, the key information of part of the first image feature is retained, making the features of the moving object more prominent and clear.
[0097] Among them, in the embodiment of the present application, the method further includes:
[0098] Performing non-linear transformation processing on the first low-dynamic image to obtain a first transformed image;
[0099] Performing channel dimension splicing processing on the first transformed image and the first low-dynamic image to obtain a first multi-channel image.
[0100] Among them, the first low-dynamic image can refer to a low-dynamic image with normal exposure.
[0101] Performing non-linear transformation processing on the first low-dynamic image to obtain a first transformed image can refer to the process of obtaining a high-dynamic range domain image by processing the first low-dynamic image. Among them, a high-dynamic range domain image (High-Dynamic Range Image, HDR) can refer to a set of technologies in computer graphics and cinematography that are used to achieve a larger exposure dynamic range (i.e., a larger difference between light and dark) than ordinary digital image technologies.
[0102] In the embodiment of the present application, the non-linear transformation processing can refer to the process of performing gamma correction on the first low-dynamic image. Among them:
[0103] Gamma correction The formula can be:
[0104] ;
[0105] Wherein, represents a parameter in the gamma transformation, which can be set to 2.2 in the embodiments of the present application; may refer to a low-dynamic-range image; may refer to the exposure time of the i-th LDR image.
[0106] Performing channel dimension splicing processing on the first transformed image and the first low-dynamic-range image to obtain the first multi-channel image may refer to a method of splicing and combining the low-dynamic-range image and the high-dynamic-range image in the channel direction, so as to obtain the first multi-channel image process. For example, the RGB channels of an image can record the corresponding color component information, and the low-dynamic-range image and the high-dynamic-range image can be regarded as units containing specific information. When the low-dynamic-range image and the high-dynamic-range image are single-channel images, the first multi-channel image can be a two-channel image, storing the data of the low-dynamic-range image and the high-dynamic-range image respectively; when the low-dynamic-range image and the high-dynamic-range image are multi-channel images, the number of channels will increase accordingly. Thus, it can provide richer inputs for subsequent steps and facilitate the task of image enhancement.
[0107] Wherein, in the embodiments of the present application, the method further includes:
[0108] Inputting the first multi-channel image into a convolutional block for convolutional processing to obtain a first initial feature;
[0109] Inputting the first initial feature into a residual block for residual processing to obtain a first image feature.
[0110] Wherein, in the image processing flow, the first multi-channel image can be first sent into a convolutional block. Inside the convolutional block, through convolutional layers, possibly combined with batch normalization layers and activation functions, etc., convolutional operations are performed on the image to extract local features and generate a first initial feature. Then, the first initial feature is input into a residual block. The residual block uses an identity mapping branch to directly transmit information, and at the same time performs operations such as convolutional transformation through a residual mapping branch to learn residual features, thereby obtaining a first image feature.
[0111] S202. Determine the position offset of the moving object between the first image feature and the second image feature, and the position offset feature of the position offset, according to the intermediate feature between the first image feature and the second image feature.
[0112] Among them, the position offset of the moving object between the first image feature and the second image feature can refer to the measurement value of the change in the spatial position of the moving object between the first image feature and the second image feature. The position offset can quantify the movement of the moving object between the first image feature and the second image feature. Among them, the position offset can first calculate the difference between the first image feature and the second image feature vector, analyze the numerical changes of each channel of the intermediate feature, then parse the position encoding information therein, determine its spatial distribution on the feature map, then use the convolution module to capture local changes, rely on the attention mechanism to determine the weight of the key area, and finally combine context information such as time series to combine global and local features to determine the position offset.
[0113] The position offset feature of the position offset can refer to the key information for accurately describing the characteristics of the position offset. This key information can include multi-dimensional quantization indexes such as direction, distance, speed, and change trend.
[0114] After determining the position offset, vector analysis can be performed on the offset trajectory to clarify the direction feature of the position offset; using the distance measurement model and combining with the spatial coordinate system, calculate the distance feature of the position offset; by performing curve fitting or trend analysis on multiple groups of continuous position offsets, excavate the change trend feature of the position offset, so as to determine the position offset feature.
[0115] S203. According to the position offset feature of the moving object between the first image feature and the second image feature, perform deformable alignment processing on the second image feature to the first image feature to obtain an aligned feature.
[0116] Among them, performing deformable alignment processing on the second image feature to the first image feature can refer to geometric twisting and alignment of the second image feature respectively based on the position offset feature of the moving object between the first image feature and the second image feature, so that it is more matched with the first image feature in terms of spatial position and feature expression. That is, it can refer to the process of using information such as direction and distance included in the position offset feature to adjust the position of the moving object in the second image feature to be consistent with the position of the corresponding object in the first image feature through geometric transformation (such as translation, rotation, scaling, etc.).
[0117] The aligned feature can refer to the feature result that is highly consistent at the spatial position and semantic expression levels after performing alignment processing on the second image feature according to the position offset feature of the moving object between the second image feature and the first image feature.
[0118] Among them, in the embodiment of the present application, after performing deformable alignment processing on the second image feature to the first image feature according to the position offset feature of the moving object between the first image feature and the second image feature to obtain an aligned feature, the method further includes:
[0119] Take the alignment feature as the second image feature, and repeatedly execute the step of capturing the moving object between the first image feature and the second image feature based on the cross-attention mechanism to determine the intermediate feature between the first image feature and the second image feature until the number of repetitions meets the preset number requirement, and then obtain the alignment feature that meets the number of repetitions.
[0120] Among them, in the embodiments of the present application, the steps of S201~S203 can be performed by a cross-attention deformable alignment module. In order to ensure the effect of obtaining the intermediate feature, S201~S203 can be repeatedly executed to obtain the intermediate feature that meets the accuracy requirement.
[0121] S204. Obtain the merged image feature according to the alignment feature and the first image feature.
[0122] Among them, the merged feature can refer to integrating the alignment feature and the first image feature into a new feature set containing richer information through a specific fusion method. Among them, this fusion method can be a weighted average-based method, which assigns weights according to the importance of each feature in expressing the overall information, and then performs weighted summation on the corresponding elements, so that the merged feature can not only retain the original details and unique information of the first image feature, but also incorporate the matching relationship between the alignment feature and the first image feature in terms of space and semantics; it can also adopt a splicing strategy to splice the two features in the channel dimension or the spatial dimension to form a feature vector or feature matrix with a higher dimension and more comprehensive information, so as to provide a more rich feature representation.
[0123] S205. Determine the target image according to the merged image feature.
[0124] Among them, the target image can refer to the reconstruction of compressing the channels of the merged image feature through a convolutional layer, so as to obtain the image. In the embodiments of the present application, the target image can be an HDR image.
[0125] The image processing method provided by the embodiments of the present application uses a cross-attention deformable alignment module to effectively align low-dynamic-range images at the feature level. Among them, the global information of non-reference features is learned through a cross-attention-based offset estimator to capture the position of the moving object, and then based on the captured position, the non-reference features are geometrically transformed through deformable convolution, effectively alleviating the problem of ghost artifacts.
[0126] Figure 3 This is the process schematic of the image processing method provided by the present application Figure 2 , as Figure 3 shown, this embodiment is in Figure 2Based on the embodiments, the steps of capturing moving objects between the first image feature and the second image feature and determining the intermediate feature between the first image feature and the second image feature based on the cross-attention mechanism are described in detail. The method includes:
[0127] S301. Determine the first window feature of the first image feature and the second window feature of the second image feature;
[0128] S302. Based on the cross-attention mechanism, capture the first moving object between the first window feature and the second window feature to obtain an initial intermediate feature;
[0129] S303. Respectively perform window merging on the first window feature and the initial intermediate feature to obtain the first window merged feature and the second window merged feature after window merging;
[0130] S304. Perform downsampling on the first window merged feature and the second window merged feature to obtain a first adjusted feature and a second adjusted feature;
[0131] S305. Based on the cross-attention mechanism, capture the second moving object between the first adjusted feature and the second adjusted feature to obtain an intermediate feature, where the moving distance of the first moving object is less than the moving distance of the second moving object.
[0132] Among them, after obtaining the first image feature and the second image feature , add an absolute position encoding that is learnable and has the same size as the second image feature to the second image feature ;
[0133] Simultaneously divide the first image feature and the second image feature with the added position encoding into several windows of size P×P to obtain the first window feature of the first image feature and the second window feature of the second image feature ;
[0134] Inside the window, capture the first moving object between the reference feature and the non-reference feature through M cascaded cross-attention layers to obtain an initial intermediate feature ; Among them, the first moving object is a small-distance moving object. A small-distance moving object can refer to an object that moves a relatively small distance in terms of spatial position within the image window range relative to the object corresponding to the second image feature. Here, "relatively small" can be determined based on pixel distance and the proportion of the image area. For example, in the pixel coordinate system of the image, the displacement of the moving object is measured in pixels. When the displacement of a moving object in the horizontal or vertical direction between two adjacent frames of images does not exceed 5 pixels, it can be considered a moving object with a relatively small moving distance; or the image is divided into several regions, and the position change of the moving object in these regions is calculated. If the moving object moves from one region to an adjacent region and the moving distance accounts for a relatively small proportion of the entire image width or height, such as not exceeding 5% - 10%, it can be considered a moving object with a relatively small moving distance.
[0135] For the initial intermediate feature and the first window feature perform window merging and restore the size to obtain the second window merged feature and the first window merged feature ;
[0136] For the obtained second window merged feature and the first window merged feature execute the Patchembedding layer for downsampling to obtain the second adjusted feature with a low size and the first adjusted feature ;
[0137] On the entire feature map, capture the second moving object between the second adjusted feature and the first adjusted feature through N cascaded cross-attention layers to obtain the intermediate feature ; Among them, the second moving object is a long-distance moving object. A long-distance moving object can refer to an object that moves a relatively large distance in terms of spatial position within the image window range relative to the object corresponding to the second image feature. Here, "relatively large" can be determined based on pixel distance and the proportion of the image area. For example, in the pixel coordinate system of the image, the displacement of the moving object is measured in pixels. When the displacement of a moving object in the horizontal or vertical direction between two adjacent frames of images exceeds 50 pixels, it can be considered a moving object with a relatively large moving distance; or the image is divided into several regions, and the position change of the moving object in these regions is calculated. If the moving object moves from one region to a region separated by multiple regions and the moving distance accounts for a relatively large proportion of the entire image width or height, such as exceeding 30% - 50%, it can be determined as a moving object with a relatively large moving distance.
[0138] After obtaining the intermediate feature, the intermediate feature can be Perform bilinear interpolation operation, perform upsampling, restore the size, and then through a layer of convolution to reduce the channel dimension to obtain features ; Using the features as input, estimate the position offset of the moving object between the first image feature and the second image feature through a convolution module, and output the position offset feature ; Using the second image feature and the position offset feature as input, use a deformable convolutional layer to perform geometric transformation on the non-reference features to obtain aligned features .
[0139] Figure 4 is the process schematic of the image processing method provided by this application Figure 3 , as Figure 4 shown, on the basis of the Figure 2 embodiment, the step of obtaining the merged image feature according to the aligned feature and the first image feature is described in detail. The method includes:
[0140] S401. Concatenate the aligned feature and the first image feature to obtain a concatenated image feature.
[0141] Among them, concatenating the aligned feature and the first image feature is a process of integrating these two different but related feature information. In the embodiment of this application, the aligned feature is obtained by adjusting the second image feature through the position offset feature of the moving object between the first image feature and the second image feature. It is more compatible with the first image feature in terms of spatial position and semantic expression. The first image feature carries the initial information of the original image. During the concatenation process, by connecting the channel dimensions of the aligned feature and the first image feature, the newly generated concatenated image feature not only contains the original details of the first image feature but also incorporates the matching relationship with the second image feature reflected by the aligned feature, thus forming a concatenated image feature with richer information and higher dimension.
[0142] S402. Determine the spatial dynamic weight of the concatenated image feature.
[0143] Among them, the spatial dynamic weight of the concatenated image feature can be determined according to the spatial attention block, where the spatial dynamic weight satisfies:
[0144]
[0145] Among them, represents the spatial attention block, which concatenates a convolutional layer and a sigmoid function.
[0146] S403. Obtain an initial merged feature based on the spatial dynamic weight and the stitched image feature;
[0147] S404. Obtain a merged image feature based on the initial merged feature and the first image feature.
[0148] Among them, multiply the spatial dynamic weight by the stitched image feature at the pixel level to suppress the content of bad exposure, and select the beneficial content (that is, the positions greater than the preset threshold, and retain and emphasize the corresponding pixel values), while for the positions less than or equal to the preset threshold, the corresponding pixel values in the stitched image feature will be suppressed, so as to reduce the influence of the content in the positions representing the bad exposure area on the subsequent processing. Thus, the stitched image feature is adjusted and optimized spatially.
[0149] Finally, use a 1×1 convolution layer to fuse the features between channels, so as to obtain a merged image feature.
[0150] Among them, in the embodiment of the present application, after obtaining the initial merged feature according to the spatial dynamic weight and the stitched image feature, the method further includes:
[0151] Use the initial merged feature as the stitched image feature, and repeat the step of determining the spatial dynamic weight of the stitched image feature until the number of repeated executions meets the preset number of times, and then obtain the merged image feature.
[0152] Among them, steps S402 - S404 can all be processed by a residual fusion module. In the embodiment of the present application, in order to ensure that the dynamic weight adaptively suppresses the content of the bad exposure area and selects the beneficial content to merge the aligned multi-exposure features, effectively alleviating the effect of fusion distortion, the steps of S402 - S404 can be repeated to improve the effect of the finally obtained merged image feature.
[0153] The image processing method provided by the embodiment of the present application generates a spatial dynamic weight through a spatial attention module, and this dynamic weight adaptively suppresses the content of the bad exposure area and selects the beneficial content to merge the aligned multi-exposure features, effectively alleviating the fusion distortion.
[0154] Figure 5 It is a schematic flowchart of a model training method provided by an embodiment of the present application. As Figure 5 shown, this method includes:
[0155] S501. Input the preprocessed image into the model to be trained.
[0156] Among them, the preprocessed image may refer to a multi-channel image obtained through non-linear transformation processing and channel dimension splicing processing. In the embodiments of the present application, the input multi-channel image may be divided into a normal exposure image, an overexposed image, and an underexposed image.
[0157] The model to be trained may include a feature extraction sub-network and a feature merging sub-network. Among them, the feature extraction sub-network may be a sub-network for performing steps S201 to S203 in the embodiments of the present application, and the feature merging sub-network may be a sub-network for performing step S204 in the embodiments of the present application.
[0158] That is, after the preprocessed image is input into the model to be trained, a predicted target image output by the model to be trained is obtained.
[0159] S502. Train the model to be trained according to the predicted target image and the labeled HDR image corresponding to the preprocessed image to obtain a high dynamic range imaging network model.
[0160] Among them, the predicted target image and the labeled HDR image corresponding to the preprocessed image can determine the loss function of the model to be trained, where the loss function satisfies:
[0161]
[0162] Among them, represents the predicted target image reconstructed by the model to be trained, is the labeled HDR image, represents a hyperparameter. represents the Sobel operator, which detects the high-frequency region of the image through the change of the gray value of the image for detail extraction. represents the tone mapping operation, where:
[0163]
[0164] Among them, is the compression parameter, which can be set to 5000.
[0165] After determining the loss function, the weight parameters of the model to be trained can be adjusted according to the loss function until the loss function of the model to be trained converges, and then a high dynamic range imaging network model is obtained.
[0166] Figure 5a is a schematic diagram of the use of the high dynamic range imaging network model provided by the embodiments of the present application. As Figure 5a shown, the preprocessed image is input into the high dynamic range imaging network model, and thus the HDR image output by the high dynamic range imaging network model can be obtained.
[0167] Figure 6Schematic diagram of another image processing method provided by an embodiment of the present application. As Figure 6 shown, the execution subject of this image processing method is a high-dynamic range imaging network model, which includes a feature extraction sub-network and a feature merging sub-network. Among them, the feature extraction sub-network includes a cross-attention deformable alignment module, and the feature merging sub-network includes a residual fusion module. Each residual fusion module includes two spatial attention fusion modules and a relu layer.
[0168] Among them, the image processing method may include:
[0169] 1. Determine three LDR images with different exposures ;
[0170] Optionally, the sizes of the three LDR images are C×H×W, where C represents the number of channels of the image, H is the height of the image, which can be represented by the number of pixels of the image in the vertical dimension; W is the width of the image, which can be represented by the number of pixels of the image in the horizontal dimension.
[0171] Optionally, the sizes of the three LDR images with different exposures can be 3×256×256; it means that each image has three channels of red (R), green (G), and blue (B), with a height of 256 pixels (px) and a width of 256 px;
[0172] 2. Perform gamma correction on the LDR image and convert it to the high-dynamic range domain , to obtain an image in the high-dynamic range domain, where the size of the image in the high-dynamic range domain is 3×256×256;
[0173] Among them, the parameter in gamma correction can be set to 2.2.
[0174] 3. Concatenate the obtained high-dynamic range domain image and the LDR image along the channel dimension to obtain a 6-channel model input , where the size is 6×256×256;
[0175] 4. Input into the high-dynamic range imaging network model;
[0176] Among them, each is first converted into features using a 3×3 convolutional layer, and the sizes are all 64×256×256;
[0177] Then use a residual block to process the extracted features to eliminate the exposure differences caused by different exposures, and obtain three shallow features , the sizes of the three shallow features are all 64×256×256;
[0178] On the low-exposure feature branch, two cascaded cross-attention deformable alignment modules are used, taking the low-exposure feature and the reference feature as inputs to obtain the aligned low-exposure feature , with a size of 64×256×256;
[0179] On the high-exposure feature branch, two cascaded cross-attention deformable alignment modules are also used, taking the high-exposure feature and the reference feature as inputs to obtain the aligned high-exposure feature , with a size of 64×256×256; the corresponding cross-attention deformable alignment modules between these two branches share parameters;
[0180] The aligned non-reference features and as well as the reference feature are concatenated along the channels to form a merged feature , with a size of 128×256×256;
[0181] Optionally, a 3×3 convolution layer is used to compress the channels of the preliminarily merged features to form a more advanced fused feature as the input of the feature merging sub-network. This feature is denoted as , with a size of 64×256×256;
[0182] In the feature merging sub-network, 3 residual fusion modules are used to gradually merge the content of the multi-exposure images to form a high-dynamic range feature. The output features (i represents the i-th module) of each residual fusion module all have a size of 64×256×256;
[0183] Inside each residual fusion module, a spatial attention fusion module, a ReLu function, a spatial attention fusion module, and a residual connection are used to effectively fuse the multi-exposure features. The output features of each spatial attention fusion module have a size of 64×256×256;
[0184] The output feature of the last residual fusion module and the shallow reference feature are pixel-wise added: ;
[0185] Optionally, the obtained feature is passed through a 3×3 convolution layer to obtain an HDR image , with a size of 3×256×256.
[0186] Among them, for the cross-attention deformable alignment module:
[0187] Input any non-reference feature and reference features , both with dimensions of 64×256×256;
[0188] Input them into the cross-attention offset estimator to estimate the position offset, that is:
[0189] Add learnable absolute position encoding features to the non-reference features: , the position encoding features have the same dimensions as the non-reference features, which are 64×256×256;
[0190] Simultaneously divide the non-reference features with position encoding and the reference features into several 8×8 size windows to obtain the features of the new non-reference image and the features of the reference image , and their dimensions are (32×32)×64×(8×8);
[0191] Inside the window, capture small moving objects between the reference features and the non-reference features through M cascaded cross-attention layers to obtain the finally output intermediate features , with dimensions of (32×32)×64×(8×8), and M can be set to 2 here;
[0192] For the non-reference intermediate features and the features of the reference image Perform window merging to restore the dimensions to obtain the features and , both with dimensions of 64×256×256;
[0193] For the intermediate non-reference features and the reference features Execute the Patch embedding layer for downsampling. The Patch embedding layer contains 3 cascaded depthwise separable convolutional layers. The important parameters of these three depthwise separable convolutional layers are:
[0194] The first layer: input_channel = 64, output_channel = 128, kernel = 3×3, stride = 2;
[0195] The second layer: input_channel = 128, output_channel = 128, kernel = 3×3, stride = 2;
[0196] The third layer: input_channel = 128, output_channel = 128, kernel = 3×3, stride = 2;
[0197] After downsampling, low - dimensional non - reference features are obtained and reference features , both with a size of 128×32×32;
[0198] On the entire feature map, N cascaded cross - attention layers are used to capture long - range moving objects between the reference features and non - reference features, and the finally output intermediate features are obtained , with a size of 128×32×32, and N can be set to 3;
[0199] Perform bilinear interpolation on the intermediate features to upsample and restore the spatial size to 128×256×256, and then pass through a 3×3 convolution layer to reduce the channel dimension, obtaining features , with a size of 64×256×256;
[0200] Using the features as the input, a CNN module is used to estimate the position offset of the moving object between the reference features and non - reference features, and the position offset features are output , with a size of 27×256×256;
[0201] Using the non - reference features and the position offset features as the input, a deformable convolution layer is used to perform geometric transformation on the non - reference features, obtaining aligned non - reference features :
[0202]
[0203] Among them, is the number and index of the convolution kernel weights, represents the weight, center position, and the k - th offset of the k - th kernel, is the weighted value of the k - th kernel, represents the sampling offset of the k - th kernel centered on , and both come from . Here, the convolution kernel size of the deformable convolution layer is 3×3. The output aligned non - reference features have a size of 64×256×256.
[0204] Among them, for the cross - attention layer:
[0205] Input any reference token: and non - reference token: , and their sizes can be (32×32)×128;
[0206] Tokens of the reference image: Generate query Q in cross-attention instead of the reference token: Generate key K and value V, and perform the multi-head cross-attention mechanism, where:
[0207]
[0208]
[0209] Where is the corresponding weight, is the number of channel dimensions of the feature, which can be 128. The feature has a size of (32×32)×128;
[0210] Perform an FFN layer on the above output feature to output intermediate features with a size of (32×32)×128. The FFN layer contains two MLP layers.
[0211] For the spatial attention fusion module:
[0212] Input any fused feature , and the size can be 64×256×256;
[0213] Use a spatial attention block to obtain spatial dynamic weights :
[0214] ;
[0215] Where represents the spatial attention block, which concatenates a 3×3 convolutional layer and a sigmoid function. has a size of 64×256×256;
[0216] Multiply the spatial dynamic weights with the fused feature pixel-wise to suppress the content with bad exposure and select the beneficial content;
[0217] Use a 1×1 convolutional layer to fuse the features in the channels to obtain the output feature with a size of 64×256×256.
[0218] Figure 7 is the structural schematic diagram of the image processing device provided by this application. As Figure 7 shown, the image processing device 70 provided in this embodiment includes:
[0219] A capture module 701, configured to capture a moving object between a first image feature and a second image feature based on a cross-attention mechanism, and determine an intermediate feature between the first image feature and the second image feature, where the first image feature is an image feature in a first multi-channel image, the second image feature is an image feature in a second multi-channel image, and the exposure degrees of the first multi-channel image and the second multi-channel image are different;
[0220] A determination module 702, configured to determine a position offset of the moving object between the first image feature and the second image feature, and a position offset feature of the position offset, according to the intermediate feature between the first image feature and the second image feature;
[0221] An alignment module 703, configured to perform deformable alignment processing on the second image feature toward the first image feature according to the position offset feature of the moving object between the first image feature and the second image feature, to obtain an aligned feature;
[0222] A merging module 704, configured to obtain a merged image feature according to the aligned feature and the first image feature;
[0223] A determination module 705, configured to determine a target image according to the merged image feature.
[0224] In a possible implementation manner, the capture module 701 may further specifically be configured to:
[0225] Determine a first window feature of the first image feature and a second window feature of the second image feature;
[0226] Capture a first moving object between the first window feature and the second window feature based on the cross-attention mechanism, to obtain an initial intermediate feature;
[0227] Obtain a first adjustment feature and a second adjustment feature according to the first window feature and the initial intermediate feature;
[0228] Capture a second moving object between the first adjustment feature and the second adjustment feature based on the cross-attention mechanism, to obtain an intermediate feature, where the moving distance of the first moving object is less than the moving distance of the second moving object.
[0229] In a possible implementation manner, the capture module 701 may further specifically be configured to:
[0230] Perform window merging on the first window feature and the initial intermediate feature respectively, to obtain a first window merged feature and a second window merged feature after window merging;
[0231] Perform downsampling processing on the first window merged feature and the second window merged feature, to obtain a first adjustment feature and a second adjustment feature.
[0232] In a possible implementation, the capture module 701 may further be specifically configured to:
[0233] Perform a non-linear transformation on the first low-dynamic range image to obtain a first transformed image;
[0234] Perform a channel dimension concatenation on the first transformed image and the first low-dynamic range image to obtain a first multi-channel image.
[0235] In a possible implementation, the capture module 701 may further be specifically configured to:
[0236] Input the first multi-channel image into a convolutional block for convolution processing to obtain a first initial feature;
[0237] Input the first initial feature into a residual block for residual processing to obtain a first image feature.
[0238] In a possible implementation, the alignment module 703 may further be specifically configured to:
[0239] Take the alignment feature as the second image feature, and repeatedly execute the step of capturing the moving object between the first image feature and the second image feature based on the cross-attention mechanism to determine the intermediate feature between the first image feature and the second image feature, until the number of repetitions meets the preset number requirement, and then obtain the alignment feature that meets the number of repetitions.
[0240] In a possible implementation, the merging module 704 may further be specifically configured to:
[0241] Perform a concatenation on the alignment feature and the first image feature to obtain a concatenated image feature;
[0242] Perform a feature merging on the concatenated image feature and the first image feature to obtain a merged image feature.
[0243] In a possible implementation, the merging module 704 may further be specifically configured to:
[0244] Determine the spatial dynamic weight of the concatenated image feature;
[0245] Obtain an initial merged feature according to the spatial dynamic weight and the concatenated image feature;
[0246] Obtain a merged image feature according to the initial merged feature and the first image feature.
[0247] In a possible implementation, the merging module 704 may further be specifically configured to:
[0248] Take the initial merged feature as the stitched image feature, and repeat the step of determining the spatial dynamic weight of the stitched image feature until the number of repetitions meets the preset number of times, and then obtain the merged image feature.
[0249] The image processing device 70 provided in this embodiment can execute the method provided in the above method embodiment, and its implementation principle and technical effect are similar, so details are not described here in this embodiment.
[0250] Figure 8 It is a schematic structural diagram of the electronic device provided in this application. As Figure 8 shown, the electronic device 80 provided in this embodiment includes: at least one processor 801 and a memory 802. Optionally, the device 80 further includes a communication component 803. Among them, the processor 801, the memory 802, and the communication component 803 are connected through a bus 804.
[0251] In a specific implementation process, at least one processor 801 executes the computer execution instructions stored in the memory 802, so that at least one processor 801 executes the above method.
[0252] For the specific implementation process of the processor 801, reference can be made to the above method embodiment, and its implementation principle and technical effect are similar, so details are not described here again in this embodiment.
[0253] In the above embodiment, it should be understood that the processor may be a central processing unit (English: Central Processing Unit, abbreviated as: CPU), or other general-purpose processors, digital signal processors (English: Digital Signal Processor, abbreviated as: DSP), application specific integrated circuits (English: Application Specific Integrated Circuit, abbreviated as: ASIC), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.
[0254] The memory may include a high-speed memory (Random Access Memory, RAM), and may also include a non-volatile memory (Non-volatile Memory, NVM), such as at least one disk memory.
[0255] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the buses in the drawings of this application are not limited to only one bus or one type of bus.
[0256] This application also provides a computer program product, including a computer program which, when executed by a processor, implements the above method.
[0257] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above method.
[0258] The above-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk or an optical disk. The readable storage medium can be any available medium accessible by a general-purpose or special-purpose computer.
[0259] An exemplary readable storage medium is coupled to the processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an Application Specific Integrated Circuit (ASIC). Of course, the processor and the readable storage medium can also exist as discrete components in a device.
[0260] The division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.
[0261] The unit described as a separate component may or may not be physically separated, and the component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed across multiple network units. Some or all of these units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0262] In addition, in each embodiment of the present invention, each functional unit may be integrated into one processing unit, may exist physically separately for each unit, or two or more units may be integrated into one unit.
[0263] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present invention. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs, etc., all kinds of media that can store program codes.
[0264] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When this program is executed, it executes the steps including the above method embodiments; and the aforementioned storage medium includes: ROM, RAM, magnetic disks, or optical discs, etc., all kinds of media that can store program codes.
[0265] Finally, it should be noted that: After considering the specification and practicing the invention disclosed herein, those skilled in the art will easily think of other implementation schemes of the present invention. The present invention aims to cover any variations, uses, or adaptive changes of the present invention. These variations, uses, or adaptive changes follow the general principles of the present invention and include common general knowledge or conventional technical means in the technical field not disclosed in the present invention. It is not limited to the precise structure described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.
Claims
1. An image processing method, characterized in that: include: Based on a cross attention mechanism, a moving object between a first image feature and a second image feature is captured, and an intermediate feature between the first image feature and the second image feature is determined, wherein the first image feature is an image feature in a first multi-channel image, the second image feature is an image feature in a second multi-channel image, and the first multi-channel image and the second multi-channel image have different exposures; determining, according to an intermediate feature between the first image feature and the second image feature, a position offset of the moving object between the first image feature and the second image feature, and a position offset feature of the position offset; According to a position offset feature of the moving object between the first image feature and the second image feature, performing a deformable alignment process on the second image feature to the first image feature to obtain an alignment feature; Obtaining a merged image feature according to the alignment feature and the first image feature; A target image is determined according to the combined image features.
2. The method according to claim 1, characterized in that The method of capturing a moving object between a first image feature and a second image feature based on a cross attention mechanism and determining an intermediate feature between the first image feature and the second image feature includes: Determine a first window feature of the first image feature and a second window feature of the second image feature; Based on the cross attention mechanism, a first moving object between the first window feature and the second window feature is captured to obtain an initial intermediate feature; Obtaining a first adjustment feature and a second adjustment feature according to the first window feature and the initial intermediate feature; Based on the cross attention mechanism, a second moving object between the first adjustment feature and the second adjustment feature is captured to obtain an intermediate feature, wherein the movement distance of the first moving object is smaller than the movement distance of the second moving object.
3. The method according to claim 2, characterized in that The obtaining of the first adjustment feature and the second adjustment feature according to the first window feature and the initial intermediate feature comprises: Performing window merging on the first window feature and the initial intermediate feature respectively to obtain a first window merging feature and a second window merging feature after the window merging; Downsampling is performed on the first window merged feature and the second window merged feature to obtain a first adjusted feature and a second adjusted feature.
4. The method according to any one of claims 1 to 3, characterized in that The method further comprises: Performing nonlinear transformation processing on the first low-dynamic image to obtain a first transformed image; The first transformed image and the first low-dynamic image are spliced in a channel dimension to obtain the first multi-channel image.
5. The method according to any one of claims 1 to 3, characterized in that: The method further comprises: Inputting the first multi-channel image into a convolution block for convolution processing to obtain a first initial feature; The first initial feature is input into the residual block for residual processing to obtain a first image feature.
6. The method according to claim 1, characterized in that After performing deformable alignment processing on the second image feature to the first image feature according to the position offset feature of the moving object between the first image feature and the second image feature to obtain the alignment feature, the method further includes: The alignment feature is used as the second image feature, and the steps of capturing the moving object between the first image feature and the second image feature based on the cross-attention mechanism and determining the intermediate feature between the first image feature and the second image feature are repeatedly performed until the number of repetitions meets the preset number of requirements, thereby obtaining the alignment feature that meets the number of repetitions.
7. The method according to claim 1, characterized in that The step of obtaining a combined image feature according to the alignment feature and the first image feature comprises: Performing splicing processing on the alignment feature and the first image feature to obtain a spliced image feature; The spliced image feature and the first image feature are subjected to feature merging processing to obtain a merged image feature.
8. The method according to claim 7, characterized in that The step of performing feature merging processing on the spliced image feature and the first image feature to obtain a merged image feature includes: Determining spatial dynamic weights of the stitched image features; Obtaining an initial merged feature according to the spatial dynamic weight and the spliced image feature; A merged image feature is obtained according to the initial merged feature and the first image feature.
9. The method according to claim 8, characterized in that After obtaining the initial merged feature according to the spatial dynamic weight and the spliced image feature, the method further includes: The initial merged feature is used as the spliced image feature, and the step of determining the spatial dynamic weight of the spliced image feature is repeatedly performed until the number of repetitions meets the preset number of times, thereby obtaining the merged image feature.
10. An electronic device, characterized in that: include: Memory, processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor performs the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
High-dynamic image reconstruction method based on neural network
CN111986106A
Image alignment method, system and device based on deep learning and storage medium
CN118365851A
RAW domain multi-exposure image fusion method and device and storage medium
CN118396868A
HDR image reconstruction method and device, equipment and storage medium
CN119379575A
Image acquisition method, apparatus and device, and non-transient computer storage medium
WO2023246392A1