Burst image registration method combining projection transformation and optical flow fine tuning
Through the combined projection transformation and optical flow fine-tuning method, combined with neural networks and occlusion mask networks, the problems of low accuracy and slow speed of Burst image registration are solved, and efficient and robust image registration is achieved, meeting the real-time requirements of mobile devices.
Patent Information
- Application Number
- CN202510327865.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-19
- Publication Date
- 2025-06-20
AI Technical Summary
The existing Burst image registration method has low registration accuracy in occlusion and light transformation areas, and when considering factors such as noise, motion blur, light changes and occlusion, the registration speed is slow, making it difficult to meet the real-time needs of mobile devices.
The combined projection transformation and optical flow fine-tuning method is used to estimate the affine transformation parameters through a neural network, and the pyramid structure is used to estimate the parameters from low-scale features and fine-tune them on high scales. Combined with the occlusion mask network, the occlusion position is predicted to prevent the impact of the occlusion area on the estimated affine parameters.
It effectively solves the occlusion problems caused by noise and dynamic objects, improves the robustness of optical flow estimation in motion areas, reduces the time for global transformation estimation, and meets the real-time registration requirements on mobile devices.
Smart Images

Figure CN120182341A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image registration technology, and particularly relates to a Burst image registration method combining projective transformation and optical flow fine-tuning. Background Art
[0002] Burst images are formed by continuously capturing multiple images of the same scene using a handheld mobile device, and can be fused into a high-resolution image after registration. Different from conventional RGB, Burst images are RAW images with a larger dynamic range. Currently, most registration algorithms for this scenario are based on optical flow alignment, and the registration accuracy is low in occlusion and illumination transformation regions.
[0003] The core task of image registration is to align one image (floating image) with another image (reference image) through geometric transformation. Burst image registration is to select one image from a group of images (8 frames or 14 frames) as the reference image, and align the remaining images with the reference image. Image registration methods are mainly divided into two categories: global spatial transformation and local spatial transformation. Global spatial transformation achieves alignment by applying a unified geometric transformation to the entire image. Common transformations include rigid transformation, affine transformation, and projective transformation, etc. These transformations can usually be represented by matrices with different degrees of freedom, and all spatial coordinates of the target image are uniformly transformed through this matrix. While local registration methods allow different regions of the target image to use different spatial transformations. This method usually realizes it by modeling a deformation field with the same size as the image, and each pixel or region in the deformation field has its independent transformation parameters, such as methods based on optical flow and splines.
[0004] Recently, deep learning-based registration has achieved great advantages. Balakrishnan et al. proposed VoxelMorph as a learning framework for deformable medical image registration. Subsequent research improved the basic network, including the Transform structure, multiple attention, etc. Mok and Chung proposed C2FViT, which learns global affine registration by leveraging the characteristics of CNNs and Transformers and a multi-resolution strategy, and is superior to CNN-based methods in terms of registration accuracy, robustness, and generalization, bringing a new direction to medical image registration research. Traditional Burst image super-resolution methods usually use optical flow techniques to extract motion information for aligning input burst frames. As a convolutional neural network-based optical flow estimation algorithm, SpyNet shows good performance in processing motion estimation between images with its simple and efficient architecture. PWCNet is also an important algorithm in the field of optical flow estimation. By constructing a pyramid structure, introducing image warping operations, and cost volume calculation, etc., it can achieve more accurate and robust optical flow estimation. However, these methods use pre-trained off-the-shelf networks in the RGB space, which leads to inaccurate alignment when applied to RAW images, resulting in negative impacts and high computational costs.
[0005] The registration accuracy of Burst images affects the subsequent super-resolution reconstruction quality. Currently, a fast registration method based on optical flow is used, without considering many problems brought by moving objects and illumination changes. When considering the influence of various factors such as noise, motion blur, illumination changes, and occlusion, the registration speed will be slow, which limits its use on mobile devices. Summary of the Invention
[0006] To solve the above technical problems, the present invention provides a Burst image registration method combining projective transformation and optical flow fine-tuning, including the following steps:
[0007] S1: Continuously capture multiple images of the same scene with a handheld mobile device to obtain a Burst image;
[0008] S2: Perform affine transformation and mask prediction on the Burst image to obtain the affine transformation parameters and mask of the Burst image;
[0009] S21: Continuously capture multiple images of the same scene with a handheld mobile device to obtain a Burst image in RAW format. Recombine the Burst image in the RAW domain into a four-channel image in RGGB mode, halving the resolution, and generating a sequence of Burst image frames to be aligned
[0010] S22: Select the first frame in Randomly select one frame as the floating image M from them; where N represents the number of image frames, and when i = 1, F i represents the highest-scale image;
[0011] S23: Input the highest-scale F1 and M1 images into the feature matching module to extract features, obtaining feature O FMM1 , and generate the highest-scale affine transformation parameter A1 according to feature O FMM1 . In the affine transformation parameter A1, the second dimension represents the estimated global transformation parameter t x , t y , θ, s x , s y , h x , h y , and generate a complete affine transformation matrix according to the global transformation parameter where t x , r y represent the translation parameters in the x and y directions respectively, θ represents the rotation parameter, s x , s y represent the scaling parameters in the x and y directions respectively, h x , h y represent the shear parameters in the x and y directions respectively;
[0012] S24: Estimate the initial mask at the highest scale through the occlusion prediction network, and after upsampling, filter and update the mask for the features of the next scale;
[0013] S25: At the next scale, double the translation parameters in the affine transformation matrix to obtain the affine transformation matrix with doubled parameters and use this matrix to warp the floating image M2;
[0014] S26: Multiply the warped image M2 by the mask and then input it together with the fixed image F2 after channel concatenation into the feature matching module to generate the affine transformation parameter A2 at this scale, and generate a complete affine transformation matrix
[0015] S27: Repeat the process of S21 - S26 for l - 1 times, and output the affine transformation matrix at the lowest scale
[0016] S3: Based on the affine transformation parameters and the mask, perform optical flow fine-tuning to achieve Burst image registration.
[0017] Advantages of the present invention:
[0018] The present invention uses a neural network to estimate affine transformation parameters for modeling camera motion during Burst photography. Specifically, through a pyramid structure, parameters are estimated from low-scale features and fine-tuned at high scales. Meanwhile, an occlusion mask network is used to predict occlusion positions, preventing the occlusion area from affecting the estimation of affine parameters. The global transformation can effectively solve the occlusion problems caused by noise and dynamic objects. Secondly, the optical flow of the occlusion area is estimated, and the foreground and background of moving objects are determined by calculating the correlation between the input frame and the current frame. In this way, it is easy to distinguish the unregistered part and estimate the motion of objects in the occlusion area.
[0019] The present invention adopts a multi-scale coarse-to-fine parameter registration network to estimate global transformation parameters. Compared with traditional methods, it only requires one inference process and does not need iterative optimization like traditional methods, thus accelerating the global transformation estimation speed. Secondly, an occlusion mask prediction network is designed to exclude the motion area during parameter estimation, preventing interference with the global transformation parameter estimation, and fine-tuning the optical flow estimated by the global transformation, improving the robustness of the optical flow estimation in the motion area. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 It is a structural diagram of the affine transformation and mask prediction process of a Burst image registration method combining projective transformation and optical flow fine-tuning according to the present invention;
[0021] Figure 2 It is a schematic structural diagram of the feature matching module of a Burst image registration method combining projective transformation and optical flow fine-tuning according to the present invention;
[0022] Figure 3 It is a schematic diagram of the registration results of different models in the BRA synthetic dataset according to the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0023] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0024] On mobile devices, due to hardware limitations, the resolution of captured images is restricted. Through Burst image super-resolution, the resolution and quality of images can be improved, enabling high-resolution photo output and meeting users' needs for high-quality photography.
[0025] Image alignment is an essential step in multi-frame image super-resolution. In burst photography image super-resolution, the image alignment method follows the video super-resolution method, using a lightweight optical flow network pre-trained in RGB space for explicit alignment, or using a deformable convolution-based method for implicit alignment. Explicit alignment provides a clear direction of motion, which provides prior information for the subsequent upsampling process. Implicit alignment requires a complex feature extraction process to estimate the convolution offset, which affects the efficiency of the model.
[0026] A Burst image registration method combining projection transformation and optical flow fine-tuning, comprising:
[0027] S1: Use a handheld mobile device to continuously take multiple images of the same scene to obtain a Burst image;
[0028] S2: Perform affine transformation and mask prediction on the Burst image to obtain the affine transformation parameters and mask of the Burst image;
[0029] S3: Fine-tune the optical flow based on affine transformation parameters and masks to achieve Burst image registration.
[0030] In traditional projective transformation modeling, the projective transformation is implemented by a 3*3 matrix H, which contains 8 degrees of freedom, namely h11, h12, h13, h21, h22, h23, h31, h32. The projective transformation can accurately describe complex spatial transformations including perspective deformation. Traditional methods can use direct optimization based on brightness consistency to iteratively optimize the similarity of distorted images and fixed images to solve the transformation parameters; by designing feature operators, such as SIFT and SURF, etc., and then iteratively optimizing the transformation parameters through methods such as RANSAC; frequency domain phase correlation method, etc. Most traditional methods need to be based on iterative optimization, and different hyperparameter settings are set for different situations. In actual use, the speed is slow and it is difficult to meet real-time requirements. The frequency domain-based method is faster, but it can only achieve good results in simple scenes and cannot achieve the accuracy of Burst photography registration. For this reason, this chapter proposes an end-to-end global parameter estimation method based on deep learning. STN can adjust images or feature maps in the form of pixels, such as translation, rotation and scaling. It provides a differentiable image sampling tool, which becomes the basis for subsequent deep learning-based methods. In burst photography scenarios, the displacement of burst images is mainly caused by the tiny movements of handheld devices. The applicant ignores the perspective components h31 and h32 of the projection transformation and uses affine transformation to achieve the desired effect. The main reasons may be the short shooting time, small jitter, and small perspective component. Secondly, more parameter estimation also increases the learning difficulty of the network, leading to model collapse. This paper uses STN to perform the affine transformation of the current frame. From a geometric perspective, the affine transformation is a non-singular linear transformation, which ensures the reversibility of the transformation.
[0031] Instead of directly estimating the affine transformation parameters, this application decomposes them into a set of linear geometric transformation matrices, namely the translation matrix T, the rotation matrix R, the scaling matrix S, and the shear matrix H. The final affine transformation matrix A can be obtained by matrix multiplication. To accelerate network convergence and collapse, this application restricts the parameters of the transformation model, reducing the search space of the model. The rotation and shear parameters are restricted between -π / 10 and π / 10, the translation parameters are restricted between -10% and +10% of the maximum spatial rate, and the scaling parameters are restricted between 0.8 and 1.2.
[0032] Preferably, perform affine transformation and mask prediction on the Burst image to obtain the affine transformation parameters and mask of the Burst image, as Figure 1 shown, including:
[0033] S21: Continuously capture multiple images of the same scene through a handheld mobile device to obtain a Burst image in RAW format. Recombine the Burst image in the RAW domain into a four-channel image in the RGGB pattern, reducing the resolution by half to generate a sequence of Burst image frames to be aligned
[0034] S22: Select the first frame in as the fixed image F, and randomly select a frame in i as the floating image M; where N represents the number of image frames. When i = 1, F i represents the highest-scale image;
[0035] S23: Input the highest-scale F1 and M1 images into the feature matching module to extract features, obtaining feature O FMM1 , and generate the highest-scale affine transformation parameter A1 based on feature O FMM1 . In the affine transformation parameter A1, the second dimension represents the estimated global transformation parameter t x , t y , θ, s x , s y , h x , h y , and generate a complete affine transformation matrix based on the global transformation parameter where t x , t y represent the translation parameters in the x and y directions respectively, θ represents the rotation parameter, s x , s y represent the scaling parameters in the x and y directions respectively, and h x , h y represent the shear parameters in the x and y directions respectively;
[0036] S24: Estimate the initial mask at the highest scale through the occlusion prediction network, and after upsampling, filter the features at the next scale and update the mask;
[0037] S25: At the next scale, double the translation parameters in the affine transformation matrix to obtain the affine transformation matrix with doubled parameters, and use this matrix to warp the floating image M2;
[0038] S26: The warped image M2 is multiplied by the mask and then concatenated with the fixed image F2 in the channel dimension and input into the feature matching module together to generate the affine transformation parameters A2 at this scale, and generate the complete affine transformation matrix
[0039] S27: Repeat the process of S21 - S26 for l - 1 times, and output the affine transformation matrix at the lowest scale
[0040] Preferably, as Figure 2 shown, the feature matching module is composed of hierarchically depth - separable convolutional blocks HDSC with shared weights stacked in series;
[0041] The input fixed image F and floating image M respectively pass through the series - connected HDSC modules and output features as R(O F ), R(O M ), and use cross - attention at the high scale to calculate the channel correlation between the fixed image and the floating image to extract the global relationship between the fixed image and the floating image, and obtain the fixed - image feature O FCA and the floating - image feature O MCA , and concatenate them in the channel dimension;
[0042] The input fixed image F and floating image M respectively pass through the series - connected HDSC modules and output features as R(O F ), R(O M ), and use cross - attention at the high scale to calculate the channel correlation between the fixed image and the floating image to extract the global relationship between the fixed image and the floating image, and obtain the fixed - image feature O FCA and the floating - image feature O MCA , including:
[0043] O FCA = R(O F )*(O F *R(O M ))
[0044] O MCA = R(O M )*(O M*R(O F ))
[0045] Among them, O FCA and O MCA respectively represent the fixed image feature and the floating image feature after capturing the global relationship. R represents the transpose operation. R(O F ), R(O M ) respectively represent flattening and transposing the features of the fixed image and the floating image. O F , O M respectively represent the features after flattening the fixed image and the floating image. * represents matrix multiplication.
[0046] Preferably, the highest-scale F1 and M1 images are respectively input into the feature matching module to extract features, and the feature O FMM is obtained, including:
[0047] O FMM1 = FMM1((M1, F1))
[0048] Among them, O FMM1 represents the highest-scale feature extracted by the feature matching module. FMM1 represents the first feature matching module. F1 and M1 respectively represent the highest-scale fixed image and floating image.
[0049] Preferably, the highest-scale affine transformation parameter A1 is generated according to the feature O FMM1 , including:
[0050] A1 = MLP1(GAP(O FMM ))
[0051] Among them, A1 represents the highest-scale affine transformation parameter. MLP1 represents the first MLP head. GAP represents the global average pooling operation. O FMM1 represents the highest-scale feature extracted by the feature matching module.
[0052] Preferably, the complete affine transformation matrix is generated according to the global transformation parameter including:
[0053]
[0054] Among them, represents the highest-scale affine transformation matrix. t x , t y respectively represent the translation parameters in the x and y directions. θ represents the rotation parameter. s x , s y respectively represent the scaling parameters in the x and y directions. h x , h y respectively represent the shear parameters in the x and y directions.
[0055] Preferably, the initial mask is estimated at the highest scale by the occlusion prediction network, and after upsampling, the features of the next scale are filtered and the mask is updated, including:
[0056] Initial mask:
[0057]
[0058] And update the mask:
[0059]
[0060] Where Mask1 and Mask i Are the occlusion mask at the highest scale and the mask at the i-th scale respectively, Sigmoid represents the normalization function, Conv represents the convolution, Respectively represent the output features of the highest scale fixed image F1 and the highest scale floating image M1 after passing through the cascaded HDSC module, Represents the highest scale affine transformation matrix after doubling the parameters, Respectively represent the output features of the i-th scale fixed image F i And the i-th scale floating image M i After passing through the cascaded HDSC module, x represents the horizontal coordinate of the pixel point, bilinearUp represents bilinear upsampling, Mask i-1 Represents the mask at scale i - 1, Represents the affine transformation matrix after doubling the parameters at scale i - 1, M i 、F i Respectively represent the floating image and the fixed image at scale i, ° represents image distortion.
[0061] Preferably, the distorted M2 image is multiplied by the binary mask Mask2 and then concatenated with the fixed image F2 in the channel and input into the feature matching module together to generate the affine transformation parameter A2 at this scale, including:
[0062]
[0063] Where A2 represents the affine transformation parameter at the second highest scale, MLP2 represents the second MLP head, FMM2 represents the second feature matching module, Represents the highest scale affine transformation matrix after doubling the parameters, M2 and F2 respectively represent the floating image and the fixed image at the second highest scale, x represents the horizontal coordinate of the pixel point, Mask2 represents the mask at the second highest scale, A 1U Represents the global affine transformation parameter after upsampling, Represents image distortion, Represents dot multiplication.
[0064] In the affine parameter estimation stage, the network parameters are Θ, and the overall loss is:
[0065]
[0066] Similar to the point-to-point loss of the optical flow network, this method calculates the error between the affine transformation parameters and the gt parameters. The former is the self-supervised loss, and the latter is the supervised loss, where is the affine transformation parameters estimated by iterative optimization through expensive traditional methods. Considering that the parameters estimated by traditional methods are not the real gt and there may be certain errors. The self-supervised loss is added, where SIM represents the similarity metric function. The MSE metric is selected when there is no occlusion, and the more robust CharbonnierLoss is selected when there is occlusion. Among them, the hyperparameters α1 = 0.4, α2 = 0.2, α 3,4,5 = 0.1, and β is 0.5.
[0067] Preferably, based on the affine transformation parameters and the mask, the optical flow is fine-tuned to achieve Burst image registration, including:
[0068] S31: Calculate the initial optical flow at the highest scale where, represents the affine transformation matrix at the highest scale, E represents the identity matrix, and I represents the unit grid corresponding to the image at the corresponding scale;
[0069] S32: Calculate the correlation between the fixed image and the warped image through Cost Volume, and calculate the fine-tuned optical flow based on the affine transformation matrix and the initial optical flow at the highest scale:
[0070]
[0071] where, represents the fine-tuned optical flow at the highest scale, ConvBlock1 represents the first convolutional block for optical flow fine-tuning, CV represents calculating the correlation between images, F1 and M1 respectively represent the fixed image and the floating image at the highest scale, and Mask1 represents the highest scale mask;
[0072] S33: At the subsequent scale, use the affine transformation matrix at the corresponding scale and fine-tune the optical flow based on the correlation between its images;
[0073]
[0074] where, represents the fine-tuned optical flow at scale i, M i 、F i respectively represent the floating image and the fixed image at scale i, represents image warping, Maski Denote the mask at scale i, Denote the affine transformation matrix at scale i, bilinearUp denotes bilinear upsampling;
[0075] S34: Repeat steps S31 - S33 to obtain the fine-tuned optical flow at the highest scale And refine it
[0076]
[0077] S35: Align the floating image at each scale with the fixed image through the optical flow at that scale to achieve Burst image registration.
[0078] Preferably, calculate the correlation between the fixed image and the warped image through Cost Volume calculation, including:
[0079]
[0080] Among them, CV(F, M) represents calculating the correlation between the fixed image F and the floating image M, d x , d y Represent the lateral and axial offsets respectively.
[0081] Preferably, align the floating image at each scale with the fixed image through the optical flow at that scale to achieve Burst image registration, including:
[0082]
[0083] Among them, φ represents the refined optical flow, M represents the floating image, x and y represent the pixel coordinates respectively, Δx and Δy represent the displacement of each pixel point, Represents image warping.
[0084] In the optical flow fine-tuning stage, point-to-point loss and self-supervised loss are adopted. In addition to real occlusion data, synthetic occlusion data is added. In the synthetic occlusion data, the real optical flow can be accurately calculated, the optical flow in the occlusion area is zero, and there is no real optical flow in the real data, and only the self-supervised loss is used to fine-tune the network. On the synthetic occlusion dataset, the loss function used is:
[0085]
[0086] Among them, the hyperparameters are α1 = 0.4, α2 = 0.2, α 3,4,5 = 0.1, r = 0.2. Only the first part of the self-supervised loss is used in the real data fine-tuning stage.
[0087] In this embodiment, the present invention uses two datasets to evaluate the performance of the proposed method and compares a series of representative methods in this field to verify the effectiveness of the method proposed in this chapter. The first dataset is the RealBSR dataset proposed by Wei et al., and the second dataset is the dataset provided by the RAW Burst Alignment and ISP Challenge track of NTIRE2024 challenges, which is abbreviated as the RBA dataset in this paper.
[0088] Registration experiments are carried out on the RealBSR dataset and the RBA dataset. First, training is performed on RealBSR, which does not contain occluded data, and then fine-tuning is performed on the RBA dataset, which contains synthetic moving objects and real occluded data. Subjective effects and objective indicators of the registration results are analyzed. To verify the superiority of the proposed model, a series of typical optical flow estimation methods are selected: PWCNet [1], GMFlow [2], CRAFT [3], LLA-Flow [4].
[0089] Figure 3 The registration results of different models in the BRA synthetic dataset are shown. The first frame, the fourth frame, and the seventh frame are extracted from the Burst sequence, as Figure 3 shown in the first row, where the first frame is used as the base frame, i.e., the fixed image, and the fourth frame and the seventh frame are used as floating images to be aligned with the fixed image respectively. Starting from the second row, each row represents the registration result of a method. The left two columns show the registration results of the fourth frame image and the first frame image, and the right two columns represent the registration results of the seventh frame image and the first frame image. In each two columns, the left represents the optical flow estimation map, and the right represents the map after the floating image is aligned with the fixed image. In summary, the method of the present invention is significantly superior to other methods, especially in the edge region of optical flow estimation.
[0090] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A Burst image registration method combining projection transformation and optical flow fine-tuning, characterized in that: include: S1: Use a handheld mobile device to continuously take multiple images of the same scene to obtain a Burst image; S2: Perform affine transformation and mask prediction on the Burst image to obtain the affine transformation parameters and mask of the Burst image; S21: Use a handheld mobile device to continuously shoot multiple images of the same scene to obtain a RAW-format Burst image, reorganize the RAW-domain Burst image into a four-channel image in RGGB mode, so that the resolution is halved, and generate a Burst image frame sequence to be aligned S22: Select The first frame in is taken as the fixed image F, A frame is randomly selected as the floating image M; where N represents the number of image frames. When i=1, F i represents the highest scale image; S23: Input the highest scale F1 and M1 images into the feature matching module to extract features and obtain feature O FMM , and according to the characteristic O FMM Generate the highest scale affine transformation parameter A1, where the second dimension represents the estimated global transformation parameter t x ,t y ,θ,s x ,s y ,h x ,h y , and generate a complete affine transformation matrix based on the global transformation parameters Among them, t x ,t y They represent the translation parameters in the x and y directions, θ represents the rotation parameter, and s x ,s y Represents the scaling parameters in the x and y directions, h x ,h y Represent the shear parameters in the x and y directions respectively; S24: Estimate the initial mask at the highest scale through the occlusion prediction network, filter the features of the next scale after upsampling, and update the mask; S25: At the next scale, the affine transformation matrix Double the translation parameter in to get the affine transformation matrix after doubling the parameter And use this matrix to distort the floating image M2; S26: The distorted image M2 is multiplied by the mask and then input into the feature matching module together with the fixed image F2 after channel splicing to generate the affine transformation parameter A2 of the scale and generate the complete affine transformation matrix S27: Repeat the S21-S26 process l-1 times and output the affine transformation matrix at the lowest scale S3: Fine-tune the optical flow based on affine transformation parameters and masks to achieve Burst image registration.
2. The method for Burst image registration combining projection transformation and optical flow fine-tuning according to claim 1, characterized in that: The feature matching module is composed of weight-sharing hierarchical depth-separable convolution blocks HDSC stacked in series; The input fixed image F and floating image M are respectively passed through the HDSC modules in series and the output features are and Flatten and transpose the last two dimensions to become R(O F )、R(O M ), and use cross attention to calculate the channel correlation between the fixed image and the floating image at a high scale to extract the global relationship between the fixed image and the floating image, and obtain the fixed image feature O after capturing the global relationship. FCA and floating image feature O MCA , and splice them in the channel dimension; Cross attention is used at a high scale to calculate the channel correlation between the fixed image and the floating image to extract the global relationship between the fixed image and the floating image, and the fixed image feature O is obtained after capturing the global relationship. FCA and floating image feature O MCA ,include: O FCA =R(O F )*(O F *R(O M )) O MCA =R(O M )*(O M *R(O F )) Among them, O FCA and O MCA They represent the fixed image features and floating image features after capturing the global relationship, R represents the transposition operation, and R(O F )、R(O M ) represent the flattened features and transposed features of the fixed image and the floating image, respectively. F , O M They represent the flattened features of the fixed image and the floating image respectively, and * represents matrix multiplication.
3. The method for Burst image registration combining projective transformation and optical flow fine-tuning according to claim 1, characterized in that: The highest scale F1 and M1 images are input into the feature matching module to extract features and obtain feature O FMM ,include: O FMM1 =FMM1((M1,F1)) Among them, O FMM1 It represents the highest scale feature extracted by the feature matching module, FMM1 represents the first feature matching module, and F1 and M1 represent the highest scale fixed image and floating image respectively.
4. The method for Burst image registration combining projection transformation and optical flow fine-tuning according to claim 1, characterized in that: According to the feature O FMM Generate the highest scale affine transformation parameters A1, including: A1=MLP1(GAP(O FMM1 )) Where A1 represents the highest scale affine transformation parameter, MLP1 represents the first MLP head, GAP represents the global average pooling operation, O FMM1 Represents the highest-scale features extracted by the feature matching module.
5. The method for Burst image registration combining projective transformation and optical flow fine-tuning according to claim 1, characterized in that: Generate a complete affine transformation matrix based on the global transformation parameters include: in, represents the highest scale affine transformation matrix, t x ,t y They represent the translation parameters in the x and y directions, θ represents the rotation parameter, and s x ,s y Represents the scaling parameters in the x and y directions, h x ,h y Represent the shear parameters in the x and y directions respectively.
6. The method for Burst image registration combining projective transformation and optical flow fine-tuning according to claim 1, characterized in that: The initial mask is estimated at the highest scale through the occlusion prediction network, and the features of the next scale are filtered and the mask is updated after upsampling, including: Initial mask: And update the mask: Among them, Mask1 and Mask i is the highest scale occlusion mask and the i-th scale mask, Sigmoid represents the normalization function, Conv represents convolution, They represent the output features of the highest scale fixed image F1 and the highest scale floating image M1 after passing through the HDSC modules in series. represents the highest scale affine transformation matrix after doubling the parameters, Represents the i-th scale fixed image F i and the i-th scale floating image M i After the HDSC modules are connected in series, the output features are: x represents the horizontal coordinate of the pixel, bilinearUp represents bilinear upsampling, and Mask i-1 represents the mask at scale i-1, represents the affine transformation matrix after doubling the parameters on scale i-1, M i 、F i represent the floating image and fixed image at scale i respectively, Indicates that the image is distorted.
7. The method for Burst image registration combining projection transformation and optical flow fine-tuning according to claim 1, characterized in that: The distorted M2 image is multiplied by the binary mask Mask2 and then input into the feature matching module together with the fixed image F2 after channel splicing to generate the scale affine transformation parameter A2. include: Where A2 represents the second highest scale affine transformation parameter, MLP2 represents the second MLP head, and FMM2 represents the second feature matching module. represents the highest scale affine transformation matrix after doubling the parameters, M2 and F2 represent the floating image and fixed image of the next highest scale, x represents the horizontal coordinate of the pixel point, Mask2 represents the next highest scale mask, and A 1U represents the global affine transformation parameters after upsampling, Indicates that the image is distorted. Represents dot product.
8. The method for Burst image registration combining projective transformation and optical flow fine-tuning according to claim 1, characterized in that: Fine-tune the optical flow based on the affine transformation parameters and masks to achieve Burst image registration, including: S31: Calculate the highest scale initial optical flow in, represents the highest scale affine transformation matrix, E represents the identity matrix, and I represents the unit grid corresponding to the corresponding scale image; S32: Calculate the correlation between the fixed image and the distorted image through the Cost Volume, and calculate the fine-tuned optical flow based on the affine transformation matrix at the highest scale and the initial optical flow: in, represents the fine-tuned optical flow at the highest scale, ConvBlock1 represents the first convolutional block of the optical flow fine-tuning, CV represents the calculation of the correlation between images, F1 and M1 represent the fixed image and floating image at the highest scale respectively, They represent the output features of the highest scale fixed image F1 and the highest scale floating image M1 after passing through the serial HDSC modules, and Mask1 represents the highest scale mask; S33: At subsequent scales, the optical flow is fine-tuned using the affine transformation matrix at the corresponding scale and based on the correlation between its images; in, represents the fine-tuned optical flow at scale i, M i 、F i represent the floating image and fixed image at scale i respectively, Represents the i-th scale fixed image F i and the i-th scale floating image M i After passing through the series of HDSC modules, the output characteristics are: Indicates image distortion, * indicates dot product, Mask i represents the mask at scale i, represents the affine transformation matrix at scale i, and bilinearUp represents bilinear upsampling; S34: Repeat steps S31-S33 to obtain the highest scale fine-tuned optical flow and refine it Among them, convBlock2 is the second convolution block used to refine the optical flow; S35: Burst image registration is achieved by aligning the floating image with the fixed image through the refined optical flow φ.
9. The method for Burst image registration combining projective transformation and optical flow fine-tuning according to claim 8, characterized in that: The correlation between the fixed image and the distorted image is calculated by Cost Volume calculation, including: Among them, CV(F,M) represents the calculation of the Cost Volume between the fixed image F and the floating image M, that is, the local correlation between the fixed image and the floating image, d x d y Represents the lateral and axial offsets respectively.
10. The method for Burst image registration combining projective transformation and optical flow fine-tuning according to claim 8, characterized in that: The floating image is aligned with the fixed image by finally refining the optical flow to achieve Burst image registration, including: Among them, φ represents the refined optical flow, M represents the floating image, x and y represent the pixel coordinates, Δx and Δy represent the displacement of each pixel, respectively. Represents image distortion.