Large-parallax image splicing method

By learning the homography matrix and implicit feature representation of large disparity images, combining codec networks and hidden neural networks, the problem of geometric and photometric difference processing in large disparity image stitching is solved, and high-quality and robust image stitching effect is achieved.

CN120198286AActive Publication Date: 2025-06-24ORIENTAL MIND (WUHAN) COMPUTING TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510671245.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-06-24
Estimated Expiration
2045-05-23

AI Technical Summary

Technical Problem

The prior art is difficult to accurately match feature points in large parallax image stitching, resulting in inaccurate estimation of transformation matrix, lack of adaptability and robustness, ignore the processing of photometric differences, resulting in misalignment, ghosting, blurring and other problems in the stitching results.

Method used

By learning the homography matrix of two large parallax images, the geometric transformation of the image is realized, and implicit feature representations are extracted to correct the photometric difference. The codec network and the hidden neural network are used for image stitching and lighting compensation.

Benefits of technology

The quality and robustness of large-parallel image stitching are improved, and the geometric and photometric differences of images can be adaptively processed, stitching traces are reduced, and the adaptability and robustness of the stitching effect are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198286A_ABST
    Figure CN120198286A_ABST
Patent Text Reader

Abstract

The invention provides a large-parallax image splicing method. The large-parallax image splicing method comprises the following steps: extracting an overlapping region of two large-parallax images through a coding and decoding network; selecting one overlapping region as a reference region based on edge detection, and learning homography matrixes of the two overlapping regions; based on the homography matrix, the non-reference large parallax images are transformed, and the transformed large parallax images are spliced and fused; and extracting illumination compensation information represented by the implicit features of the large-parallax image after the preliminary splicing and fusion, and performing illumination compensation on the preliminary splicing image based on the illumination compensation information to obtain a final splicing and fusion image. According to the method, the homography matrixes of the two large-parallax images are learned to realize effective splicing of the images, and on the basis, the implicit feature representation of the large-parallax image pair is further extracted to correct the luminosity difference of the images so as to obtain a high-quality large-parallax image splicing fusion result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing, and more particularly, to a method for stitching large parallax images. Background Art

[0002] Stitching large parallax images is an important task in the field of computer vision, which aims to seamlessly fuse multiple images with significant viewing angle differences into a panoramic image. However, in practical applications, due to the influence of various factors such as shooting angles, lighting conditions, and camera parameters, there are often obvious geometric and photometric differences between the images to be stitched, resulting in problems such as misalignment, ghosting, and blurring in the stitching results.

[0003] Most traditional image stitching methods are based on feature matching and transformation matrix estimation, and these methods can achieve good results when the parallax is small. However, for large parallax image stitching, it is often difficult for these methods to accurately match feature points, resulting in inaccurate transformation matrix estimation, which in turn affects the stitching effect. In addition, these methods usually require manual setting of thresholds and parameters, and different adjustments are needed for different images and scenarios, lacking adaptability and robustness.

[0004] In recent years, deep learning methods have made remarkable progress in the field of image processing, and have also provided new ideas for large parallax image stitching. Some deep learning-based image stitching methods learn the transformation relationship between images by training neural networks, achieving relatively accurate stitching. However, most of these methods rely on a large amount of labeled data for supervised learning, and in practical applications, it is very difficult to obtain labeled large parallax image data. And many methods usually only focus on the geometric transformation between images, while ignoring the processing of photometric differences, resulting in obvious stitching traces in the areas with inconsistent lighting in the stitching results. Summary of the Invention

[0005] Aiming at the technical problems existing in the prior art, the present invention provides a method for stitching large parallax images, which realizes effective stitching of images by learning the homography matrix of two large parallax images, and further extracts the implicit feature representation of the large parallax image pair on this basis to correct the photometric differences of the images and obtain a higher-quality large parallax image stitching and fusion result.

[0006] The present invention provides a method for stitching large parallax images, including: Inputting two large parallax images into a first encoder-decoder network to obtain two overlapping regions, where the two overlapping regions refer to the regions where the two large parallax images overlap each other; Based on edge detection, selecting one of the two overlapping regions as a reference region and the other overlapping region as a non-reference region; Input the two overlapping regions into a second encoding and decoding network to obtain a homography matrix, which represents the transformation relationship from the non-reference region to the reference region; Based on the homography matrix, transform the first large-disparity image, and fuse the aligned regions and non-overlapping regions in the transformed first large-disparity image at the image pixel level to obtain an aligned first large-disparity image, where the first large-disparity image refers to the large-disparity image where the non-reference region is located, the aligned region refers to the overlapping region in the transformed first large-disparity image, and the non-overlapping region refers to the region outside the overlapping region of the first large-disparity image; Stitch the aligned first large-disparity image with the second large-disparity image to obtain a preliminary stitched image, where the second large-disparity image refers to the large-disparity image where the reference region is located; Input the two overlapping regions into an implicit neural network to output illumination compensation information, and perform illumination compensation on the preliminary stitched image based on the illumination compensation information; Fuse the illumination-compensated stitched image with the preliminary stitched image to obtain a final stitched image.

[0007] A method for stitching large-disparity images provided by the present invention extracts the overlapping regions of two large-disparity images through an encoding and decoding network; selects one of the overlapping regions as a reference region based on edge detection, and learns the homography matrix of the two overlapping regions; based on the homography matrix, transforms the non-reference large-disparity image, and stitches and fuses the transformed large-disparity image; extracts the illumination compensation information represented by the implicit features of the preliminarily stitched and fused large-disparity image, and performs illumination compensation on the preliminary stitched image based on the illumination compensation information to obtain a final stitched and fused image. The present invention realizes effective stitching of images by learning the homography matrix of two large-disparity images. On this basis, the implicit feature representation of the large-disparity image pair is further extracted to correct the photometric differences of the images, and a higher-quality large-disparity image stitching and fusion result is obtained. Description of the Drawings

[0008] Figure 1 It is a flowchart of a method for stitching large-disparity images provided by an embodiment of the present invention; Figure 2 It is an architecture diagram of the first encoding and decoding network provided by an embodiment of the present invention; Figure 3 It is an architecture diagram of the second encoding and decoding network provided by an embodiment of the present invention; Figure 4 It is an architecture diagram of the implicit neural network provided by an embodiment of the present invention; Figure 5 It is an architecture diagram of the feature fusion network provided by an embodiment of the present invention; Figure 6This is the overall network architecture diagram for large parallax image stitching provided by the embodiments of the present invention; Figure 7 This is a schematic structural diagram of a large parallax image stitching system provided by the embodiments of the present invention. Specific embodiments

[0009] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. Additionally, the technical features in each embodiment or individual embodiment provided by the present invention can be combined with each other arbitrarily to form a feasible technical solution. This combination is not restricted by the order of steps and / or the structural composition mode, but must be based on what can be achieved by those of ordinary skill in the art. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.

[0010] In response to the urgent needs of the prior art, the present invention provides a large parallax image stitching method, which should be able to adaptively learn the transformation relationship between images and simultaneously consider the processing of geometric and photometric differences to achieve high-quality large parallax image stitching. Specifically, through the method of implicit feature representation and homography matrix learning, where the homography matrix learning mainly deals with the geometric transformation of the stitched images, and the implicit feature representation mainly learns the photometric differences between the two images, unsupervised large parallax image stitching is achieved, improving the adaptability and robustness of the stitching effect and providing strong technical support for related applications in the field of computer vision.

[0011] Figure 1 This is a flowchart of a large parallax image stitching method provided by the present invention. As Figure 1 shown, the method includes: Step 1, input two large parallax images into the first encoder-decoder network to obtain two overlapping regions, where the two overlapping regions refer to the regions where the two large parallax images overlap each other.

[0012] It can be understood that for two large parallax RGB images with different illuminations, the overlapping regions of the two large parallax images are extracted through the first encoder-decoder network to obtain two overlapping regions, including the first overlapping region in the first large parallax image overlapping with the second region in the second large parallax image, and these two overlapping regions have the most identical content.

[0013] Among them, the first-stage decoding network includes three encoders and three decoders. Each of the encoders includes a densely connected bottleneck structure residual block and a spatial attention block, and each of the decoders includes a densely connected bottleneck structure residual block and a weighted spatial attention block; inputting two large-disparity images into the first-stage decoding network to obtain two overlapping regions includes: Input the two large-disparity images into the first-stage decoding network respectively. Through the three bottleneck structure residual blocks in the encoder, extract the local features of three scales of each large-disparity image respectively, and aggregate the local features of three scales through the three spatial attention blocks in the encoder by downsampling; Restore the aggregated local features of three scales through the three residual blocks in the decoder, and perform weighted fusion on the local features of three scales by upsampling through three weighted spatial attention blocks, and output the overlapping region attention map of each large-disparity image; According to each large-disparity image and the corresponding overlapping region attention map, calculate the overlapping region of each large-disparity image.

[0014] See Figure 2 , which is the network structure diagram of the first-stage decoding network. The first-stage decoding network processes large-disparity image pairs through an unsupervised densely connected encoding and decoding network. The main part of this first-stage decoding network is an encoding and decoding network based on U-Net, and obtains the overlapping region attention map by extracting and fusing two large-disparity images with different illuminations. The first-stage decoding network specifically includes three encoders and three decoders. The basic blocks of each encoder and each decoder are composed of a densely connected bottleneck structure residual block and a spatial attention block. The specific implementation of the unsupervised densely connected encoding and decoding network is as follows: Encoder part: The basic block of each encoder is composed of a densely connected bottleneck structure residual block and a spatial attention. Each bottleneck structure residual block extracts the local features of the large-disparity image through a series of convolution operations and skip connections, and then aggregates them by the spatial attention block. The three encoders are in a multi-scale progressive relationship, and downsample the feature map to reduce the spatial resolution and computational complexity.

[0015]

[0016] Among them, represents the output feature map of the i-th layer encoder, and DenseBlock represents the operation of the densely connected bottleneck structure residual block.

[0017] Decoder part: The basic block of each decoder is composed of a densely connected bottleneck structure residual block and a weighted spatial attention. Each bottleneck structure residual block restores the local features of the image through a series of operations, and adopts a weighted spatial attention mechanism to fuse the features transmitted from the encoder.

[0018] The three decoders have a step-by-step restoration relationship. By upsampling the feature maps, the spatial resolution of the image is gradually restored.

[0019]

[0020] Among them, represents the output feature map of the i-th layer decoder, DenseBlock represents the operation of the bottleneck structure residual block with dense connections, and Fusion represents the weighted spatial attention fusion operation, where the output features of the current decoder and the features of the corresponding size of the encoder are fused. The decoder outputs the attention map of the overlapping region of two large-disparity images.

[0021] Overlapping region estimation: Multiply the overlapping region attention map by the input large-disparity image and add them to obtain the estimation of the overlapping region of the large-disparity image.

[0022]

[0023] Among them, represents the estimated overlapping region, represents the input large-disparity image, and A represents the overlapping region attention map. The overlapping regions of the two large-disparity images are estimated according to this formula.

[0024] Among them, for the training of the first encoding and decoding network, first, a training set and a test set of large-disparity image pairs are constructed. Specifically, the ELA algorithm is used to splice the images from the UDISD dataset to collect a large number of real spliced images. Ignore those images whose extrapolated area is less than 10% of the entire image to ensure that the spliced images have sufficient disparity and irregular boundaries. A content-aware warping algorithm is used to generate a rich variety of different mesh deformation matrices from these real spliced images, and these meshes will be used in the subsequent rectification process. Apply the inverse matrix of the mesh deformation to warp the real rectangular image to the synthetic spliced image, and 5705 available samples are retained from more than 60000 samples. Manually screen the synthetic images and eliminate those with obvious distortions. 653 distortion-free samples are selected from the initially collected large number of real spliced images and added to the training set to improve the diversity and generalization ability of the dataset. Generally speaking, the dataset includes 5839 samples for training and 519 samples for testing. The resolution of each image in the dataset is 512×384.

[0025] Input the training set images into Figure 2The first codec network for estimating the overlapping region is trained until the loss function reaches a preset convergence condition. Among them, the structure of the codec network model for estimating the overlapping region includes three stages. The basic blocks of the encoder and decoder are both composed of a bottleneck structure residual block with dense connections and a spatial attention block.

[0026] The input large-disparity image pair first undergoes 3×3 convolution for channel upsampling and preliminary feature extraction to obtain a tensor of H×W×C. Then, it traverses the three-stage encoder part. That is, in the first stage, this tensor passes through the first dense connection bottleneck structure residual block of the encoder to obtain a tensor of the same size; in the second stage, this tensor passes through the second residual block for feature extraction and then undergoes downsampling to obtain a tensor of 1 / 2H×1 / 2W×2C; in the third stage, the downsampled tensor passes through the third residual block for feature extraction and then undergoes downsampling to obtain a tensor of 1 / 4H×1 / 4W×4C. Similarly, the decoder also has three similar stages. First, a tensor of shape 1 / 4H×1 / 4W×4C passes through a residual block for decoding, then undergoes upsampling, and then passes through the residual block of the next stage for decoding. In the middle of the encoding and decoding, skip connections are made to fuse the features of the encoding and decoding. Specifically, the output of the third-stage encoder is upsampled through 3×3 convolution and then concatenated with the output of the second-stage encoder. The concatenated tensor undergoes 1×1 convolution to reduce the channels, and then the output tensor is connected with the output tensor of the third-stage decoder through a residual connection. In the last stage, the output image is regarded as the initial overlapping region attention map. Then, this attention map is multiplied by the input image and then connected with the input image through a residual connection to obtain the final estimation of the overlapping region of the large-disparity image.

[0027] During the training of the first codec network, the loss of the first codec network is calculated based on the L2 loss function and the frequency-domain loss function, which is used to measure the difference between the overlapping region estimation and the true overlapping region and guide the training of the first codec network. The calculation formula for the loss of the first codec network is:

[0028] Among them, represents the total loss of the first codec network, represents the L2 loss function, represents the frequency-domain loss function, represents the overlapping region, represent two large-disparity images respectively.

[0029] Among them, regarding the loss function, in the unsupervised large-disparity image stitching method of the present invention, in order to ensure the quality and consistency of the generated images at each stage, multiple loss functions are applied during network training, including mean square error loss, frequency domain loss, and content alignment loss. The different loss functions are introduced as follows: 1. L2 loss (Mean square error loss) The L2 loss, also known as the mean square error loss, is used to measure the difference between the model's predicted value and the true value. It works by calculating the sum of the squares of the differences between the predicted value and the true value, giving a greater penalty to large errors, thereby encouraging the model to predict outputs closer to the true values.

[0030] For the estimation of the overlapping region of the large-disparity image pair, the L2 loss can be expressed as:

[0031] Where represents the L2 loss, represents the predicted pixel value of the model at position , represents the pixel value of the true overlapping region at position .

[0032] 2. Frequency domain loss The frequency domain loss is calculated in the frequency domain of the image, which takes into account the overall structure and texture information of the image. By penalizing the differences between the predicted image and the true image in the frequency domain, the frequency domain loss helps the model better understand the structural relationships between images, thereby more accurately estimating the overlapping region.

[0033] The frequency domain loss based on the Fourier transform can be expressed as:

[0034] Where represents the frequency domain loss, and represent the values of the Fourier transforms of the predicted image and the true image at the frequency domain position respectively. Here represents the modulus of the complex number.

[0035] Through the processing of the above unsupervised densely connected encoder-decoder network, the overlapping regions of two large-disparity images with different illuminations are effectively estimated, providing a basis for subsequent image fusion or reconstruction tasks.

[0036] Step 2, select one of the overlapping regions as the reference region and the other overlapping region as the non-reference region based on edge detection.

[0037] It is understandable that after obtaining the estimation of the overlapping regions of the two large-disparity images, by using the information of vertical edges and horizontal edges, an overlapping region is selected from the overlapping regions of the two large-disparity images as the reference region for subsequent processing, and the overlapping region of the other large-disparity image is used as the non-reference region.

[0038] In a possible embodiment of the present invention, the method of selecting one overlapping region as the reference region and the other overlapping region as the non-reference region from the two overlapping regions based on edge detection includes: performing gamma transformation on the two overlapping regions respectively to obtain the two overlapping regions after gamma transformation; detecting edge pixel points of the two overlapping regions after gamma transformation respectively based on Sobel operator edge detection; counting the number of edge pixel points of each overlapping region; taking the overlapping region with more edge pixel points as the reference region and the overlapping region with fewer edge pixel points as the non-reference region.

[0039] Specifically, in order to effectively improve the distinguishability of the overlapping regions, the estimation maps of the overlapping regions of the two output large-disparity images are first subjected to gamma transformation with a set value of 2. This means that pixels with lower gray values (usually corresponding to the dark parts or regions with rich details in the image) will be relatively amplified, while pixels with higher gray values (bright parts) will increase more slowly. This transformation can not only enhance the overall contrast of the image, making the information in the image more abundant and hierarchical, but also is particularly conducive to sharpening and highlighting the edges of the overlapping regions.

[0040] Gamma transformation: For the overlapping region estimation images of the large-disparity image pair obtained in step 1 1 and 2 are respectively subjected to gamma transformation to enhance the contrast of the image and make the edges of the overlapping regions more prominent. The formula of gamma transformation is as follows:

[0041] Wherein, represents the overlapping region image, is the gamma value, is the overlapping region image after transformation.

[0042] After completing the Gamma transformation of the overlapping region of two large-disparity images to enhance contrast and highlight edge features, the Sobel operator is then used for edge detection. Specifically, the horizontal Sobel operator and the vertical Sobel operator are respectively applied to the Gamma-transformed image. By means of convolution operations, the horizontal gradient and the vertical gradient of the image at each pixel point are calculated, obtaining two results, which respectively represent the edge intensities of the original image in the horizontal and vertical directions. The Otsu threshold segmentation method is applied to the calculated gradient image. This method finds an optimal threshold by analyzing the grayscale histogram of the image, marks the pixels with gradient values higher than this set threshold as edges, and the pixels with values lower than the threshold are regarded as non-edges, thereby obtaining a binary edge image. This step ensures that only significant edges are retained while noise and minor variations are suppressed.

[0043] Specifically, the Sobel operator is selected for edge detection on the overlapping regions of the two Gamma-transformed images, the edge pixel points of each overlapping region are extracted, the numbers of edge pixel points in the two overlapping regions are respectively counted, the counting results are compared, and the overlapping region with a larger number of edge pixel points is set as the reference region, and the other overlapping region is set as the non-overlapping region. Visual inspection and necessary post-processing, such as morphological operations (dilation, erosion, etc.), are performed on the determined reference region to remove isolated points or refine the edges to ensure the accuracy and coherence of the reference region.

[0044] In a possible embodiment of the present invention, based on the Sobel operator edge detection, the horizontal edge pixel points and the vertical edge pixel points of the two overlapping regions after Gamma transformation are respectively detected, including: Calculating the horizontal gradient value of each pixel point in each overlapping region and the longitudinal gradient value ; According to the horizontal gradient value of each pixel point and the longitudinal gradient value , calculating the gradient magnitude of each pixel point; Based on the gradient magnitude of each pixel point, determining whether each pixel point is an edge pixel point.

[0045] It can be understood that for each pixel point in the overlapping region, calculating the horizontal gradient value of each pixel point in each overlapping region and the longitudinal gradient value , including: Obtaining the horizontal gradient operator Gx and the vertical gradient operator Gy; Convolving the horizontal gradient operator with each pixel point to obtain the horizontal gradient value of each pixel point ; Convolve the vertical gradient operator with each pixel to obtain the longitudinal gradient value of each pixel .

[0046] Among them, the horizontal gradient operator is as follows: ; The vertical gradient operator is as follows: ; The horizontal gradient value Gx(x, y) of each pixel is: , and the longitudinal gradient value of each pixel is: .

[0047] The calculation of the gradient magnitude of each pixel according to the horizontal gradient value and the longitudinal gradient value of each pixel includes: ; Among them, is the gradient magnitude of each pixel.

[0048] The determination of whether each pixel is an edge pixel based on the gradient magnitude of each pixel includes: When the gradient magnitude of the pixel is greater than the preset magnitude threshold, the pixel is an edge pixel; otherwise, the pixel is a non-edge pixel.

[0049] Detect the edge pixels in the two overlapping regions respectively through Sobel operator edge detection, select the overlapping region with more edge pixels as the reference region, and use the overlapping region with fewer edge pixels as the non-reference region.

[0050] Step 3: Input the two overlapping regions into the second encoder-decoder network to obtain a homography matrix, and the homography matrix represents the transformation relationship from the non-reference region to the reference region.

[0051] It can be understood that in order to effectively stitch two large-disparity images, the homography matrix of the reference region is learned through the second encoder-decoder network, so as to transform the non-reference region and align the non-reference region with the reference region in terms of geometric content. The specific implementation steps are as follows: Input the two overlapping regions (one is the reference region and the other is the non-reference region) into the multi-scale encoder-decoder network (the second encoder-decoder network). The output of the second encoder-decoder network is a multi-grid homography matrix, which describes the transformation relationship from the non-reference region to the reference region. The multi-grid homography matrix can be regarded as a collection of multiple local homography matrices, and each local homography matrix is responsible for the transformation of a grid region.

[0052] Among them, reference can be made to Figure 3 , which is the architecture diagram of the second-stage decoding network. Its working principle is as follows: First, the two estimated overlapping regions are subjected to feature extraction through convolutional residual blocks to obtain a tensor of H×W×C. Then, this tensor is flattened into a shape of N×C (N = H×W), and then these two N×C tensors are multiplied matrix-wise to calculate the correlation between the resulting feature vectors. By comparing the distances between the feature point descriptors on each pair of images, the feature point with the smallest distance to each feature point is selected as the matching point. The multiplied N×N tensor is reshaped into H×W×N, and a homography matrix is obtained through an MLP module composed of three consecutive convolutional layers and two fully connected layers. To prevent the network model from overfitting, a Dropout layer is added before each fully connected layer, and the dropout probability value is set to 0.5. The output of the MLP module is 8 real numbers. By applying the direct linear transformation (DLT) algorithm to the obtained 8 real numbers, a 3×3 form homography matrix between the two images can be obtained.

[0053] Step 4: Based on the homography matrix, transform the first large-disparity image, and fuse the pixel points of the aligned region and the non-overlapping region in the transformed first large-disparity image to obtain the aligned first large-disparity image, where the first large-disparity image refers to the large-disparity image where the non-reference region is located, the aligned region refers to the overlapping region in the transformed first large-disparity image, and the non-overlapping region refers to the region outside the overlapping region in the first large-disparity image.

[0054] It can be understood that by using the homography matrix, the non-reference region is transformed to the same perspective as the reference region to obtain the aligned region. In the present invention, by using the homography matrix, the non-reference large-disparity image is transformed to obtain the transformed non-reference large-disparity image. Among them, the overlapping region of the transformed non-reference large-disparity image is called the aligned region. Due to the perspective difference, at the junction of the aligned region and the non-overlapping region, a smooth transition method is adopted to fuse the image pixel points to eliminate possible seams or inconsistencies. Finally, the aligned large-disparity image is updated.

[0055] Among them, the overall architecture of the second-stage decoding network is the same as that of the first-stage decoding network in Step 1, and is realized through downsampling, upsampling, and feature fusion. This multi-scale feature can better capture the detailed information and context information of the image, thereby improving the accuracy and quality of image reconstruction.

[0056] Among them, when performing unsupervised learning on the second-stage decoding network, its loss function is composed of a content alignment loss and a smooth transition loss, which is used to measure the difference between the stitched perspective image and the real perspective image, and guide the training of the second-stage decoding network. The calculation formula of its loss function is:

[0057] Among them, represents the total loss of the second - stage decoding network, represents the content alignment loss function, represents the smooth transition loss function, represents the stitched image.

[0058] For the content alignment loss function and the smooth transition loss function, they are introduced as follows: 1. Content alignment loss In the encoding - decoding network for learning the homography matrix, the content alignment loss ensures that the transformed non - reference region is highly consistent with the reference region in content. By minimizing the content alignment loss, the model can learn a more accurate homography matrix, thereby achieving precise alignment between images.

[0059] The content alignment loss can be calculated by comparing the pixel differences between the reference region and the non - reference region transformed by the multi - grid homography matrix. The formula is as follows:

[0060] Among them, represents the content alignment loss, N and M respectively represent the height and width of the image, represents the pixel value of the reference region at position , represents the pixel value of the non - reference region transformed by the multi - grid homography matrix at the corresponding position. This formula calculates the squared difference of all pixel points between the reference region and the transformed region, and obtains the final content alignment loss through summation and averaging.

[0061] 2. Smooth transition loss The smooth transition loss ensures that at the junction of the aligned region and the non - overlapping region, the image pixel points can transition smoothly, thereby reducing visual artifacts and abruptness during fusion. By minimizing the smooth transition loss, the model can learn a more natural image fusion method, making the aligned image more visually coherent and consistent.

[0062] Based on the smooth transition loss using the gradient difference of pixel values, the formula is as follows:

[0063] Among them, represents the smooth transition loss, N represents the number of pixel points at the junction, and B represents the set of pixel points at the junction of the aligned region and the non - overlapping region. Represents the aligned image, Represents the image of the non-overlapping region, and respectively represent the gradient calculations of the image in the horizontal and vertical directions. This formula calculates the gradient differences of the pixels at the junction in the horizontal and vertical directions, and obtains the final smooth transition loss through summation and averaging.

[0064] Step 5: Stitch the first largest disparity image after alignment with the second largest disparity image to obtain a preliminary stitched image, where the second largest disparity image refers to the large disparity image where the reference region is located.

[0065] It can be understood that through the selection of the reference region in step 2 and the extraction of the homography matrix, the non-reference large disparity image is subjected to an alignment transformation through the homography matrix, and the aligned large disparity image and the reference large disparity image are stitched to form a preliminary stitched image.

[0066] Step 6: Input the two overlapping regions into the implicit neural network to output illumination compensation information, and perform illumination compensation on the preliminary stitched image based on the illumination compensation information.

[0067] It can be understood that the two overlapping regions are input into the implicit neural representation network. By training the implicit neural network to learn a photometric adjustment function, this function can adaptively adjust the brightness of the image according to the photometric characteristics of the image. The implicitly represented features after photometric adjustment and the aligned image are input into the convolutional neural network for feature fusion and image reconstruction.

[0068] In a possible embodiment of the present invention, inputting the two overlapping regions into the implicit neural network to output illumination compensation information and performing illumination compensation on the preliminary stitched image based on the illumination compensation information includes: Respectively obtain the coordinate positions and content windows of each pixel point in each overlapping region; Input the position and content window of each pixel point in the two overlapping regions into the implicit neural network to output illumination compensation information.

[0069] See Figure 4 , which is the architecture diagram of the implicit neural network. The implicit neural representation reconstructs or generates an image by learning a mapping function from continuous coordinates (such as spatial coordinates) to image features or attributes. First, the aligned region of the input preliminary stitched large disparity image is converted from the RGB color gamut to the HSV color gamut, and the enhancement process is redefined by mapping the two-dimensional coordinates of the image with uneven brightness to its illumination component (i.e., V).

[0070] The position coordinates are obtained by generating uniformly distributed coordinate points in the height and width directions respectively and combining these points into a three-dimensional coordinate grid, where each point represents a specific position in the image or feature map. The content window is obtained by performing a convolution operation on the padded input image using a specially designed convolution kernel (i.e., taking the pixel values of the position coordinates and the coordinates in the surrounding area), thereby extracting a fixed-size image region corresponding to each position. The prepared position coordinates and the values of their content windows are input into the hidden neural network. The hidden neural network first processes the content window features and the position coordinates separately, and the output network combines the information of both for the final prediction or regression. Among them, these three parts are all through an nn.Linear linear layer, which is used to perform a linear transformation on the input. The activation function of the layer is the sine function, but it is changed to the Sigmoid function in the last layer to adapt to different output requirements. At the same time, it also accepts a w0 parameter to control the frequency of the sine function, which helps the network capture features of different frequencies. During the training process, different weight decay strategies are applied. The parameters of the network are divided into different groups, and each group of parameters has its corresponding weight decay value, with default values of 0.1, 0.0001, and 0.001. This strategy helps prevent the network from overfitting and improves the generalization ability of the model.

[0071] The illumination compensation information of the image is output through the hidden neural network. Based on the illumination compensation information, the preliminary stitched image in step 5 is subjected to illumination compensation. The steps of illumination compensation include: Convert the initial stitched image from the RGB color gamut to the HSV color gamut, and extract the V component in the HSV color gamut; Based on the illumination compensation information and the V component, calculate the corrected illumination intensity, and use the corrected illumination intensity as the new V component; Among them, the calculation formula for the corrected illumination intensity is: ; The new V component is , is the corrected illumination intensity, is the illumination compensation information; Fuse the new V component with the H component and the S component in the HSV color gamut to obtain the stitched image after illumination compensation, and finally transform the stitched image from the HSV color gamut back to the RGB color gamut. Obtain the stitched image after illumination compensation.

[0072] Step 7, fuse the stitched image after illumination compensation with the preliminary stitched image to obtain the final stitched image.

[0073] Among them, the spliced image after illumination compensation and the preliminary spliced image are input into a feature fusion network. Through three residual blocks in the feature fusion network, the spliced image after illumination compensation and the preliminary spliced image are subjected to residual connection to obtain a final spliced image.

[0074] See Figure 5 , which is the architecture diagram of the feature fusion network. The spliced image after illumination compensation in step 6 and the preliminary spliced image output in step 5 are fed into Net4 together. Through 3 ResBlock residual blocks, the extracted result is subjected to residual connection with the spliced image estimated, compensated, and corrected for illumination information in step 5 to obtain a final fused spliced image.

[0075] See Figure 6 , which is the overall network architecture of the large-disparity image splicing method of the present invention. Its working principle is as follows: Two large-disparity images are input into the first encoder-decoder network to output two overlapping regions. One of the overlapping regions is selected as the reference region, and the other overlapping region is the non-reference region. The reference region and the non-reference region are input into the second encoder-decoder network to output a homography matrix. Based on the homography matrix, the non-reference large-disparity image is transformed to the same perspective as the reference large-disparity image, and the pixel points in the overlapping region and the non-overlapping region are smoothly transitioned to obtain a preliminary spliced image. The two overlapping regions are input into a hidden neural network to output illumination compensation information. Based on the illumination compensation information, the preliminary spliced image is subjected to illumination compensation. Finally, the spliced image after illumination compensation and the preliminary spliced image are spliced and fused to form a final spliced and fused large-disparity image.

[0076] See Figure 7 , which is a large-disparity image splicing system provided by the present invention. The system includes: A first acquisition module 701, configured to input two large-disparity images into the first encoder-decoder network to obtain two overlapping regions, where the two overlapping regions refer to the regions where the two large-disparity images overlap each other; A determination module 702, configured to determine one of the overlapping regions as the reference region and the other overlapping region as the non-reference region based on edge detection from the two overlapping regions; A second acquisition module 703, configured to input the two overlapping regions into the second encoder-decoder network to obtain a homography matrix, where the homography matrix represents the transformation relationship from the non-reference region to the reference region; A transformation module 704 is configured to transform the first large-disparity image based on the homography matrix, and fuse the image pixel points in the aligned region and the non-overlapping region in the transformed first large-disparity image to obtain the aligned first large-disparity image, where the first large-disparity image refers to the large-disparity image where the non-reference region is located, the aligned region refers to the overlapping region in the transformed first large-disparity image, and the non-overlapping region refers to the region outside the overlapping region of the first large-disparity image; An image stitching module 705 is configured to stitch the aligned first large-disparity image with the second large-disparity image to obtain a preliminary stitched image, where the second large-disparity image refers to the large-disparity image where the reference region is located; and fuse the stitched image after illumination compensation with the preliminary stitched image to obtain a final stitched image; An illumination compensation module 706 is configured to input the two overlapping regions into a hidden neural network, output illumination compensation information, and perform illumination compensation on the preliminary stitched image based on the illumination compensation information.

[0077] It can be understood that a large-disparity image stitching system provided by the present invention corresponds to the large-disparity image stitching method provided by the foregoing embodiments. The relevant technical features of the large-disparity image stitching system can refer to the relevant technical features of the large-disparity image stitching method, which will not be elaborated here.

[0078] A large-disparity image stitching method provided by an embodiment of the present invention extracts the overlapping regions of two large-disparity images through an encoding and decoding network; selects one of the overlapping regions as a reference region based on edge detection, and learns the homography matrix of the two overlapping regions; based on the homography matrix, transforms the non-reference large-disparity image, and stitches and fuses the transformed large-disparity image; extracts the illumination compensation information represented by the implicit features of the preliminary stitched and fused large-disparity image, and performs illumination compensation on the preliminary stitched image based on the illumination compensation information to obtain a final stitched and fused image. The present invention realizes effective stitching of images by learning the homography matrix of two large-disparity images, and further extracts the implicit feature representation of the large-disparity image pair on this basis to correct the photometric difference of the images and obtain a higher-quality large-disparity image stitching and fusion result.

[0079] It should be noted that in the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0080] Those skilled in the art will understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0081] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0082] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0083] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0084] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications to these embodiments once they know the basic inventive concept. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0085] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. A large parallax image stitching method, characterized in that, Including: Input two large-disparity images into the first encoder-decoder network to obtain two overlapping regions, where the two overlapping regions refer to the regions where the two large-disparity images overlap with each other; Based on edge detection, select one of the two overlapping regions as the reference region and the other overlapping region as the non-reference region; Input the two overlapping regions into the second encoder-decoder network to obtain a homography matrix, and the homography matrix represents the transformation relationship from the non-reference region to the reference region; Based on the homography matrix, transform the first large-disparity image, and fuse the pixel points of the aligned region and the non-overlapping region in the transformed first large-disparity image to obtain the first large-disparity image after alignment, where the first large-disparity image refers to the large-disparity image where the non-reference region is located, the aligned region refers to the overlapping region in the transformed first large-disparity image, and the non-overlapping region refers to the region outside the overlapping region of the first large-disparity image; Stitch the first large-disparity image after alignment with the second large-disparity image to obtain a preliminary stitched image, where the second large-disparity image refers to the large-disparity image where the reference region is located; Input the two overlapping regions into the hidden neural network to output illumination compensation information, and perform illumination compensation on the preliminary stitched image based on the illumination compensation information; Fuse the stitched image after illumination compensation with the preliminary stitched image to obtain the final stitched image.

2. The large parallax image stitching method according to claim 1, wherein The first encoder-decoder network includes three encoders and three decoders. Each encoder includes a densely connected bottleneck structure residual block and a spatial attention block, and each decoder includes a densely connected bottleneck structure residual block and a weighted spatial attention block; The step of inputting two large-disparity images into the first encoder-decoder network to obtain two overlapping regions includes: Input the two large-disparity images into the first encoder-decoder network respectively, extract the local features of three scales of each large-disparity image through the three bottleneck structure residual blocks in the encoder, and aggregate the local features of three scales by downsampling through the three spatial attention blocks in the encoder; Restore the aggregated local features of three scales through the three bottleneck structure residual blocks in the decoder, and perform weighted fusion on the local features of three scales by upsampling through the three weighted spatial attention blocks to output the overlapping region attention map of each large-disparity image; Calculate the overlapping region of each large-disparity image according to each large-disparity image and the corresponding overlapping region attention map.

3. The large parallax image stitching method according to claim 2, wherein The step of calculating the overlapping region of each large-disparity image according to each large-disparity image and the corresponding overlapping region attention map includes: Among them, represents the overlapping region of the large parallax image, represents the input large parallax image, and A represents the overlapping region attention map.

4. The large parallax image stitching method according to claim 1 or 2, characterized in that When training the first encoder-decoder network, calculate the loss of the first encoder-decoder network based on the L2 loss function and the frequency domain loss function, and the calculation formula is: Among them, represents the loss of the first-stage decoding network, represents the L2 loss function, represents the frequency-domain loss function, represents the overlapping region, respectively represent two large-disparity images.

5. The method for stitching large parallax images according to claim 1, wherein The step of selecting one overlapping region as the reference region and the other overlapping region as the non-reference region based on edge detection from the two overlapping regions includes: Perform gamma transformation on the two overlapping regions respectively to obtain the two overlapping regions after gamma transformation; Based on Sobel operator edge detection, detect the edge pixel points of the two overlapping regions after gamma transformation respectively; Count the number of edge pixel points in each of the overlapping regions; Take the overlapping region with more edge pixel points as the reference region, and the overlapping region with fewer edge pixel points as the non-reference region.

6. The method for stitching large parallax images according to claim 5, characterized in that, The above-mentioned based on Sobel operator edge detection, detecting the horizontal edge pixel points and vertical edge pixel points of the two overlapping regions after gamma transformation respectively, includes: Calculate the horizontal gradient value Gx(x, y) and the vertical gradient value Gy(x, y) of each pixel point in each of the overlapping regions; Calculate the gradient magnitude of each pixel point according to the horizontal gradient value Gx(x, y) and the vertical gradient value Gy(x, y) of each pixel point; Based on the gradient magnitude of each pixel point, determine whether each pixel point is an edge pixel point.

7. The large parallax image stitching method according to claim 6, characterized in that The above-mentioned calculating the horizontal gradient value Gx(x, y) and the vertical gradient value Gy(x, y) of each pixel point in each overlapping region, includes: Obtain the horizontal gradient operator Gx and the vertical gradient operator Gy; Convolve the horizontal gradient operator with each pixel point to obtain the horizontal gradient value Gx(x, y) of each pixel point; Convolve the vertical gradient operator with each pixel point to obtain the vertical gradient value Gy(x, y) of each pixel point; Among them, the horizontal gradient operator is: ; The vertical gradient operator is: ; Among them, the horizontal gradient value Gx(x, y) of each pixel is as follows: , and the vertical gradient value Gy(x, y) of each pixel is as follows: ; According to the horizontal gradient value of each pixel point and the vertical gradient value , calculating the gradient magnitude of each pixel point, including: ; Among them, is the gradient magnitude of each pixel point; The above-mentioned based on the gradient magnitude of each pixel point, determining whether each pixel point is an edge pixel point, includes: When the gradient magnitude of the pixel point is greater than the preset magnitude threshold, the pixel point is an edge pixel point, otherwise, the pixel point is a non-edge pixel point.

8. The large parallax image stitching method according to claim 1, wherein Calculate the loss of the second encoder-decoder network based on the content alignment loss function and the smooth transition loss function, and the calculation formula is: Among them, represents the loss of the first-stage decoding network, represents the content alignment loss function, represents the smooth transition loss function, represents the large-disparity image after alignment.

9. The method for stitching large parallax images according to claim 1, characterized in that The above-mentioned inputting the two overlapping regions into the hidden neural network to output the illumination compensation information, includes: Obtain the coordinate position and content window of each pixel point in each of the overlapping regions respectively; Input the coordinate position and content window of each pixel point of the two overlapping regions into the hidden neural network to output the illumination compensation information, and the content window is the neighborhood of the pixel point; The above-mentioned performing illumination compensation on the preliminary stitching image based on the illumination compensation information, includes: Convert the initial stitching image from the RGB color gamut to the HSV color gamut, and extract the V component in the HSV color gamut; Based on the illumination compensation information and the V component, calculate the corrected illumination intensity, and take the corrected illumination intensity as the new V component; Among them, the calculation formula of the corrected illumination intensity is: ; The new V component is , is the corrected light intensity, is the light compensation information; Fuse the new V component with the H component and S component in the HSV color gamut, and restore it to the RGB color gamut to obtain the stitched image after illumination compensation.

10. The large parallax image stitching method according to claim 1, wherein The above-mentioned fusing the stitched image after illumination compensation with the preliminary stitched image to obtain the final stitched image, includes: Input the spliced image after illumination compensation and the preliminary spliced image into the feature fusion network. After three residual blocks in the feature fusion network, perform residual connection on the spliced image after illumination compensation and the preliminary spliced image to obtain the final spliced image.

Citation Information

Patent Citations

  • Uneven light compensation method

    CN105405110A

  • System and methods for correcting overlapping digital images of a panorama

    CN110651275A

  • Medical image segmentation method based on deep learning

    CN112150428A

  • Image splicing method and device based on deep learning

    CN116612006A

  • Wheat leaf disease monitoring method based on Internet of Things

    CN119516271A