Image splicing method for reconstructing joints based on feature difference perception

By using feature difference-perceptual reconstruction seams in image stitching technology, a two-stage network model is built for image registration and fusion, the problem of poor image fusion effect in the existing technology is solved, and high-quality and stable image stitching is achieved.

CN120235751APending Publication Date: 2025-07-01SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510242056.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

When existing image stitching technology deals with complex textures or images with large brightness differences, it is difficult to retain the original features of the image, resulting in poor fusion effect and problems of ghosting, blurring and edge discontinuity.

Method used

The image stitching method based on feature difference perception reconstruction seams is adopted, and an end-to-end efficient image stitching process is achieved by building a two-stage network model, including a deep semantic consistent image registration network and a feature difference perception reconstruction seams.

Benefits of technology

It significantly improves the quality and stability of image fusion, avoids ghosting, blurring and splicing traces, and improves the accuracy and visual effect of image splicing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120235751A_ABST
    Figure CN120235751A_ABST
Patent Text Reader

Abstract

The invention discloses an image splicing method for reconstructing a joint based on feature difference perception, and the method comprises the following steps: carrying out the preprocessing of an obtained image data set, dividing a training data set and a test data set, and enabling every two images with an overlapping region to be a pair of data; building an image splicing network model for reconstructing the seam based on the feature difference perception, wherein the image splicing network model comprises an image registration network based on depth semantic consistency and an image fusion network for reconstructing the seam based on the feature difference perception, which are connected in sequence; training an image registration network based on depth semantic consistency by using the training data set to obtain an optimal image registration network, and outputting an image registration result of the training set; training the image fusion network by using the image registration output of the training data set to obtain an optimal image fusion network; and inputting the test data set into the obtained optimal image registration network and the image fusion network to complete the splicing task of the test images.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to an image stitching method for reconstructing seams based on feature difference perception. Background Art

[0002] In the current era of accelerating digitalization, image stitching technology plays a crucial role in many fields such as cultural heritage protection, medical image processing, and geographical mapping. Its aim is to synthesize multiple images with overlapping regions into a complete and coherent image, thereby broadening the field of vision and providing a more comprehensive information display. This process mainly consists of key steps such as image acquisition, image registration, and image fusion. Image fusion, as one of the core steps of image stitching, is responsible for processing the pixel information in the overlapping regions between images to achieve seamless connection and transition. If the fusion effect is not good, it will cause problems such as ghosting, blurring, and discontinuous edges in the stitched image, seriously affecting the quality and usability of the final stitched image.

[0003] Among the existing image fusion algorithms, directly fusing the pixels in the overlapping regions of multiple images is a relatively simple method. Its advantages are fast calculation speed and simple and easy-to-understand algorithm; however, the disadvantages are also very obvious. Since the differences and structural information between images are not considered, it is easy to cause a decrease in the contrast of the fusion region and loss of details. Especially when processing images with complex textures or large brightness differences, it is impossible to effectively retain the original features of the images and it is difficult to obtain an ideal fusion effect. The seam generation algorithm is another common image fusion approach. Traditional seam generation methods usually determine the stitching relationship between images based on feature point matching and geometric transformation, and then generate seams. However, such methods have many disadvantages. They highly rely on the quality of feature point detection. In cases of low texture, low light, or image blurring, it is difficult to accurately extract high-quality feature points, resulting in inaccurate seam generation and ultimately affecting the success rate and quality of image stitching. Summary of the Invention

[0004] In order to overcome the disadvantages and deficiencies of the existing technology, the present invention provides an image stitching method for reconstructing seams based on feature difference perception. The present invention can effectively overcome the problems of difficult feature point detection and inaccurate seam generation, significantly improve the quality and stability of image fusion, and is applicable to application scenarios that require high-precision image stitching.

[0005] In order to achieve the above objectives, the present invention adopts the following technical solutions:

[0006] An image stitching method for reconstructing seams based on feature difference perception, comprising the following:

[0007] Preprocess the obtained image dataset and divide it into a training dataset and a test dataset;

[0008] Build an image stitching network model based on feature difference-aware reconstruction of seams. The image stitching network model includes an image registration network based on deep semantic consistency and an image fusion network based on feature difference-aware reconstruction of seams, which are connected in sequence. Use the training dataset to train the image registration network based on deep semantic consistency, obtain the optimal image registration network, and output the image registration results of the training set.

[0009] Use the image registration output of the training dataset to train the image fusion network and obtain the optimal image fusion network.

[0010] Input the test dataset into the obtained optimal image registration network and image fusion network to complete the stitching task of the test images.

[0011] Furthermore, the image registration network includes a deep semantic feature extraction module and a homography prediction module.

[0012] Furthermore, the deep semantic feature extraction module includes a Layer0 layer, a layer1 layer, a layer2 layer, and a layer3 layer connected in sequence. The outputs of the layer1 layer, layer2 layer, and layer3 layer constitute the deep semantic feature I. s 。

[0013] Furthermore,

[0014] The Layer0 layer includes a convolutional layer and a max pooling layer.

[0015] The layer1 layer includes one Bottleneck1 basic module and two Bottleneck2 basic modules.

[0016] The layer2 layer includes one Bottleneck1 basic module and three Bottleneck2 basic modules.

[0017] The layer3 layer includes one Bottleneck1 basic module and five Bottleneck2 basic modules.

[0018] Furthermore, the Bottleneck1 basic module includes a first branch and a second branch.

[0019] The first branch is composed of three convolutional layers connected in sequence. The second branch includes a convolutional layer. The output results of the first branch and the second branch are added and then passed through the ReLU activation function to obtain the output of the Bottleneck1 basic module.

[0020] Furthermore, the BottleNeck2 basic module includes three convolutional layers connected in sequence. The outputs of the three convolutional layers are added to the input of the BottleNeck2 basic module, and then passed through the ReLU activation function to obtain the output of the BottleNeck2 basic module.

[0021] Furthermore, the image fusion network includes a feature difference perception module and a seam generator.

[0022] Furthermore,

[0023] The feature difference perception module performs feature difference perception and reconstruction on the result after image registration. The feature difference perception module consists of three multi-branch convolutions, and uses convolution kernels and channel numbers of different sizes to perform feature perception on the registered image, and outputs feature maps F b1 、F b2 、F b3 , and then for the feature vectors at each spatial position, a difference operation is performed in the channel dimension to obtain a difference perception feature map.

[0024] Then the difference feature maps are fused, and F diff12 、F diff13 / F diff23 are concatenated together along the channel dimension to obtain a new feature difference perception feature map F conbined_diff .

[0025] Furthermore, the specific processing process of the seam generator is as follows:

[0026] First, through two bilinear interpolation upsampling operations, the feature difference perception feature map F conbined_diff is restored to a size that matches the registered image;

[0027] Then, two reparameterized convolution operations are performed to obtain a seam mask;

[0028] The seam mask is subjected to image weighted fusion to obtain the final stitching result.

[0029] Furthermore, the two reparameterized convolutional layers use convolutional kernels with a size of 3×3, the stride is set to 1, and after the convolutional operation, they pass through a batch normalization layer to normalize the features, accelerate the network convergence and improve the stability of the features, and then pass through the ReLU activation function to increase the non-linear expression ability of the network;

[0030] Then, the Sigmoid function is used for seam prediction to obtain the final seam mask output.

[0031] Compared with the prior art, the present invention has the following advantages and beneficial effects:

[0032] (1) The two-stage network model constructed by the present invention realizes an end-to-end efficient image stitching process through the orderly cooperation of image registration and image fusion. During the training process, each stage can optimize the parameters related to its own task specifically, avoiding the complexity and instability of the overall network training, improving the training efficiency and convergence speed of the model, and thus being able to quickly respond and complete the image stitching task in practical applications.

[0033] (2) The image registration network based on deep semantic consistency designed by the present invention can effectively alleviate the problem of large subsequent image registration errors caused by poor feature extraction effects in complex scenes. By using an advanced deep learning architecture, deeply mining the semantic information of images, and through the comprehensive analysis of multi-scale features and the understanding at the semantic level, it can accurately extract representative and stable features, thereby improving the accuracy of image registration and laying a solid foundation for subsequent high-quality image stitching.

[0034] (3) The image fusion network based on feature difference-aware reconstruction of seams designed by the present invention fully considers the feature difference information between images. During the fusion process, instead of simply performing pixel-level averaging or conventional weighted operations, it reconstructs the seam area specifically through the fine perception and analysis of feature differences. This method can better adapt to the changes between different images, effectively avoiding problems such as ghosting, blurring, and obvious stitching traces caused by image differences, making the fused image more natural and smooth visually, and greatly improving the quality and visual effect of image stitching.

[0035] (4) The overall image stitching method proposed by the present invention has good generality and generalization ability. Compared with many existing stitching technologies designed for specific scenes or image types, the present invention can better handle various unknown image data and complex and changeable practical application scenarios, reducing the workload of re-adjusting and optimizing the model due to data differences or scene changes, and lowering the application cost. Brief Description of the Drawings

[0036] Figure 1 is a schematic diagram of the image stitching method based on feature difference-aware reconstruction of seams of the present invention;

[0037] Figure 2(a) is the original image to be stitched, Figure 2(b) is the stitching effect diagram of the deep learning method based on the traditional seam, and Figure 2(c) is the stitching effect diagram of the method described in this embodiment;

[0038] Figure 3 is the working flow chart of the present invention. Detailed Embodiments

[0039] The present invention will be further described in detail below in conjunction with embodiments and the accompanying drawings, but the embodiments of the present invention are not limited thereto.

[0040] Embodiment

[0041] This embodiment provides an image stitching method for reconstructing seams based on feature difference perception. The flowchart is as Figure 1 and Figure 3 shown. Combining with specific numerical values of examples, it includes the following steps:

[0042] Taking an image dataset containing more than ten thousand pairs of images with a resolution of 512*512*3 as an example, the following processing is carried out:

[0043] Step 1: Preprocess the obtained image dataset. Divide the dataset into a training dataset and a test dataset according to a ratio of 8:2. Every two images with an overlapping area are a pair of data;

[0044] Step 2: Build an image stitching network model based on feature difference perception for reconstructing seams. The image stitching network model includes an image registration network based on deep semantic consistency and an image fusion network based on feature difference perception for reconstructing seams, which are connected in sequence;

[0045] Step 3: Use the training dataset to train the image registration network based on deep semantic consistency. During the training process, continuously minimize the unsupervised loss based on content alignment constraints to obtain the optimal image registration network, and output the image registration results of the training set. Specifically, calculate the unsupervised loss based on content alignment constraints according to the similarity of the reconstructed area after image registration. When the loss reaches the minimum value or reaches the maximum number of iteration steps, the optimal image registration network is obtained.

[0046] Step 4: Use the image registration output of the training dataset to train the image stitching network model based on feature difference perception for reconstructing seams. Through forward propagation, obtain the seam position parameters required in image fusion, and use the structure and energy constraint loss function for generating the seam as the model loss. Finally, obtain the optimal image fusion network;

[0047] Step 5: Input the test dataset into the optimal network models obtained in steps 3 and 4 to complete the stitching task of the test images.

[0048] Furthermore, for the specific construction process of the image registration network based on deep semantic consistency in step 2, the image registration network includes a deep semantic feature extraction module and a homography prediction module. Specifically:

[0049] Step 2.1: For the input image I∈R 512×512×3 , use the ResNet-style structure to extract the deep semantics of the image. First, construct the Bottleneck1 basic module and the Bottleneck2 basic module respectively as the basic modules of the deep semantic feature extraction module:

[0050] BottleNeck1 Basic Module:

[0051] Four parameters C, W, C1, S. C represents the number of input channels; W represents the input size, i.e., length × width; C1 represents the number of convolutional kernels, which is also the number of output channels; S represents the convolutional stride.

[0052] The said BottleNeck1 basic module includes two branches. The first branch consists of three convolutional layers, specifically:

[0053] Convolutional layer 1: C1 convolutional kernels, with a size of 1*1 and a stride of S. The input feature map is convolved through C1 1*1 convolutional kernels, then passes through the BN layer (Batch Normalization), and uses the ReLU activation function.

[0054] Convolutional layer 2: C1 convolutional kernels, with a size of 3*3 and a stride of 1. The output of convolutional layer 1 is convolved through C1 3*3 convolutional kernels, then passes through the BN layer (Batch Normalization), and uses the ReLU activation function.

[0055] Convolutional layer 3: C1*4 convolutional kernels, with a size of 1*1 and a stride of 1. The output of convolutional layer 2 is convolved through C1*4 1*1 convolutional kernels, then passes through the BN layer (Batch Normalization).

[0056] The second branch:

[0057] Convolutional layer: C1*4 convolutional kernels, with a size of 1*1 and a stride of S. The input feature map is convolved through C1*4 1*1 convolutional kernels, then passes through the BN layer (Batch Normalization).

[0058] The results of the two branches are added and then pass through the ReLU activation function to obtain the output of the BottleNeck1 basic module.

[0059] For the input of (W, W, C), after passing through the BottleNeck1 basic module, the output is (W / S, W / S, C1*4).

[0060] BottleNeck2 Basic Module:

[0061] Four parameters C, W. C represents the number of input channels; W represents the input size, i.e., length × width.

[0062] It includes three convolutional layers connected in sequence, specifically:

[0063] Convolutional layer 1: C / 4 convolutional kernels, with a size of 1*1, a stride of 1. The input feature map is convolved by C / 4 1*1 convolutional kernels, then passes through a BN layer (Batch Normalization), and the ReLU activation function is used.

[0064] Convolutional layer 2: C / 4 convolutional kernels, with a size of 3*3, a stride of 1. The output of convolutional layer 1 is convolved by C / 4 3*3 convolutional kernels, then passes through a BN layer (Batch Normalization), and the ReLU activation function is used.

[0065] Convolutional layer 3: C convolutional kernels, with a size of 1*1, a stride of 1. The output of convolutional layer 2 is convolved by C 1*1 convolutional kernels, then passes through a BN layer (Batch Normalization).

[0066] The outputs of the three convolutional layers are added to the input of BottleNeck2 and then passed through the ReLU activation function to obtain the output of BottleNeck2. For an input of (W, W, C), the output after passing through BottleNeck2 is (W, W, C).

[0067] The deep semantic feature extraction module is constructed from the BottleNeck1 basic module and the BottleNeck2 basic module. This module consists of four layers from Layer0 to Layer3. Combining the outputs of its layer1, layer2, and layer3, the output feature I is obtained s , and the specific composition and hierarchical structure are as follows:

[0068] Layer0: It consists of a convolutional layer and a max pooling layer. Among them:

[0069] Convolutional layer: 64 convolutional kernels, with a size of 7*7, a stride of 2. The input image is convolved by 64 7*7 convolutional kernels, then undergoes batch normalization processing, and the ReLU activation function is used for activation.

[0070] Pooling layer: A 3*3 max pooling layer is used with a stride of 2 to downsample the output of the convolutional layer.

[0071] Taking the input I ∈ R 512×512×3 as an example, after passing through Layer0, the output I1 ∈ R 128×128×64 .

[0072] Layer1: It consists of one BottleNeck1 basic module and two BottleNeck2 basic modules.

[0073] Bottleneck1 Basic Module: Parameters C = 64, W = 128, C1 = 64, S = 1.

[0074] Bottleneck2 Basic Module: Parameters C = 64, W = 128.

[0075] Bottleneck2 Basic Module: Parameters C = 64, W = 128.

[0076] Taking the input I1 ∈ R 128×128×64 as an example, after passing through Layer1, the output is I2 ∈ R 128×128×256 .

[0077] Layer2: Consists of one Bottleneck1 basic module and three Bottleneck2 basic modules.

[0078] Bottleneck1 Basic Module: Parameters C = 256, W = 128, C1 = 128, S = 2.

[0079] Bottleneck2 Basic Module: Parameters C = 512, W = 64.

[0080] Bottleneck2 Basic Module: Parameters C = 512, W = 64.

[0081] Bottleneck2 Basic Module: Parameters C = 512, W = 64.

[0082] Taking the input I2 ∈ R 128×128×256 as an example, after passing through Layer2, the output is I3 ∈ R 64×64×512 .

[0083] Layer3: Consists of one Bottleneck1 basic module and five Bottleneck2 basic modules.

[0084] Bottleneck1 Basic Module: Parameters C = 512, W = 64, C1 = 256, S = 2.

[0085] Bottleneck2 Basic Module: Parameters C = 1024, W = 32.

[0086] Bottleneck2 Basic Module: Parameters C = 1024, W = 32.

[0087] Bottleneck2 Basic Module: Parameters C = 1024, W = 32.

[0088] Bottleneck2 Basic Module: Parameters C = 1024, W = 32.

[0089] Bottleneck2 basic module: parameter C = 1024, W = 32.

[0090] Taking the input I3 ∈ R 64×64×512 as an example, after passing through Layer3, the output is I4 ∈ R 32×32×1024 .

[0091] Finally, the outputs of layer1, layer2, and layer3 form the deep semantic feature I s .

[0092] The deep semantic feature predicts the spatial transformation matrix of the image, which is completed by a homography prediction module.

[0093] The specific structure of the homography prediction module is as follows:

[0094] When the homography transformation is applied globally, it is usually parameterized as the movement of four vertices, and then the DLT method is used to solve it into a 3*3 matrix.

[0095] First, multiple consecutive convolutional layers are used for feature compression. The first convolutional layer uses a 3×3 convolutional kernel, a stride of 2, and the number of output channels is C / 2, followed by normalization and activation processing. The second convolutional layer uses a 3×3 convolutional kernel, a stride of 1, and the number of output channels is C / 4, also followed by normalization and activation processing. These convolutional compression layers gradually reduce the number of channels and spatial dimensions of the feature map, realizing feature aggregation and dimensionality reduction, and reducing the computational burden of the subsequent fully connected layers.

[0096] The feature map after convolutional compression is flattened into a one-dimensional vector and then connected to multiple fully connected layers. For example, the first fully connected layer has 512 neurons, passes through the ReLU activation function, and then the second fully connected layer has 256 neurons, also passing through the ReLU activation function.

[0097] Finally, a regression layer is connected, and its output dimension is 8 parameters related to the four-point offset regression. The predicted values of the four-point offset are directly output through a linear activation function, and then used for the calculation and adjustment of the subsequent homography matrix. Among them, the final output 8 parameters correspond to the two-dimensional image homography transformation matrix:

[0098]

[0099] where u1 = (x1, y1) T , that is, the homography transformation matrix. In the two-dimensional space, four points determine a plane, with a total of 8 parameters;

[0100] Furthermore, the image fusion network for reconstructing the seam based on feature difference perception consists of a feature difference perception module and a seam generator. The specific construction process is as follows:

[0101] The structural feature difference perception module performs feature difference perception and reconstruction on the result after image registration. This module consists of three multi-branch convolutions, which use convolution kernels and channel numbers of different sizes to perform feature perception on the registered image and output feature maps F b1 、F b2 、F b3 , where the subscript represents three different branches. Then, for the feature vectors at each spatial position (i.e., the same position corresponding to the height and width dimensions), a differential operation is performed in the channel dimension to obtain the difference perception feature map, such as:

[0102] F diff12 =F b1 -F b2 , F diff13 、F diff23 Similarly.

[0103] Then these differential feature maps are fused, and F diff12 、F diff13 / F diff23 are concatenated together along the channel dimension to obtain a new feature difference perception feature map F conbined_diff .

[0104] Through the above feature difference operations performed in the shallow feature extractor, it is possible to better focus on the key feature difference parts in the image, providing a more valuable feature basis for the entire image fusion network that reconstructs seams based on feature difference perception.

[0105] The feature difference perception feature map F conbined_diff is input into the seam generator to predict accurate seams. The specific processing process of the seam generator is as follows:

[0106] First, it undergoes two bilinear interpolation upsampling operations to restore the feature map to a size that matches the registered image. Then, two reparameterized convolution operations are performed to further adjust the distribution and expression of the features and enhance the ability to generate seam masks. Both convolutional layers use convolution kernels with a size of 3×3 and a stride of 1. After the convolution operation, it passes through a batch normalization layer to normalize the features, accelerate network convergence, and improve the stability of the features. Then, it passes through the ReLU activation function to increase the non-linear expression ability of the network.

[0107] Then, the Sigmoid function is used for seam prediction, and the function expression is as follows:

[0108]

[0109] Among them, x represents the input eigenvalue. When the feature map processed through the previous series of processes is input into the Sigmoid function, the function will calculate each pixel value in the feature map and convert it into a value between 0 and 1. This process can transform the feature map into output information with probability significance, where the value of each pixel reflects the likelihood of the pixel belonging to the seam area. The soft-coded mask output information is obtained. After obtaining the preliminary seam mask, it is multiplied by the aligned image mask to eliminate invalid area information, thereby obtaining the final seam mask output.

[0110] Finally, the seam mask output is directly used for image weighted fusion to obtain the final stitching result.

[0111] Furthermore, the unsupervised loss based on content alignment constraints is used to measure the difference between the reference image and the target image after the homography transformation. The calculation formula is:

[0112]

[0113] where I A , I B represent the input image pair, E is the all-ones matrix, H is the homography matrix, is the image space transformation operation. ||·||1 represents the Manhattan distance. During the training process, the goal of the network is to minimize this loss function L c .

[0114] Furthermore, the optimal image registration result of the training dataset is input to train the image stitching network model for reconstructing the seam based on feature difference perception to generate the structure and energy constraint loss function of the seam as the model loss. The calculation formula of this loss function is as follows:

[0115] L seam =||I ωA -I ωA ||2

[0116] where ||·||2 represents the Euclidean distance, that is, the square root of the sum of the squares of the elements. The pixel difference map reflects the difference degree of each pixel position in the overlapping area of the two images.

[0117] In this way, the stitching result is as similar as possible to image A at the boundary of the overlapping area with image A; at the boundary of the overlapping area with image B, it is as similar as possible to image B. Ensure that the start and end points of the stitching seam are in the correct positions, that is, on the boundary of the overlapping area, making the stitching look more natural.

[0118] Figure 2(a) is the original image to be spliced, Figure 2(b) is a splicing effect diagram based on the deep learning method of traditional seams, and Figure 2(c) is a splicing effect diagram using the method described in this embodiment. Although the deep learning method based on traditional seams can achieve basic splicing when splicing images, it has many defects. There are obvious traces of splicing, such as unnatural color transition, which affects the overall visual effect of the image and reduces the authenticity and integrity of the image. The method of reconstructing seams based on feature difference perception in this embodiment achieves seamless fusion, smooth color transition, smooth and natural visual effects, and improves the image's viewing and information communication capabilities. The method of the present invention overcomes the limitations of traditional seam deep learning methods, and achieves improvements in the accuracy and visual effects of image splicing, expanding a broader space for the practical application of image splicing technology.

[0119] The present invention proposes an image stitching method based on feature difference perception and reconstruction of seams. This method uses the powerful feature learning ability of deep neural networks to focus on perceiving and analyzing the feature differences of images. Through a carefully designed network structure, the feature information of the image can be accurately extracted, and the seams can be reconstructed based on the feature differences. Compared with traditional methods, the present invention can effectively overcome the problem of inaccurate seam generation, significantly improve the quality and stability of image fusion, and bring new breakthroughs and application prospects to the development of image stitching technology.

[0120] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be equivalent replacement methods and are included in the protection scope of the present invention.

Claims

1. An image stitching method based on feature difference perception and reconstruction of seams, characterized in that: These include: Preprocess the acquired image data set and divide it into training data set and test data set; Build an image stitching network model based on feature difference perception and reconstruction seams. The image stitching network model includes a sequentially connected image registration network based on deep semantic consistency and an image fusion network based on feature difference perception and reconstruction seams; use the training data set to train the image registration network based on deep semantic consistency to obtain the optimal image registration network, and output the image registration result of the training set; The image fusion network is trained using the image registration output of the training data set to obtain the optimal image fusion network; The test data set is input into the obtained optimal image registration network and image fusion network to complete the task of stitching the test images.

2. The image stitching method according to claim 1, characterized in that: The image registration network includes a deep semantic feature extraction module and a homography prediction module.

3. The image stitching method according to claim 2, characterized in that: The deep semantic feature extraction module includes Layer0, Layer1, Layer2 and Layer3 layers connected in sequence; the outputs of Layer1, Layer2 and Layer3 constitute the deep semantic feature I s .

4. The image stitching method according to claim 3, characterized in that: The Layer0 layer includes a convolution layer and a maximum pooling layer; The Layer1 layer includes a BottleNeck1 basic module and two BottleNeck2 basic modules; The Layer2 layer includes a BottleNeck1 basic module and three BottleNeck2 basic modules; The Layer3 layer includes one BottleNeck1 basic module and five BottleNeck2 basic modules.

5. The image stitching method according to claim 4, characterized in that: The BottleNeck1 basic module includes a first branch and a second branch; The first branch is composed of three convolutional layers connected in sequence; the second branch includes a convolutional layer, and the output results of the first branch and the second branch are added and then subjected to a ReLU activation function to obtain the output of the BottleNeck1 basic module.

6. The image stitching method according to claim 5, characterized in that: The BottleNeck2 basic module includes three convolutional layers connected in sequence, the outputs of the three convolutional layers are added to the input of the BottleNeck2 basic module, and then the output of the BottleNeck2 basic module is obtained through a ReLU activation function.

7. The image stitching method according to any one of claims 1 to 6, characterized in that: The image fusion network includes a feature difference perception module and a seam generator.

8. The image stitching method according to claim 7, characterized in that: include: The feature difference perception module performs feature difference perception and reconstruction on the image registration result. The feature difference perception module is composed of three multi-branch convolutions, which use convolution kernels of different sizes and channel numbers to perform feature perception on the registered image and output feature map F b1 、F b2 、F b3 , then for each feature vector at each spatial position, a differential operation is performed on the channel dimension to obtain a difference perception feature map, Then the differential feature maps are fused and F is transformed along the channel dimension. diff12 、F diff13 / F diff23 Spliced ​​together, we get a new feature difference perception feature map F conbined_diff .

9. The image stitching method according to claim 8, characterized in that: The specific processing process of the seam generator is as follows: First, after two bilinear interpolation upsampling operations, the feature difference perception feature map F conbined_diff Restore to a size that matches the registered image; Then two re-parameterized convolution operations are performed to obtain the seam mask; The seam masks are weighted fused to obtain the final stitching result.

10. The image stitching method according to claim 1, characterized in that: The two re-parameterized convolutional layers use a convolution kernel of size 3×3 and a step size of 1. After the convolution operation, they pass through a batch normalization layer to normalize the features, accelerate network convergence and improve feature stability, and then pass through a ReLU activation function to increase the nonlinear expression ability of the network. The Sigmoid function is then used to perform seam prediction and obtain the final seam mask output.