Image splicing method and device and storage medium
Through the stitching network, correlation analysis network and fusion network in the stitching model, the precise fusion of deep semantic features and shallow detail features is achieved, which solves the problem of poor image stitching quality in traditional methods and improves image clarity and computational efficiency.
Patent Information
- Application Number
- CN202511247910.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-03
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2045-09-03
AI Technical Summary
Existing image stitching methods find it difficult to effectively integrate deep semantic information with shallow detail features, resulting in problems such as blurred edges and decreased texture contrast in the stitched images, affecting image clarity and subsequent analysis and application effects.
A stitching model, including a stitching network, a correlation analysis network, and a fusion network, is used to extract deep semantic features and shallow detail features, combine them with contextual information, perform precise geometric registration and fusion of deep and shallow features, and generate high-quality fused images.
It significantly improves the naturalness of the transition of the stitching edges, preserves the clarity of local details, and improves the global coherence and computational efficiency of the stitched image, adapting to the application needs of mobile devices and complex scenes.
Smart Images

Figure CN120725864A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of image stitching, and in particular to an image stitching method, device, and storage medium. Background Art
[0002] With the rapid development of computer vision, remote sensing monitoring, virtual reality, and other fields, the demand for large-scale, high-resolution images is becoming increasingly urgent. Image stitching technology, a key means of acquiring such images, can overcome the limitations of a single image's field of view by fusing multiple overlapping images into a complete panoramic image. This provides important support for applications such as scene understanding, target detection, and geographic information collection.
[0003] However, existing image stitching methods often suffer from insufficient registration accuracy, blurred edges, and loss of detail when faced with complex scenes (such as those with dynamic objects, dramatic lighting changes, and areas of repeated texture). Traditional image stitching methods struggle to effectively integrate deep semantic information with shallow detail features, inevitably leading to loss of high-frequency detail. This leads to blurred edges and decreased texture contrast in the stitched image, compromising image clarity and the effectiveness of subsequent analytical applications. Summary of the Invention
[0004] The main technical problem solved by this application is to provide an image stitching method to solve the problem that traditional image stitching methods are difficult to take into account the effective integration of deep semantic information and shallow detail features, which will inevitably cause loss of high-frequency details, resulting in blurred edges and decreased texture contrast in the stitched images, affecting the image clarity and subsequent analysis and application effects.
[0005] To solve the above technical problems, a technical solution adopted in this application is to provide an image stitching method, comprising the steps of:
[0006] fusing paired or grouped images to be spliced into a fused image through a splicing model; the splicing model includes a splicing network, a correlation analysis network, and a fusion network connected in sequence;
[0007] The stitching network receives the image to be stitched, and is used to extract deep semantic features and shallow detail features from the image to be stitched to obtain a stitching feature map;
[0008] The correlation analysis network receives the images to be stitched and the stitching feature map, mines contextual information in the images to be stitched, determines the minimum circumscribed rectangle size of the images to be stitched, and fuses features of the images to be stitched and the stitching feature map to obtain a registration feature map;
[0009] The fusion network receives the registration feature map, encodes and decodes the registration feature map, reduces the size of the registration feature map during encoding to increase the feature expression of the registration feature map, increases the size of the registration feature map during decoding, and fuses the deep semantic features and shallow detail features in the registration feature map to obtain a fused image.
[0010] The present application also provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the image stitching method when executing the computer program.
[0011] The present application also provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the image stitching method are implemented.
[0012] The beneficial effects of this application are as follows: In this application, the deep semantic features (such as target outlines and scene structures) and shallow detail features (such as texture and color) of the images to be stitched are extracted through the stitching network, providing a rich feature foundation for subsequent fusion and effectively avoiding the loss of stitching information caused by a single feature. The correlation analysis network can more accurately judge the overlapping relationship and spatial position between images by mining the contextual association information between images. Combined with the dynamic determination of the minimum bounding rectangle size, it significantly reduces the registration deviation. The fusion network encoding stage enhances the feature expression capability by reducing the feature map size, and the decoding stage achieves a refined fusion of deep and shallow features by restoring the size, which not only ensures the global coherence of the stitched image, but also retains the clarity of local details and significantly improves the naturalness of the transition of the stitching edge. The three networks work together to realize the end-to-end image stitching process of "efficient feature extraction-precise geometric registration-seamless fusion optimization", effectively solving the problems of high computational complexity, low registration accuracy, and poor stitching quality of traditional methods, and adapting to the application needs of mobile devices and complex scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 is a flow chart according to an embodiment of the present application;
[0014] Figure 2 is a structural diagram according to an embodiment of the present application;
[0015] Figure 3 is a schematic diagram of the structure of a splicing network according to an embodiment of the present application;
[0016] Figure 4 is a schematic structural diagram of a first structured residual block according to an embodiment of the present application;
[0017] Figure 5is a schematic structural diagram of a second structured residual block according to an embodiment of the present application;
[0018] Figure 6 is a schematic diagram of the structure of a correlation analysis network according to an embodiment of the present application;
[0019] Figure 7 is a schematic diagram of the structure of a converged network according to an embodiment of the present application;
[0020] Figure 8 is a schematic diagram of fusion results according to an embodiment of the present application;
[0021] Figure 9 is another schematic diagram of a fusion result according to an embodiment of the present application;
[0022] Figure 10 is another schematic diagram of a fusion result according to an embodiment of the present application;
[0023] Figure 11 is another schematic diagram of a fusion result according to an embodiment of the present application;
[0024] Figure 12 is another schematic diagram of a fusion result according to an embodiment of the present application;
[0025] Figure 13 is another schematic diagram of a fusion result according to an embodiment of the present application;
[0026] Figure numerals: 1. Splicing network, 11. Splicing preprocessing module, 12. Splicing feature extraction module, 121. Dilated convolution unit, 122. Depthwise separable convolution unit, 123. Attention unit, 124. Adjustment convolution unit, 125. Addition unit, 126. Activation unit, 127. Additional adjustment convolution unit, 13. Attention processing module, 14. Sampling module, 2. Correlation analysis network, 21. Context-related module, 22. Related feature extraction module, 23. Fully connected module, 24. Control point offset module, 25. Direct linear transformation module, 26. Splicing domain transformation module, 3. Fusion network, 31. Initial processing module, 32. Encoding module, 33. Decoding module. DETAILED DESCRIPTION
[0027] To facilitate understanding of the present application, the present application is described in more detail below with reference to the accompanying drawings and specific embodiments. The accompanying drawings provide preferred embodiments of the present application. However, the present application can be implemented in many different forms and is not limited to the embodiments described in this specification. Rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of the present application.
[0028] It should be noted that, unless otherwise defined, all technical and scientific terms used in this specification have the same meanings as those commonly understood by those skilled in the art to which this application belongs. The terms used in this specification are for the purpose of describing specific embodiments only and are not intended to limit this application. The term "and / or" as used in this specification includes any and all combinations of one or more of the relevant listed items.
[0029] Figure 1 and Figure 2 An embodiment of the image stitching method of the present application is shown, comprising:
[0030] Step S1: fusing paired or grouped images to be stitched into a fused image through a stitching model; the stitching model includes a stitching network 1, a correlation analysis network 2, and a fusion network 3 connected in sequence;
[0031] Step S2: The stitching network 1 receives the image to be stitched, and is used to extract deep semantic features and shallow detail features from the image to be stitched to obtain a stitching feature map;
[0032] Step S3: The correlation analysis network 2 receives the image to be stitched and the stitching feature map, mines the contextual information in the image to be stitched, determines the minimum bounding rectangle size of the image to be stitched, and fuses the features of the image to be stitched and the stitching feature map to obtain a registration feature map;
[0033] Step S4: The fusion network 3 receives the registration feature map, and the fusion network 3 encodes and decodes the registration feature map. The size of the registration feature map is reduced during encoding to increase the feature expression of the registration feature map, and the size of the registration feature map is increased during decoding. The deep semantic features in the registration feature map are fused with the shallow detail features to obtain a fused image.
[0034] In this application, the stitching network 1 extracts deep semantic features (such as target outlines and scene structures) and shallow detail features (such as texture and color) of the images to be stitched, providing a rich feature foundation for subsequent fusion and effectively avoiding the loss of stitching information caused by a single feature. The correlation analysis network 2 can more accurately judge the overlapping relationship and spatial position between images by mining the contextual association information between images. Combined with the dynamic determination of the minimum bounding rectangle size, it significantly reduces the registration deviation. The fusion network 3 enhances the feature expression capability by reducing the size of the feature map during the encoding stage and achieves a refined fusion of deep and shallow features by restoring the size during the decoding stage. This ensures the global coherence of the stitched image while retaining the clarity of local details and significantly improving the natural transition of the stitching edges. The three networks work together to realize an end-to-end image stitching process of "efficient feature extraction-precise geometric registration-seamless fusion optimization", effectively solving the problems of high computational complexity, low registration accuracy, and poor stitching quality of traditional methods, and adapting to the application needs of mobile devices and complex scenes.
[0035] Stitching networks 1 can be configured in pairs or groups, with each network processing a single image to be stitched. Through multi-scale parallel branching and feature grouping, Stitching Network 1 avoids the information loss caused by channel dimensionality reduction in traditional methods. Compared to single-scale feature extraction schemes, this architecture reduces computational complexity by approximately 35% while maintaining feature expressiveness, making it more suitable for lightweight training in unsupervised learning where data annotation is lacking.
[0036] In some embodiments, as Figure 3 As shown, the stitching network 1 includes a stitching preprocessing module 11, multiple stitching feature extraction modules 12, an attention processing module 13 and multiple groups of sampling modules 14 connected in sequence. The stitching preprocessing module 11 is used to reduce the size of the image to be stitched and expand the number of channels of the image to be stitched to generate a first feature map. The stitching feature extraction module 12 is used to extract features in the first feature map and generate multiple intermediate feature maps of different sizes and numbers of channels; the attention processing module 13 is used to capture the global dependency between channels in the intermediate feature map and generate an attention feature map. The sampling module 14 is used to upsample the attention feature map to generate multiple sampling feature maps, and fuse the upsampled feature map with the intermediate feature map corresponding to its number of channels to generate a stitching feature map.
[0037] In some embodiments, after the image to be stitched is input into the stitching network 1, the stitching preprocessing module 11 first performs preliminary processing on the image to be stitched. The stitching preprocessing module 11 includes a convolution layer and a maximum pooling layer. The convolution layer and the maximum pooling layer are used to reduce the size of the image to be stitched, and expand the number of channels of the image to be stitched to generate a first feature map.
[0038] Specifically, a 512×512 pixel image to be stitched is input into the stitching network 1. The image is first processed using a 3×3 normal convolutional layer and a max pooling layer. This operation compresses the image to half its original size and expands the number of channels to 16, resulting in a first feature map with a size and channel representation of (1 / 2, 16). The stitching preprocessing module 11 effectively reduces the spatial dimensionality of the stitched image, significantly reducing subsequent computational effort. It also extracts basic image features, laying the foundation for deep feature extraction. The stitching preprocessing module 11 automatically downsamples the original high-resolution image (e.g., 4K, 8K) to a standard size of 512×512 pixels and records the image's length and width scaling factors (e.g., 1 / 8, 1 / 16). This operation significantly reduces the computational complexity of feature extraction and network inference, resulting in a lightweight model that reduces video memory usage by 90% while increasing inference speed by 5-8 times.
[0039] In some embodiments, as Figure 3 As shown, after the splicing preprocessing module 11 generates the first feature map, the first feature map is input into the splicing feature extraction module 12, and the splicing feature extraction module 12 extracts features from the first feature map to generate multiple intermediate feature maps of different sizes and channel numbers.
[0040] In some embodiments, the splicing feature extraction module 12 includes four, each splicing feature extraction module 12 outputs an intermediate feature map of different size and number of channels, the intermediate feature maps output by the four splicing feature extraction modules 12 may be respectively the first intermediate sub-feature map, the second intermediate sub-feature map, the third intermediate sub-feature map and the fourth intermediate sub-feature map, the intermediate feature maps output by the four splicing feature extraction modules 12 may be respectively the first intermediate sub-feature map, the second intermediate sub-feature map, the third intermediate sub-feature map and the fourth intermediate sub-feature map, the first feature map serves as the input of the first splicing feature extraction module 12, the output of the first splicing feature extraction module 12 is the first intermediate sub-feature map, the first intermediate sub-feature map serves as the input of the second splicing feature extraction module 12, the output of the second splicing feature extraction module 12 is the second intermediate sub-feature map, the second intermediate sub-feature map serves as the input of the third splicing feature extraction module 12, the output of the third splicing feature extraction module 12 is the third intermediate sub-feature map, the third intermediate sub-feature map serves as the input of the fourth splicing feature extraction module 12, and the output of the fourth splicing feature extraction module 12 is the fourth intermediate sub-feature map. The corresponding sizes and channel numbers of the first intermediate sub-feature map, the second intermediate sub-feature map, the third intermediate sub-feature map, and the fourth intermediate sub-feature map are (1 / 4, 32), (1 / 8, 64), (1 / 16, 128), and (1 / 32, 256), respectively.
[0041] In some embodiments, as Figure 4As shown, the splicing feature extraction module 12 includes at least two residual blocks connected in sequence and a maximum pooling layer. The residual block includes a dilated convolution unit 121, a depthwise separable convolution unit 122, an attention unit 123, a regulated convolution unit 124, an addition unit 125, and an activation unit 126 connected in sequence. The dilated convolution unit 121 is used to expand the number of channels to broaden the feature expression space; the depthwise separable convolution unit 122 is used to extract spatial and channel features; the attention unit 123 is used to select the local optimal coverage range of cross-channel interactions, capture key dependencies between channels, and focus on discriminative features; the regulated convolution unit 124 is used to adjust the number of channels; the addition unit 125 is used to fuse features, and the activation unit 126 is a rectified linear unit (ReLU) function.
[0042] In some embodiments, as Figure 5 As shown, the extended convolution unit 121 includes a convolution layer (Conv) 1×1, a normalization (BN) layer, and a linear rectification (Rectified Linear Unit, Relu) function set in sequence. The extended convolution unit 121 is used to double the number of channels of the input image to broaden the feature expression space.
[0043] The depthwise separable convolution unit 122 comprises two depthwise offset convolution (DSconv) layers and is used to extract spatial and channel features. Introducing the depthwise separable convolution unit 122 within the residual block significantly reduces the number of parameters and computational cost by decoupling the spatial and channel dimensions, while maintaining feature extraction capabilities. This frees up more computing resources for subsequent network layers. The parameter count of a conventional convolution layer is calculated as follows: assuming the number of input channels is Cin, the number of output channels is Cout, and the convolution kernel size is K×K, the parameter count is Cin×Cout×K×K. The depthwise separable convolution unit 122 decomposes the standard convolution operation into spatial convolution (channel-by-channel convolution) and channel-by-channel convolution (point-by-point convolution), resulting in a parameter count of Cin×K×K+Cin×Cout. Taking Cin=32, Cout=64, and K=3 as an example, the number of parameters for ordinary convolution is 32×64×3×3=18432, and the number of parameters for the depthwise separable convolution unit 122 is 32×3×3+32×64=2304, which is a reduction of about 87.5%. This significantly reduces the computing cost and improves computing efficiency, meeting the demand for lightweight models in unsupervised learning scenarios.
[0044] In some embodiments, attention unit 123 is based on efficient channel attention (ECA). By adaptively selecting the local optimal coverage of cross-channel interactions, attention unit 123 accurately captures key dependencies between channels and focuses on discriminative features. Attention unit 123 avoids the dimensionality reduction loss of traditional attention, dynamically weighting channel features at minimal computational cost (only increasing the number of parameters by 0.1%), focusing on discriminative features (such as edges, corners, and other key splicing features), and improving the robustness of feature extraction in unsupervised scenarios. This enables the adaptive selection of the local optimal coverage of cross-channel interactions, accurately capturing key dependencies between channels, and enabling the residual block to focus on extracting and enhancing discriminative features.
[0045] In some embodiments, the convolution adjustment unit 124 includes a 1×1 convolution layer (Conv) and a batch normalization (BN) layer. The convolution adjustment unit 124 is used to adjust the number of channels of the image to a suitable dimension.
[0046] In some embodiments, the adding unit 125 is used to achieve feature fusion and enhance the network learning ability.
[0047] Taking the residual block in the first spliced feature extraction module 12 as an example, the first feature map serves as the input to the first spliced feature extraction module 12. The addition unit 125 adds and fuses the input first feature map with the first feature map processed by the dilated convolution unit 121, the depthwise separable convolution unit 122, the attention unit 123, and the regulated convolution unit 124. The addition result is then processed by the Rectified Linear Unit (ReLU) function in the activation unit 126, which outputs the activated feature map. This structure effectively avoids the vanishing gradient problem while adjusting the channel dimension, improving network training stability.
[0048] In some embodiments, the residual block further includes an additional adjustment convolution unit 127, which is used to adjust the number of channels of the output image of the previous residual block, including a 1×1 convolution layer (Conv) and a batch normalization (BN) layer.
[0049] Taking the residual block in the first splicing feature extraction module 12 as an example, the addition unit 125 adds and fuses the first feature map processed by the additional adjustment convolution unit 127 with the first feature map processed by the extended convolution unit 121, the depth-separable convolution unit 122, the attention unit 123 and the adjustment convolution unit 124. The addition result is then processed by the linear rectification function (Rectified Linear Unit, Relu) to output the activated feature map.
[0050] In some embodiments, the residual block without the additional adjustment convolution unit 127 is used as the first structured residual block, and the residual block with the additional adjustment convolution unit 127 is used as the second structured residual block. The number of residual blocks in the four splicing feature extraction modules 12 is 2, 3, 2, and 2, respectively. Among these residual blocks, the first structured residual block is used as the front residual block, and the second structured residual block is used as the last residual block. This can effectively avoid the gradient vanishing problem and improve the stability of network training.
[0051] Traditional residual blocks use standard convolution, which has a large number of parameters and redundant computation. This application introduces a depth-wise separable convolution unit 122, which reduces the number of parameters of a single residual block by about 80% compared to standard convolution by decoupling the spatial and channel dimensions, significantly improving computational efficiency. Combined with an efficient channel attention mechanism (which only increases the number of parameters by 0.1%), it dynamically enhances key dependencies between channels with almost no increase in computational cost, avoiding the feature representation limitations of traditional attention (such as SE) caused by the fixed channel interaction range. This combination enables the feature extraction module to capture discriminative features (such as edges and corners) at a lower computational cost in unsupervised scenarios, supports the design of deeper network structures, and improves model robustness.
[0052] In the above, multiple splicing feature extraction modules 12 in the splicing network 1 can generate multiple intermediate feature maps of different sizes and channel numbers. After the intermediate feature maps are generated, the intermediate feature maps are processed by the attention processing module 13 in the splicing network 1 to obtain attention feature maps.
[0053] In some embodiments, the attention processing module 13 adopts an efficient multi-scale attention (EMA) mechanism, and the size and number of channels of the intermediate feature map are kept unchanged during the processing of the attention processing module 13.
[0054] In some embodiments, the attention processing module 13 employs a multi-scale parallel sub-network design, where a 1×1 convolutional branch focuses on capturing global dependencies between channels, while a 3×3 convolutional branch captures local texture information, effectively avoiding information loss caused by channel dimensionality reduction. Furthermore, by dynamically aggregating features at different scales through cross-dimensional interactions, a richer semantic representation is formed, automatically mining correspondences between images under unsupervised conditions and enhancing feature responses in overlapping regions.
[0055] While retaining channel information, the attention processing module 13 efficiently captures multi-scale spatial semantic features and enhances pixel-level context modeling through cross-space interaction, effectively improving the stability and accuracy of the model's capture of image feature information. The attention processing module 13 fuses the output features of parallel branches through a cross-space learning method to effectively capture long-range dependencies at the pixel level. Specifically, the attention processing module 13 encodes global spatial information through two-dimensional global average pooling and generates a spatial attention map through cross-dimensional interaction, thereby highlighting the semantic information of key areas (such as overlapping areas). This design enables the model to automatically discover correspondences between images under unsupervised conditions, such as optimizing splicing alignment by enhancing the feature responses of overlapping areas. In addition, the attention processing module 13 further improves the uniform distribution of spatial semantic features and reduces noise interference through a feature grouping strategy (dividing the channel dimension into multiple sub-feature groups).
[0056] The stitching network 1 upsamples the attention feature map through multiple sampling modules 14 to generate multiple sampled feature maps. The upsampled feature maps are then fused with the intermediate feature maps corresponding to their number of channels to generate a stitched feature map. The stitched feature map organically combines deep semantic features with shallow detail features, enabling the stitching network 1 to accurately capture image detail texture information while accurately grasping the overall structural outline, providing comprehensive and accurate feature information support for subsequent image registration.
[0057] In some embodiments, as Figure 3 As shown, the sampling module 14 includes an upsampling unit (Upsample), a fusion unit (Concat), and a residual block, which are arranged in sequence. The upsampling unit is used to upsample the attention feature map, and the fusion unit is used to fuse the attention feature map with the intermediate feature map to ultimately generate a spliced feature map. For example, after the attention feature map is upsampled by the upsampling unit, the size and number of channels of the attention feature map are (1 / 16, 128). The image output by the upsampling unit is input into the fusion unit. At this time, the third intermediate sub-feature map is also input into the fusion unit. The size and number of channels of the third intermediate sub-feature map are also (1 / 16, 128). The fusion unit then inputs the image obtained by fusing the attention feature map and the third intermediate sub-feature map into the first residual unit or the second residual unit, which then processes the input image. After processing, the processed image is input into the next sampling module 14. After upsampling, the upsampling unit of the next sampling module 14 obtains an image with a size and number of channels of (1 / 8, 64). This image is fused with the second intermediate sub-feature map, which also has a size and number of channels of (1 / 8, 64). Thus, a spliced feature map with a size and number of channels of (1 / 8, 64) is obtained.
[0058] Single-scale features cannot take into account both detail information (high-resolution features) and semantic information (low-resolution features). Directly using deep features may lose key registration details such as edges, while shallow features lack global structural information. In this application, the sampling module 14 is used to perform channel splicing and fusion of deep semantic features (structural contours) and shallow detail features (texture edges), which can organically integrate the features of the image at different resolutions and generate a spliced feature map containing multi-scale information. This solves the problem of "detail loss" or "semantic ambiguity" of traditional single-scale features. After fusing multi-scale features, the feature matching accuracy is improved by 22% compared to the single-scale solution.
[0059] The stitching feature map contains both texture details (such as pixel-level color differences) and structural semantics (such as regional shape features) of the image. This allows the model to capture both the image's detailed textures and the overall structural semantics, providing more comprehensive input for subsequent contextual analysis. The sampling module 14 adapts to scenarios where objects of varying sizes are encountered in unsupervised image stitching, improving the generalization capability of matching objects of varying sizes through multi-scale feature matching.
[0060] In some embodiments, the stitching feature map and the image to be stitched (including the image to be stitched a and the image to be stitched b) are input into the correlation analysis network 2, and the correlation analysis network 2 processes the stitching feature map and the image to be stitched to obtain the registration feature map (including the registration feature map a and the registration feature map b) and the mask (including mask a and mask b).
[0061] In some embodiments, as Figure 6 As shown, the correlation analysis network 2 includes a context-related module 21, a relevant feature extraction module 22, a fully connected module 23, a control point offset module 24, a direct linear transformation module 25, and a splicing domain transformation module 26, which are connected in sequence. The context-related module 21 is based on the directional feature interaction mechanism and accurately captures the spatial correspondence clues in the splicing feature map by separating the feature dependencies in the horizontal and vertical directions. The relevant feature extraction module 22 is used to extract features in the splicing feature map, the fully connected module 23 is used to adjust the feature vector of the splicing feature map, the control point offset module 24 is used to control the offset of the pixel points in the splicing feature map, and the direct linear transformation module 25 is used to transform the splicing feature map and generate a distorted feature map. After the distorted feature map is fused with one of the images to be spliced, it is spliced with the other image to be spliced through the splicing domain transformation module 26 to generate a registration feature map and a mask.
[0062] In some embodiments, as Figure 6 As shown, the context-related module 21 includes a cluster center loss (CCL) layer and a convolution layer. The cluster center loss layer performs horizontal and vertical correlation calculations on the input concatenated feature map:
[0063] Horizontal feature flow generation: Calculate the dot product similarity of feature vectors along the row dimension to generate a feature flow matrix that describes the horizontal displacement trend;
[0064] Vertical feature flow generation: The cosine similarity of the feature vectors is calculated along the column dimension to generate a feature flow matrix that describes the vertical displacement trend.
[0065] The concatenated feature map is input to the context-related module 21. The size and number of channels of the feature flow matrix output by the cluster center loss layer in the context-related module 21 are (1 / 8, 2). The feature flow matrix output by the cluster center loss layer is then expanded through a 1×1 convolution layer to a first related feature map with a size and number of channels of (1 / 8, 64).
[0066] The core of image registration is the pixel-level spatial correspondence, and horizontal / vertical displacement is the basic component of geometric transformation. The feature interaction methods in the prior art (such as point-by-point calculation) do not show separation direction dependence and are difficult to capture long-distance contextual associations (such as row and column feature matching across image boundaries). In this application, the contextual association information between image pairs is deeply mined through the context-related module 21. The context-related module 21 can carefully analyze the feature interaction relationship of the image pair in the horizontal and vertical directions, and generate horizontal feature streams and vertical feature streams by calculating the feature correlation of the row / column dimensions, thereby accurately describing the displacement trend of the image pair in the horizontal / vertical direction, providing structured context clues for subsequent offset regression; reducing the feature mismatch problem caused by complex geometric transformations in unsupervised learning, especially for image pairs with rotation and scaling, the directional feature flow can enhance the correspondence modeling of local areas. The context-related module 21 shows the analysis of horizontal / vertical feature interactions and generates a structured directional feature flow. Compared with traditional point-by-point feature interaction methods, this mechanism reduces the registration error for rotated and scaled scenes by 18%, and has significant advantages in modeling long-distance contextual associations across image boundaries (such as accurately capturing the horizontal offset distribution of a row of pixels in the left image in the right image).
[0067] In some embodiments, as Figure 6As shown, the relevant feature extraction module 22 includes three residual blocks and a maximum pooling layer connected in sequence. The residual block has the same structure as the residual block in the above-mentioned splicing feature extraction module 12. The first relevant feature map is input into the relevant feature extraction module 22, and the relevant feature extraction module 22 gradually compresses the size of the input first relevant feature map to 1 / 16, 1 / 32, and 1 / 64, and expands the number of channels to 128, 256, and 512. For example, the relevant feature extraction module 22 expands the number of channels from 64 to 128 through 1×1 convolution, strengthens the discriminative features through depthwise separable convolution units and attention units, and then maintains the number of channels unchanged through 1×1 convolution, and cooperates with residual connections to prevent gradient vanishing; the maximum pooling layer performs 2×2 kernel downsampling to compress the size of the first relevant feature map from 1 / 8 to 1 / 16.
[0068] In some embodiments, global average pooling is performed on the image output by the relevant feature extraction module 22, compressing the spatial dimension to 1×1 to obtain a 512-dimensional feature vector. This vector passes through the fully connected (FC) module 23 and the control point offset module 24 (Control Point Offset), regressing the control point offsets of the 12×12 grid to generate a second relevant feature map. The 512-dimensional feature vector in the fully connected module 23 changes as follows: 512 → 256 → 288 (12×12×2), where the hidden layer is activated using a linear rectified unit (Relu) function. In the control point offset module 24, the predicted offset (Δx, Δy) of each grid vertex represents the displacement of that point in the original image coordinate system relative to the reference image (unit: pixel).
[0069] This application divides the image into several grid cells. This decomposes the image into local regions, with each region independently regressing offsets, supporting more flexible local geometric transformation modeling. The local control mechanism enables the model to adapt to complex scenarios (such as multi-plane stitching and curved images). Each grid can independently learn transformation parameters such as scaling, rotation, and translation, improving registration accuracy. The nonlinear mapping capability of the fully connected layer captures the complex relationship between control point offsets and features (for example, offsets in densely textured areas require more refined adjustments), avoiding the reliance of traditional methods (such as manual feature matching) on specific scenarios. This allows for accurate prediction of the offset of each control point, thus achieving precise image registration.
[0070] In some embodiments, the direct linear transformation module 25 calculates the local homography matrix H from the offset using direct linear transformation (DLT) for each 12×12 grid (a total of 13×13 control points) in the second correlation feature map. The mathematical expression of the homography transformation is:
[0071] ;
[0072] in, is the source image coordinate, is the target image coordinate; is an element in the local homography matrix H; the direct linear transformation module solves the following overdetermined equations by least squares:
[0073] ;
[0074] Convert to matrix form . The optimal solution is obtained by finding the eigenvector corresponding to the minimum singular value through singular value decomposition, where T represents transpose. Thus, the distorted feature map can be obtained through the direct linear transformation module 25.
[0075] In the prior art, the homography matrix is a standard mathematical tool for describing plane transformations. The direct linear transformation method ensures the geometric consistency of local transformations through least squares fitting. Direct stitching may lead to redundant or missing image boundaries, and the output size needs to be dynamically adjusted through the minimum bounding rectangle. In this application, the direct linear transformation module 25 accurately models the geometric transformation (such as rotation angle and scaling) of each grid through the local homography matrix, significantly improving the registration accuracy of non-planar scenes compared to the global transformation. The minimum bounding rectangle mechanism automatically adapts to the actual size of the stitched image, avoiding information cropping or invalid padding caused by fixed-size output in traditional methods, thereby achieving accurate registration of the entire image and ensuring that the image is highly aligned in geometric position and structure. This improves the practicality and flexibility of unsupervised stitching.
[0076] Traditional global homography assumes that the image is a planar rigid body and cannot handle multi-plane or non-rigid deformations (such as curved objects). This application decomposes the image into independent local areas through grid division and local homography transformation. Each grid regresses the control point offset (Δx, Δy) through a fully connected layer, supporting flexible transformations such as scaling, rotation, and translation. The direct linear transformation module 25 proportionally amplifies the predicted offset result according to the pre-recorded scaling ratio (for example, if the scaling ratio is 1 / 8, the Δx and Δy coordinate values are multiplied by 8), thereby accurately mapping the registration parameters of the low-resolution space back to the original high-resolution image space. This strategy not only ensures that the model effectively captures global structure and local details, but also achieves fast and low-resource registration of high-resolution images without losing stitching accuracy through a resolution layered processing mechanism, successfully solving the performance bottleneck of traditional methods when processing large-size images. In multi-plane stitching scenarios, the registration error (peak signal-to-noise ratio (PSNR) and structural similarity (SSIM)) of this application is reduced by 31% and 25% compared with the global homography method, significantly improving the geometric alignment accuracy in complex scenarios.
[0077] In some embodiments, a distortion feature map is obtained by the direct linear transformation module 25. After the distortion feature map is fused with an image to be stitched, it is input into the stitching domain transformation module 26 together with another image to be stitched. After being processed by the stitching domain transformation module 26, the stitching domain transformation module 26 outputs a pair of alignment feature maps and their corresponding masks.
[0078] In some embodiments, the stitching domain transformation module 26 is used to calculate the minimum bounding rectangle of the image to be stitched, and determine the size and boundary of the image to be stitched. Specific steps: apply the homography matrix Hi to each grid, calculate the coordinates of the four corner points of the grid in the stitching coordinate system; collect all corner point coordinates, solve the extreme values of the horizontal and vertical axes (Xmin, Xmax, Ymin, Ymax), and determine the minimum bounding rectangle; map each grid area to the minimum bounding rectangle coordinate system through bilinear interpolation to complete the whole image registration. After processing by the stitching domain transformation module 26, the stitching domain transformation module 26 outputs a pair of registration feature maps and their corresponding masks. The minimum bounding rectangle boundary is automatically calculated by the stitching domain transformation module 26 layer, avoiding information cropping or invalid filling caused by traditional fixed-size output, and adapting to flexible stitching requirements in multiple scenarios such as video surveillance, drone inspection, and medical microscopy.
[0079] The stitching network is trained using the following loss function. To minimize the matching error in the overlapping regions while constraining the geometric deformation energy in the non-overlapping regions, the objective function L is a combination of pixel consistency loss L1 and mesh deformation loss L2:
[0080] ;
[0081] in, 、 are the weight coefficients corresponding to different losses, which can be 1 and 5 respectively.
[0082] The pixel consistency loss is expressed as:
[0083] ;
[0084] Among them, i and j represent the pixel coordinates of the overlapping area, represents the overlapping area of the registered feature maps, represents the twist operation, represents the target image, Represents the reference image, N represents the total number of pixels in the overlapping area of the registration feature map. During image registration, one image remains unchanged and the other is distorted for registration. The image that remains unchanged is called the reference image, and the image used for distortion is called the target image.
[0085] The mesh deformation loss is expressed as:
[0086] ;
[0087] in, and Indicates the number of horizontal and vertical edges, which are 12 and 12 respectively. and Represent the width and height of the registration feature map respectively, and denote the sets of horizontal and vertical edges respectively, represents an edge vector in a set of horizontal or vertical edge vectors, and denote the horizontal and vertical unit vectors respectively.
[0088] Represents the Rectified Linear Unit (Relu) function, P represents the total number of adjacent edges, is a 0-1 mask, with non-overlapping areas set to 1. and Indicates adjacent edges.
[0089] In some embodiments, as Figure 7 As shown, a pair of registration feature maps are input into the fusion network 3, which includes an initial processing module 31, multiple encoding modules 32, and multiple decoding modules 33 corresponding to the encoding modules 32. The initial processing module 31 is used to fuse the registration feature maps, and the encoding module 32 is used to halve the size of the registration feature maps to enhance the feature expression of the registration feature maps. The decoding module 33 is used to increase the size of the registration feature maps in the corresponding decoding process and fuse the deep semantic features in the registration feature maps with the shallow detail features to obtain a fused image.
[0090] In some embodiments, as Figure 7 As shown, the initial processing module 31 includes a fusion unit and two residual blocks. The fusion unit fuses the features in a pair of registered feature maps. The registered feature maps are concatenated in the channel dimension (e.g., 512×512×3 → 512×512×6). The first residual block is the second residual unit, and the second residual block is the first residual unit. The first residual block expands the number of channels from 6 to 12 (channel expansion) through 1×1 convolution. Two depthwise separable convolution units and the ECA attention mechanism strengthen the feature representation. The dimension is then increased to 16 through 1×1 convolution. The result is added to the registered feature map and the output is (1 / 1, 16). The second residual block maintains the number of channels at 16, further enhancing the feature representation.
[0091] In some embodiments, as Figure 7As shown, there are four encoding modules 32, each comprising a maximum pooling layer with a 2×2 kernel, which is used to halve the size of the registered feature map. The first residual block is a second-structure residual block, and the second is a first-structure residual block. The first residual block expands the number of channels from 6 to 12 (channel expansion) through a 1×1 convolution. Feature representation is enhanced through two depthwise separable convolutional units and the ECA attention mechanism. The dimension is then increased to 16 through a 1×1 convolution, and the result is summed with the registered feature map to output (1 / 1, 16). The second residual block maintains the number of channels at 16, further enhancing feature representation.
[0092] The input and output sizes and channel numbers of the four encoding modules 32 are shown in Table 1:
[0093] Table 1 Input and output size and number of channels of the encoding module
[0094] Encoding module Input size Output size Channel number changes 1 1 / 1,16 1 / 2,32 16→32→32 2 1 / 2,32 1 / 4,64 32→64→64 3 1 / 4,64 1 / 8,128 64→128→128 4 1 / 8,128 1 / 16,256 128→256→256
[0095] The four encoding modules 32 achieve multi-scale feature capture through step-by-step dimensionality reduction and channel expansion. For example, the 256-channel registration feature map at 1 / 16 scale encodes global semantic information (such as object outlines), while the 32-channel registration feature map at 1 / 2 scale preserves local details (such as texture edges).
[0096] Unsupervised learning relies on lightweight models to reduce training complexity. Traditional convolutions are computationally complex and prone to falling into local optima. Image fusion requires preserving both local details (such as pixel-level texture) and global semantics (such as object outlines), requiring hierarchical feature extraction to capture multi-granular information. Therefore, in fusion network 3, the downsampling path serves as the encoding module 32, utilizing depthwise separable convolutional units for multi-scale feature extraction. This lightweight convolution operation effectively captures local structural features of the image while reducing computational complexity. Pooling layers are then added to achieve a step-by-step reduction in spatial resolution, gradually stripping away redundant detail information, ultimately forming a high-level feature representation that contains the core semantic information of the registered feature map.
[0097] In some embodiments, as Figure 7 As shown, the decoding module 33 includes an upsampling layer, a fusion layer, and two residual blocks connected in sequence. The upsampling layer performs bilinear interpolation with a magnification of 2. Of the two residual blocks, the first residual block is a second-structured residual block, which reduces the number of channels, while the second residual block is a first-structured residual block, which maintains the same number of channels. The number of decoding modules 33 corresponds to that of encoding modules 32, and the last decoding module 33 also includes a convolutional layer.
[0098] Traditional deconvolution (transposed convolution) is prone to producing checkerboard artifacts in unsupervised scenarios. Bilinear interpolation provides a smoother upsampling foundation. Detailed information (such as color gradients) contained in shallow features must be gradually restored through lightweight convolution operations. In fusion network 3, the upsampling path serves as the decoding module 33, combining bilinear interpolation upsampling with depthwise separable convolution units to achieve layer-by-layer restoration of spatial dimensions. Bilinear interpolation avoids deconvolution artifacts while supplementing high-frequency details with lightweight operations (reducing computational effort by 40%), resulting in a 19% improvement in pixel-level semantic consistency in reconstructed overlapping regions. Depthwise separable convolution units restore the size while supplementing high-frequency details such as edges and textures at a low computational cost. Upsampling scales the feature map, providing a spatial foundation for detail reconstruction. Subsequent depthwise separable convolution units refine the enlarged feature map to supplement high-frequency details.
[0099] The input and output sizes and channel numbers of the four decoding modules 33 are shown in Table 2:
[0100] Table 2 Input and output size and number of channels of the decoding module
[0101] Decoding module Input size Skip connection size Output size Channel number changes 1 1 / 16,256 1 / 8,128 1 / 8,128 256→384→128 2 1 / 8,128 1 / 4,64 1 / 4,64 128→192→64 3 1 / 4,64 1 / 2,32 1 / 2,32 64→96→32 4 1 / 2,32 1 / 1,16 1 / 1,16 32→48→16→3
[0102] In this application, the upsampling layer replaces deconvolution with bilinear interpolation (O(N2)) to eliminate the checkerboard artifacts common in unsupervised learning (traditional deconvolution produces jagged artifacts at the edges, while bilinear interpolation maintains smooth transitions).
[0103] In some embodiments, the decoding module 33 employs a symmetrical structure, gradually restoring the size of the registration feature map through upsampling and cross-layer skip connections. The decoding module 33 is connected to the encoding module 32 via skip connections, meaning that the registration feature maps of the same scale as the encoding stage are connected by the number of channels. For example, when the decoding size is 1 / 8 layer, the encoding output size is connected to 1 / 8 layer. The skip connection performs channel-wise splicing and fusion on the registration feature maps of corresponding levels. The shallow features of the encoding module (high resolution, containing detailed information) and the deep features of the decoding module (low resolution, containing semantic information) are spliced by the channel dimension. The fused feature maps are then fed into the subsequent decoding layer for joint processing.
[0104] In traditional codec structures, deep semantic features lack support for shallow details during decoding, leading to problems such as blurred edges and texture loss. Unsupervised learning lacks real label supervision, and feature representation capabilities need to be enhanced through cross-layer information complementarity. In this application, the decoding module and the encoding module are connected through jump connections, and the shallow detail features (such as color gradients) of the corresponding levels in the encoding module are channel-wise spliced and fused with the deep semantic features (such as object boundaries) of the decoding module to form a composite feature representation containing multi-scale information. This feature fusion strategy effectively solves the problem of detail information loss in traditional codec structures and significantly enhances the model's ability to recover image edges, textures and other details.
[0105] In some embodiments, the fusion network finally reconstructs the overlapping areas of the registered feature map pairs. The pixel details and semantic consistency of the overlapping areas are accurately restored by using the registered feature map generated after decoding and the reconstruction result with the same size as the input image to be spliced. For non-overlapping areas, the original image area after registration is directly used as the retained part to avoid errors introduced in the reconstruction process. Finally, through pixel-level fusion operations, the retained non-overlapping areas and the reconstructed overlapping areas are seamlessly spliced to generate a high-quality fused image with a unified color space and geometric consistency. This method effectively improves the detail fidelity and visual smoothness of image fusion while ensuring semantic consistency.
[0106] There are conflicts in the pixel values of multiple images in the overlapping area (such as exposure differences and visual color differences), which require network reconstruction to achieve semantic consistency; the original pixels in the non-overlapping area already have geometric consistency, and directly retaining them can avoid the accumulation of errors introduced by reconstruction.
[0107] In this application, the overlapping area reconstruction result is generated by the decoded full-size registration feature map (the same size as the input image), that is, the reconstructed image (including RGB three channels) with the same resolution as the input image is output by the 1×1 convolution layer, and the reconstruction target is focused on the overlapping area of the registered image pair (located by the mask matrix); the non-overlapping area directly uses the original image area after registration: the coordinate mapping relationship is determined, and the non-overlapping original image blocks are extracted; the final fused image is formed by combining the reconstructed overlapping area and the original non-overlapping area.
[0108] Specifically, the fusion network uses masks to reconstruct overlapping regions and preserve non-overlapping regions. Based on the generated masks, masks M1 and M2 are constructed for the registration feature maps, respectively. Pixels in overlapping regions are marked as 1, and pixels in non-overlapping regions are marked as 0. The intersection region is precisely extracted through the mask operation M = M1 × M2, while the respective non-overlapping region masks M1-M and M2-M are calculated. This mask separation strategy divides the image domain into three mutually exclusive regions: a shared overlap region (M = 1), a region unique to the first image (M1-M = 1), and a region unique to the second image (M2-M = 1).
[0109] During the reconstruction phase, to ensure that the reconstruction process focuses on key areas, this application designs a mask constraint mechanism: reconstruction loss constraints are only imposed on overlapping areas (M=1), achieving dual alignment at the pixel level and semantic level.
[0110] For non-overlapping regions, this application adopts an information-preserving strategy: the corresponding regions (I1·(M1-M) and I2·(M2-M)) are directly extracted from the original registered images to avoid the loss of details that may be introduced by the reconstruction process.
[0111] The final fused image I _final Generated by the following formula:
[0112] I _final =I _rec ・M+I1・(M1-M)+I2・(M2-M)
[0113] Among them, I _rec Represents the original registered image, I_rec・M represents the overlapping area corresponding to the original registered image, which together with the non-overlapping area of I1 (I1・(M1-M)) and the non-overlapping area of I2 (I2・(M2-M)) constitute the final fused image I _final .
[0114] This achieves semantically consistent reconstruction of overlapping regions while preserving the original information in non-overlapping regions. This region separation processing mechanism not only solves the color conflict problem in overlapping regions caused by traditional methods, but also preserves the texture details in non-overlapping regions.
[0115] The traditional stitching method directly superimposes overlapping areas, which can easily lead to exposure differences or color discontinuities. This application locates the overlapping areas through a mask matrix, uses an autoencoder to reconstruct pixel details (such as RGB three-channel consistency adjustment), and directly retains the original pixels in the non-overlapping areas to avoid the accumulation of reconstruction errors. The color brightness unevenness of the fused image (measured by standard deviation) is reduced by 28%, and the incidence of artifacts in the overlapping area is reduced by 62% compared with the traditional method. The generated fused image has higher visual smoothness and semantic consistency. The image to be stitched is converted into a fused image such as Figures 8-13 shown.
[0116] The fusion network is trained using the following loss function. The fusion network fuses the processed registration feature maps, and the pixel-level reconstruction loss L is used in the fusion stage of the registration feature maps. pixel and semantic-level perception loss L perceptual The combination strategy constructs the total loss function L by weighted summation z, realizing multi-granularity constraints in overlapping areas. The total loss function theoretically guarantees the visual smoothness and semantic consistency of the fused image through the dual constraints of pixel-level accuracy and semantic-level rationality, providing a solid optimization foundation for unsupervised image stitching. z The specific formula is as follows:
[0117] L z =L pixel +aL perceptual
[0118] Among them, a=0.2 to balance pixel-level accuracy and semantic consistency.
[0119] Pixel reconstruction loss L pixel It is expressed as follows:
[0120] ;
[0121] Among them, M is the overlapping area mask, Q is the number of pixels in the overlapping area, I1 and I2 are the registered feature maps a and b respectively.
[0122] Perceptual loss L perceptual It is expressed as follows:
[0123] ;
[0124] in, Represents the output feature map of the convolutional layer.
[0125] The above stitching model can be used to fuse pairs or groups of images to be stitched into a fused image. Before applying the stitched images, the stitching model can be trained to improve the stitching accuracy of the stitching model. The training method is as follows:
[0126] First, the videos of 400 application scenes under different shooting angles are split frame by frame, and two frames of images with random time intervals are extracted. Image sub-blocks with overlapping areas are generated through random cropping operations; then the image resolution is uniformly adjusted to 512×512 pixels to construct image pair samples. Finally, a dataset containing 40,000 pairs of images is formed, and the training set, validation set, and test set are divided in the same 8:1:1 ratio. This construction method effectively covers complex factors such as perspective changes and lighting differences in real scenes, and provides the model with training samples close to practical applications. Ensure that the model has strong generalization ability in unsupervised learning. Compared with the model trained with traditional synthetic datasets, the registration success rate of this application in real scene tests is improved by 34%.
[0127] In this application, the deep semantic features (such as target outlines and scene structures) and shallow detail features (such as texture and color) of the images to be stitched are extracted through the stitching network, providing a rich feature basis for subsequent fusion and effectively avoiding the loss of stitching information caused by a single feature. The correlation analysis network can more accurately judge the overlapping relationship and spatial position between images by mining the contextual association information between images. Combined with the dynamic determination of the minimum bounding rectangle size, the registration deviation is greatly reduced. The fusion network encoding stage enhances the feature expression capability by reducing the size of the feature map, and the decoding stage achieves a refined fusion of deep and shallow features by restoring the size, which not only ensures the global coherence of the stitched image, but also retains the clarity of local details and significantly improves the naturalness of the transition of the stitching edge. The three networks work together to realize the end-to-end image stitching process of "efficient feature extraction-accurate geometric registration-seamless fusion optimization", effectively solving the problems of high computational complexity, low registration accuracy, and poor stitching quality of traditional methods, and adapting to the application needs of mobile devices and complex scenes.
[0128] The above are merely embodiments of the present application and are not intended to limit the patent scope of the present application. Any equivalent structural transformations made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. An image stitching method, characterized in that: Including steps: fusing paired or grouped images to be spliced into a fused image through a splicing model; the splicing model includes a splicing network, a correlation analysis network, and a fusion network connected in sequence; The stitching network receives the image to be stitched, and is used to extract deep semantic features and shallow detail features from the image to be stitched to obtain a stitching feature map; The correlation analysis network receives the images to be stitched and the stitching feature map, mines contextual information in the images to be stitched, determines the minimum circumscribed rectangle size of the images to be stitched, and fuses features of the images to be stitched and the stitching feature map to obtain a registration feature map; The fusion network receives the registration feature map, encodes and decodes the registration feature map, reduces the size of the registration feature map during encoding to increase the feature expression of the registration feature map, increases the size of the registration feature map during decoding, and fuses the deep semantic features and shallow detail features in the registration feature map to obtain a fused image.
2. The image stitching method according to claim 1, wherein: The stitching network includes a stitching preprocessing module, multiple stitching feature extraction modules, an attention processing module and multiple groups of sampling modules connected in sequence. The stitching preprocessing module is used to reduce the size of the image to be stitched and expand the number of channels of the image to be stitched to generate a first feature map. The stitching feature extraction module is used to extract features in the first feature map and generate multiple intermediate feature maps with different sizes and channel numbers; the attention processing module is used to capture the global dependency between channels in the intermediate feature map and generate an attention feature map. The sampling module is used to upsample the attention feature map to generate multiple sampling feature maps, and fuse the upsampled feature map with the intermediate feature map corresponding to its channel number to generate the stitching feature map.
3. The image stitching method according to claim 2, wherein: The splicing feature extraction module includes at least two residual blocks and a maximum pooling layer connected in sequence, and the residual block includes an extended convolution unit, a depth-separable convolution unit, an attention unit, an adjustment convolution unit, an addition unit and an activation unit connected in sequence; the extended convolution unit is used to expand the number of channels to broaden the feature expression space; the depth-separable convolution unit is used to extract spatial and channel features, the attention unit is used to select the local optimal coverage range of cross-channel interaction, capture the key dependencies between channels, and focus on discriminative features; the adjustment convolution unit is used to adjust the number of channels; the addition unit is used to fuse features.
4. The image stitching method according to claim 3, wherein: The residual block also includes an additional adjustment convolution unit, which is used to adjust the number of channels of the output image of the previous residual block; the residual block without the additional adjustment convolution unit is used as the first structural residual block, and the residual block with the additional adjustment convolution unit is used as the second structural residual block. There are 4 splicing feature extraction modules, and the number of residual blocks in the 4 splicing feature extraction modules is 2, 3, 2 and 2 respectively; in the same splicing feature extraction module, the residual block on the front side adopts the first structural residual block, and the last residual block adopts the second structural residual block.
5. The image stitching method according to claim 4, characterized in that: The correlation analysis network includes a context-related module, a relevant feature extraction module, a fully connected module, a control point offset module, a direct linear transformation module and a splicing domain transformation module, which are connected in sequence. The context-related module is based on a directional feature interaction mechanism and accurately captures spatial correspondence clues in the splicing feature map by separating the feature dependencies in the horizontal and vertical directions. The relevant feature extraction module is used to extract features in the splicing feature map, the fully connected module is used to adjust the feature vector of the splicing feature map, the control point offset module is used to control the offset of the pixel points in the splicing feature map, and the direct linear transformation module is used to transform the splicing feature map and generate a distorted feature map. After the distorted feature map is fused with one of the images to be spliced, it is spliced with the other image to be spliced through the splicing domain transformation module to generate the registration feature map and mask.
6. The image stitching method according to claim 5, characterized in that: The registration feature map is divided into several grid cells. In the direct linear transformation module, a direct linear transformation is used for each grid cell to calculate the local homography matrix H from the offset. The mathematical expression of the homography transformation is: ; in, is the source image coordinate, is the target image coordinate; is the element in the local homography matrix H; The direct linear transformation module solves the following overdetermined equations by least squares: ; Convert to matrix form , , the optimal solution is obtained by finding the eigenvector corresponding to the minimum singular value through singular value decomposition, and T represents transpose.
7. The image stitching method according to claim 6, characterized in that: The fusion network includes an initial processing module, multiple encoding modules and multiple decoding modules corresponding to the encoding modules, which are connected in sequence. The initial processing module is used to fuse the registration feature map. The encoding module is used to halve the size of the registration feature map to enhance the feature expression of the registration feature map. The decoding module is used to increase the size of the registration feature map in the corresponding decoding and fuse the deep semantic features with the shallow detail features in the registration feature map to obtain the fused image.
8. The image stitching method according to claim 7, wherein: The fusion network reconstructs the overlapping area of the registered feature map pair, generates a reconstruction result with the same size of the registered feature map and the input image to be stitched after decoding, restores the pixel details and semantic consistency of the overlapping area, and directly uses the original image area after registration as the retained part for the non-overlapping area. Finally, through pixel-level fusion operation, the retained non-overlapping area and the reconstructed overlapping area are seamlessly spliced to generate a fused image with a unified color space and geometric consistency.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the image stitching method according to any one of claims 1 to 8 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the image stitching method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Shelf commodity image splicing method and system
CN115115522A
Image splicing method based on pyramid structure super-resolution network
CN115841422A
Image splicing method, system and device based on deep learning and medium
CN116934592A
Low-light image splicing method and device, medium and equipment
CN118350990A
Image splicing method and device based on homography estimation
CN118552402A
Cited By
Image fusion system
CN121685754A