Image splicing method and device, storage medium and computer equipment
By introducing multimodal prior information on instance-level semantics and geometric structure, a twisted network integrating instance- and structure-aware attention mechanisms is constructed. This solves the alignment and structure preservation problems of image stitching in large parallax scenarios, improving the geometric accuracy and visual realism of image stitching.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-15
- Publication Date
- 2026-03-10
AI Technical Summary
When dealing with complex scenes with large parallax, existing image stitching technologies struggle to maintain the stability of the global structure while ensuring alignment accuracy, leading to artifacts such as misalignment and ghosting, as well as unnatural distortions.
By introducing multimodal prior information on instance-level semantics and geometric structure, a warped network integrating instance- and structure-aware attention mechanisms is constructed to perform image warping alignment and fusion, generating high-quality stitched images.
It significantly improves the geometric accuracy and visual realism of image stitching in complex scenes, resolves the contradiction between content alignment and structure preservation in scenes with large parallax, and achieves high-quality image stitching.
Smart Images

Figure CN121639451A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of image stitching, and in particular, to an image stitching method and device, a storage medium and a computer device. BACKGROUND
[0002] Image stitching is a key technology in computer vision, aiming to generate a seamless panoramic image with a large field of view by merging multiple images with narrow fields of view. With the continuous progress of computer vision technology, image stitching has been widely applied in various fields such as remote sensing, virtual reality, medical imaging, etc.
[0003] However, when dealing with complex scenes with large parallax, especially when the scene contains non-planar structures or there is a large amount of camera translation, image stitching still faces great challenges. The core problem lies in how to balance the accurate alignment of the overlapping area and the protection of the global structure, avoiding misalignment, ghosting and other artifacts caused by improper alignment, while avoiding unnatural distortion caused by excessive alignment. Although there are many methods trying to solve these problems, how to ensure the alignment accuracy while effectively maintaining the stability of the global structure is still a difficult point in image stitching technology.
[0004] In related technologies, existing traditional image stitching techniques are usually based on local deformation models such as bi-homographic deformation and mesh methods, which improve the stitching effect by estimating the local transformation of image blocks. Some methods based on geometric structure preservation, such as line-point consistency and geometric structure preservation, preserve the geometric shapes in the scene to reduce projection distortion and improve the stitching effect. However, these methods usually rely on low-level feature extraction, resulting in poor performance in complex scenes or texture-poor situations. SUMMARY
[0005] The embodiments of the present disclosure at least provide an image stitching method, device, storage medium and computer device, by introducing multi-modal prior information containing instance-level semantics and geometric structure, and constructing a distortion network integrating instance and structure perception attention mechanism, effectively solving the contradiction between content alignment and structure preservation in large parallax scenes, thereby significantly improving the geometric accuracy and visual reality of image stitching in complex scenes.
[0006] The embodiments of the present disclosure provide an image stitching method, comprising: obtaining a set of images to be stitched, the set of images to be stitched comprising at least two images to be stitched; for each image to be stitched in the set of images to be stitched, generating multi-modal prior information, wherein the multi-modal prior information comprises an instance ID mapping graph, a binary mask graph and an edge feature graph; Based on the multimodal prior information, the image group to be stitched is distorted and aligned using a distortion network to obtain the target image group; The target image group is processed by image fusion based on the image fusion model to obtain the target stitched image.
[0007] This disclosure provides an image stitching device, including: An image acquisition module is used to acquire a group of images to be stitched together, wherein the group of images to be stitched together includes at least two images to be stitched together; The information generation module is used to generate multimodal prior information for each image in the image group to be stitched; wherein, the multimodal prior information includes an instance ID mapping map, a binary mask map, and an edge feature map; The image processing module is used to perform distortion alignment processing on the image group to be stitched through a distortion network based on the multimodal prior information to obtain the target image group; The image fusion module is used to perform image fusion processing on the target image group based on the image fusion model to obtain the target stitched image.
[0008] This disclosure provides a computer device, including a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the computer device is running, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, they perform an image stitching method as described in any of the above possible embodiments.
[0009] This disclosure provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the image stitching method as described in any of the possible embodiments above.
[0010] The image stitching method, apparatus, storage medium, and computer device provided in this disclosure acquire a group of images to be stitched and generate multimodal prior information including instance ID mapping, binary masks, and edge features. Based on this, a twisted network integrating instance and structure-aware attention mechanisms is used for twist alignment processing. Finally, a fusion network is used to generate the stitching result (i.e., the target stitched image). Thus, by introducing instance-level information and geometric structural features with clear semantic distinguishing capabilities, and constructing a twisted network integrating instance and structure-aware attention mechanisms, this disclosure fully considers the semantic information and structural features in the image, effectively solving the fundamental contradiction between content alignment and structure preservation in high parallax scenes, and significantly improving the geometric accuracy and visual realism of image stitching in complex scenes.
[0011] To make the above-mentioned objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0012] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings referenced in the embodiments will be briefly described below. These drawings are incorporated in and constitute a part of this specification. They illustrate embodiments conforming to this disclosure and, together with the specification, serve to explain the technical solutions of this disclosure. It should be understood that the following drawings only show some embodiments of this disclosure and should not be considered as limiting the scope. Those skilled in the art can obtain other related drawings based on these drawings without creative effort.
[0013] Figure 1 A flowchart of an image stitching method provided by an embodiment of this disclosure is shown; Figure 2 A flowchart of an image distortion alignment processing method provided by an embodiment of this disclosure is shown; Figure 3 A flowchart of an image stitching method based on a multi-dimensional attention mechanism provided in an embodiment of this disclosure is shown; Figure 4 A flowchart of an image transformation method based on multi-feature fusion provided in an embodiment of this disclosure is shown; Figure 5 A flowchart illustrating a training method for a twisted network provided in an embodiment of this disclosure is shown; Figure 6 A flowchart illustrating a specific training method for a twisted network provided in an embodiment of this disclosure is shown; Figure 7 A schematic diagram of the structure of an image stitching device provided in an embodiment of this disclosure is shown; Figure 8 A schematic diagram of another image stitching device provided in an embodiment of this disclosure is shown; Figure 9 A schematic diagram of the structure of a computer device provided in an embodiment of this disclosure is shown. Detailed Implementation
[0014] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. The components of the embodiments of this disclosure described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this disclosure provided in the accompanying drawings is not intended to limit the scope of the claimed disclosure, but merely represents selected embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.
[0015] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0016] In this document, the term "and / or" merely describes a relationship, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Furthermore, the term "at least one" in this document means any combination of at least two of any one or more elements. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.
[0017] Image stitching, a fundamental technology in computer vision, aims to construct wide-field-of-view panoramas by seamlessly combining multiple narrow-field-of-view images. Despite significant progress in this field, critical and persistent challenges remain when dealing with scenes exhibiting large parallax caused by non-planar structures or extensive camera panning. The core of this challenge lies in the inherent tension between two competing objectives: achieving precise content alignment in overlapping areas while maintaining global structural protection to ensure a natural and distortion-free result; misalignment can lead to undesirable artifacts such as ghosting, while overly aggressive alignment can cause unrealistic distortions, severely degrading the visual quality of the panorama.
[0018] Research has revealed that various local deformation models have been introduced into related methods to alleviate these limitations. For example, biisomorphic deformation can estimate multiple isomorphisms in different planar regions; mesh-based methods, such as the "project as far as possible" algorithm, can compute local transformations of image patches; and there are methods utilizing line-point consistency or geometry preservation techniques, focusing on preserving scene geometry and reducing projection distortion. However, the reliance of these methods on handcrafted, low-level features limits their robustness in complex, texture-sparse scenes.
[0019] With the development of deep learning, a new paradigm for image stitching has emerged. Unsupervised end-to-end frameworks (such as UDIS and its successors) can learn transform fields from data, providing a novel approach to stitching. These methods typically combine global isomorphism and local refinement strategies, optimizing the stitching effect through photometric consistency loss. However, these photometric-based optimization methods often neglect the semantic structure of the scene, potentially leading to unnatural distortions in complex scenes, especially when dealing with structural objects like buildings. In such cases, the network may unreasonably distort these objects to reduce pixel errors.
[0020] Based on the above research, this disclosure provides an image stitching method, apparatus, storage medium, and computer device. By acquiring a group of images to be stitched and generating multimodal prior information including instance ID mapping, binary mask, and edge features, the method performs distortion alignment processing through a distortion network that integrates instance and structure-aware attention mechanisms, and finally generates the stitching result (i.e., the target stitched image) through a fusion network.
[0021] In this embodiment, by introducing instance-level information and geometric structural features with clear semantic distinguishing capabilities, and constructing a twisted network that integrates instance and structure-aware attention mechanisms, the semantic information and structural features in the image are fully considered. This effectively solves the fundamental contradiction between content alignment and structure preservation in large parallax scenes, and significantly improves the geometric accuracy and visual realism of image stitching in complex scenes.
[0022] To facilitate understanding of this embodiment, the executing entity of the image stitching method provided in this disclosure will first be described in detail. The executing entity of the image stitching method provided in this disclosure is a computer device. This computer device can be a terminal device or a server. The terminal device can also be a mobile device, a user terminal, a terminal, a handheld device, a computing device, an in-vehicle device, a wearable device, etc. The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud storage, big data, and artificial intelligence platforms. Optionally, this method can also be applied to an implementation environment composed of computer devices and servers.
[0023] The image stitching method provided in the embodiments of this application will be described in detail below with reference to the accompanying drawings. See also Figure 1 The diagram shows a flowchart of an image stitching method provided in this embodiment of the present disclosure. The method includes the following steps S101 to S104: S101, Obtain the image group to be stitched.
[0024] Here, the image set to be stitched includes at least two images, typically taken by a camera from different perspectives with some overlap. These images can be captured by a camera, retrieved from a storage device, or obtained from other image sources. For example, when you want to stitch multiple photos of different parts of a large scene into a complete panorama, these photos constitute the image set to be stitched. Each image records information about a specific part of the large scene, and through subsequent stitching processing, this scattered information can be integrated to form a more comprehensive presentation of the scene.
[0025] S102, generate multimodal prior information for each image in the image group to be stitched.
[0026] It is understandable that multimodal prior information refers to a collection of various types of information extracted from the image to be stitched before performing image stitching and related tasks. This information can provide key guidance and assistance for subsequent processing. It can be used to guide key operations such as alignment and fusion during the image stitching process, ensuring the accuracy and high quality of the stitching results. It is composed of a combination of various types of information, including instance ID mapping maps, binary mask maps, and edge feature maps.
[0027] An instance ID map is a mapping diagram that records the unique identifiers of different instances (such as objects, people, etc.) in an image. For example, in an image containing multiple different objects, the instance ID map can assign a unique ID to each object to accurately identify and distinguish them during the stitching process. A binary mask is an image with only two values (usually 0 and 1). It is used to mark specific regions in an image. For example, when stitching images, it may be of interest only to a specific region. A binary mask can mark that region as 1 and other regions as 0, so that subsequent processing only operates on that region of interest. An edge feature map is an image that extracts the edge information of objects in an image. It can highlight the outlines of objects in an image. For example, in an image containing buildings and the sky, the edge feature map can clearly outline the outlines of the buildings, helping to better align the edges of objects in different images during the stitching process.
[0028] For example, in order to achieve more accurate object alignment and targeted region processing in the subsequent image stitching process, thereby improving the quality and accuracy of the stitched image, the following steps (1) to (2) may be included when generating the multimodal prior information corresponding to each image to be stitched: (1) For each image to be stitched in the image group to be stitched, the image to be stitched is segmented using an image segmentation model to generate an instance ID mapping map and a binary mask map corresponding to the image to be stitched. (2) For each image to be stitched in the image group, the edge detection algorithm is used to perform convolution processing on the image to be stitched to generate an edge feature map corresponding to the image to be stitched.
[0029] As we can understand, an image segmentation model is a model built based on deep learning or traditional image processing algorithms. It is primarily used to divide an image into multiple semantically meaningful regions and classify and label these different regions. In instance segmentation, it can not only distinguish different object categories in an image (such as people, vehicles, buildings, etc.) but also assign a unique identifier to each independent object instance. Taking an image of a city street containing multiple pedestrians and vehicles as an example, the image segmentation model can identify each pedestrian and each car as an independent instance and assign them different IDs. The instance ID mapping map is a visual representation of this identification relationship. It associates each instance in the image with its unique ID. In subsequent image stitching, when processing multiple images containing the same or similar instances, the instance ID mapping map can accurately find the corresponding instances, achieving precise instance-level alignment.
[0030] For example, based on the segmentation results of an image segmentation model, given a set of k binary masks generated by the image segmentation model, an instance ID map can be generated by assigning a unique integer index to each mask and overlaying them onto a canvas. For each pixel location, its ID value is defined as: .
[0031] Understandably, binary masks are generated based on instance segmentation results. They highlight areas of interest by marking specific instance regions in the image as 1 and the rest as 0. For example, if only vehicles are of interest in an image, the binary mask will set the vehicle region to 1 and the background and other non-vehicle regions to 0. This allows subsequent processing to target only the vehicle region, improving efficiency and focus.
[0032] Specifically, edge detection algorithms are a class of algorithms used to identify the edges of objects in an image. Their basic principle is to extract and analyze local features of the image through convolution operations. Convolution processing uses a specific convolution kernel to slide across the image, weighting and summing the pixel values of each local region to obtain the feature representation of that region. Commonly used convolution kernels in edge detection include the Sobel operator and the Prewitt operator. These operators can highlight areas in the image where pixel values change drastically, and the edges of objects are usually where pixel values change abruptly. Taking an image containing a circular object as an example, after convolution processing using an edge detection algorithm, the generated edge feature map will clearly show the outline of the circle, distinguishing the circular object from the background. In image stitching, edge feature maps can help to more accurately find the corresponding edges of objects in different images, achieving precise edge-based alignment. For example, when stitching two images containing similar objects, by comparing their edge feature maps, matching points of the object edges can be found, thus accurately aligning the two images together and making the stitched image more natural and seamless.
[0033] S103, based on the multimodal prior information, the image group to be stitched is distorted and aligned using a distortion network to obtain the target image group.
[0034] Specifically, the warping network is a specially designed neural network model constructed in this disclosure, and its training steps will be discussed in detail later. Here, the warping network can deform and adjust each image to be stitched in the image group based on multimodal prior information, so that the same objects or regions in different images can be aligned. For example, when the same building exists in two images to be stitched, but the position and shape of the building in the images differ due to different shooting angles, the warping network can distort and deform the two images based on information such as instance ID mapping map, binary mask map, and edge feature map, so that the position and shape of the building in the two images are consistent, thereby achieving alignment. The image group obtained after warping and alignment processing is the target image group, at which point the images in the target image group have been initially aligned in terms of content.
[0035] Here, the warping network can include instance and structure-aware attention modules. These modules, during image warping alignment, comprehensively consider instance information (such as instance ID mappings of different objects) and structural information (such as object edge contours and region distribution) within the image. They highlight features that have a key impact on the stitching process and suppress irrelevant or interfering information, thereby more accurately guiding image warping and achieving high-quality alignment at both the instance and structural levels between different images. (Refer to...) Figure 2 As shown, the process of distorting and aligning the image groups to be stitched together may specifically include the following steps S201~205: S201, Extract the visual feature pyramid of each image to be stitched together.
[0036] Specifically, a visual feature pyramid is a multi-level feature representation structure, primarily obtained by downsampling and feature extraction operations at different scales on the input image. Starting with the original image, operations such as convolution and pooling are used to progressively reduce the image resolution, while simultaneously extracting corresponding features at each scale. For example, high-resolution layers can capture detailed information about the image, such as object textures and small parts; while low-resolution layers can obtain global information about the image, such as the overall shape of objects and the general layout of the scene. Thus, by constructing a visual feature pyramid, the content of an image can be represented more comprehensively and at multiple levels.
[0037] S202, for each image to be stitched, the instance ID mapping map corresponding to the image to be stitched is converted into a dense semantic feature map through an embedding layer; and the dense semantic feature map is fused with the visual feature pyramid of the image to be stitched to generate an instance-aware feature map corresponding to the image to be stitched.
[0038] Understandably, since convolutional networks cannot directly process discrete integer IDs, this disclosure proposes to address this problem by employing an embedding layer. Here, the embedding layer is a neural network layer used to convert discrete identification information (such as instance IDs) into continuous vector representations. Essentially, it is a learnable lookup table that can retrieve the corresponding N-dimensional vector (N being the dimension of the embedding space) for each pixel ID.
[0039] Specifically, the instance ID map records the unique identifier of each instance in the image. After processing by the embedding layer, these discrete IDs are transformed into dense vectors with semantic information, namely, dense semantic feature maps. These feature maps can more effectively express the semantic characteristics of instances. Then, the dense semantic feature maps are fused with the visual feature pyramid. The fusion process can employ operations such as concatenation or addition, so that the fused instance-aware feature map contains both the visual hierarchical information of the image and the semantic information of the instances. For example, in an image containing multiple animals and plants, the instance-aware feature map can distinguish different animal and plant instances while preserving their respective visual features, thus enabling more accurate identification and manipulation of each instance in subsequent processing.
[0040] S203, for each image to be stitched, the binary mask image and edge feature map corresponding to the image to be stitched are stitched together, and the feature of the stitched structure map is extracted by the guided encoder to generate a multi-scale structure guided feature map.
[0041] Understandably, a binary mask is used to mark specific regions in an image, while an edge feature map highlights the edge information of objects in the image. Concatenating these two creates a structure map that comprehensively reflects the image's structural information. Here, the guided encoder is a neural network structure capable of multi-level feature extraction from the input structure map. Through different levels of convolution and pooling operations, the guided encoder can progressively extract low- to high-level features from the structure map, generating multi-scale guided structure feature maps. These feature maps describe the image's structural information at different scales. For example, at low scales, they may focus on detailed edge features, while at high scales, they can grasp the general outline and regional distribution of objects, providing structural guidance for subsequent image alignment.
[0042] S204, the instance-aware feature map and the structure-guided feature map corresponding to each image in the image group to be stitched are input to the instance and structure-aware attention module to generate an enhanced feature map.
[0043] Specifically, the instance- and structure-aware attention module receives instance-aware feature maps and structure-guided feature maps from each image to be stitched, and adaptively calculates attention weights based on the information in these feature maps. Specifically, it analyzes the importance of different instance and structural features in image alignment, assigning higher attention weights to key features and lower weights to irrelevant or interfering features. Through this attention mechanism, the module can highlight features that significantly impact image alignment and suppress irrelevant information. For example, when stitching two images containing similar buildings, the module focuses on the features of the building instances and their edge structure information, assigning higher weights to these features to generate an enhanced feature map. This enhanced feature map integrates key information from instances and structures, possessing stronger representational capabilities and more accurately guiding the image warping alignment operation, enabling better matching and fusion of instances and structures from different images.
[0044] For example, refer to Figure 3 As shown, this disclosure proposes an image stitching method based on a multi-dimensional attention mechanism. When performing feature processing and alignment on the image groups to be stitched to achieve high-quality image stitching, the method may specifically include the following steps S2041~S2043: S2041, using the instance and the structure-aware attention module, interactive processing is performed on the structure-guided feature maps corresponding to each image to be stitched to generate a spatial attention map.
[0045] Specifically, the structure-guided feature map is a feature representation that integrates multi-scale structural features such as object edge information and region distribution information in an image. It characterizes the structural properties of an image from different levels and perspectives. During the interactive processing of the structure-guided feature maps corresponding to the images to be stitched, the instance and structure-aware attention module analyzes the spatial distribution and correlation of structural features in different images. For example, when stitching two images containing similar architectural scenes, the module identifies the spatial correspondence of structural elements such as building edges and windows in the two images. Through this interactive analysis, a spatial attention map can be generated. Each element value in the map reflects the importance of structural features at the corresponding spatial location, providing spatial attention guidance for subsequent feature processing, directing attention to spatial regions that play a key role in image stitching.
[0046] In some possible implementations, a structural attention head can be used during the structural attention generation process. This head is a sub-network that converts encoded structural features into a refined spatial attention map, and it consists of two convolutional layers. First, a 1x1 convolution projects the high-dimensional guiding features into a lower-dimensional space, followed by batch normalization and ReLU activation. This reduces channel redundancy and aligns the spatial distribution of the guiding features with the main visual flow. During the interactive processing of the structural guiding feature maps corresponding to the images to be stitched, the instance and structure-aware attention modules analyze the spatial distribution and correlation of structural features in different images. For example, when stitching two images containing similar architectural scenes, the module identifies the spatial correspondence of structural elements such as building edges and windows in the two images. Subsequently, a 3x3 convolution aggregates local spatial information and compresses the channel dimension to one dimension, thereby generating the final attention map precursor.
[0047] Furthermore, after obtaining the attention map precursor, it can be fed into a soft gating system using a scaled sigmoid. Here, instead of using the standard sigmoid function, this disclosure applies a custom scale-sigmoid activation function to the attention map precursor. The reason for this is that the standard sigmoid function outputs in the range [0,1], which can lead to aggressive feature suppression (values close to 0) in regions with weak or potentially inaccurate structural priors (e.g., caused by segmentation errors). This suppression can be harmful because these regions may still contain valuable visual context. The scale-sigmoid, however, maps the output to a more constrained range, where a hyperparameter controls the minimum attenuation factor. This transformation is defined in a specific functional form, and this design choice acts as a soft gating mechanism, ensuring that even in regions with the lowest structural confidence, the original visual features are attenuated by at most a factor of [1-a], rather than being completely suppressed. The transformation is defined as follows: .
[0048] Furthermore, the generated soft attention map is then used to modulate the instance-aware features of the input. This process can be achieved through additive interactions within a residual framework, aiming to enhance features of structurally salient regions without suppressing the underlying visual context. To optimally integrate this enhanced information and further refine the representation, the modulated features undergo a final refinement stage. This stage primarily involves a 1x1 refined convolution that linearly reassembles the feature channels. This is followed by a Convolutional Block Attention Module (CBAM) that recalibrates the feature map through sequential channel and spatial attention to suppress noise and highlight the most discriminative information. The entire process ultimately produces a highly refined feature representation that incorporates structural information.
[0049] In this way, through the above interactive analysis and subsequent special activation function processing, a spatial attention map can be generated. Each element value in the map reflects the importance of structural features at the corresponding spatial location, providing spatial attention guidance for subsequent feature processing. This guides the focus on spatial regions that play a key role in image stitching, making the guidance mechanism more resilient to noise or imperfect prior information. While emphasizing structurally significant regions, it maintains the integrity of the underlying feature map.
[0050] S2042, using the instance and structure-aware attention module, the spatial attention map is used to perform cross-image weighted modulation on the input instance-aware feature map to generate a modulated feature map.
[0051] Specifically, an instance-aware feature map is a feature representation that integrates semantic information and visual features of different instances (such as people, vehicles, and other independent objects) in an image. It can distinguish and represent various instances in an image. After obtaining the spatial attention map, the instance and structure-aware attention modules perform cross-image weighted modulation on the instance-aware feature map based on the weight information in the spatial attention map.
[0052] For example, features at different spatial locations in the instance-aware feature map can be weighted according to their corresponding weights in the spatial attention map. Spatial locations with higher weights will have their corresponding instance features enhanced, while those with lower weights will have their features relatively suppressed. For instance, if the spatial attention map shows that the structural features of a certain region are important for stitching, then the instance-aware features corresponding to that region will be enhanced, thus generating a modulated feature map. The modulated feature map retains the semantic information of the instances while incorporating the importance of spatial structure, making the instance features between different images more spatially correlated and consistent.
[0053] S2043, using the instance and structure-aware attention module, perform channel and spatial attention recalibration processing on the modulated feature map, and output the enhanced feature map.
[0054] Understandably, the modulated feature map contains rich instance and structural information, but the importance of features in different channels and spatial locations may vary. Channel and spatial attention recalibration aims to further optimize these features. In the channel dimension, the module analyzes the contribution of each channel's features to image stitching, assigning higher weights to important channels to enhance their feature representation; for less important channels, the weights are reduced. In the spatial dimension, the previously generated spatial attention map can be combined to further refine the features at spatial locations.
[0055] Thus, through this dual attention recalibration of channel and space, the embodiments of this disclosure can make the features in the enhanced feature map more focused and effective, more accurately reflect the key image information between different images, and provide more discriminative and representative feature representations.
[0056] S205, based on the enhanced feature map, determine the transformation relationship between each image to be stitched in the image group to be stitched, and perform distortion alignment processing on each image to be stitched in the image group to be stitched based on the transformation relationship.
[0057] Understandably, enhanced feature maps integrate semantic information of instances, structural information of images, and key features optimized by attention mechanisms, enabling them to accurately describe the similarities and differences between different images. Using enhanced feature maps, transformation relationships between images to be stitched can be determined through feature matching algorithms (such as deep learning-based feature matching methods or traditional feature point matching algorithms). These transformation relationships include geometric transformation parameters such as translation, rotation, and scaling. For example, by comparing the positions and feature similarities of corresponding feature points in the enhanced feature maps, information such as relative positional changes and angular rotations between images can be calculated. After determining the transformation relationships, each image in the stitching group is warped and aligned according to these parameters, ensuring accurate matching and fusion of corresponding content in different images, ultimately achieving high-quality image stitching.
[0058] In some possible embodiments, when faced with multiple images to be stitched together, each with complex scenes and different perspectives, significant geometric differences may exist between the images, such as different shooting angles, changes in the position and shape of objects within the frame, etc. Therefore, relying solely on a single alignment method may be insufficient to achieve high-quality image stitching. Thus, a combination of global and local transformation methods can be used to precisely adjust the positional relationships between the images to achieve seamless stitching. Here, the transformation relationship between the images to be stitched in this disclosure is jointly determined by the global homography matrix and the local thin-plate spline deformation field. The global homography matrix describes the projection transformation relationship between two planes, mapping points on one plane to another through linear transformation. It can handle overall rotation, translation, scaling, and perspective transformations between images, and is suitable for preliminary, overall image alignment. The local thin-plate spline deformation field is a local deformation model based on thin-plate spline interpolation theory. It can perform finer deformation adjustments on local areas of the image, capturing subtle local changes such as local deformation of objects and distortion of textures.
[0059] Specifically, refer to Figure 4 As shown, an image transformation method based on multi-feature fusion provided in this disclosure may include the following steps S2051~S2053: S2051, determine the reference image and at least one target image in the image group to be stitched.
[0060] Understandably, in a group of images to be stitched together, the reference image serves as the alignment benchmark, providing a fixed standard for aligning subsequent images. When selecting a reference image, factors such as image quality, sharpness, and the amount of key information it contains are typically considered. For example, an image with complete content, no obvious occlusions, and rich features is chosen as the reference image. The target images are the other images that need to be aligned and stitched with the reference image. They may differ from the reference image in various aspects, requiring subsequent transformation operations to achieve matching with the reference image.
[0061] S2052, based on the enhanced feature map, determine the global homography matrix that aligns each of the target images to the reference image; and perform coarse alignment of each of the target images based on the global homography matrix and using the reference image as a reference.
[0062] Specifically, enhanced feature maps are feature representations containing rich image information obtained after previous feature extraction and processing. They integrate various features such as semantic and structural information of the image, and can more accurately reflect the similarity and differences between images. Using enhanced feature maps, corresponding feature points between the target image and the reference image can be found through feature matching algorithms (such as deep learning-based feature point matching algorithms). Then, the global homography matrix is calculated based on the relationship between these corresponding points. The global homography matrix contains transformation parameters such as rotation, translation, and scaling between images. By applying this matrix to transform the target image, the target image can be roughly aligned to the position and pose of the reference image, achieving coarse image alignment.
[0063] S2053, Based on the enhanced feature map, predict the local thin-plate spline deformation field for fine deformation of each target image after coarse alignment.
[0064] Furthermore, after coarse alignment, some local differences may still exist between the target image and the reference image, such as local texture distortion or minor deformation of objects. At this point, the enhanced feature map can be used again to predict the local thin-plate spline deformation field using a specific algorithm (such as a machine learning-based deformation field prediction algorithm). The local thin-plate spline deformation field defines the deformation pattern of each local region in the image. Based on the local feature differences reflected in the enhanced feature map, it can finely adjust the coarsely aligned target image, enabling local regions of the target image to match the corresponding regions of the reference image more accurately, thereby achieving high-quality image stitching.
[0065] S104, Based on the image fusion model, perform image fusion processing on the target image group to obtain the target stitched image.
[0066] Understandably, an image fusion model is a model that can integrate information from multiple images to generate a high-quality stitched image. It can integrate the target image after the distortion of the target image group by the distortion network, eliminating seams and discontinuities between the images. For example, when stitching multiple photos, there may be differences in lighting and color between different photos. The image fusion model can also adjust and optimize the pixel values to make the stitched image more uniform and natural in lighting and color, ultimately obtaining a seamless, high-quality stitched image that fully presents the information of the entire scene.
[0067] The following is combined with Figure 5 The training steps S501 to S504 of the twisted network constructed in this application are described in detail: S501, acquire multiple complete images, crop each complete image according to preset rules, and obtain multiple cropped images corresponding to each complete image.
[0068] Understandably, training a twisted network first requires acquiring multiple complete images. These complete images refer to the original images without any deformation or special processing, serving as the raw material for subsequent operations. Next, each complete image can be cropped according to preset rules, which can be set according to actual needs, such as uniform cropping or cropping at a specific ratio. Through this cropping, multiple cropped images corresponding to each complete image can be obtained, covering the entirety of the original complete image.
[0069] S502, for each cropped image corresponding to a complete image, determine one of them as the reference image, and perform random image deformation processing on the other cropped images to obtain multiple sample images.
[0070] Furthermore, for each cropped image corresponding to a complete image, one needs to be selected as the baseline image. The baseline image serves as a reference throughout the training process, providing a standard for the deformation and comparison of other images. After determining the baseline image, random image deformation processing, such as rotation, scaling, translation, and distortion, can be applied to the other cropped images to simulate various image deformations that may occur in real-world scenarios. This processing yields multiple sample images, which, together with the baseline image, constitute the data pairs required for subsequent training.
[0071] S503: Based on the baseline image and sample image corresponding to each complete image, construct a training sample subset, and construct the training sample set from all training sample subsets.
[0072] Understandably, each training sample subset contains a baseline image and its corresponding deformed sample image, with a specific correspondence between them; that is, the sample image is obtained by deforming the baseline image in a certain way. All training sample subsets are combined to form the training sample set, which serves as the data foundation for the entire training process. It contains a large number of image pairs, and based on this training sample set, the network can learn image restoration rules under different deformation conditions.
[0073] S504, Obtain the twisted network to be trained, and perform unsupervised training on the twisted network to be trained based on each set of training sample subsets to obtain the trained image processing model.
[0074] A twisted network to be trained is obtained, and unsupervised training is performed on the twisted network based on each subset of training samples to obtain the trained image processing model. Specifically, refer to... Figure 6 As shown, the training process may include the following steps S5041~S5044: S5041, for each training sample subset, the baseline image is used as the reference image, and each sample image is used as the target image, and input into the twisted network to be trained.
[0075] Specifically, for each subset of training samples, a reference image is used as the reference image, and each sample image is used as the target image, both input into the warping network to be trained. The reference image provides a standard for the network, while the target image is the image to be processed and restored. The task of the warping network is to appropriately deform and adjust the target image based on the information from the reference image, aligning it with the reference image. For the specific implementation steps of the warping network, please refer to steps S201-S204 above, which will not be repeated here.
[0076] S5042, Obtain the distorted target image output by the distorted network and aligned to the reference image.
[0077] It is understandable that the distortion network can predict the corresponding deformation parameters by performing a series of distortions on the input target image (see steps S201~S204 above), and then deform the target image according to these parameters to finally obtain a distorted target image aligned with the reference image.
[0078] S5043, construct the composite loss function.
[0079] Understandably, the composite loss function is a key metric for evaluating the performance of twisted networks. Here, the composite loss function includes the photometric alignment loss function, the boundary consistency loss function, and the smoothness regularization loss function, and its expression can be represented as: ; ; ; ; ; in, Represented as the optical alignment loss function; Represented as the smoothness regularization loss function; Represented as the smoothness regularization loss function; , and It is a hyperparameter that balances the contributions of each component; H represents the global homography matrix; Represented as the target image; Represented as a reference image; It represents the hyperparameter of the optical alignment loss function; TPS represents the fine deformation after passing through the local thin-plate spline deformation field; Represented as a binary mask image; This is represented as a boundary region mask determined by a binary mask image; Residual motion fields are represented as loss penalty predictions; and These represent the spatial gradients in the horizontal and vertical directions, respectively.
[0080] Specifically, the photometric alignment loss function is constructed based on the pixel differences between the distorted target image and the reference image within the effective overlap region. The effective overlap region refers to the part where the distorted target image and the reference image overlap spatially. By calculating the pixel differences within this region, the degree of alignment between the distorted target image and the reference image in terms of brightness, color, and other photometric aspects can be measured. The boundary consistency loss function is constructed based on the pixel differences between the distorted target image and the reference image within the boundary region determined by the binary mask image of the reference image. By calculating the pixel differences between the distorted target image and the reference image within the boundary region, it can be ensured that the distorted image remains consistent with the reference image at the boundary, avoiding boundary mismatch problems. The smoothness regularization loss function is constructed based on the spatial gradient of the local thin-plate spline deformation field predicted by the distortion network. The spatial gradient reflects the rate of change of the deformation field in different directions. By constructing the smoothness regularization loss function, the changes in the deformation field can be constrained, making the deformation process smoother and avoiding unreasonable deformation.
[0081] S5044, Based on the calculation result of the composite loss function, the parameters of the twisted network are updated through the backpropagation algorithm until the performance of the twisted network converges or reaches the preset training rounds, thus obtaining the trained twisted network.
[0082] Here, backpropagation is an algorithm used to calculate the gradients of parameters in a neural network. It propagates gradient information back from the output layer to the input layer based on the value of the loss function, thereby calculating the contribution of each parameter to the loss function. Based on this gradient information, optimization algorithms (such as stochastic gradient descent) are used to update the parameters of the warped network, gradually reducing the value of the loss function. This process is repeated until the performance of the warped network converges or a preset number of training epochs is reached, resulting in a well-trained warped network. When the network's performance converges, it means that the network has learned good image deformation and restoration capabilities; reaching the preset number of training epochs prevents the training process from becoming too long and wasting computational resources.
[0083] The image stitching method, apparatus, storage medium, and computer device provided in this disclosure introduce instance-level information and geometric structural features with clear semantic distinguishing capabilities, and construct a twisted network that integrates instance and structure-aware attention mechanisms. This fully considers the semantic information and structural features in the image, effectively solves the fundamental contradiction between content alignment and structure preservation in large parallax scenes, and significantly improves the geometric accuracy and visual realism of image stitching in complex scenes.
[0084] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
[0085] Based on the same inventive concept, this disclosure also provides an image stitching device corresponding to the image stitching method. Since the principle of the device in this disclosure for solving the problem is similar to the image stitching method described above in this disclosure, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.
[0086] Reference Figure 7 The diagram shown is a schematic representation of an image stitching device 700 provided in an embodiment of this disclosure. The device includes: Image acquisition module 701 is used to acquire a group of images to be stitched together, wherein the group of images to be stitched together includes at least two images to be stitched together; The information generation module 702 is used to generate multimodal prior information for each image in the image group to be stitched; wherein, the multimodal prior information includes an instance ID mapping map, a binary mask map, and an edge feature map; Image processing module 703 is used to perform distortion alignment processing on the image group to be stitched through a distortion network based on the multimodal prior information to obtain the target image group; The image fusion module 704 is used to perform image fusion processing on the target image group based on the image fusion model to obtain the target stitched image.
[0087] In some possible embodiments, the information generation module 702 is specifically used for: For each image in the image group to be stitched, an image segmentation model is used to perform instance segmentation processing on the image to be stitched, generating an instance ID mapping map and a binary mask map corresponding to the image to be stitched. For each image in the image group to be stitched, an edge detection algorithm is used to perform convolution processing on the image to be stitched to generate an edge feature map corresponding to the image to be stitched.
[0088] In some possible embodiments, the image processing module 703 is specifically used for: Extract the visual feature pyramid of each image to be stitched together; For each image to be stitched, the instance ID mapping map corresponding to the image to be stitched is converted into a dense semantic feature map through an embedding layer; and the dense semantic feature map is fused with the visual feature pyramid of the image to be stitched to generate an instance-aware feature map corresponding to the image to be stitched. For each image to be stitched, the binary mask image and edge feature map corresponding to the image to be stitched are stitched together, and the feature map of the stitched structure is extracted by the guided encoder to generate a multi-scale structure guided feature map; The instance-aware feature map and the structure-guided feature map corresponding to each image in the image group to be stitched are input into the instance and structure-aware attention module to generate an enhanced feature map; Based on the enhanced feature map, the transformation relationship between each image to be stitched in the image group is determined, and the distortion alignment process is performed on each image to be stitched in the image group based on the transformation relationship.
[0089] In some possible embodiments, the image processing module is further configured to: Using the example and the structure-aware attention module, the structure-guided feature maps corresponding to each image to be stitched are interactively processed to generate a spatial attention map. Using the instance and structure-aware attention module, the spatial attention map is used to perform cross-image weighted modulation on the input instance-aware feature map to generate a modulated feature map; Using the example and the structure-aware attention module, the modulated feature map is subjected to channel and spatial attention recalibration processing to output the enhanced feature map.
[0090] In some possible embodiments, the image processing module 703 is further configured to: Determine a reference image and at least one target image in the group of images to be stitched together; Based on the enhanced feature map, a global homography matrix is determined to align each of the target images to the reference image; and based on the global homography matrix, each of the target images is roughly aligned with the reference image. Based on the enhanced feature map, a local thin-plate strip deformation field is predicted for fine deformation of each target image after coarse alignment. Based on the global homography matrix and the local thin-plate spline deformation field, the transformation relationship between each image to be stitched in the image group is determined.
[0091] In some possible embodiments, refer toFigure 8 As shown, the device further includes: The image cropping module 705 is used to acquire multiple complete images, crop each complete image according to a preset rule, and obtain multiple cropped images corresponding to each complete image, wherein the multiple cropped images can cover the entire part of the complete image; The image deformation module 706 is used to determine one of the cropped images corresponding to each complete image as the reference image, and to perform random image deformation processing on the other cropped images to obtain multiple sample images. The sample construction module 707 is used to construct a training sample subset based on the reference image and sample image corresponding to each complete image, and to construct the training sample set from all training sample subsets. The model training module 708 is used to acquire the twisted network to be trained and perform unsupervised training on the twisted network to be trained based on each set of training sample subsets to obtain the trained image processing model.
[0092] In some possible embodiments, the model training module 708 is specifically used for: For each training sample subset, the baseline image is used as the reference image, and each sample image is used as the target image, and input into the twisted network to be trained. Obtain the distorted target image output by the distorting network and aligned to the reference image; A composite loss function is constructed, comprising a photometric alignment loss function, a boundary consistency loss function, and a smoothness regularization loss function. The photometric alignment loss function is constructed based on the pixel differences between the distorted target image and the reference image within the effective overlapping region. The boundary consistency loss function is constructed based on the pixel differences between the distorted target image and the reference image within the boundary region determined by the binary mask image of the reference image. The smoothness regularization loss function is constructed based on the spatial gradient of the local thin-plate strip deformation field predicted by the distorting network. Based on the calculation results of the composite loss function, the parameters of the twisted network are updated through the backpropagation algorithm until the performance of the twisted network converges or reaches the preset number of training rounds, thus obtaining a trained twisted network.
[0093] Based on the same technical concept, this disclosure also provides a computer device. (See also...) Figure 9The diagram shown is a structural schematic of a computer device 900 provided in an embodiment of this disclosure, including a processor 901, a memory 902, and a bus 903. The memory 902 stores execution instructions and includes a main memory 9021 and an external memory 9022. The main memory 9021, also called internal memory, is used to temporarily store computational data in the processor 901, as well as data exchanged with external memory 9022 such as a hard disk. The processor 901 exchanges data with the external memory 9022 through the main memory 9021.
[0094] In this embodiment, the memory 902 is specifically used to store application code that executes the solution of this application, and its execution is controlled by the processor 901. That is, when the computer device 900 is running, the processor 901 communicates with the memory 902 through the bus 903, causing the processor 901 to execute the application code stored in the memory 902, thereby executing the method described in any of the foregoing embodiments.
[0095] The memory 902 may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.
[0096] Processor 901 may be an integrated circuit chip with signal processing capabilities. The aforementioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), a Graphics Processing Unit (GPU), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor can be a microprocessor or any conventional processor.
[0097] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the computer device 900. In other embodiments of this application, the computer device 900 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0098] This disclosure also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the image stitching method described in the above-described method embodiments. The storage medium can be a volatile or non-volatile computer-readable storage medium.
[0099] This disclosure also provides a computer program product carrying program code. The program code includes instructions that can be used to execute the steps of the image stitching method described in the above method embodiments. For details, please refer to the above method embodiments, which will not be repeated here.
[0100] The aforementioned computer program product can be implemented through hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically embodied in a computer storage medium; in another optional embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.
[0101] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and devices described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. In the several embodiments provided in this disclosure, it should be understood that the disclosed systems and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some communication interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.
[0102] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs. Furthermore, the functional units in the various embodiments of this disclosure may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0103] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0104] Finally, it should be noted that the above-described embodiments are merely specific implementations of this disclosure, used to illustrate the technical solutions of this disclosure, and not to limit it. The protection scope of this disclosure is not limited thereto. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this disclosure. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this disclosure, and should all be covered within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be determined by the protection scope of the claims.
Claims
1. An image stitching method, characterized by, The method comprises the following steps: acquiring a set of images to be stitched, the set of images to be stitched comprising at least two images to be stitched; generating multi-modal prior information for each image to be stitched in the set of images to be stitched; wherein the multi-modal prior information comprises an instance ID mapping graph, a binary mask graph and an edge feature graph; performing warping alignment processing on the set of images to be stitched based on the multi-modal prior information through a warping network to obtain a set of target images; performing image fusion processing on the set of target images based on an image fusion model to obtain a target stitched image.
2. The method of claim 1, wherein, The generation of the multi-modal prior information comprises: performing instance segmentation processing on each image to be stitched in the set of images to be stitched using an image segmentation model to generate an instance ID mapping graph and a binary mask graph corresponding to the image to be stitched; performing convolution processing on each image to be stitched in the set of images to be stitched using an edge detection algorithm to generate an edge feature graph corresponding to the image to be stitched.
3. The method of claim 2, wherein, The warping network comprises an instance and structure perception attention module; the warping alignment processing on the set of images to be stitched through the warping network comprises: extracting a visual feature pyramid of each image to be stitched; for each image to be stitched, converting the instance ID mapping graph corresponding to the image to be stitched into a dense semantic feature graph through an embedding layer; and fusing the dense semantic feature graph with the visual feature pyramid of the image to be stitched to generate an instance perception feature graph corresponding to the image to be stitched; for each image to be stitched, splicing the binary mask graph and the edge feature graph corresponding to the image to be stitched, and performing feature extraction on the spliced structure graph through a guided encoder to generate a multi-scale structure guided feature graph; inputting the instance perception feature graph and the structure guided feature graph corresponding to each image to be stitched in the set of images to be stitched into the instance and structure perception attention module to generate an enhanced feature graph; determining a transformation relationship between each image to be stitched in the set of images to be stitched based on the enhanced feature graph, and performing warping alignment processing on each image to be stitched in the set of images to be stitched based on the transformation relationship.
4. The method of claim 3, wherein, The inputting of the instance perception feature graph and the structure guided feature graph corresponding to each image to be stitched in the set of images to be stitched into the instance and structure perception attention module comprises: performing interactive processing on the structure guided feature graph corresponding to each image to be stitched through the instance and structure perception attention module to generate a spatial attention graph; performing cross-image weighting modulation on the input instance perception feature graph using the spatial attention graph through the instance and structure perception attention module to generate a modulated feature graph; performing channel and spatial attention recalibration processing on the modulated feature graph through the instance and structure perception attention module to output the enhanced feature graph.
5. The method of claim 3, wherein, The determination of the transformation relationship between each image to be stitched in the set of images to be stitched based on the enhanced feature graph comprises: determining a reference image and at least one target image in the set of images to be stitched; determine a global homography matrix for aligning each of the target images to the reference image based on the enhanced feature map; and perform coarse alignment of each of the target images based on the global homography matrix and the reference image as a reference; predict a local thin plate spline deformation field for fine deformation of each of the coarsely aligned target images based on the enhanced feature map; determine a transformation relationship between each of the to-be-stitched images in the to-be-stitched image group based on the global homography matrix and the local thin plate spline deformation field.
6. The method of claim 4, wherein, The warping network is trained by the following steps: obtain a plurality of complete images, and obtain a plurality of cropped images corresponding to each complete image by cropping each complete image according to a preset rule, wherein the plurality of cropped images can cover all parts of the complete image; for each complete image, determine one of the cropped images as a reference image, and perform random image deformation processing on the other cropped images to obtain a plurality of sample images; based on the reference image and the sample images corresponding to each complete image, respectively, construct a training sample subset, and construct all training sample subsets into a training sample set; obtain a to-be-trained warping network, and perform unsupervised training on the to-be-trained warping network based on each training sample subset to obtain the trained image processing model.
7. The method of claim 6, wherein, The unsupervised training of the to-be-trained warping network based on each training sample subset comprises: for each training sample subset, input the reference image as a reference image and each sample image as a target image into the to-be-trained warping network; obtain a warped target image output by the warping network and aligned to the reference image; construct a composite loss function; wherein the composite loss function comprises a photometric alignment loss function, a boundary consistency loss function, and a smoothness regularization loss function; the photometric alignment loss function is constructed based on the pixel difference between the warped target image and the reference image in the effective overlap region; the boundary consistency loss function is constructed based on the pixel difference between the warped target image and the reference image in the boundary region determined by the binary mask image of the reference image; and the smoothness regularization loss function is constructed based on the spatial gradient of the local thin plate spline deformation field predicted by the warping network; update the parameters of the warping network through a back propagation algorithm based on the calculation result of the composite loss function until the performance of the warping network converges or a preset training round is reached, to obtain the trained warping network.
8. An image stitching apparatus characterized by comprising: comprise: an image acquisition module configured to acquire a to-be-stitched image group, wherein the to-be-stitched image group comprises at least two to-be-stitched images; an information generation module configured to generate multi-modal prior information for each to-be-stitched image in the to-be-stitched image group, wherein the multi-modal prior information comprises an instance ID mapping image, a binary mask image, and an edge feature map; an image processing module configured to perform warping alignment processing on the to-be-stitched image group based on the multi-modal prior information through a warping network to obtain a target image group; An image fusion module is configured to perform image fusion processing on the target image group based on an image fusion model to obtain a target spliced image.
9. A storage medium having stored thereon a computer program, characterized in that The computer program, when executed by a processor, implements the method of any one of claims 1 to 7.
10. A computer device comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, The processor, when executing the computer program, implements the method of any one of claims 1 to 7.