Remote sensing image splicing method and device and electronic equipment

By using preset semantic segmentation models and feature matching algorithms to process remote sensing images, the stitching misalignment caused by moving objects in urban aerial images is solved, and high-quality large-field image stitching is achieved.

CN120070199APending Publication Date: 2025-05-30AEROSPACE INFORMATION RES INST CAS
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510095448.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the quality problems such as splicing misalignment and tearing caused by changes in the position of moving objects such as vehicles in urban aerial images, especially when flying at low altitudes.

Method used

The remote sensing image is processed using a preset semantic segmentation model, a mask image is generated, and pixel matching pairing is determined through a feature matching algorithm, and image registration and fusion are performed to obtain high-quality stitching images.

Benefits of technology

The quality of the stitching image is improved, the negative impact of moving objects on the stitching results is reduced, and high-precision stitching of dynamic scenes is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070199A_ABST
    Figure CN120070199A_ABST
Patent Text Reader

Abstract

The invention discloses a splicing method and device of remote sensing images and electronic equipment, and relates to the field of remote sensing image processing, and the splicing method comprises the steps: collecting a remote sensing image video of a target area, extracting multiple frames of remote sensing images from the remote sensing image video, processing each frame of remote sensing image by adopting a preset semantic segmentation model, and obtaining a splicing result; the method comprises the steps of obtaining a mask image corresponding to each frame of remote sensing image, extracting features of each frame of mask image in a mask image pair formed by two continuous frames of mask images in each group, obtaining a feature image set of each frame of mask image, determining a pixel matching pair of the mask image pair by adopting a preset feature matching algorithm based on the feature image set, and obtaining a pixel matching pair of the mask image pair based on the pixel matching pair. And registering the mask image pairs, and fusing all the registered mask image pairs to obtain a spliced shadow of the target area. The technical problem that the quality of the spliced image is poor due to the fact that the collected image has repeated textures and moving objects in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of remote sensing image processing, and more particularly, to a method and apparatus for stitching remote sensing images, and an electronic device. Background Art

[0002] With the rapid progress of unmanned aerial vehicle (UAV) technology, UAV aerial images have become one of the main means of obtaining geographical information. Compared with remote sensing images, UAV aerial photography is carried out in a low-altitude environment, with a shorter acquisition cycle, lower cost, and higher flexibility. UAVs can acquire images at different positions for a specific target area. However, due to the limitations of flight altitude and camera focal length, the coverage of each image is small, belonging to small field-of-view images, and cannot fully capture the entire target area. To comprehensively observe the target area, image stitching is required to generate large field-of-view images.

[0003] There is a wide range of demands for the large field-of-view images generated by stitching urban aerial images. First of all, it can comprehensively evaluate infrastructure, which is very important for safety and infrastructure improvement. Secondly, the stitched images can assist urban construction, evaluate traffic conditions, and ensure reasonable road construction. Moreover, these images can be used for land planning. Therefore, researching aerial image stitching algorithms has high social and economic value. However, due to the frequent changes in the positions of elements such as vehicles in urban aerial images due to movement, image stitching poses great challenges.

[0004] Most of the current stitching algorithms are applicable to images collected by UAVs in a static environment, but in fact, the images are captured dynamically, especially when flying at low altitude. Considering that large changes in object positions may cause low-altitude aerial images to be affected by moving foregrounds during stitching, resulting in quality problems such as stitching misalignment and tearing. In related technologies, generally, a UAV aerial image stitching method based on semantic segmentation and feature point matching algorithms is adopted. A semantic segmentation network is introduced to separate the foreground and background of the image and obtain foreground semantic information. At the same time, a classic feature point matching algorithm is used to extract feature points. By comparing the feature point information with the foreground semantic information, foreground feature points can be deleted to achieve feature point matching. However, the current algorithms have not been able to well solve the problem of stitching misalignment during background stitching caused by the dynamic foregrounds of moving objects. Therefore, stitching urban remote sensing data containing a large number of moving objects to obtain a panoramic view remains a difficult problem.

[0005] In related technologies, the FCN (Fully Convolutional Network) semantic segmentation network is generally used for semantic segmentation to extract foreground objects. However, objects with large position changes in the images captured by drones account for a relatively small proportion in the images, belonging to the small target segmentation task. The FCN semantic segmentation network is difficult to accurately segment small targets. Moreover, the ORB (Oriented FAST and Rotated BRIEF) algorithm is usually used to extract feature points of two images. However, the ORB algorithm is based on a detector for matching. In order to enhance affine invariance, general feature operators such as ORB have complex operations and large computational amounts, and the algorithm has poor robustness, and the registration effect for images with repeated textures is not ideal. And there are a large number of repeated textures such as roads in urban remote sensing images, and the current feature matching algorithms have poor effects on feature point matching.

[0006] In response to the above problems, no effective solution has been proposed yet. Summary of the Invention

[0007] An embodiment of the present invention provides a method and device for stitching remote sensing images, and an electronic device, so as to at least solve the technical problem in related technologies that due to the existence of repeated textures and moving objects in the collected images, the quality of the stitched images is poor.

[0008] According to one aspect of an embodiment of the present invention, a method for stitching remote sensing images is provided, including: collecting a remote sensing image video of a target area, and extracting multiple frames of remote sensing images from the remote sensing image video; processing each frame of remote sensing image with a preset semantic segmentation model to obtain a mask image corresponding to each frame of remote sensing image, where the model structure of the preset semantic segmentation model at least includes: multiple layers of hybrid dilated convolutional layers, and an attention module is connected after each layer of hybrid dilated convolutional layer; for each pair of mask images composed of two consecutive frames of mask images, extracting the features of each frame of mask image in the pair of mask images to obtain a feature map set of each frame of mask image, and based on the feature map set, using a preset feature matching algorithm to determine the pixel matching pairs of the pair of mask images, where the feature map set includes: feature maps of different resolutions; based on the pixel matching pairs, registering the pair of mask images, and fusing all the registered pairs of mask images to obtain a stitched image of the target area.

[0009] Further, before processing each frame of remote sensing image using a preset semantic segmentation model to obtain a mask map corresponding to each frame of remote sensing image, it further includes: determining a first preset number of target convolutional layer sets from the semantic segmentation model structure, where the semantic segmentation model structure includes: a second preset number of convolutional layers, and the first preset number is less than the second preset number; embedding dilated convolutions on each target convolutional layer in the target convolutional layer set to obtain a first preset number of hybrid dilated convolutional layers, where the dilation rates of the dilated convolutions embedded on each target convolutional layer are different; connecting an attention module after each hybrid dilated convolutional layer to obtain an initial semantic segmentation model.

[0010] Further, after obtaining the initial semantic segmentation model, it further includes: collecting a historical image set, where the historical image set includes: multiple positive sample images containing target objects and multiple negative sample images not containing target objects; annotating the target objects on each positive sample image to obtain annotation data; training the initial semantic segmentation model based on the historical image set and the annotation data until the loss value determined based on a preset loss function is less than a preset loss threshold to obtain a preset semantic segmentation model, where the preset loss function includes a preset penalty weight, and the preset penalty weight is used to assign weights to positive sample images; the loss value is determined by the preset loss function based on the predicted values output by the initial semantic segmentation model and the true values indicated by the annotation data.

[0011] Further, the step of processing each frame of remote sensing image using a preset semantic segmentation model to obtain a mask map corresponding to each frame of remote sensing image includes: extracting a feature representation vector of the remote sensing image using a convolutional structure and transmitting the feature representation vector to a hybrid dilated convolutional layer connected to the last convolutional layer, where the convolutional structure is a structure composed of multiple convolutional layers; using the hybrid dilated convolutional layer to insert holes in the convolutional kernel to process the received feature representation vector to obtain an expanded feature representation vector, and transmitting the expanded feature representation vector to an attention module connected after the hybrid dilated convolutional layer; using the attention module to divide the received expanded feature representation vector into N groups according to the channel dimension to obtain N groups of sub-feature representation vectors for each channel dimension, where N is a positive integer; generating a segmentation prediction map based on the N groups of sub-feature representation vectors for each channel dimension, where each pixel value on the segmentation prediction map represents the probability value corresponding to each pixel position; based on the segmentation prediction map, comparing each probability value with a preset probability threshold to obtain a mask map corresponding to the remote sensing image.

[0012] Further, the step of generating a segmentation prediction map based on N groups of sub-feature representation vectors for each channel dimension includes: for each channel dimension, converting each group of sub-feature representation vectors into a query matrix, a key matrix, and a value matrix; determining an attention weight matrix based on the query matrix and the key matrix, and multiplying the value matrix by the attention weight matrix to obtain a weighted feature map; generating an attention feature vector for the channel dimension based on all the weighted feature maps corresponding to the channel dimension; fusing the attention feature vectors for all channel dimensions to obtain a fused feature vector; and performing convolution processing on the fused feature vector to generate a segmentation prediction map.

[0013] Further, the feature map set at least includes: a first feature map, a second feature map, and a third feature map. The resolution of the first feature map is less than that of the second feature map, and the resolution of the second feature map is less than that of the third feature map. The step of determining pixel matching pairs of the mask map pair based on the feature map set using a preset feature matching algorithm includes: serializing the first feature map of each frame of the mask map pair in the mask map pair to obtain a first two-dimensional matrix; determining a first matching probability matrix of the first two-dimensional matrix, and determining a coarse-grained matching region of the mask map pair using a minimum bounding rectangle based on the first matching probability matrix, where the coarse-grained matching region includes: the coarse-grained regions on each frame of the mask map in the mask map pair; mapping all the coarse-grained matching regions to the second feature map of each frame of the mask map in the mask map pair to obtain a first mapped feature map, and serializing the first mapped feature map of each frame of the mask map in the mask map pair to obtain a second two-dimensional matrix; determining a second matching probability matrix of the second two-dimensional matrix, and determining a fine-grained matching region of the mask map pair using a minimum bounding rectangle based on the second matching probability matrix, where the fine-grained matching region includes: the fine-grained regions on each frame of the mask map in the mask map pair, and the fine-grained region is smaller than the coarse-grained region; mapping all the fine-grained matching regions to the third feature map of each frame of the mask map in the mask map pair to obtain a second mapped feature map; determining a third matching probability matrix based on the second mapped feature map of each frame of the mask map, and determining the position corresponding to the maximum value in the third matching probability matrix as the matching position; and determining the pixel matching pairs of the mask map pair based on the matching position.

[0014] Further, the step of registering the mask map pair based on the pixel matching pairs includes: determining a homography matrix of the mask map pair based on the pixel matching pairs; and performing a spatial transformation on any one of the mask maps in the mask map pair based on the homography matrix to complete the registration of the mask map pair.

[0015] Further, the step of fusing all the registered mask map pairs to obtain a stitched image of the target region includes: determining the motion mask information of each frame of the mask map; and fusing all the registered mask map pairs using a seam fusion algorithm based on the motion mask information to obtain a stitched image of the target region.

[0016] According to another aspect of the embodiments of the present invention, there is also provided a splicing device for remote sensing images, including: an acquisition unit, configured to acquire a remote sensing image video of a target area and extract multiple frames of remote sensing images from the remote sensing image video; a processing unit, configured to process each frame of remote sensing image by using a preset semantic segmentation model to obtain a mask image corresponding to each frame of remote sensing image, wherein the model structure of the preset semantic segmentation model at least includes: multiple layers of hybrid dilated convolutional layers, and an attention module is connected after each layer of hybrid dilated convolutional layer; a determination unit, configured to, for each pair of mask images composed of two consecutive frames of mask images in a group, extract the features of each frame of mask image in the pair of mask images to obtain a feature map set of each frame of mask image, and based on the feature map set, use a preset feature matching algorithm to determine the pixel matching pairs of the pair of mask images, wherein the feature map set includes: feature maps with different resolutions; a registration unit, configured to register the pair of mask images based on the pixel matching pairs and fuse all the registered pairs of mask images to obtain a spliced image of the target area.

[0017] Further, the splicing device further includes: a first determination module, configured to determine a set of first preset number of target convolutional layers from the semantic segmentation model structure before processing each frame of remote sensing image by using the preset semantic segmentation model to obtain a mask image corresponding to each frame of remote sensing image, wherein the semantic segmentation model structure includes: a second preset number of convolutional layers, and the first preset number is less than the second preset number; a first embedding module, configured to embed dilated convolutions on each layer of the target convolutional layers in the set of target convolutional layers to obtain a set of first preset number of hybrid dilated convolutional layers, wherein the dilation rates of the dilated convolutions embedded on each layer of the target convolutional layers are different; a first connection module, configured to connect an attention module after each layer of hybrid dilated convolutional layer to obtain an initial semantic segmentation model.

[0018] Further, the splicing device further includes: a first acquisition module, configured to acquire a set of historical images after obtaining the initial semantic segmentation model, wherein the set of historical images includes: multiple positive sample images containing a target object and multiple negative sample images not containing the target object; a first annotation module, configured to annotate the target object on each positive sample image to obtain annotation data; a first training module, configured to train the initial semantic segmentation model based on the set of historical images and the annotation data until the loss value determined based on a preset loss function is less than a preset loss threshold to obtain a preset semantic segmentation model, wherein the preset loss function includes a preset penalty weight, and the preset penalty weight is used to assign a weight to the positive sample images; the loss value is determined by the preset loss function based on the predicted value output by the initial semantic segmentation model and the true value indicated by the annotation data.

[0019] Further, the processing unit includes: a first extraction module, configured to extract a feature representation vector of a remote sensing image by using a convolutional structure and transmit the feature representation vector to a hybrid dilated convolutional layer connected to the last convolutional layer, where the convolutional structure is a structure composed of multiple convolutional layers; a first insertion module, configured to insert holes into the convolutional kernel by using the hybrid dilated convolutional layer to process the received feature representation vector to obtain an expanded feature representation vector and transmit the expanded feature representation vector to an attention module connected after the hybrid dilated convolutional layer; a first splitting module, configured to split the received expanded feature representation vector into N groups according to the channel dimension by using the attention module to obtain N groups of sub-feature representation vectors for each channel dimension, where N is a positive integer; a first generation module, configured to generate a segmentation prediction map based on the N groups of sub-feature representation vectors for each channel dimension, where each pixel value on the segmentation prediction map represents a probability value corresponding to each pixel position; a first comparison module, configured to compare each probability value with a preset probability threshold based on the segmentation prediction map to obtain a mask map corresponding to the remote sensing image.

[0020] Further, the first generation module includes: a first transformation sub-module, configured to transform each group of sub-feature representation vectors into a query matrix, a key matrix, and a value matrix for each channel dimension; a first determination sub-module, configured to determine an attention weight matrix based on the query matrix and the key matrix and multiply the value matrix by the attention weight matrix to obtain a weighted feature map; a first generation sub-module, configured to generate an attention feature vector for the channel dimension based on all the weighted feature maps corresponding to the channel dimension; a first fusion sub-module, configured to fuse the attention feature vectors for all channel dimensions to obtain a fused feature vector; a first processing sub-module, configured to perform convolutional processing on the fused feature vector to generate a segmentation prediction map.

[0021] Further, the set of feature maps includes at least: a first feature map, a second feature map, and a third feature map. The resolution of the first feature map is less than that of the second feature map, and the resolution of the second feature map is less than that of the third feature map. The determination unit includes: a first processing module for serializing the first feature map of each frame of the mask image pair to obtain a first two-dimensional matrix; a second determination module for determining a first matching probability matrix of the first two-dimensional matrix, and based on the first matching probability matrix, using the minimum bounding rectangle to determine the coarse-grained matching region of the mask image pair, where the coarse-grained matching region includes: the coarse-grained regions on each frame of the mask image pair; a first mapping module for mapping all the coarse-grained matching regions to the second feature map of each frame of the mask image pair to obtain a first mapped feature map, and serializing the first mapped feature map of each frame of the mask image pair to obtain a second two-dimensional matrix; a third determination module for determining a second matching probability matrix of the second two-dimensional matrix, and based on the second matching probability matrix, using the minimum bounding rectangle to determine the fine-grained matching region of the mask image pair, where the fine-grained matching region includes: the fine-grained regions on each frame of the mask image pair, and the fine-grained regions are smaller than the coarse-grained regions; a second mapping module for mapping all the fine-grained matching regions to the third feature map of each frame of the mask image pair to obtain a second mapped feature map; a fourth determination module for determining a third matching probability matrix based on the second mapped feature map of each frame of the mask image pair, and determining the position corresponding to the maximum value in the third matching probability matrix as the matching position; a fifth determination module for determining the pixel matching pairs of the mask image pair based on the matching position.

[0022] Further, the registration unit includes: a sixth determination module for determining the homography matrix of the mask image pair based on the pixel matching pairs; a first transformation module for performing a spatial transformation on any one of the mask images in the mask image pair based on the homography matrix to complete the registration of the mask image pair.

[0023] Further, the registration unit further includes: a seventh determination module for determining the motion mask information of each frame of the mask image; a first fusion module for fusing all the registered mask image pairs using the suture fusion algorithm based on the motion mask information to obtain the stitched image of the target region.

[0024] According to another aspect of the embodiments of the present invention, there is also provided a computer program product, including a non-volatile computer-readable storage medium storing a computer program, where the computer program, when executed by a processor, implements the remote sensing image stitching method of any one of the above.

[0025] According to another aspect of the embodiments of the present invention, an electronic device is further provided, including one or more processors and a memory. The memory is used to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the stitching method of remote sensing images as described in any one of the above.

[0026] In the present invention, a video of remote sensing images of a target area is collected, and multiple frames of remote sensing images are extracted from the video of remote sensing images. A preset semantic segmentation model is used to process each frame of remote sensing image to obtain a mask image corresponding to each frame of remote sensing image. For each pair of mask images composed of two consecutive frames of mask images in a group, features of each frame of mask image in the pair of mask images are extracted to obtain a set of feature maps of each frame of mask image. Based on the set of feature maps, a preset feature matching algorithm is used to determine pixel matching pairs of the pair of mask images. Based on the pixel matching pairs, the pair of mask images is registered, and all registered pairs of mask images are fused to obtain a stitched image of the target area, thereby solving the technical problem in the related art that due to the existence of repeated textures and moving objects in the collected images, the quality of the stitched image is poor.

[0027] In the present invention, through a semantic segmentation model that combines hybrid dilated convolution and attention mechanism, features in remote sensing images can be extracted more accurately, thereby improving the accuracy of the generated mask images. Moreover, using feature maps of different resolutions for feature matching can effectively improve the robustness and accuracy of the matching, thereby reducing errors in the registration and fusion processes and improving the quality of the stitched image. In this way, not only can static remote sensing images be processed, but also it can be applied to video streams containing small moving target objects to achieve high-precision stitching of dynamic scenes, achieving the technical effect of accurately stitching remote sensing images to obtain high-quality large-field-of-view images. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:

[0029] Figure 1 is a flowchart of an optional method for stitching remote sensing images according to an embodiment of the present invention;

[0030] Figure 2 is a schematic diagram of an optional stitching process of urban remote sensing images based on detector-free matching according to an embodiment of the present invention;

[0031] Figure 3 is a schematic diagram of an optional framework of the stitching process of urban remote sensing images according to an embodiment of the present invention;

[0032] Figure 4Schematic diagram of an optional remote sensing image stitching device according to an embodiment of the present invention;

[0033] Figure 5 Hardware structure block diagram of an electronic device (or mobile device) for a remote sensing image stitching method according to an embodiment of the present invention. Detailed implementation manners

[0034] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0035] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order different from those illustrated or described here. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0036] It should be noted that the relevant information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) collected and involved in the present invention are all information and data authorized by the user or fully authorized by all parties. And the collection, storage, use, processing, transmission, provision, disclosure and application and other processing of the relevant data all comply with the relevant laws, regulations and standards of the relevant regions, take necessary confidentiality measures, do not violate public order and good customs, and provide corresponding operation entrances for users to choose to authorize or refuse. For example, an interface is set between the present system and relevant users or institutions. Before obtaining relevant information, a request for obtaining needs to be sent to the aforementioned users or institutions through the interface, and after receiving the consent information feedback from the aforementioned users or institutions, the relevant information is obtained.

[0037] In order to solve the problems of the influence of moving objects and poor feature point matching effect during the process of remote sensing image mosaicking, the present invention proposes a method for urban remote sensing image mosaicking based on detector-free matching, which can cope with the challenges of moving objects in urban remote sensing image mosaicking problems.

[0038] In the present invention, a semantic segmentation algorithm suitable for small target segmentation is proposed to process images, so as to accurately identify and generate a mask representing moving vehicles. Moreover, a detector-free coarse-to-fine feature matching algorithm for image mosaicking is proposed, which can improve the matching accuracy of feature points, and combine the moving mask information to match image features, which can avoid the interference of moving objects on feature point matching.

[0039] The present invention will be described in detail below in conjunction with each embodiment.

[0040] Embodiment 1

[0041] According to an embodiment of the present invention, an embodiment of a method for mosaicking remote sensing images is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0042] Figure 1 is a flowchart of an optional method for mosaicking remote sensing images according to an embodiment of the present invention, as Figure 1 shown, the method includes the following steps:

[0043] Step S101, collect a remote sensing image video of a target area, and extract multiple frames of remote sensing images from the remote sensing image video.

[0044] In an embodiment of the present invention, remote sensing devices such as drones can be used to continuously shoot a target area (for example, a certain city) to form video data (i.e., a remote sensing image video). For example, if the target area is an urban block, a drone can be used to fly over the area and continuously shoot images of the block to form a video. Then, several image frames are extracted from the remote sensing image video as the input for image mosaicking. These frames are key frames in the video image sequence, which can be one frame extracted every certain number of frames, or frames automatically selected based on the degree of change in image content.

[0045] Step S102, process each frame of remote sensing image with a preset semantic segmentation model to obtain a mask map corresponding to each frame of remote sensing image, where the model structure of the preset semantic segmentation model at least includes: multiple layers of hybrid dilated convolutional layers, and an attention module is connected after each layer of hybrid dilated convolutional layer.

[0046] In the embodiments of the present invention, since the objects with relatively large position changes in the urban images captured by the drone account for a relatively small proportion in the images, which belong to the small target segmentation task, and the current semantic segmentation network (such as the FCN network) is difficult to accurately segment small targets. Therefore, in this embodiment, a semantic segmentation model suitable for small target segmentation (i.e., the preset semantic segmentation model) is constructed to process each frame of remote sensing image, which can accurately segment small targets and generate a mask representing moving objects (such as moving vehicles).

[0047] In the embodiments of the present invention, the preset semantic segmentation model can identify different objects or scene elements in the image and assign different labels to them. The structure of the preset semantic segmentation model at least includes: multiple layers of hybrid dilated convolutional layers, and each layer is connected with an attention module behind it. Among them, the hybrid dilated convolutional layer is a convolutional layer embedding dilated convolution, which can expand the receptive field by increasing the blank between convolutional kernels without increasing additional parameters. The hybrid dilated convolution refers to using different dilation rates in a single-layer convolution to capture feature information of different scales. The attention mechanism is used in deep learning to highlight important feature information while suppressing unimportant parts. In this embodiment, the attention module is designed to enhance the segmentation performance of small targets and improve the overall segmentation accuracy by analyzing and focusing on small target information.

[0048] In the embodiments of the present invention, the mask map corresponding to each frame of remote sensing image output by the preset semantic segmentation model is a binary black-and-white map, that is, an image with the same size as the input image but marked with the category to which each pixel belongs. For example, if a vehicle is correctly identified, the area corresponding to the vehicle in the mask map will be marked.

[0049] Step S103, for each pair of mask maps composed of two consecutive frames of mask maps in each group, extract the features of each frame of mask map in the pair of mask maps to obtain a set of feature maps for each frame of mask map, and based on the set of feature maps, use a preset feature matching algorithm to determine the pixel matching pairs of the pair of mask maps, where the set of feature maps includes: feature maps of different resolutions.

[0050] In the embodiments of the present invention, since the ORB algorithm needs to perform image feature matching based on a detector, the operation is relatively complex, and general feature operators such as ORB have a large amount of calculation in order to enhance affine invariance. Therefore, the ORB algorithm generally performs poorly in terms of matching speed and has poor effects on high-resolution images. So, in this embodiment, a detector-free coarse-to-fine feature matching algorithm (i.e., the preset feature matching algorithm) is proposed, which can improve the matching accuracy of feature points and combine the motion mask information to match the image features, and can avoid the interference of moving objects on the feature point matching.

[0051] In an embodiment of the present invention, a feature matching algorithm from coarse-grained to fine-grained without a detector can be used to perform feature matching on two consecutive frames of mask images. Specifically, two consecutive frames of mask images can be determined as a pair of mask images, and then feature extraction is performed on each frame of the mask images in each pair of mask images in turn to obtain a set of feature maps for each frame of the mask image. After that, a preset feature matching algorithm is used to process the feature maps with different resolutions in the set of feature maps to obtain pixel matching pairs of the pair of mask images.

[0052] In an embodiment of the present invention, ResNet-18 (i.e., a residual network with 18 layers of depth) can be used as the backbone network of the feature extraction model to extract the features of the mask image through the feature extraction model. In this embodiment, the feature extraction model can include feature extraction structures at different levels and can obtain feature maps containing rich information at different layers (for example, three feature maps with resolutions of 1 / 8, 1 / 4, and 1 / 2 of the original image).

[0053] In an embodiment of the present invention, the pixel matching pairs of the pair of mask images refer to the pixels corresponding to the same scene part in the two frames of mask images. By determining the matching pairs through the feature matching algorithm, it can be ensured that the two images can be aligned when spliced.

[0054] Step S104, based on the pixel matching pairs, register the pair of mask images, and fuse all the registered pairs of mask images to obtain a spliced image of the target area.

[0055] In an embodiment of the present invention, all the pixel matching pairs on the obtained pair of mask images can be used to register the two frames of mask images indicated by the pair of mask images, and then all the registered pairs of mask images are fused to generate a complete, continuous, non-misaligned and non-torn urban remote sensing spliced image, which can comprehensively reflect the geographical information and dynamic changes of the target area.

[0056] In summary, through the semantic segmentation model that combines hybrid dilated convolution and attention mechanism, the features in the remote sensing image can be extracted more accurately, thereby improving the accuracy of the generated mask image. Moreover, using the set of feature maps with different resolutions for feature matching can effectively improve the robustness and accuracy of the matching, thereby reducing errors in the registration and fusion processes and improving the quality of the spliced image. In this way, not only static remote sensing images can be processed, but also it can be applied to video streams containing small moving target objects to achieve high-precision splicing of dynamic scenes, achieving the technical effect of accurately splicing remote sensing images and obtaining high-quality large-field-of-view images, and further solving the technical problem in the related art that due to the existence of repeated textures and moving objects in the collected images, the quality of the spliced image is poor.

[0057] To accurately segment small target objects, a semantic segmentation model is constructed. In the method for stitching remote sensing images provided in the first embodiment of this application, a first preset number of target convolutional layer sets are determined from the semantic segmentation model structure, where the semantic segmentation model structure includes: a second preset number of convolutional layers, and the first preset number is less than the second preset number; at each target convolutional layer in the target convolutional layer set, a dilated convolution is embedded to obtain a first preset number of hybrid dilated convolutional layers, where the dilation rates of the dilated convolutions embedded at each target convolutional layer are different; an attention module is connected after each hybrid dilated convolutional layer to obtain an initial semantic segmentation model.

[0058] In the embodiment of the present invention, to achieve accurate semantic segmentation of small targets in urban remote sensing images, an optimized semantic segmentation model design method is proposed. This method can improve the performance of the model in small target segmentation tasks by embedding dilated convolutions and attention mechanisms. Specifically: First, a semantic segmentation model structure can be constructed, which includes a second preset number (e.g., 15 layers) of convolutional layers. For example, when constructing the model, a network structure including 15 convolutional layers is designed. Convolutional layers are basic modules in deep learning networks for extracting features from images. Then, a first preset number (e.g., 8 layers, and this first preset number is less than the second preset number) of target convolutional layer sets are determined from the semantic segmentation model structure. For example, in a model structure with 15 convolutional layers, the middle 8 convolutional layers are selected as the target convolutional layer set for subsequent embedding of dilated convolutions and connection of attention modules.

[0059] After that, dilated convolutions with different dilation rates (e.g., the dilation rates are set to r = 1, 2, 3, 5) can be embedded at each target convolutional layer in the target convolutional layer set to obtain a first preset number of hybrid dilated convolutional layers. For example, dilated convolutions with dilation rates set to 2, 3, and 5 can be used on different target convolutional layers respectively.

[0060] Here, dilated convolution is a special convolution technique. By adding "holes" (i.e., skip connections within the convolution kernel) to the convolution kernel, it can expand the receptive field of the model without increasing additional parameters, enabling the model to capture more extensive context information. Moreover, dilated convolutional layers with different dilation rates can extract features from different scales, which can further enhance the segmentation ability of the model.

[0061] After that, an attention module can be connected after each layer of the hybrid dilated convolutional layer to obtain an initial semantic segmentation model. Through the attention module, the focusing ability of the model in processing local features can be enhanced. In this embodiment, the attention module can divide the features output by each layer of the hybrid dilated convolutional layer into 64 groups along the channels, so as to dynamically allocate attention weights, enabling the model to pay more attention to the features of small targets when processing them, improving the accuracy of segmentation, and at the same time reducing the excessive attention to the background or large targets to avoid resource waste.

[0062] In the embodiment of the present invention, through the improvement of the semantic segmentation model structure, an initial semantic segmentation model capable of efficiently performing semantic segmentation of small targets in urban remote sensing images is obtained. This model can accurately identify and segment small targets such as moving vehicles.

[0063] In this embodiment, on a specific semantic segmentation model structure, by embedding dilated convolution and an attention mechanism, accurate segmentation of small targets in urban remote sensing images is achieved. The introduction of dilated convolution expands the receptive field of the model, enabling the model to more comprehensively understand the image content from different scales, especially those smaller or targets with large position changes. The addition of the attention module further improves the sensitivity of the model to the features of small targets, enabling the model to focus more on key information when processing these targets, thereby improving the accuracy of segmentation.

[0064] To improve the accuracy of the semantic segmentation model, in the method for stitching remote sensing images provided in the first embodiment of this application, a historical image set is collected, where the historical image set includes: multiple positive sample images containing target objects and multiple negative sample images not containing target objects; the target objects on each positive sample image are labeled to obtain labeled data; based on the historical image set and the labeled data, the initial semantic segmentation model is trained until the loss value determined based on a preset loss function is less than a preset loss threshold to obtain a preset semantic segmentation model, where the preset loss function includes a preset penalty weight, and the preset penalty weight is used to assign weights to the positive sample images; the loss value is determined by the preset loss function based on the predicted value output by the initial semantic segmentation model and the true value indicated by the labeled data.

[0065] In the embodiment of the present invention, a historical image set can be collected first. The historical image set includes multiple positive sample images containing target objects (such as vehicles) and multiple negative sample images not containing target objects. Each image can be an urban remote sensing image taken by a drone in a historical time period. Among them, a positive sample image refers to an image in which there are small targets to be segmented (such as moving vehicles), and a negative sample image refers to an image in which these small targets are not included.

[0066] In the embodiments of the present invention, the target objects on each positive sample image can be labeled to obtain labeled data. In this embodiment, the boundary of the target object can be delimited manually or using a labeling tool in the image, and a class label can be assigned to it to generate labeled data, so that the specific shape, size, and position of the small target can be guided for the model to learn based on the labeled data, so as to accurately identify and segment these small targets in subsequent predictions.

[0067] In the embodiments of the present invention, the initial semantic segmentation model can be trained based on the historical image set and the labeled data. During the training process, the model will learn the feature representations of small targets from the positive sample images, and at the same time learn the feature representations of background textures from the negative sample images. Through such training, the model can gradually form the ability to effectively identify and segment small targets. And, during the training process, a preset loss function (such as the cross-entropy loss function) can be used to evaluate the difference between the model prediction and the labeled data until the loss value determined based on the preset loss function is less than a preset loss threshold (which can be set according to the actual situation) to obtain the trained preset semantic segmentation model. Here, the loss value is determined by the preset loss function based on the predicted values (i.e., the predicted object positions and classes) output by the initial semantic segmentation model and the true values (i.e., the true positions and classes of the objects) indicated by the labeled data. The loss value gradually decreases as the training progresses until it is less than the preset loss threshold, at which point the model considers that it has achieved sufficient segmentation performance.

[0068] In the embodiments of the present invention, since the positive and negative samples in the collected historical image set are unbalanced, a preset penalty weight can be introduced into the preset loss function to assign a weight to the positive sample image. For example, the penalty weight value in the loss function is 2, so that the errors of the model in segmenting small targets will be given more attention, thereby prompting the model to show higher accuracy in small target segmentation.

[0069] In this embodiment, the trained preset semantic segmentation model can effectively and accurately identify and segment small targets in urban remote sensing images, such as moving vehicles. Even in the case where the target proportion is small and the texture is complex, it can maintain a high segmentation accuracy and stability. The optimized training process of the model, especially the introduction of a preset penalty weight in the loss function, ensures the robustness and accuracy of the model in dealing with small target segmentation tasks, provides a high-quality motion mask for subsequent image stitching, helps to avoid the influence of moving objects on the stitching process, and improves the overall quality of the stitched image.

[0070] In order to improve the accuracy of determining the mask map corresponding to the remote sensing image, in the method for stitching remote sensing images provided in Embodiment 1 of the present application, a convolutional structure is used to extract the feature representation vector of the remote sensing image, and the feature representation vector is transmitted to a hybrid dilated convolutional layer connected to the last convolutional layer, where the convolutional structure is a structure composed of multiple convolutional layers; the hybrid dilated convolutional layer inserts holes in the convolutional kernel to process the received feature representation vector, obtaining an expanded feature representation vector, and transmitting the expanded feature representation vector to an attention module connected after the hybrid dilated convolutional layer; the attention module divides the received expanded feature representation vector into N groups according to the channel dimension, obtaining N groups of sub-feature representation vectors for each channel dimension, where N is a positive integer; based on the N groups of sub-feature representation vectors for each channel dimension, a segmentation prediction map is generated, where each pixel value on the segmentation prediction map represents the probability value corresponding to each pixel position; based on the segmentation prediction map, each probability value is compared with a preset probability threshold to obtain the mask map corresponding to the remote sensing image.

[0071] In an embodiment of the present invention, after the remote sensing image is input into a preset semantic segmentation model, the convolutional structure (including multiple convolutional layers) of the preset semantic segmentation model can extract the feature representation vector of the remote sensing image and transmit the feature representation vector to a hybrid dilated convolutional layer connected to the last convolutional layer. Here, the convolutional structure is a basic component for image feature extraction in deep learning, composed of multiple convolutional layers. Each convolutional layer performs a mathematical transformation on the features of the input image through a convolutional kernel to extract specific feature information, such as edges, textures, and shapes. In this embodiment, first, the convolutional structure composed of multiple convolutional layers is used to extract features from each frame of the remote sensing image, obtaining the multi-dimensional feature representation vector of the image. Then, these feature vectors are transmitted to the hybrid dilated convolutional layer to further enhance the depth and breadth of the feature representation.

[0072] In an embodiment of the present invention, the hybrid dilated convolutional layer can insert holes in the convolutional kernel to process the received feature representation vector, obtaining an expanded feature representation vector, and transmitting the expanded feature representation vector to an attention module connected after the hybrid dilated convolutional layer. Here, dilated convolution is a special convolutional layer technology. By inserting "holes" (i.e., zero padding) in the convolutional kernel, the receptive field can be expanded without increasing additional parameters, enabling the network to capture a larger range of context information. The hybrid dilated convolutional layer uses dilated convolutions with different dilation rates on multiple convolutional layers to obtain feature information at different scales, thereby enhancing the model's ability to recognize small targets. After being processed by the hybrid dilated convolutional layer, the original feature representation vector is expanded, containing more scene details and semantic information. Then, these expanded feature representation vectors are transmitted to the attention module to further optimize feature extraction.

[0073] In the embodiment of the present invention, the attention module can divide the received dilated feature representation vector into N groups (N is a positive integer, such as 64 groups) according to the channel dimension (i.e., the three-channel dimension of RGB (Red, Green, Blue)), obtaining N groups of sub-feature representation vectors for each channel dimension. In this embodiment, the attention module can dynamically adjust the attention of the model to different parts during processing. The attention module can first decompose the dilated feature representation vector according to the channel dimension into N groups of sub-feature representation vectors. The selection of N (for example, 64) depends on the architecture of the model and the requirements of the task. This decomposition helps the model to focus more on those feature channels that may contain small target information, while reducing the processing of irrelevant features, improving the segmentation efficiency and accuracy.

[0074] In the embodiment of the present invention, by processing the N groups of sub-feature representation vectors corresponding to all channel dimensions, the probability value corresponding to each pixel position (represented by the pixel value on the segmentation prediction map) can be obtained, thereby generating a segmentation prediction map. The segmentation prediction map is the result of the model performing semantic segmentation on the image based on the learned deep features. The probability value of each pixel position reflects the possibility that the position belongs to a small target (such as a moving vehicle). By using the features optimized by the attention module as the input, the model can generate a prediction map with the same size as the input image, but with the probability of small targets marked at each pixel position.

[0075] In the embodiment of the present invention, based on the segmentation prediction map, by comparing each probability value with a preset probability threshold, the mask map corresponding to the remote sensing image can be obtained. Here, the mask map is the result of threshold processing on the segmentation prediction map and can be a binary image. Among them, 1 represents the pixel positions with high probability, that is, the parts that the model considers very likely to belong to small targets, while 0 represents the parts with low probability or background. By setting a reasonable probability threshold (for example, 0.5), the model compares the probability value of each pixel position on the segmentation prediction map, thereby obtaining a mask map that clearly highlights small targets.

[0076] In this embodiment, the combined use of the hybrid dilated convolution layer and the attention module not only enhances the model's ability to capture small target features, but also optimizes the efficiency of feature processing, enabling the model to effectively extract and highlight small targets, such as moving vehicles, from complex urban remote sensing images. The generated mask map not only accurately depicts the contours of small targets, but also provides precise moving object information for subsequent image stitching, reducing the stitching misalignment caused by moving objects and improving the coherence and visual effect of the stitched image.

[0077] To improve the accuracy of generating the segmentation prediction map, in the remote sensing image stitching method provided in the first embodiment of this application, for each channel dimension, each group of sub-feature representation vectors is transformed into a query matrix, a key matrix, and a value matrix; based on the query matrix and the key matrix, an attention weight matrix is determined, and the value matrix is multiplied by the attention weight matrix to obtain a weighted feature map; based on all the weighted feature maps corresponding to the channel dimension, an attention feature vector of the channel dimension is generated; the attention feature vectors of all channel dimensions are fused to obtain a fused feature vector; the fused feature vector is subjected to convolution processing to generate a segmentation prediction map.

[0078] In the embodiment of the present invention, each group of sub-feature representation vectors corresponding to each channel dimension can be transformed into a query matrix, a key matrix, and a value matrix. Here, the query matrix (Query Matrix), the key matrix (Key Matrix), and the value matrix (Value Matrix) respectively correspond to the "query", "reference", and "response" of the model when processing features. In this embodiment, for N groups of sub-feature representation vectors of each channel dimension, each group of vectors is transformed into these three matrices for subsequent attention weight calculation. This transformation converts multi-dimensional feature information into a form suitable for processing by the attention mechanism, preparing for subsequent weighting operations.

[0079] In the embodiment of the present invention, multiplying the query matrix and the key matrix (i.e., dot product) can obtain an attention weight matrix. Each element in the attention weight matrix represents the similarity or correlation between a query vector in the Query matrix and all vectors in the Key matrix. The similarity can be calculated by dot product, cosine similarity, or other similarity measurement methods. And after obtaining the attention weight matrix, the attention weight matrix can be transformed into a probability distribution through the Softmax function (i.e., the normalized exponential function). In this way, it can be ensured that each query vector at each position can assign a weight to each Key vector according to its similarity to all Key vectors, and the sum of these weights is 1. Then, multiplying the obtained attention weight matrix by the Value matrix and performing a weighted sum on each vector in the Value matrix can obtain a new feature vector (i.e., a weighted feature map), that is, based on the query vectors in the Query matrix, all vectors in the Value matrix are weighted and fused, thereby generating a feature representation that emphasizes important information more strongly.

[0080] In an embodiment of the present invention, an attention feature vector for the channel dimension can be generated based on all weighted feature maps corresponding to the channel dimension. That is, by aggregating all weighted feature maps, an attention feature vector can be generated for each channel dimension. This vector synthesizes the information of all weighted feature maps, further emphasizes the features of small targets, and provides a clearer and deeper feature representation for the model. Then, the attention feature vectors of all channel dimensions are fused to obtain a fused feature vector. Among them, the fusion process involves integrating the attention feature vectors obtained from all channel dimensions to form a more comprehensive and integrated feature representation. This fusion takes into account multi-scale and multi-dimensional feature information, which helps the model better understand and locate small targets and improve the accuracy of segmentation.

[0081] In an embodiment of the present invention, the fused feature vector is processed through a convolutional layer to generate a segmentation prediction map. The role of the convolutional layer is to extract more advanced features from the fused feature vector and convert these features into pixel-level predictions, that is, the segmentation prediction map. Each pixel value on this prediction map represents the probability that the pixel belongs to a small target (such as a moving vehicle). For example, the feature map represented by the fused feature vector can be restored to the same size as the input image through an upsampling layer, and then the processed feature map is passed through one or more convolutional layers to generate a segmentation prediction map. The last convolutional layer usually has a Softmax function to convert the probability distribution of each pixel point into a class label.

[0082] In some optional embodiments, after generating the segmentation prediction map, some post-processing steps can be performed, such as applying thresholds, morphological operations (erosion, dilation), etc., to further optimize and refine the segmentation result.

[0083] In this embodiment, the in-depth processing of the fused feature vector and the generation of the segmentation prediction map are realized, which can improve the model's ability to capture and express the features of small targets and make the segmentation prediction more accurate.

[0084] Optionally, the set of feature maps at least includes: a first feature map, a second feature map, and a third feature map. The resolution of the first feature map is less than that of the second feature map, and the resolution of the second feature map is less than that of the third feature map. In order to improve the accuracy of determining pixel matching pairs of the mask map pair, in the remote sensing image stitching method provided in the first embodiment of the present application, the first feature map of each frame of the mask map pair is serialized to obtain a first two-dimensional matrix; the first matching probability matrix of the first two-dimensional matrix is determined, and based on the first matching probability matrix, the minimum bounding rectangle is used to determine the coarse-grained matching region of the mask map pair, where the coarse-grained matching region includes: the coarse-grained regions on each frame of the mask map pair; all the coarse-grained matching regions are mapped onto the second feature map of each frame of the mask map pair to obtain a first mapped feature map, and the first mapped feature map of each frame of the mask map pair is serialized to obtain a second two-dimensional matrix; the second matching probability matrix of the second two-dimensional matrix is determined, and based on the second matching probability matrix, the minimum bounding rectangle is used to determine the fine-grained matching region of the mask map pair, where the fine-grained matching region includes: the fine-grained regions on each frame of the mask map pair, and the fine-grained regions are smaller than the coarse-grained regions; all the fine-grained matching regions are mapped onto the third feature map of each frame of the mask map pair to obtain a second mapped feature map; based on the second mapped feature map of each frame of the mask map, the third matching probability matrix is determined, and the position corresponding to the maximum value in the third matching probability matrix is determined as the matching position; based on the matching position, the pixel matching pairs of the mask map pair are determined.

[0085] In the embodiment of the present invention, feature maps with different resolutions can be extracted from the mask map (i.e., the first feature map (such as a feature map with a resolution of 1 / 8 of the original image), the second feature map (such as a feature map with a resolution of 1 / 4 of the original image), and the third feature map (such as a feature map with a resolution of 1 / 2 of the original image)), where the resolution of the first feature map is less than that of the second feature map, and the resolution of the second feature map is less than that of the third feature map). Then, for the first feature map, a coarse-grained field of view (i.e., a coarse-grained matching region) can be obtained through the matching probability matrix and the minimum bounding rectangle, and the model field of view is focused on the region where feature points are concentrated. Specifically:

[0086] The first feature map (i.e., one-dimensional sequence feature vector) of each frame of the mask image pair can be serialized to obtain a first two-dimensional matrix. The first feature map usually refers to the low-resolution feature map extracted from the image, which contains the global information of the image but not too many details. Serialization processing refers to converting the one-dimensional sequence feature vectors corresponding to two frames of mask images into a two-dimensional matrix for subsequent mathematical operations and matching probability calculations. The first two-dimensional matrix is the result of this serialization process, which presents the feature information in two frames of mask images in matrix form for subsequent matching probability calculations. Then, the similarity between the feature information indicated by each feature position in the first two-dimensional matrix is calculated, such as using dot product or cosine similarity, to obtain a matrix (i.e., the first matching probability matrix), and each element in this matching probability matrix represents the possibility of matching at the corresponding feature position. After that, by analyzing the first matching probability matrix, the method of the minimum bounding rectangle can be used to determine the coarse-grained matching regions in the mask image pair. These regions contain the regions where the feature points in a pair of mask images are relatively concentrated, providing the range and focus for subsequent fine-grained matching.

[0087] In the embodiment of the present invention, on the second feature map with higher resolution, the regions focused by the coarse-grained field of view are globally feature-serialized and input into the linear attention network to calculate the confidence matrix, and then global matching is obtained through threshold and mutual nearest neighbor screening, specifically:

[0088] All the coarse-grained matching regions obtained on the first feature map can be first mapped to the second feature map of each frame of the mask image pair to obtain a first mapped feature map, and the first mapped feature map of each frame of the mask image pair is serialized to obtain a second two-dimensional matrix. The second feature map is a feature map with higher resolution, which contains more detailed information. By mapping the coarse-grained matching regions to the second feature map, a first mapped feature map can be obtained, and this mapping process helps to focus on the more detailed feature information of potential matching points. Then, the first mapped feature map can be serialized to generate a second two-dimensional matrix, which contains more detailed information than the first two-dimensional matrix and provides a finer basis for fine-grained matching. After that, the similarity between the feature information indicated by each feature position in the second two-dimensional matrix is calculated to obtain a second matching probability matrix, which reflects the matching relationship between the second feature maps in the mask image pair. However, this matrix is calculated at a finer feature level, so it can provide more accurate matching information. Then, the minimum bounding rectangle is used to determine the fine-grained matching regions, which can further narrow the matching range, improve the matching accuracy, and ensure that accurate matching points can be found in the regions with dense details.

[0089] In the embodiment of the present invention, all the fine-grained matching regions obtained on the second feature map are mapped into the windows of the highest-resolution feature map (i.e., the third feature map) to obtain local matches. However, in the current detector-free feature matching algorithm, the local match only performs one-way calculation, that is, only calculates the similarity between the window center vector of the image I to be stitched and all the feature vectors in the corresponding window of the image I to be stitched, losing the information of the remaining vectors in the window of the image I to be stitched. Therefore, this embodiment introduces a positioning optimization strategy in feature matching to calculate the corresponding position of the maximum value of the matching probability matrix on the basis of global matching, that is, the sub-pixel level position, so as to map the obtained sub-pixel level matching position to the position on the original resolution of the two-frame mask images, and obtain the final sub-pixel level accurate matching relationship, specifically: a The window center vector and the image I to be stitched b The similarity between all the feature vectors in the corresponding window of the image I to be stitched, losing the information of the remaining vectors in the window of the image I to be stitched. Therefore, this embodiment introduces a positioning optimization strategy in feature matching to calculate the corresponding position of the maximum value of the matching probability matrix on the basis of global matching, that is, the sub-pixel level position, so as to map the obtained sub-pixel level matching position to the position on the original resolution of the two-frame mask images, and obtain the final sub-pixel level accurate matching relationship, specifically: a Therefore, this embodiment introduces a positioning optimization strategy in feature matching to calculate the corresponding position of the maximum value of the matching probability matrix on the basis of global matching, that is, the sub-pixel level position, so as to map the obtained sub-pixel level matching position to the position on the original resolution of the two-frame mask images, and obtain the final sub-pixel level accurate matching relationship, specifically:

[0090] Map all the fine-grained matching regions to the third feature map of each frame of the mask image pair to obtain the second mapped feature map. The third feature map is the feature map with the highest resolution, which contains the most detailed image information. By mapping the fine-grained matching regions to the third feature map, the second mapped feature map can be obtained, thus ensuring that the final match is performed at the finest feature level, improving the accuracy and robustness of the match. Then, based on the second mapped feature map of each frame of the mask image, determine the third matching probability matrix, and determine the position corresponding to the maximum value in the third matching probability matrix as the matching position. The third matching probability matrix is the result of calculating the matching probability at the most detailed feature level. By analyzing this matrix, the most likely matching position between each pair of feature points, that is, the position corresponding to the maximum matching probability value, can be found. After that, by converting the position corresponding to the maximum value in the third matching probability matrix back to the pixel coordinates, the accurate matching pairs of pixel points in the mask image pair (i.e., the pixel-level matching pairs) can be obtained.

[0091] In this embodiment, through the matching process from coarse-grained to fine-grained, the accuracy and robustness of feature matching are improved. From the global matching at low resolution, gradually focusing on the local matching at high resolution, the range of subsequent matching is determined by the method of the minimum bounding rectangle at each stage, ensuring the efficiency and accuracy of the matching process. The finally obtained pixel matching pairs provide high-quality matching results for the complex urban remote sensing image stitching task, overcome the challenges brought by moving objects and repetitive textures, and improve the overall quality and visual effect of the stitched image.

[0092] In order to improve the accuracy of registering the mask image pair, in the remote sensing image stitching method provided in the first embodiment of the present application, based on the pixel matching pairs, determine the homography matrix of the mask image pair; based on the homography matrix, perform a spatial transformation on any one of the mask images in the mask image pair to complete the registration of the mask image pair.

[0093] In an embodiment of the present invention, a homography matrix of a mask image pair can be determined based on pixel matching pairs. The homography matrix is used to describe the corresponding relationship of pixel points between two images due to camera position changes or rotations, and is the key to image registration and stitching. Specifically, by using pixel matching pairs, that is, the coordinates of corresponding pixel points found in the mask image pair, the MAGSAC++ algorithm (i.e., the improved geometric verification random field sampling consistency algorithm) or other robust estimation methods can be used to estimate the homography matrix. This matrix describes the mapping relationship of pixel points between two images, including translation, rotation, scaling, and perspective changes, etc. Then, based on the homography matrix, a spatial transformation is performed on any mask image in the mask image pair to complete the registration of the mask image pair.

[0094] In an embodiment of the present invention, spatial transformation refers to performing mathematical operations on an image using a homography matrix to achieve spatial alignment between two images. In this embodiment, any mask image in the mask image pair can be selected, and the above-estimated homography matrix is applied to perform a spatial transformation on this image in a reverse mapping or forward mapping manner, so that it is spatially aligned with another image. This transformation usually involves resampling of pixels to ensure that the corresponding points of the transformed image and the reference image are accurately matched at the pixel level. After registration, the corresponding feature points (such as feature points in the static background) in the two images will be at the same position, providing an accurate geometric alignment basis for the final image stitching.

[0095] In order to improve the quality of the stitched image, in the remote sensing image stitching method provided in Embodiment 1 of this application, the motion mask information of each frame of mask image is determined; based on the motion mask information, the stitched image of the target area is obtained by fusing all registered mask image pairs using the stitched line fusion algorithm.

[0096] In an embodiment of the present invention, the motion mask information of each frame of mask image can be determined first. The motion mask information refers to a binary mask generated by a small target semantic segmentation algorithm, which is used to distinguish moving objects (such as vehicles, pedestrians) from the static background in the image. These mask information clearly indicate which parts should be protected or adjusted to prevent motion artifacts or tearing phenomena during the stitching process. Then, based on the motion mask information, the stitched line fusion algorithm is used to fuse all registered mask image pairs to obtain the stitched image of the target area.

[0097] In an embodiment of the present invention, the seam fusion algorithm is an algorithm for image stitching. It fuses two or more images based on the seams at the image boundaries, ensuring a natural transition at the stitching location and avoiding obvious boundaries between images. In this embodiment, the seam fusion algorithm can be improved by incorporating motion mask information. Specifically, in the registered mask image pair, the regions of moving objects defined by the motion mask information can be first identified. Then, when performing image fusion, special fusion strategies can be adopted for these regions, such as using weighted average, bilinear interpolation, or content-based fusion methods, to reduce the impact of moving objects on the stitching result. For static background regions, more conventional seam fusion methods, such as direct pixel replication or simple weighted fusion, can be used.

[0098] In this embodiment, by accurately identifying and processing the regions of moving objects, stitching errors caused by moving objects, such as tearing, misalignment, or artifacts, can be effectively avoided, while maintaining high-quality stitching of static background regions. This fusion method not only considers the geometric accuracy of image registration but also the actual changes in image content, ensuring visual consistency and natural transition in the stitched image. The final stitched image of the target region not only has the characteristics of high resolution and wide field of view but also minimizes the negative impact of moving objects on the stitching result, providing a more accurate and reliable data basis for advanced applications of urban remote sensing images, such as intelligent traffic management, urban planning, and environmental monitoring.

[0099] A more detailed description will be given below in conjunction with another alternative specific implementation.

[0100] In an embodiment of the present invention, a method for stitching urban remote sensing images based on detector-free matching is proposed to address the challenge of moving objects in urban remote sensing image stitching.

[0101] In an embodiment of the present invention, the accurate motion mask generated by the small target semantic segmentation algorithm can reduce the interference of moving vehicles on feature matching, ensuring the accuracy and reliability of feature matching. Moreover, through the detector-free feature matching algorithm for urban remote sensing image stitching, the accuracy of feature matching is improved in the case of repeated textures in complex urban remote sensing images, thereby enhancing the stitching accuracy and giving it significant advantages in practical applications. Then, based on the results of the feature matching algorithm, the urban remote sensing images are spatially transformed through a homography matrix, mapping the points in the image to the corresponding points in another image with which it performs feature matching to complete image registration. Finally, based on the seam fusion algorithm and combining the obtained motion mask information, the final stitched image is obtained.

[0102] Figure 2 is a schematic diagram of an alternative detector-free matching-based urban remote sensing image stitching process according to an embodiment of the present invention, asFigure 2 As shown, a motion mask can be generated by semantic segmentation for small targets first, then, detector-free coarse-to-fine feature matching is performed, and then homography is used to estimate the spatial transformation to complete image registration. Finally, image fusion is performed based on the seam fusion algorithm to complete stitching.

[0103] Figure 3 It is a schematic diagram of an optional urban remote sensing image stitching process framework according to an embodiment of the present invention. As Figure 3 shown, for complex urban remote sensing images, a motion mask can be generated through small target semantic segmentation, cyclic hybrid dilated convolution, and morphological erosion and dilation post-processing. Then, detector-free matching is performed on two consecutive frames of images. Specifically, feature extraction can be performed on the images first, and then coarse-grained matching is performed on the extracted low-resolution feature maps, that is, global feature matching is performed through the matching probability matrix and the minimum bounding rectangle of the field of view, and then on higher-resolution feature maps, fine-grained matching is performed on the coarse-grained matching regions to achieve local bijective optimization and obtain pixel-level matching pairs. Then, based on the feature matching results (i.e., the obtained pixel-level matching pairs), the MAGSAC++ algorithm is used to estimate the homography matrix between the two frames of images to perform image registration on the two frames of images using the homography matrix, and then the seam fusion algorithm is used for image fusion to obtain the global aerial stitched image.

[0104] In this embodiment, by processing the image with a semantic segmentation algorithm for small target segmentation that embeds dilated convolution and lightweight attention mechanism, accurate segmentation of small targets can be achieved, the corresponding region information can be obtained, and a mask representing moving vehicles can be generated, which can avoid interference of moving vehicles on feature matching. And, through a detector-free coarse-to-fine feature matching algorithm, interleaved self-attention and cross-attention are used to obtain the focused area, and the image features are matched in combination with the motion mask information. Then, spatial transformation of the urban remote sensing image is performed through the homography matrix, and image registration can be accurately completed.

[0105] The following is a detailed description in combination with another embodiment.

[0106] Embodiment 2

[0107] A stitching device for remote sensing images provided in this embodiment includes multiple implementation units, and each implementation unit corresponds to each implementation step in Embodiment 1 above.

[0108] Figure 4 It is a schematic diagram of an optional stitching device for remote sensing images according to an embodiment of the present invention. As Figure 4 shown, the stitching device may include: a collection unit 40, a processing unit 41, a determination unit 42, and a registration unit 43.

[0109] Among them, the acquisition unit 40 is used to acquire the remote sensing image video of the target area and extract multiple frames of remote sensing images from the remote sensing image video;

[0110] The processing unit 41 is used to process each frame of the remote sensing image by using a preset semantic segmentation model to obtain a mask image corresponding to each frame of the remote sensing image. Among them, the model structure of the preset semantic segmentation model at least includes: multiple layers of hybrid dilated convolutional layers, and an attention module is connected after each layer of the hybrid dilated convolutional layer;

[0111] The determination unit 42 is used to extract the features of each frame of the mask image in the mask image pair composed of each group of two consecutive frames of mask images to obtain a feature map set of each frame of the mask image, and based on the feature map set, use a preset feature matching algorithm to determine the pixel matching pairs of the mask image pair. Among them, the feature map set includes: feature maps with different resolutions;

[0112] The registration unit 43 is used to register the mask image pair based on the pixel matching pairs and fuse all the registered mask image pairs to obtain a stitched image of the target area.

[0113] The above stitching device, through a semantic segmentation model that combines hybrid dilated convolution and an attention mechanism, can more accurately extract the features in the remote sensing image, thereby improving the accuracy of the generated mask image. Moreover, using the feature map set with different resolutions for feature matching can effectively improve the robustness and accuracy of the matching, thereby reducing errors in the registration and fusion processes and improving the quality of the stitched image. In this way, it can not only process static remote sensing images but also be applied to video streams containing small moving target objects, realizing high-precision stitching of dynamic scenes, achieving the technical effect of accurately stitching remote sensing images to obtain high-quality large-field images, and further solving the technical problem in the related art that the quality of the stitched image is poor due to the presence of repeated textures and moving objects in the acquired images.

[0114] Optionally, the stitched device further includes: a first determination module, which is used to determine a first preset number of target convolutional layer sets from the semantic segmentation model structure before processing each frame of the remote sensing image by using the preset semantic segmentation model to obtain a mask image corresponding to each frame of the remote sensing image. Among them, the semantic segmentation model structure includes: a second preset number of convolutional layers, and the first preset number is less than the second preset number; a first embedding module, which is used to embed dilated convolutions on each layer of the target convolutional layer set to obtain a first preset number of hybrid dilated convolutional layers, where the dilation rates of the dilated convolutions embedded on each layer of the target convolutional layer are different; a first connection module, which is used to connect an attention module after each layer of the hybrid dilated convolutional layer to obtain an initial semantic segmentation model.

[0115] Optionally, the splicing device further includes: a first acquisition module, configured to acquire a historical image set after obtaining an initial semantic segmentation model, where the historical image set includes: multiple positive sample images containing a target object and multiple negative sample images not containing the target object; a first annotation module, configured to annotate the target object on each positive sample image to obtain annotation data; a first training module, configured to train the initial semantic segmentation model based on the historical image set and the annotation data until the loss value determined based on a preset loss function is less than a preset loss threshold, to obtain a preset semantic segmentation model, where the preset loss function includes a preset penalty weight, and the preset penalty weight is used to assign a weight to the positive sample image; the loss value is determined by the preset loss function based on the predicted value output by the initial semantic segmentation model and the true value indicated by the annotation data.

[0116] Optionally, the processing unit includes: a first extraction module, configured to extract a feature representation vector of a remote sensing image by using a convolutional structure and transmit the feature representation vector to a hybrid dilated convolutional layer connected to the last convolutional layer, where the convolutional structure is a structure composed of multiple convolutional layers; a first insertion module, configured to insert holes in the convolution kernel by using the hybrid dilated convolutional layer to process the received feature representation vector to obtain an expanded feature representation vector and transmit the expanded feature representation vector to an attention module connected after the hybrid dilated convolutional layer; a first division module, configured to divide the received expanded feature representation vector into N groups according to the channel dimension by using the attention module to obtain N groups of sub-feature representation vectors for each channel dimension, where N is a positive integer; a first generation module, configured to generate a segmentation prediction map based on the N groups of sub-feature representation vectors for each channel dimension, where each pixel value on the segmentation prediction map represents a probability value corresponding to each pixel position; a first comparison module, configured to compare each probability value with a preset probability threshold based on the segmentation prediction map to obtain a mask map corresponding to the remote sensing image.

[0117] Optionally, the first generation module includes: a first conversion sub-module, configured to convert each group of sub-feature representation vectors into a query matrix, a key matrix, and a value matrix for each channel dimension; a first determination sub-module, configured to determine an attention weight matrix based on the query matrix and the key matrix and multiply the value matrix by the attention weight matrix to obtain a weighted feature map; a first generation sub-module, configured to generate an attention feature vector for the channel dimension based on all the weighted feature maps corresponding to the channel dimension; a first fusion sub-module, configured to fuse the attention feature vectors for all channel dimensions to obtain a fused feature vector; a first processing sub-module, configured to perform convolutional processing on the fused feature vector to generate a segmentation prediction map.

[0118] Optionally, the set of feature maps includes at least: a first feature map, a second feature map, and a third feature map. The resolution of the first feature map is less than that of the second feature map, and the resolution of the second feature map is less than that of the third feature map. The determination unit includes: a first processing module for serializing the first feature map of each frame of the mask image pair to obtain a first two-dimensional matrix; a second determination module for determining a first matching probability matrix of the first two-dimensional matrix and, based on the first matching probability matrix, using a minimum bounding rectangle to determine a coarse-grained matching region of the mask image pair, where the coarse-grained matching region includes: a coarse-grained region on each frame of the mask image pair; a first mapping module for mapping all the coarse-grained matching regions to the second feature map of each frame of the mask image pair to obtain a first mapped feature map and serializing the first mapped feature map of each frame of the mask image pair to obtain a second two-dimensional matrix; a third determination module for determining a second matching probability matrix of the second two-dimensional matrix and, based on the second matching probability matrix, using a minimum bounding rectangle to determine a fine-grained matching region of the mask image pair, where the fine-grained matching region includes: a fine-grained region on each frame of the mask image pair, and the fine-grained region is smaller than the coarse-grained region; a second mapping module for mapping all the fine-grained matching regions to the third feature map of each frame of the mask image pair to obtain a second mapped feature map; a fourth determination module for determining a third matching probability matrix based on the second mapped feature map of each frame of the mask image pair and determining the position corresponding to the maximum value in the third matching probability matrix as the matching position; a fifth determination module for determining a pixel matching pair of the mask image pair based on the matching position.

[0119] Optionally, the registration unit includes: a sixth determination module for determining a homography matrix of the mask image pair based on the pixel matching pair; a first transformation module for performing a spatial transformation on any one of the mask images in the mask image pair based on the homography matrix to complete the registration of the mask image pair.

[0120] Optionally, the registration unit further includes: a seventh determination module for determining the motion mask information of each frame of the mask image; a first fusion module for fusing all the registered mask image pairs using a stitching line fusion algorithm based on the motion mask information to obtain a stitched image of the target region.

[0121] The above-mentioned stitching device may further include a processor and a memory. The above-mentioned acquisition unit 40, processing unit 41, determination unit 42, registration unit 43, etc. are all stored in the memory as program units, and the processor executes the above program units stored in the memory to implement corresponding functions.

[0122] The above-mentioned processor includes a kernel, which retrieves corresponding program units from the memory. One or more kernels can be set. By adjusting the kernel parameters, the mask image pairs are registered based on pixel matching pairs, and all registered mask image pairs are fused to obtain the mosaic image of the target area.

[0123] The above-mentioned memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash RAM (flash RAM). The memory includes at least one memory chip.

[0124] The present invention also provides a computer program product, which is adapted to execute a program initialized with the following method steps when executed on a data processing device: collecting a remote sensing image video of a target area, extracting multiple frames of remote sensing images from the remote sensing image video, processing each frame of the remote sensing image with a preset semantic segmentation model to obtain a mask image corresponding to each frame of the remote sensing image. For each pair of consecutive frames of mask images, extracting the features of each frame of the mask image in the mask image pair to obtain a feature map set of each frame of the mask image, and based on the feature map set, using a preset feature matching algorithm to determine the pixel matching pairs of the mask image pair. Based on the pixel matching pairs, the mask image pair is registered, and all registered mask image pairs are fused to obtain the mosaic image of the target area.

[0125] According to another aspect of the embodiments of the present invention, there is also provided a computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the method for mosaicking remote sensing images in any one of the above.

[0126] According to another aspect of the embodiments of the present invention, there is also provided an electronic device, including one or more processors and a memory. The memory is used to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors implement the above-mentioned method for mosaicking remote sensing images.

[0127] Figure 5 It is a hardware structure block diagram of an electronic device (or mobile device) for the method of mosaicking remote sensing images according to an embodiment of the present invention. As Figure 5 shown, the electronic device may include one or more processors (for example, Figure 5The processors 502a, 502b, ……, 502n, etc. in it, these processors may include but are not limited to processing devices such as microprocessor MCUs or programmable logic devices FPGAs), a memory 504 for storing data. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be included as one of the ports of the I / O interface), a network interface, a keyboard, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 5 The structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the electronic device may further include more or fewer components than Figure 5 shown in it, or have a different configuration from Figure 5 shown in it.

[0128] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.

[0129] The embodiments or examples of the present disclosure are not exhaustive. They are only schematic of some embodiments or examples and do not serve as specific limitations on the protection scope of the present disclosure. Without contradiction, each step in a certain embodiment or example can be implemented as an independent embodiment, and the steps can be combined arbitrarily. For example, the solution after removing some steps in a certain embodiment or example can also be implemented as an independent embodiment, and the order of the steps in a certain embodiment or example can be exchanged arbitrarily. In addition, the optional ways or optional examples in a certain embodiment or example can be combined arbitrarily; furthermore, the embodiments or examples can be combined arbitrarily. For example, some or all of the steps of different embodiments or examples can be combined arbitrarily, and a certain embodiment or example can be combined arbitrarily with the optional ways or optional examples of other embodiments or examples.

[0130] In the above embodiments of the present invention, the descriptions of the respective embodiments have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0131] In the several embodiments provided by the present invention, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces. The indirect couplings or communication connections of units or modules can be in electrical or other forms.

[0132] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed over multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0133] In addition, each functional unit in various embodiments of the present invention may be integrated into a processing unit, may exist physically alone for each unit, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0134] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs, etc., which can store program codes.

[0135] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

Claims

1. A remote sensing image stitching method, characterized in that: include: Collecting remote sensing image videos of the target area, and extracting multiple frames of remote sensing images from the remote sensing image videos; Each frame of the remote sensing image is processed using a preset semantic segmentation model to obtain a mask image corresponding to each frame of the remote sensing image, wherein the model structure of the preset semantic segmentation model at least includes: multiple layers of mixed void convolution layers, and each layer of the mixed void convolution layer is connected to an attention module; For each mask image pair consisting of two consecutive frames of the mask images, extract the features of the mask image of each frame in the mask image pair to obtain a feature image set of the mask image of each frame, and based on the feature image set, use a preset feature matching algorithm to determine a pixel matching pair of the mask image pair, wherein the feature image set includes: feature images of different resolutions; Based on the pixel matching pairs, the mask image pairs are registered, and all the registered mask image pairs are fused to obtain a spliced ​​image of the target area.

2. The splicing method according to claim 1, characterized in that: Before using the preset semantic segmentation model to process each frame of the remote sensing image to obtain a mask image corresponding to each frame of the remote sensing image, the method further includes: Determining a first preset number of target convolutional layer sets from a semantic segmentation model structure, wherein the semantic segmentation model structure includes: a second preset number of convolutional layers, the first preset number being less than the second preset number; Embedding a dilated convolution on each target convolution layer in the target convolution layer set to obtain the first preset number of mixed dilated convolution layers, wherein the dilated convolution rates of the dilated convolutions embedded on each target convolution layer are different; The attention module is connected after each layer of the mixed hole convolution layer to obtain an initial semantic segmentation model.

3. The splicing method according to claim 2, characterized in that: After obtaining the initial semantic segmentation model, it also includes: Collecting a historical image set, wherein the historical image set includes: a plurality of positive sample images containing a target object and a plurality of negative sample images not containing the target object; Annotating the target object on each positive sample image to obtain annotation data; Based on the historical image set and the annotated data, the initial semantic segmentation model is trained until the loss value determined based on a preset loss function is less than a preset loss threshold, so as to obtain the preset semantic segmentation model, wherein the preset loss function includes a preset penalty weight, and the preset penalty weight is used to assign a weight to the positive sample image; the loss value is determined by the preset loss function based on the predicted value output by the initial semantic segmentation model and the true value indicated by the annotated data.

4. The splicing method according to claim 2, characterized in that: The step of processing each frame of the remote sensing image using a preset semantic segmentation model to obtain a mask image corresponding to each frame of the remote sensing image includes: A convolution structure is used to extract a feature representation vector of the remote sensing image, and the feature representation vector is transmitted to the mixed hole convolution layer connected to the last convolution layer, wherein the convolution structure is a structure composed of multiple convolution layers; Using the hybrid atrous convolution layer to insert holes in the convolution kernel to process the received feature representation vector to obtain an expanded feature representation vector, and transmitting the expanded feature representation vector to the attention module connected to the hybrid atrous convolution layer; Using the attention module to divide the received expanded feature representation vector into N groups according to the channel dimension, to obtain N groups of sub-feature representation vectors of each channel dimension, where N is a positive integer; Generate a segmentation prediction map based on the N groups of sub-feature representation vectors of each channel dimension, wherein each pixel value on the segmentation prediction map represents a probability value corresponding to each pixel position; Based on the segmentation prediction map, each of the probability values ​​is compared with a preset probability threshold to obtain the mask map corresponding to the remote sensing image.

5. The splicing method according to claim 4, characterized in that: The step of generating a segmentation prediction map based on the N groups of sub-feature representation vectors of each channel dimension comprises: For each of the channel dimensions, converting each group of the sub-feature representation vectors into a query matrix, a key matrix, and a value matrix; Determine an attention weight matrix based on the query matrix and the key matrix, and multiply the value matrix by the attention weight matrix to obtain a weighted feature map; Based on all the weighted feature maps corresponding to the channel dimension, generating an attention feature vector of the channel dimension; Fusing the attention feature vectors of all the channel dimensions to obtain a fused feature vector; The fused feature vector is convolved to generate the segmentation prediction map.

6. The splicing method according to claim 1, characterized in that: The feature map set includes at least: a first feature map, a second feature map, and a third feature map, the resolution of the first feature map is smaller than the resolution of the second feature map, and the resolution of the second feature map is smaller than the resolution of the third feature map. Based on the feature map set, a preset feature matching algorithm is used to determine the pixel matching pairs of the mask image pair, including: Performing serialization processing on the first feature image of each frame of the mask image in the mask image pair to obtain a first two-dimensional matrix; Determine a first matching probability matrix of the first two-dimensional matrix, and based on the first matching probability matrix, determine a coarse-grained matching area of ​​the mask image pair using a minimum circumscribed rectangle, wherein the coarse-grained matching area includes: a coarse-grained area on the mask image of each frame in the mask image pair; Mapping all the coarse-grained matching areas onto the second feature map of each frame of the mask image in the mask image pair to obtain a first mapping feature map, and performing serialization processing on the first mapping feature map of each frame of the mask image in the mask image pair to obtain a second two-dimensional matrix; Determine a second matching probability matrix of the second two-dimensional matrix, and determine a fine-grained matching region of the mask image pair using a minimum circumscribed rectangle based on the second matching probability matrix, wherein the fine-grained matching region includes: a fine-grained region on the mask image of each frame in the mask image pair, and the fine-grained region is smaller than the coarse-grained region; Mapping all the fine-grained matching areas onto the third feature map of each frame of the mask image in the mask image pair to obtain a second mapping feature map; Determine a third matching probability matrix based on the second mapping feature map of each frame of the mask image, and determine a position corresponding to a maximum value in the third matching probability matrix as a matching position; Based on the matching positions, the matching pairs of pixels of the mask image pair are determined.

7. The splicing method according to claim 1, characterized in that: The step of registering the mask image pair based on the pixel matching pair comprises: Determining a homography matrix of the mask image pair based on the pixel matching pairs; Based on the homography matrix, spatial transformation is performed on any of the mask images in the mask image pair to complete the registration of the mask image pair.

8. The splicing method according to claim 1, characterized in that: The step of fusing all the registered mask image pairs to obtain a spliced ​​image of the target area includes: Determining motion mask information of the mask image of each frame; Based on the motion mask information, a seam fusion algorithm is used to fuse all the registered mask image pairs to obtain the spliced ​​image of the target area.

9. A remote sensing image stitching device, characterized in that: include: An acquisition unit, used for acquiring remote sensing image videos of a target area and extracting multiple frames of remote sensing images from the remote sensing image videos; A processing unit, used to process each frame of the remote sensing image using a preset semantic segmentation model to obtain a mask image corresponding to each frame of the remote sensing image, wherein the model structure of the preset semantic segmentation model at least includes: multiple layers of mixed void convolution layers, each layer of the mixed void convolution layer is connected to an attention module; A determination unit is used to extract features of the mask images of each frame in the mask image pair for each group of two consecutive frames of the mask images, obtain a feature map set of the mask images of each frame, and determine pixel matching pairs of the mask image pair based on the feature map set and using a preset feature matching algorithm, wherein the feature map set includes: feature maps of different resolutions; A registration unit is used to register the mask image pairs based on the pixel matching pairs, and fuse all the registered mask image pairs to obtain a spliced ​​image of the target area.

10. An electronic device, characterized in that: It includes one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the remote sensing image stitching method described in any one of claims 1 to 8.

Citation Information

Cited By

  • Dynamic image sequence splicing optimization method, system, equipment, medium and product

    CN122115237A