Unsupervised low-altitude parallax image stitching method based on linear-nonlinear control grid
Patent Information
- Application Number
- CN202610974142.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-01
- Publication Date
- 2026-09-22
AI Technical Summary
[0003]目前,现有方法均在常规视角下获取图像,存在丰富且适量纹理特征信息,但针对于低空无人机视角,纹理特征信息稀疏且较弱,现有方法难以生成高质量的拼接图像
[0020]有益效果:本发明提供了一种基于线性-非线性控制网格的无监督低空视差图像拼接方法,通过设计的无监督低空视差图像拼接损失函数,对构建的基于线性-非线性控制网格的低空视差图像拼接模型进行训练,获取最优低空视差图像拼接,以实现基于线性-非线性控制网格的无监督低空视差图像拼接。本发明提出了无监督图像拼接框架,即在无真值标签的情况下通过最优代价评估的方式,实现对待拼接图像的拼接效果的评估与优化,提升了场景鲁棒性。通过线性偏移预测模块与非线性偏移预测模块以及图像对齐模块,实现网格偏移预测与对齐,具有更高精度的图像对齐结果。基于构建无监督低空视差图像拼接损失函数,通过自适应宽度区域融合提供了更加平滑的融合效果,结合了像素集分配的代价评估器在获得最优代价评估的前提下,保证了感知目标的完整性。
Smart Images

Figure CN122798620A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of low-altitude intelligent and sensing detection technology, and in particular to an unsupervised low-altitude parallax image stitching method based on linear-nonlinear control grids. Background Technology
[0002] In the fields of UAVs and low-altitude vision, numerous visual perception solutions exist. However, limitations in equipment size and cost prevent the installation of larger-scale imaging devices. Therefore, high-quality, wide-field-of-view image stitching methods are needed to compensate for imaging limitations and improve the perception range. Existing image stitching methods often rely on manual feature extraction, thus requiring a high degree of texture richness in the image. In contrast, deep learning-based image stitching methods often depend on high-quality training sets, and excellent stitching results are built upon rich and clearly labeled datasets.
[0003] Currently, existing methods acquire images from conventional perspectives, which possess rich and adequate texture features. However, from the perspective of low-altitude drones, texture features are sparse and weak, making it difficult for existing methods to generate high-quality stitched images. Traditional methods not only have high requirements for texture quality but also suffer from excessive computational complexity, making real-time performance impossible with the computing power of conventional drone equipment. Furthermore, the difficulty in obtaining ground truth data often limits image stitching datasets for supervised learning, resulting in poor training performance in unsupervised learning and even lower stitching quality in some scenarios. Summary of the Invention
[0004] This invention provides an unsupervised low-altitude parallax image stitching method based on linear-nonlinear control grids to overcome the aforementioned technical problems.
[0005] To achieve the above objectives, the technical solution of the present invention is as follows: An unsupervised low-altitude parallax image stitching method based on linear-nonlinear control grids specifically includes the following steps: S1: Obtain the image to be stitched from a low-altitude perspective; the image to be stitched includes a reference image and a target image; S2: Construct a low-altitude parallax image stitching model based on linear-nonlinear control grid; The model includes an input layer, a feature representation module built with ResNet50, a linear offset prediction module, a nonlinear offset prediction module, an image alignment module, and an artifact removal fusion network based on optimal cost evaluation. The input layer is used to input the images to be stitched into the feature representation module; The feature representation module is used to extract semantic feature information of the target image in the image to be stitched and obtain a semantic feature map; and the semantic feature information includes at least texture and contour information; and the semantic feature map includes a first-scale feature map and a second-scale feature map with different scales. The linear offset prediction module is used to perform a linear transformation operation on the first-scale feature map to obtain the linear homography transformation matrix; The nonlinear offset prediction module is used to interpolate the second-scale feature map based on the radial basis function interpolation algorithm to obtain a nonlinear mesh deformation field; and to obtain a coupled deformation field based on the linear homography transformation matrix and the nonlinear mesh deformation field; the coupled deformation field is used to perform coupled deformation on the target image to obtain an optimized target image; The image alignment module is used to obtain the misaligned and overlapping regions corresponding to the target image and the reference image. An artifact elimination fusion network based on optimal cost evaluation is used to perform disparity bridging and optimal cost fusion processing on the misaligned overlapping regions to obtain the final stitched image. S3: With the goal of minimizing the constructed unsupervised low-altitude parallax image stitching loss function, the constructed low-altitude parallax image stitching model is trained using the image to be stitched to obtain the optimal low-altitude parallax image stitching model, so as to realize the unsupervised low-altitude parallax image stitching process based on linear-nonlinear control grid.
[0006] Furthermore, the linear homography transformation matrix obtained in S2 is:
[0007] In the formula: Represents the linear homography transformation matrix; This represents the coordinates of the k-th pixel in the reference image; This represents the coordinates of the corresponding pixel in the target image; This represents the intermediate result of the linear homography transformation matrix regression process, and the intermediate result includes a normalized point set and a homogeneous coordinate representation.
[0008] Furthermore, the method for obtaining the coupled deformation field in S2 is as follows: Define the grid corresponding to the second-scale feature map as follows ; Based on the radial basis function interpolation algorithm, according to the grid Interpolation is performed on the second-scale feature map to obtain the nonlinear mesh deformation field:
[0009] In the formula: Represents the coordinates of the grid control points in the nonlinear grid deformation field; Indicates the interpolation weights of the radial basis functions; Represents the Gaussian function; This represents the coordinates of the pixel point to be interpolated in the second-scale feature map. Represents a grid The Middle The position coordinates of each grid control point; The coupled deformation field is obtained from the linear homography transformation matrix and the nonlinear mesh deformation field as follows:
[0010] In the formula: This indicates that for any point in the target image The coupled deformation field; Represents any point in the target image The processing results of the corresponding nonlinear mesh deformation field.
[0011] Furthermore, the artifact elimination fusion network based on optimal cost evaluation in S2 includes a cost estimator and a fusion decoder; The cost estimator is used to construct an initial cost map based on pixel color differences in the misaligned overlapping areas of the aligned image, and the formula for constructing the initial cost map is as follows:
[0012]
[0013]
[0014] In the formula: This represents the Manhattan distance between pixels in the overlapping region; This represents the pixels in the misaligned and overlapping region that correspond to the optimized target image. This represents the pixel in the misaligned and overlapping area that corresponds to the reference image. Indicates pixel channel; This represents the result of normalizing the Manhattan distance of the pixels; This represents the minimum Manhattan distance between pixels in the overlapping region. This represents the maximum Manhattan distance between pixels in the overlapping region; Indicates the inverse value after normalization; Based on the initial cost map, the cost redistribution of pixels within the target image and the reference image in the misaligned and overlapping regions is optimized, specifically as follows: Obtain a pixel within the target image and the reference image for optimization. The color difference between pixels spatially adjacent to it is determined, and the magnitude of the color difference is compared with a preset difference threshold. Then, it is compared with a certain pixel... Pixels that are spatially adjacent and whose color difference is less than a preset difference threshold are grouped into a preset set of pixels representing the same perceptual target. And improve pixel points based on empirical values. It is adjacent to and belongs to the same pixel set The value between pixels; conversely, confirming that they do not belong to the same set of perceived target pixels. And reduce pixel count based on empirical values. It is adjacent to and belongs to the same pixel set The cost value between pixels; and the cost value is the color difference value between pixels; The fusion decoder is used to perform dimensionality-upgrading fusion of the semantic feature map obtained by the feature representation module based on a preset inverse pyramid structure to obtain a fused semantic feature map with the same resolution as the original input reference image, so as to obtain the splicing overlap region features of the fused semantic feature map and the target image; and to confirm the fusion boundary region by combining the cost value reallocated by the cost estimator. The fusion boundary region is the seam region in the splicing overlap region features corresponding to the cost value being lower than the preset cost threshold; the gradient change of the cost map of the fusion boundary region is obtained by the cost estimator during training to dynamically adjust the width of the fusion boundary region, so as to achieve low-altitude parallax image splicing.
[0015] Furthermore, the expression for the unsupervised low-altitude parallax image stitching loss function constructed in S3 is as follows:
[0016]
[0017]
[0018]
[0019] In the formula: This represents the strong alignment loss of the image; , These represent the reference image and the target image to be stitched together, respectively. This represents the homography matrix obtained through the linear offset prediction module. The inverse matrix; Indicates the learning parameters; Represents the target image Optimized target image obtained based on coupled deformation location; The cost-optimal constraint function represents the evaluation of the optimality of the fusion region selection. Indicates correspondence The merging boundary region of pixel positions; Indicates in The luminance difference between the reference image and the target image at a pixel location is the RGB Euclidean distance between pixels. Indicates the learning parameters; Indicates correspondence The merging boundary region of pixel positions; Indicates in The luminance difference between the reference image and the target image at a pixel location is the RGB Euclidean distance between pixels. Indicates in The luminance difference between the reference image and the target image at a pixel location is the RGB Euclidean distance between pixels. Indicates correspondence The merging boundary region of pixel positions; This indicates the smoothness loss used to assess the merging boundary region; This represents the brightness gradient between adjacent pixels in the stitched image; Indicates the learning parameters; This represents the loss function for unsupervised low-altitude parallax image stitching.
[0020] Beneficial Effects: This invention provides an unsupervised low-altitude parallax image stitching method based on linear-nonlinear control grids. By using a designed unsupervised low-altitude parallax image stitching loss function, a constructed low-altitude parallax image stitching model based on linear-nonlinear control grids is trained to obtain the optimal low-altitude parallax image stitching, thus achieving unsupervised low-altitude parallax image stitching based on linear-nonlinear control grids. This invention proposes an unsupervised image stitching framework, which evaluates and optimizes the stitching effect of the images to be stitched through optimal cost evaluation in the absence of ground truth labels, improving scene robustness. Through linear offset prediction, nonlinear offset prediction, and image alignment modules, grid offset prediction and alignment are achieved, resulting in higher-precision image alignment results. Based on the constructed unsupervised low-altitude parallax image stitching loss function, adaptive width region fusion provides a smoother fusion effect. Combined with a pixel set allocation cost estimator, the integrity of the perceived target is guaranteed while obtaining the optimal cost evaluation. Attached Figure Description
[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a flowchart of the unsupervised low-altitude parallax image stitching method based on linear-nonlinear control grids according to the present invention; Figure 2 This is a schematic diagram of the structure of the low-altitude parallax image stitching model based on linear-nonlinear control grid constructed in this embodiment; Figure 3 This is a schematic diagram of the linear-nonlinear control grid in this embodiment; Figure 4 This is a schematic diagram of the optimal cost distribution in this embodiment; Figure 5 This is a low-altitude parallax image stitching result from an application example in this embodiment. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0024] This embodiment provides an unsupervised low-altitude parallax image stitching method based on linear-nonlinear control grids, such as... Figure 1 As shown, the specific steps include: S1: Obtain the image to be stitched from a low-altitude perspective; The images to be stitched together include a target image (left image) and a reference image (right image); S2: Construct a low-altitude parallax image stitching model based on linear-nonlinear control grid; like Figure 2 As shown, the model includes an input layer, a feature representation module constructed from ResNet50, a linear offset prediction module, a nonlinear offset prediction module, an image alignment module, and an artifact removal fusion network based on optimal cost evaluation. The input layer is used to input the images to be stitched into the feature representation module; The feature representation module is used to extract semantic feature information of the target image in the image to be stitched and obtain a semantic feature map; and the semantic feature information includes at least texture and contour information; and the semantic feature map includes a first scale feature map of different sizes, namely a set of feature maps of 1 / 16 size and a second scale feature map of 1 / 8 size. In this embodiment, a 1 / 16 feature map group is input into the linear migration prediction module to execute the linear migration prediction process, and a control point grid is created with the four corner points of the first-scale feature map as control points; a 1 / 8 feature map group is input into the nonlinear migration prediction module to execute the nonlinear prediction process, and a 12-point control point grid is created based on the second-scale feature map. For a control point grid of size 12, this embodiment performs linear and nonlinear grid offset regression predictions separately. During the prediction process, the linear offset prediction module and the nonlinear offset prediction module share prediction information through cross-attention, ultimately obtaining the first-stage linear control point grid offset prediction result and the second-stage nonlinear control point grid offset prediction result, as shown below. Figure 3 As shown; The linear migration prediction module is used to perform a linear transformation operation on the first-scale feature map to obtain a linear homography transformation matrix; specifically, the one-stage linear control point grid migration prediction is calculated and the linear transformation homography matrix is obtained through the Direct Linear Transformation (DLT) method.
[0025] In the formula: Represents the linear homography transformation matrix; This represents the coordinates of the k-th pixel in the reference image; This represents the coordinates of the corresponding pixel in the target image; This represents the intermediate result of the linear homography transformation matrix regression process, and the intermediate result includes a normalized point set and a homogeneous coordinate representation.
[0026] The nonlinear offset prediction module is used to interpolate the second-scale feature map based on the radial basis function interpolation algorithm to obtain a nonlinear mesh deformation field; and to obtain a coupled deformation field based on the linear homography transformation matrix and the nonlinear mesh deformation field; the coupled deformation field is used to perform coupled deformation on the target image to obtain an optimized target image; The method for obtaining the coupled deformation field in this embodiment is as follows: Define the grid corresponding to the second-scale feature map as follows In this embodiment, a two-dimensional grid is used to record control points, emphasizing the spatial distribution and topological relationship of control points in the grid. Based on the radial basis function interpolation algorithm, according to the grid Interpolation is performed on the second-scale feature map to obtain the nonlinear mesh deformation field:
[0027] In the formula: Represents the coordinates of the grid control points in the nonlinear grid deformation field; Indicates the interpolation weights of the radial basis functions; Describe the Gaussian function and ; This represents the coordinates of the pixel point to be interpolated in the second-scale feature map. Represents a grid The Middle The position coordinates of each grid control point; The coupled deformation field is obtained from the linear homography transformation matrix and the nonlinear mesh deformation field as follows:
[0028] In the formula: This indicates that for any point in the target image The coupled deformation field; Represents any point in the target image The processing results of the corresponding nonlinear mesh deformation field.
[0029] Furthermore, this embodiment also includes strict constraints on mesh deformation to ensure that the mesh does not cause image distortion due to overfit alignment. These strict constraints include: Mesh constraints: The width and height deformation of the mesh cannot exceed 50% of its side length; Mesh Constraints: In the two-stage nonlinear mesh migration, during the coupled deformation of the target image through the nonlinear migration prediction module, the angle between adjacent horizontal edges between meshes is no greater than the angle between the top and bottom edges of the mesh during the one-stage linear mesh migration, and the angle between adjacent vertical edges between meshes is no greater than the angle between the left and right adjacent edges of the mesh during the one-stage linear mesh migration. This embodiment applies over-deformation penalties to the mesh constraints during model learning to preserve the mesh shape, and the penalty weights are adaptively adjusted during learning. By applying mesh deformation constraints and penalties, this embodiment avoids overfitting of image alignment and maintains the normal line and surface structure of the stitched image.
[0030] The image alignment module is used to obtain the misaligned and overlapping regions corresponding to the optimized target image and the reference image; the artifact elimination fusion network based on optimal cost evaluation is used to perform disparity bridging and optimal cost fusion processing on the misaligned and overlapping regions to obtain the final stitched image. Specifically, the artifact removal fusion network based on optimal cost evaluation includes a cost estimator and a fusion decoder. In this embodiment, the artifact removal fusion network based on optimal cost evaluation performs cost evaluation on the misaligned overlapping regions of the aligned image to achieve optimal cost fusion while avoiding misalignment artifacts. In this embodiment, the cost estimator is used to construct an initial cost map for the misaligned overlapping regions of the aligned image based on pixel color differences, and the formula for constructing the initial cost map is:
[0031]
[0032]
[0033] In the formula: This represents the Manhattan distance between pixels in the overlapping region; This represents the pixels in the misaligned and overlapping region that correspond to the optimized target image. This represents the pixel in the misaligned and overlapping area that corresponds to the reference image. Indicates pixel channel; This represents the result of normalizing the Manhattan distance of the pixels; This represents the minimum Manhattan distance between pixels in the overlapping region. This represents the maximum Manhattan distance between pixels in the overlapping region; Indicates the inverse value after normalization; Specifically, in this embodiment, the initial cost map is constructed using the Manhattan distance between corresponding coordinate pixels in the misaligned overlapping region of two images. The pixel distance is then normalized, and the absolute value of the normalized distance is subtracted from 1 to obtain the final cost of that point. Here, cost refers to the color difference between corresponding pixels in the two images, which is the cost incurred when performing pixel line segmentation on the two pixels. The closer the RGB values of the two pixels, the higher the cost difference. The smaller the value, The larger the pixel, the less likely it is that the two pixels should be separated. This embodiment optimizes the cost redistribution of pixels within the target image and reference image in the misaligned overlapping region based on the initial cost map, to obtain the optimal cost distribution map as shown below. Figure 4 As shown, specifically: obtaining a pixel point within the target image and the reference image. The color difference between pixels spatially adjacent to it is determined, and the magnitude of the color difference is compared with a preset difference threshold. Then, it is compared with a certain pixel... Pixels that are spatially adjacent and whose color difference is less than a preset difference threshold are grouped into a preset set of pixels representing the same perceptual target. And improve pixel points based on empirical values. It is adjacent to and belongs to the same pixel set The value between pixels; conversely, confirming that they do not belong to the same set of perceived target pixels. And reduce pixel count based on empirical values. It is adjacent to and belongs to the same pixel set The cost value between pixels; and the cost value is the color difference value between pixels; In this embodiment, the cost estimator explicitly constructs an initial cost map for the misaligned and overlapping regions after image alignment based on pixel color differences, and redistributes costs according to the corresponding pixel assignments in the cost map. For example, when a pixel... Belonging to a certain set of perceived target pixels At that time, the pixel It is adjacent to and belongs to the same pixel set The cost between pixels will increase; when a pixel Not belonging to a certain set of perceived target pixels At that time, the pixel It is adjacent to and belongs to the same pixel set The cost between pixels will decrease, and the weight allocation parameters of the cost estimator will be adaptively updated during learning. In this embodiment, pixels with similar colors and spatial proximity are defined as belonging to the same perceptual target: when the color difference between a pixel and its neighboring pixels in a certain direction (such as its left) is small (less than a certain preset threshold), it is considered to belong to the same perceptual target, and therefore the cost of that point is increased. The difference between this and the initial cost map is that the initial cost map is constructed by the pixel color difference between two aligned images, while the cost redistribution is calculated between the pixel values of each image and its neighboring pixels, which makes it easier to comprehensively consider the differences between images and within images during segmentation.
[0034] The fusion decoder is used to perform dimensionality-upgrading fusion of the semantic feature map obtained by the feature representation module based on a preset inverse pyramid structure to obtain a fused semantic feature map with the same resolution as the original input reference image, so as to obtain the splicing overlap region features of the fused semantic feature map and the target image; and to confirm the fusion boundary region by combining the cost value reallocated by the cost estimator. The fusion boundary region is the seam region in the splicing overlap region features corresponding to the cost value being lower than the preset cost threshold; the gradient change of the cost map of the fusion boundary region is obtained by the cost estimator during training to dynamically adjust the width of the fusion boundary region, so as to achieve low-altitude parallax image splicing. Specifically, in this embodiment, the fusion decoder gradually performs residual upscaling and fusion of low-dimensional feature maps using an inverted pyramid structure. This inverted pyramid structure contains at least three sequentially connected upscaling and fusion modules, and each upscaling and fusion module includes: an upsampling layer (such as bilinear interpolation or transposed convolution) to upscale the spatial resolution of the input low-dimensional feature map to the size of the previous layer; and a residual connection unit to add the upsampled feature map element-wise with the feature map of the same scale from the corresponding layer of the feature representation module to achieve residual upscaling and fusion. By repeating the above operations layer by layer, a fused semantic feature map with the same resolution as the original input reference image is finally output. The cost estimator is based on a learnable deep neural network (such as U-N). The system is constructed using either a fusion semantic feature map or a fully convolutional network. It takes the overlapping region features of the fused semantic feature map and the target image as input and outputs a cost map of the same size as the overlapping region. Based on the cost map provided by the cost estimator, and considering the characteristics of the cost map (the cost value of each pixel in the cost map represents the "cost" of image cutting or fusion at that location: a higher cost value indicates complex texture, rich edges, or significant parallax, making cutting or fusion prone to artifacts; a lower cost value indicates smooth texture and small parallax, suitable as a stitching seam region, i.e., a fusion boundary region), the fusion boundary region is located at the boundary between the sets of perceived target pixels to ensure the integrity of the perceived target and avoid object artifacts. Furthermore, The fusion boundary region is not a single pixel line seam, but rather employs adaptive width fusion. This means fusion occurs within the fusion boundary region itself. Instead of using fixed pixel line stitching, the fusion width is dynamically adjusted based on the gradient changes in the cost map on both sides of the defined fusion boundary region. For example, for each pixel position within the fusion boundary region, the fusion width extends outwards along a direction perpendicular to the seam line, centered on that pixel. The width of this extension is determined by the cost value at that position—the lower the cost value, the larger the extension width (e.g., 1-5 pixels); when the cost value approaches a threshold, the extension width approaches 1 pixel. Within the extended fusion boundary region, a weighted average fusion algorithm is used. The pixel weight is determined by both the Euclidean distance from the seam line and the cost value, resulting in higher weights for pixels closer to the seam line and gradually decreasing weights for pixels farther from the seam line, achieving a smooth transition. Subsequently, by establishing a smoothness loss function to evaluate the fusion boundary region, the smoothness of the fusion boundary region is optimized by splicing smoothness constraints during training to eliminate the sense of disjointness in the fusion region of the spliced image. Specifically, within the fusion boundary region and its neighborhood, the brightness gradient difference between adjacent pixels is calculated and the gradient difference is minimized to promote a smooth brightness change within the fusion region and eliminate the sense of disjointness. S3: With the goal of minimizing the constructed unsupervised low-altitude parallax image stitching loss function, the constructed low-altitude parallax image stitching model is trained using the images to be stitched. During the training process, the quality of a certain set of images is evaluated by the geometric relationship between the source image, the aligned image, and the stitched image, and backpropagation is performed to optimize and update the model parameters to obtain the optimal low-altitude parallax image stitching model, so as to realize the unsupervised low-altitude parallax image stitching process based on linear-nonlinear control grid.
[0035] Specifically, the unsupervised low-altitude parallax image stitching loss function constructed in this embodiment is:
[0036]
[0037]
[0038]
[0039] In the formula: This represents the strong alignment loss of the image; , These represent the reference image and the target image to be stitched together, respectively. This represents the homography matrix obtained through the linear offset prediction module. The inverse matrix; Indicates the learning parameters; Represents the target image Optimized target image obtained based on coupled deformation location; The cost-optimal constraint function represents the evaluation of the optimality of the fusion region selection. Indicates correspondence The merging boundary region of pixel positions; Indicates in The luminance difference between the reference image and the target image at a pixel location is the RGB Euclidean distance between pixels. In this embodiment, the RGB Euclidean distance between pixels is... The calculation formula is ; Indicates the learning parameters; Indicates correspondence The merging boundary region of pixel positions; Indicates in The luminance difference between the reference image and the target image at a pixel location is the RGB Euclidean distance between pixels. Indicates in The luminance difference between the reference image and the target image at a pixel location is the RGB Euclidean distance between pixels. Indicates correspondence The merging boundary region of pixel positions; This indicates the smoothness loss used to assess the merging boundary region; This represents the brightness gradient between adjacent pixels in the stitched image; Indicates the learning parameters; This represents the unsupervised low-altitude parallax image stitching loss function. In this embodiment, since the inverse operation of the two-stage nonlinear deformation cannot be solved, mesh deformation constraints are only performed in the coupled deformation field. That is, the effectiveness of coupled deformation and homography transformation is evaluated through the designed image strong alignment loss. In this embodiment, the optimality of the fusion region selection is evaluated through the cost-optimal constraint function. In this embodiment, the smoothness of the fusion region is evaluated through the smoothness loss.
[0040] The application example results in this embodiment are as follows: Figure 5 As shown, when stitching sparse textured images acquired from the perspective of low-altitude UAVs, the unsupervised low-altitude parallax image stitching method based on linear-nonlinear control grids proposed in this embodiment still exhibits good image stitching capabilities: it maintains high visual quality for targets such as houses in rich textured areas of the scene, with no significant misalignment; and it also shows good continuity in water and sky boundary areas. This indicates that the method described in this embodiment can still complete the stitching task with limited information guidance in sparse textured scenes, improving the discrimination power for effective features from the perspective of low-altitude UAVs and the robustness in preserving the shape of low-confidence regions, laying the foundation for subsequent intelligent visual perception of mid- and high-altitude unmanned equipment and low-cost deployment in actual detection scenarios.
[0041] Compared with the prior art, the beneficial effects of the method described in this embodiment are as follows: (1) Compared with existing deep learning methods, the method described in this embodiment not only achieves unsupervised learning, but also maintains excellent stitching effect in sparse feature scenes from the perspective of low-altitude UAVs, compared with existing unsupervised learning image stitching methods.
[0042] (2) The two-stage mesh offset prediction and image alignment have higher accuracy alignment results compared with single linear or nonlinear alignment methods. In addition, the method described in this embodiment has a more complete shape preservation method and computational efficiency compared with existing two-stage alignment methods.
[0043] (3) Based on the optimal cost evaluation, the artifact elimination fusion network achieves seamless region fusion and provides a better smoothing effect in overlapping regions. After evaluation in multiple scenarios, it achieves a double advantage in both visual effect and objective evaluation.
[0044] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. An unsupervised low-altitude parallax image stitching method based on linear-nonlinear control grids, characterized in that, Specifically, the following steps are included: S1: Obtain the image to be stitched from a low-altitude perspective; the image to be stitched includes a reference image and a target image; S2: Construct a low-altitude parallax image stitching model based on linear-nonlinear control grid; The model includes an input layer, a feature representation module built with ResNet50, a linear offset prediction module, a nonlinear offset prediction module, an image alignment module, and an artifact removal fusion network based on optimal cost evaluation. The input layer is used to input the images to be stitched into the feature representation module; The feature representation module is used to extract semantic feature information of the target image in the image to be stitched and obtain a semantic feature map; and the semantic feature information includes at least texture and contour information; and the semantic feature map includes a first-scale feature map and a second-scale feature map with different scales. The linear offset prediction module is used to perform a linear transformation operation on the first-scale feature map to obtain the linear homography transformation matrix; The nonlinear offset prediction module is used to interpolate the second-scale feature map based on the radial basis function interpolation algorithm to obtain a nonlinear mesh deformation field; and to obtain a coupled deformation field based on the linear homography transformation matrix and the nonlinear mesh deformation field; the coupled deformation field is used to perform coupled deformation on the target image to obtain an optimized target image; The image alignment module is used to obtain the misaligned and overlapping regions corresponding to the target image and the reference image. An artifact elimination fusion network based on optimal cost evaluation is used to perform disparity bridging and optimal cost fusion processing on the misaligned overlapping regions to obtain the final stitched image. S3: With the goal of minimizing the constructed unsupervised low-altitude parallax image stitching loss function, the constructed low-altitude parallax image stitching model is trained using the image to be stitched to obtain the optimal low-altitude parallax image stitching model, so as to realize the unsupervised low-altitude parallax image stitching process based on linear-nonlinear control grid.
2. The unsupervised low-altitude parallax image stitching method based on linear-nonlinear control grids according to claim 1, characterized in that, The linear homography transformation matrix obtained in S2 is: In the formula: Represents the linear homography transformation matrix; This represents the coordinates of the k-th pixel in the reference image; This represents the coordinates of the corresponding pixel in the target image; This represents the intermediate result of the linear homography transformation matrix regression process, and the intermediate result includes a normalized point set and a homogeneous coordinate representation.
3. The unsupervised low-altitude parallax image stitching method based on linear-nonlinear control grids according to claim 2, characterized in that, The method for obtaining the coupled deformation field in S2 is as follows: Define the grid corresponding to the second-scale feature map as follows ; Based on the radial basis function interpolation algorithm, according to the grid Interpolation is performed on the second-scale feature map to obtain the nonlinear mesh deformation field: In the formula: Represents the coordinates of the grid control points in the nonlinear grid deformation field; Indicates the interpolation weights of the radial basis functions; Represents the Gaussian function; This represents the coordinates of the pixel point to be interpolated in the second-scale feature map. Represents a grid The Middle The position coordinates of each grid control point; The coupled deformation field is obtained from the linear homography transformation matrix and the nonlinear mesh deformation field as follows: In the formula: This indicates that for any point in the target image The coupled deformation field; Represents any point in the target image The processing results of the corresponding nonlinear mesh deformation field.
4. The unsupervised low-altitude parallax image stitching method based on linear-nonlinear control grids according to claim 3, characterized in that, The artifact elimination fusion network based on optimal cost evaluation in S2 includes a cost estimator and a fusion decoder. The cost estimator is used to construct an initial cost map based on pixel color differences in the misaligned overlapping areas of the aligned image, and the formula for constructing the initial cost map is as follows: In the formula: This represents the Manhattan distance between pixels in the overlapping region; This represents the pixels in the misaligned and overlapping region that correspond to the optimized target image. This represents the pixel in the misaligned and overlapping area that corresponds to the reference image. Indicates pixel channel; This represents the result of normalizing the Manhattan distance of the pixels; This represents the minimum Manhattan distance between pixels in the overlapping region. This represents the maximum Manhattan distance between pixels in the overlapping region; Indicates the inverse value after normalization; Based on the initial cost map, the cost redistribution of pixels within the target image and the reference image in the misaligned and overlapping regions is optimized, specifically as follows: Obtain a pixel within the target image and the reference image for optimization. The color difference between pixels spatially adjacent to it is determined, and the magnitude of the color difference is compared with a preset difference threshold. Then, it is compared with a certain pixel... Pixels that are spatially adjacent and whose color difference is less than a preset difference threshold are grouped into a preset set of pixels representing the same perceptual target. And improve pixel points based on empirical values. It is adjacent to and belongs to the same pixel set The cost between pixels; Conversely, confirm that it does not belong to the same set of perceived target pixels. And reduce pixel count based on empirical values. It is adjacent to and belongs to the same pixel set The cost value between pixels; and the cost value is the color difference value between pixels; The fusion decoder is used to perform dimensionality-upgrading fusion of the semantic feature map obtained by the feature representation module based on a preset inverted pyramid structure to obtain a fused semantic feature map with the same resolution as the original input reference image, so as to obtain the splicing overlap region features of the fused semantic feature map and the target image; and to confirm the fusion boundary region by combining the cost value reallocated by the cost estimator. The fusion boundary region is the seam region in the splicing overlap region features corresponding to the cost value being lower than the preset cost threshold; the gradient change of the cost map of the fusion boundary region is obtained by the cost estimator during training, and the width of the fusion boundary region is dynamically adjusted to achieve low-altitude parallax image splicing.
5. The unsupervised low-altitude parallax image stitching method based on linear-nonlinear control grids according to claim 4, characterized in that, The expression for the unsupervised low-altitude parallax image stitching loss function constructed in S3 is as follows: In the formula: This represents the strong alignment loss of the image; , These represent the reference image and the target image to be stitched together, respectively. This represents the homography matrix obtained through the linear offset prediction module. The inverse matrix; Indicates the learning parameters; Represents the target image Optimized target image obtained based on coupled deformation location; The cost-optimal constraint function represents the evaluation of the optimality of the fusion region selection. Indicates correspondence The merging boundary region of pixel positions; Indicates in The luminance difference between the reference image and the target image at a pixel location is the RGB Euclidean distance between pixels. Indicates the learning parameters; Indicates correspondence The merging boundary region of pixel positions; Indicates in The luminance difference between the reference image and the target image at a pixel location is the RGB Euclidean distance between pixels. Indicates in The luminance difference between the reference image and the target image at a pixel location is the RGB Euclidean distance between pixels. Indicates correspondence The merging boundary region of pixel positions; This indicates the smoothness loss used to assess the merging boundary region; This represents the brightness gradient between adjacent pixels in the stitched image; Indicates the learning parameters; This represents the loss function for unsupervised low-altitude parallax image stitching.