A panoramic image stitching method and system fusing color tone correction and POI anchoring
Patent Information
- Application Number
- CN202611278128.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-21
- Publication Date
- 2026-10-09
AI Technical Summary
[0005]针对上述存在的技术不足,本发明的目的是提出一种融合色调校正与POI锚定的全景图像拼接方法,旨在解决现有技术中对单张图像进行全局色调补偿,尤其是在多图全景拼接中多个重叠区域色差与拼接缝残差相互冲突的条件下,无法实现多重叠区色差一致性与拼接缝平滑性的技术问题
[0046]通过将POI地理坐标投影至多视角图像并提取局部结构描述子进行匹配,形成跨视图POI锚点数据,使同一POI在不同视角下的图像位置和语义信息建立可靠对应,为后续配准和色调校正提供空间约束,减少仅依赖图像特征时因弱纹理或重复纹理导致的误匹配,提升全景拼接中几何对齐的稳定性。
Smart Images

Figure CN122887672A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to a panoramic image stitching method and system that integrates tone correction and POI anchoring. Background Technology
[0002] Currently, applications such as street view mapping, urban digitization, building facade recording, and indoor navigation typically require stitching together multiple images from different perspectives and shooting poses to create a panoramic image, and mapping the geographic coordinates and semantic categories of Points of Interest (POIs) to a panoramic coordinate system. These multi-view images not only exhibit geometric transformations such as rotation, translation, and perspective, but also show significant differences in brightness and chromaticity due to variations in exposure time, white balance, shadows, and light source direction. Some applications also provide POI data, which can serve as auxiliary information for cross-view spatial correspondence and local structural constraints. Existing processing paths generally begin by extracting and matching features from the multi-view images, then estimating the projection transformation parameters between images based on the shooting pose; next, global tone compensation is applied to individual images or the initial stitched result to attempt to eliminate brightness and chromaticity differences between images; finally, a panoramic image is generated through multi-band fusion or weighted fusion, and the POI locations are mapped to panoramic coordinates based on geometric transformations. For simple scenarios with few overlapping areas and relatively consistent lighting conditions, this processing path of first fixing geometric registration and then performing global tone compensation can achieve acceptable results with a small amount of computation. However, its limitations gradually become apparent in scenarios where multiple images overlap, local color differences, and POI anchoring errors coexist.
[0003] For example, in the panoramic fusion of multi-view images, if a single image has multiple overlapping regions with several adjacent images, these overlapping regions may exhibit conflicting brightness and chromaticity deviations due to differences in local illumination, shadows, and color temperature. For instance, one overlapping region might appear cooler, while another might appear darker. Existing methods often apply global gain and bias to a single image for tone compensation. However, this type of global compensation only adjusts the average response of the entire image and cannot simultaneously meet the different requirements for color difference consistency from multiple overlapping regions. When some overlapping regions need increased brightness while others need decreased brightness, global tone compensation may improve color gradation breaks in one region but increase the seam residual in another. Furthermore, existing processing paths typically optimize the tone compensation field separately after fixing geometric transformation parameters. POI constraints and tone compensation are independent of each other, failing to incorporate POI position drift, overlapping region color difference, and seam residual into a unified optimization model. Geometric registration errors and POI position drift cannot be compensated for through subsequent geometric fine-tuning, resulting in simultaneous color gradation breaks and local misalignment near the seam. Because the color difference constraints of multiple overlapping areas and the smoothness constraints of the stitching seams conflict with each other when processed in a unified manner, global tone compensation of a single image cannot take into account both the consistency of local color difference and the smoothness of the stitching seams. This ultimately results in visible brightness abrupt changes, color band boundaries, or stitching seam afterimages in the panoramic image.
[0004] Therefore, there is an urgent need for a panoramic image stitching method that integrates tone correction and POI anchoring. Even when multiple images have multiple overlapping areas and the brightness and color differences between different images conflict with each other, it is still possible to incorporate POI position drift, color difference in overlapping areas, and stitching seam residuals into a unified process. This enables multi-image tone compensation and geometric parameter collaborative optimization, thereby improving the color consistency and stitching seam smoothness of panoramic images in overlapping areas, and enhancing the binding accuracy and reliability of POI positions in panoramic coordinates. Summary of the Invention
[0005] To address the aforementioned technical shortcomings, the present invention aims to propose a panoramic image stitching method that integrates tone correction and POI anchoring. This method addresses the technical problem in existing technologies where global tone compensation for a single image is insufficient, especially in multi-image panoramic stitching where color differences in multiple overlapping areas conflict with the residual difference at the stitching seam. This makes it impossible to achieve color difference consistency and stitching seam smoothness in these areas.
[0006] To solve the above-mentioned technical problems, the present invention adopts the following technical solution: The present invention provides a panoramic image stitching method that integrates tone correction and POI anchoring.
[0007] The panoramic image stitching method that integrates tone correction and POI anchoring includes:
[0008] Step S10: Obtain panoramic stitching source data, and based on the panoramic stitching source data, perform cross-view anchor point construction task using POI projection anchoring mechanism, and output cross-view POI anchor point data;
[0009] Step S20: Based on the cross-view POI anchor data, the Huber robust credibility assessment mechanism is used to perform anchor filtering and initial registration tasks, and output robust anchor evaluation data;
[0010] Step S30: Based on the robust anchor point evaluation data, the initial tone correction task is performed using the thin plate spline local tone correction mechanism, and the initial tone correction data is output.
[0011] Step S40: Based on the initial tone correction data, perform a joint geometry and tone optimization task using the POI-constrained Levenberg-Marquardt alternating optimization mechanism, and output the joint optimization data;
[0012] Step S50: Based on the joint optimization data, the panoramic fusion and POI location binding tasks are performed using the POI confidence-weighted Laplacian pyramid fusion mechanism, and the panoramic stitching result data is output.
[0013] Preferably, step S10, which involves acquiring panoramic stitching source data, performing a cross-view anchor point construction task based on the panoramic stitching source data using a POI projection anchoring mechanism, and outputting cross-view POI anchor point data, specifically includes:
[0014] Step S101: Obtain multi-view images, shooting poses and camera intrinsic parameter matrices corresponding to each multi-view image, and POI data. The POI data includes POI geographic coordinates, POI semantic categories, and POI reference structure descriptors. Combine the multi-view images, shooting poses, camera intrinsic parameter matrices, and POI data to form the panoramic stitching source data.
[0015] Step S102: Based on the panoramic stitching source data, the geographic coordinates of the POI are projected onto the corresponding multi-view image according to the shooting pose and the camera intrinsic parameter matrix to obtain the initial projection position of the POI; the local structure descriptor of the image is extracted in the preset neighborhood of the initial projection position of the POI, and the local structure descriptor of the image is matched with the corresponding POI reference structure descriptor to obtain the initial POI anchor point data.
[0016] Step S103: Based on the initial POI anchor point data, determine the visibility markers according to whether the projection positions of each POI in the initial POI anchor point data are located within the effective image plane and whether they are occluded. Combine the POI projection positions visible in at least two multi-view images, the image local structure descriptor, the POI semantic category, the visibility markers, and the panoramic stitching source data to form the cross-view POI anchor point data.
[0017] Preferably, step S20, which involves performing anchor point screening and initial registration tasks based on the cross-view POI anchor point data using the Huber robust credibility assessment mechanism and outputting robust anchor point assessment data, specifically includes:
[0018] Step S201: Based on the cross-view POI anchor point data, determine the reprojection error, semantic consistency, occlusion degree and local structural stability of each POI anchor point, and combine the reprojection error, semantic consistency, occlusion degree and local structural stability to form anchor point quality feature data;
[0019] Step S202: Based on the anchor point quality feature data, the reprojection error, semantic consistency, occlusion degree, and local structural stability are fused using preset weights to obtain an anchor point credibility score; POI anchor points with an anchor point credibility score greater than a preset high credibility threshold are included in a high credibility anchor point set, POI anchor points with an anchor point credibility score less than a preset rejection threshold are rejected, and the anchor point credibility score and the high credibility anchor point set are combined to form anchor point screening data;
[0020] Step S203: Based on the anchor point screening data, the anchor point confidence score is used as the registration weight of the corresponding POI anchor point. The Huber robust kernel is used to minimize the weighted projection error of the high-confidence anchor point set to obtain the initial image transformation parameters. The cross-view POI anchor point data, the anchor point confidence score, the high-confidence anchor point set and the initial image transformation parameters are combined to form the robust anchor point evaluation data.
[0021] Preferably, in step S202, the anchor point confidence score is calculated based on the anchor point quality feature data, and the anchor point confidence score is:
[0022]
[0023] in, For the first The first multi-view image Anchor credibility score for each POI anchor point For multi-view image indexing, For POI indexing, For the Sigmoid function, The reprojection error is expressed in pixels. For the sake of dimensionless semantic consistency. The degree of occlusion is dimensionless. The local structural stability is dimensionless. This is a dimensionless preset bias coefficient. These are preset weighting coefficients that match the dimensions of the reprojection error. , and The preset weight coefficients are dimensionless; the registration weights and credibility categories of each POI anchor point are determined using the anchor point credibility scores to obtain the anchor point screening data.
[0024] Preferably, step S30, which involves performing an initial tone correction task based on the robust anchor point evaluation data using a thin-plate spline local tone correction mechanism and outputting the initial tone correction data, specifically includes:
[0025] Step S301: Obtain the cross-view POI anchor data, the high-confidence anchor set, the anchor confidence score, and the initial image transformation parameters from the robust anchor evaluation data; obtain multi-view images and POI semantic categories from the cross-view POI anchor data; establish a coordinate mapping from each multi-view image to the corresponding reference view using the initial image transformation parameters; in the YCbCr color space, use robust color weights composed of distance weights, gradient weights, saturation weights, and structure weights to calculate the weighted color statistics of each POI neighborhood in each color channel within the high-confidence anchor set, and obtain POI neighborhood color statistics data.
[0026] Step S302: Based on the color statistics of the POI neighborhood, determine the saturation anomaly ratio and local structural integrity of each POI neighborhood. According to the anchor confidence score, the saturation anomaly ratio, the local structural integrity, and the POI semantic category, determine the reference view from the multi-view images corresponding to anchors belonging to the same POI within the high-confidence anchor set. Establish a local linear color correspondence relationship between each POI neighborhood in the corresponding multi-view image and the reference view based on the weighted color statistics of each POI neighborhood in the POI neighborhood color statistics, and obtain local tone parameter data. The local tone parameter data includes the local gain and local bias of each color channel.
[0027] Step S303: Based on the local tone parameter data, using the position of each POI anchor point in the high-confidence anchor point set in the corresponding multi-view image as the interpolation node, spatial interpolation is performed on the local gain and the local bias using a thin plate spline kernel to obtain an initial tone compensation field; the initial tone compensation field is used to perform tone correction on the multi-view image to obtain a tone-corrected image; the tone-corrected image, the initial tone compensation field, and the robust anchor point evaluation data are combined to form the initial tone correction data.
[0028] Preferably, step S40, which involves performing a joint geometry and tone optimization task based on the initial tone correction data using a Levenberg-Marquardt alternating optimization mechanism with POI constraints, and outputting the joint optimization data, specifically includes:
[0029] Step S401: Obtain the cross-view POI anchor point data from the initial tone correction data, and obtain the corresponding POI projection position from the cross-view POI anchor point data; within the preset search window of the tone-corrected image, calculate the image matching response and semantic matching response of the candidate positions, and determine the POI repositioning position based on the image matching response and the semantic matching response; determine the local occlusion degree at the POI repositioning position, and determine the POI repositioning confidence based on the image matching response, the semantic matching response, and the local occlusion degree; combine the POI repositioning position, the POI repositioning confidence, and the initial tone correction data to form POI constraint initialization data;
[0030] Step S402: Based on the POI constraint initialization data, construct a joint optimization objective including POI position drift term, overlapping area color difference term, seam residual term, and tone field smoothing term. The above four terms constrain the deviation between the repositioned POI position and the corresponding POI projection position, the overlapping area color difference, the seam continuity, and the tone compensation field smoothness in sequence. Combine the above four terms with preset weights to obtain the joint optimization objective data.
[0031] Step S403: Based on the joint optimization target data, using the initial image transformation parameters and the initial tone compensation field as initial values, the Levenberg-Marquardt method is used to update one type of parameters while fixing another type of parameters until the joint optimization target satisfies the preset convergence condition, thereby obtaining the optimized image transformation parameters and the optimized tone compensation field; the optimized image transformation parameters, the optimized tone compensation field, the POI repositioning location, the POI repositioning confidence, and the initial tone correction data are combined to form the joint optimization data.
[0032] Preferably, step S50, which involves performing panoramic fusion and POI location binding tasks based on the joint optimization data using a POI confidence-weighted Laplacian pyramid fusion mechanism, and outputting panoramic stitching result data, specifically includes:
[0033] Step S501: Obtain the optimized image transformation parameters, the optimized tone compensation field, the POI relocation location, the POI relocation confidence, and the cross-view POI anchor point data from the joint optimization data; and obtain the multi-view images, POI geographic coordinates, and the set of visible POIs in each multi-view image from the cross-view POI anchor point data; generate the final tone compensation image corresponding to each multi-view image based on the optimized image transformation parameters and the optimized tone compensation field, and determine the pixel position of the final tone compensation image. Final fusion weights at the point:
[0034]
[0035] in, For the first The final tone-compensated image at pixel location The dimensionless final fusion weights at the point are For multi-view image indexing, For POI indexing, The dimensionless pixel position to the first Distance weights for the boundaries of the final tone-compensated image. The gradient magnitude weights at the pixel positions are dimensionless. The dimensionless preset POI fusion weight enhancement coefficient. For the first The set of visible POIs in a multi-view image. For the first The first multi-view image Dimensionless POI relocation reliability for each POI The first one, in pixels Relocate the POI of each POI. The pixel position is expressed in pixels. The preset POI influence radius in pixels. This indicates that the summation is performed on the POIs in the set of visible POIs. It is an exponential function. The L2 norm is used; the final tone-compensated image and the final fusion weights are combined to form panoramic fusion preparation data;
[0036] Step S502: Based on the panoramic fusion preparation data, construct a Laplacian pyramid for the final tone compensation image, construct a Gaussian pyramid of the corresponding level for the final fusion weights, perform weighted fusion and reconstruction according to the corresponding level to obtain a panoramic image; calculate the image residuals on both sides of the stitching seam in the panoramic image, and combine the panoramic image and the image residuals to form panoramic fusion intermediate data;
[0037] Step S503: Based on the panoramic fusion intermediate data, when the image residual is greater than the preset stitching seam residual threshold, adjust the final fusion weight or perform local offset correction on the optimized tone compensation field; obtain the POI repositioning confidence in each multi-view image for the same POI, weight the projection position of the POI geographic coordinates in each multi-view image according to the POI repositioning confidence, convert the weighting result to panoramic coordinates, and obtain the POI position bound to the panoramic coordinates; combine the panoramic image and the POI position bound to the panoramic coordinates to form the panoramic stitching result data.
[0038] This invention also provides a panoramic image stitching system that integrates tone correction and POI anchoring, comprising:
[0039] The POI anchor point construction module is used to acquire panoramic stitching source data, and based on the panoramic stitching source data, execute the cross-view anchor point construction task using the POI projection anchoring mechanism, and output the cross-view POI anchor point data.
[0040] The robust anchor evaluation module is used to perform the anchor selection and initial registration tasks based on the cross-view POI anchor data and the Huber robust credibility evaluation mechanism, and output the robust anchor evaluation data.
[0041] The initial tone correction module is used to perform the initial tone correction task based on the robust anchor point evaluation data and the thin plate spline local tone correction mechanism, and output the initial tone correction data.
[0042] The joint optimization module is used to perform the geometry and tone joint optimization task based on the initial tone correction data and using the Levenberg-Marquardt alternating optimization mechanism with the POI constraint, and output the joint optimization data.
[0043] The panoramic fusion module is used to perform the panoramic fusion and POI location binding task based on the joint optimized data and the POI confidence-weighted Laplacian pyramid fusion mechanism, and output the panoramic stitching result data.
[0044] The present invention also provides a computer program product, the computer program product including a panoramic image stitching program that integrates tone correction and POI anchoring, the panoramic image stitching program that integrates tone correction and POI anchoring implements the above method when executed by a processor.
[0045] The beneficial effects of this invention are as follows:
[0046] By projecting POI geographic coordinates onto multi-view images and extracting local structural descriptors for matching, cross-view POI anchor point data is formed, enabling a reliable correspondence between the image position and semantic information of the same POI under different views. This provides spatial constraints for subsequent registration and tone correction, reduces mismatches caused by weak or repetitive textures when relying solely on image features, and improves the stability of geometric alignment in panoramic stitching.
[0047] The Huber robust credibility assessment mechanism is used to fuse the reprojection error, semantic consistency, occlusion degree and local structural stability of anchor points, select high credibility anchor points and assign corresponding registration weights, suppress the interference of low quality anchor points on the initial image transformation parameter estimation, so that the initial registration results still maintain high accuracy in the presence of local occlusion and semantic noise, and provide reliable initial values for subsequent tone correction and joint optimization.
[0048] By using a thin-plate spline local tone correction mechanism to spatially interpolate the local gain and bias of the high-confidence anchor point neighborhood, a tone compensation field that varies with position is generated. This can simultaneously meet the different needs of brightness and chromaticity adjustment in different overlapping areas of the same image, avoiding the global tone compensation from being incomplete in multiple overlapping areas. This reduces brightness abrupt changes and color band boundaries near the stitching seam, and improves the color consistency of the panoramic image. Attached Figure Description
[0049] Figure 1 This is a flowchart illustrating the first embodiment of a panoramic image stitching method that integrates tone correction and POI anchoring according to the present invention.
[0050] Figure 2 This is a cross-view POI anchor point reliability distribution map of a first embodiment of a panoramic image stitching method and system that integrates tone correction and POI anchoring according to the present invention.
[0051] Figure 3 This is a schematic diagram of tonal conflicts in multiple overlapping areas and direct stitching seams in a first embodiment of a panoramic image stitching method and system that integrates tone correction and POI anchoring according to the present invention.
[0052] Figure 4 This is a thin-plate spline spatial continuous tone compensation gain field diagram of a first embodiment of a panoramic image stitching method and system that integrates tone correction and POI anchoring according to the present invention.
[0053] Figure 5 This is a comprehensive color difference distribution map of the overlapping area before and after tone correction in the first embodiment of a panoramic image stitching method and system that integrates tone correction and POI anchoring according to the present invention.
[0054] Figure 6 This is a geometric and tonal alternation joint optimization convergence curve of the first embodiment of the panoramic image stitching method and system that integrates tone correction and POI anchoring of the present invention.
[0055] Figure 7 This is a residual distribution diagram of the stitching seam before and after the Laplacian pyramid fusion of a panoramic image stitching method and system that integrates tone correction and POI anchoring according to the first embodiment of the present invention. Detailed Implementation
[0056] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.
[0057] Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0058] Example 1: As Figure 1 The diagram shown is a flowchart of a first embodiment of a panoramic image stitching method that integrates tone correction and POI anchoring according to the present invention.
[0059] In the first embodiment, the panoramic image stitching method that integrates tone correction and POI anchoring includes:
[0060] Step S10: Obtain panoramic stitching source data, and based on the panoramic stitching source data, perform cross-view anchor point construction task using POI projection anchoring mechanism, and output cross-view POI anchor point data;
[0061] The panoramic stitching source data consists of multi-view images, the corresponding shooting poses and camera intrinsic matrix for each image, and POI data. The multi-view images are acquired from imaging devices at different locations or orientations, capturing the same scene. These images exhibit differences in viewpoint, scale variations, and local occlusion; overlapping areas may also show inconsistencies in brightness and color. The shooting pose describes the imaging relationship of each image in a unified world coordinate system, while the camera intrinsic matrix is used to transform world coordinates to image plane coordinates. POI data includes geographic coordinates, semantic categories, and reference structure descriptors, where the reference structure descriptors represent relatively stable texture or edge distributions near the points of interest. A pinhole camera projection mechanism is used to project the geographic coordinates of the points of interest image by image onto the corresponding multi-view images based on the shooting pose and camera intrinsic matrix, obtaining the initial projection positions. Subsequently, local structure descriptors are extracted from the preset neighborhood of the initial projection positions and matched with the corresponding reference structure descriptors to form initial POI anchor points. For each initial anchor point, a visibility marker is determined based on whether its projection position lies within the effective image plane and whether it is occluded by foreground objects. When the same point of interest is determined to be visible in at least two multi-view images, the projection location, local structure descriptor, semantic category, visibility marker, and corresponding source data in these images are organized into a cross-view point of interest anchor group. This anchor group associates the observations of the same geographic entity in different views, providing a set of candidate correspondences with geographic priors for subsequent anchor selection and initial registration.
[0062] Cross-view point of interest (POI) anchor data projects POIs, which were originally isolated by geographic coordinates, onto multiple images, forming anchors with initial image plane positions, local structural features, semantic categories, and visibility markers. Visibility markers eliminate projections located outside the effective image plane or those that are occluded, preventing these invalid projections from participating in error calculations in subsequent steps. Local structural descriptors record relatively stable texture or edge information within the neighborhood of the POI, allowing subsequent steps to locate and assess reliability based on this structural information even if there are brightness variations between different images. Cross-view POI anchor groups organize the projections of the same POI across multiple images into a single group, enabling subsequent anchor selection and initial registration to directly utilize geographic prior constraints, rather than relying solely on local image feature matching results. For scenarios with multiple overlapping regions, a stable anchor distribution helps subsequent steps distinguish the location and extent of different overlapping regions, reducing location errors in overlapping regions caused by feature mismatches. These anchors also provide cross-view location constraints for subsequent tone correction and stitching processing, enabling geometric and tone joint optimization to establish constraints at geographically referenced locations.
[0063] Traditional methods for constructing cross-view correspondences typically involve first extracting feature points from the image and performing descriptor matching, then using geometric verification to eliminate false matches. This approach is prone to generating insufficient or numerous false matches in areas with weak or repetitive textures, or where perspective distortion exists, potentially leading to local shifts in subsequent registration. When there are significant differences in brightness and chromaticity across multiple viewpoints, feature descriptors become highly sensitive to lighting conditions, and the same feature may generate different descriptions in different images, further increasing the likelihood of false matches. This approach introduces the geographic coordinates of points of interest as a priori, obtains initial positions through pinhole projection, extracts locally stable structural descriptors within the neighborhood, and constructs anchor point groups based on visibility judgments. This provides the correspondences with the assistance of geographic priors and semantic information. Subsequent filtering does not require re-identifying reliable correspondences from a large number of feature matches without prior knowledge, thus reducing the impact of weak and repetitive textures on the quality of the initial correspondences and providing a more stable spatial reference for handling color differences and seams in overlapping areas.
[0064] In a panoramic acquisition of a city street, multiple images covered intersections, street facades, and traffic signs. The geographic coordinates of a shop sign point of interest (POI) were first projected onto three of these images. In two of the images, the projected point was within the effective image plane and not obscured by trees. In the third image, the projected point was invisible due to vehicle obstruction. Edge and text structure descriptors near the sign were extracted for the first two images, and their visibility was marked as visible. The third image was marked as invisible. Ultimately, this POI formed a cross-view anchor point group containing two anchor points. Subsequent filtering steps could determine the consistency of the POI across different views based on these two anchor points and their structure descriptors, rather than relying solely on local feature points that might mismatch in areas of repeated text on the sign. Furthermore, these two anchor points, located near two different overlapping areas, provide cross-view positioning constraints for subsequent tone correction and stitching processing.
[0065] Step S20: Based on the cross-view POI anchor data, the Huber robust credibility assessment mechanism is used to perform anchor filtering and initial registration tasks, and output robust anchor evaluation data;
[0066] After cross-view point of interest (POI) anchor point data enters S20, each anchor point needs to undergo simultaneous calculations of reprojection error, semantic consistency, occlusion level, and local structural stability. Reprojection error measures the deviation between the observed position of the anchor point in each view and the position projected based on geographic coordinates and camera parameters; semantic consistency is reflected by the similarity between local image semantic features and the semantic reference features of the POI; occlusion level is determined by depth information or multi-view... Figure 1The occlusion confidence score is calculated based on consistency; local structural stability is obtained through normalized cross-correlation between the image gradient map and the local stable structure descriptor. These four indicators are combined to form anchor point quality feature data, which is then fused into an anchor point confidence score according to preset weights. Anchor points with confidence scores higher than a preset high-confidence threshold are added to the high-confidence anchor point set, anchor points lower than a preset removal threshold are directly removed, and anchor points in between are retained but with reduced weights. The high-confidence anchor point set is used to construct a weighted projection error energy function, where a robust kernel function limits the influence of anchor points with large errors, and Mahalanobis distance characterizes the uncertainty of positional errors in different directions. Minimizing this energy function yields initial image transformation parameters. Cross-view interest point anchor data, anchor point confidence scores, the high-confidence anchor point set, and the initial image transformation parameters together constitute robust anchor point evaluation data.
[0067] The selected set of high-confidence anchor points meets stringent conditions in terms of reprojection error, semantic consistency, occlusion degree, and local structural stability, thus more closely approximating the actual ground feature correspondences. The energy function no longer treats all anchor points equally but instead uses confidence scores as weights, concentrating optimization resources on constraints with higher alignment quality. The robust kernel function further limits the interference of individual anchor points with large errors on the overall registration, ensuring the stability of initial image transformation parameters in multi-view overlapping regions. These initial transformation parameters provide the foundation for S30 to establish coordinate mappings from each image to the reference view, allowing initial tone correction to be performed only if the images are correctly aligned. For multi-overlapping region scenes, stable initial transformation parameters prevent subsequent color difference and seam constraints from being incorrectly allocated due to geometric initial value deviations, thus preserving a more reasonable initial state for unified optimization.
[0068] Traditional methods often rely on a single metric in the anchor point selection stage, such as reprojection error or feature matching score. A single metric is insufficient to reflect whether anchor points are unreliable due to semantic confusion, local occlusion, or unstable texture structure. This is especially problematic when there are brightness differences and occlusion in multi-view images, easily leading to the incorrect retention of high-error anchor points or the accidental deletion of true anchor points. If the initial registration is dominated by anchor points with gross errors, subsequent color difference constraints and seam constraints in multi-overlapping regions may be established at incorrect spatial locations, resulting in local color level breaks that cannot be correctly attributed and handled. S20 incorporates reprojection error, semantic consistency, occlusion degree, and local structural stability into the credibility evaluation simultaneously, and constructs an optimization objective through a robust kernel function, ensuring good convergence results even with some erroneous anchor points in the initial registration. This multi-evidence selection method reduces the dependence on image matching quality, making the initial transformation parameters more suitable for subsequent joint optimization with color difference constraints.
[0069] In the panoramic acquisition of urban streets described in step S10, the points of interest (POIs) of shop signs form two visible anchor points, located near two different overlapping areas. In step S20, reprojection error, semantic consistency, occlusion degree, and local structural stability are calculated for these two anchor points. One anchor point is located in the overlapping area of the street facade and intersection images. This area has clear texture, semantic features highly consistent with the sign reference features, low occlusion degree, and stable local structure; therefore, its confidence score is higher than the preset high confidence threshold, and it is included in the high confidence anchor point set. The other anchor point is located in the overlapping area of the street facade and traffic sign images. Due to partial tree branch occlusion in this area, the occlusion degree is relatively high, but the reprojection error and semantic consistency are still within acceptable ranges. Its confidence score is between the high confidence threshold and the removal threshold, so it is retained but its weight is reduced. Furthermore, the third image, which is invisible due to vehicle occlusion in step S10, does not generate anchor points and is therefore not included in this step's calculation. Using the high confidence anchor point set and the confidence weights of the retained anchor points, a weighted projection error energy function is constructed and optimized using a robust kernel function to obtain the initial image transformation parameters. This parameter enables the overlapping areas between the street facade image, intersection image, and traffic sign image to be initially aligned geometrically, providing a reliable coordinate mapping basis for the local tone correction of the thin plate spline of S30.
[0070] Step S30: Based on the robust anchor point evaluation data, the initial tone correction task is performed using the thin plate spline local tone correction mechanism, and the initial tone correction data is output.
[0071] When the high-confidence anchor point set, along with the confidence score and initial image transformation parameters, enters the tone correction stage, all multi-view images have already been aligned to the same reference view through coordinate mapping. The correction work is first carried out in the color space composed of the luminance channel and two chrominance channels, statistically calculating weighted color quantities for the neighborhood of each anchor point. The statistical weights are jointly determined by four terms: distance, gradient, saturation, and structure. The distance term reduces the contribution of pixels far from the center of the interest point; the gradient term suppresses the perturbation of statistics by edges and high-frequency regions; the saturation term prevents highly saturated pixels from dominating the results; and the structure term favors pixels consistent with the local stable structure descriptor. After the neighborhood statistics of the same interest point in different views are completed, a local linear color correspondence is established based on these statistics, and local gain and local bias are estimated channel by channel. These parameters only reflect the color differences near the neighborhood of each interest point and do not yet cover the entire image. Subsequently, using the positions of each interest point anchor point in the high-confidence anchor point set in the corresponding multi-view images as interpolation nodes, a thin-plate spline kernel is used to spatially interpolate the local gain and local bias, respectively, to obtain an initial tone compensation field that continuously varies on the image plane. The compensation field is used to perform tone correction on multi-view images, generating tone-corrected images, which are then combined with the initial tone compensation field and robust anchor point evaluation data to form initial tone correction data.
[0072] After tone correction, the brightness and chromaticity differences in the neighborhood of the point of interest and near overlapping regions of the image have been initially suppressed. The initial tone compensation field records the spatially continuously varying gain and bias distribution. Since the compensation field changes with spatial location, different overlapping regions in the same image can obtain different gains and biases based on their respective local statistics, thus differentiating conflicting color difference requirements of multiple overlapping regions in the initial stage. This intermediate state provides initial values for the geometric and tone joint optimization in step S40, allowing the optimization process to handle geometric and color constraints simultaneously without eliminating large overall color shifts. Instead, it can handle residual color differences and seam residuals on a relatively consistent basis. This avoids the joint optimization from failing to converge due to excessive overall color shifts in the initial stage and reduces oscillations caused by mutual interference between geometric and tone parameters. The initial tone compensation field also provides spatial priors for subsequent optimization, making the optimization process more inclined to preserve local tone differences rather than degenerating into a globally uniform adjustment.
[0073] Traditional methods for global tone compensation of a single image typically use a single global gain and bias to stretch the entire image. This approach cannot handle situations where different overlapping regions within the same image are affected by varying lighting or color temperatures, easily leading to improvements in color difference in one overlapping region while exacerbating color gradation breaks in another. Global compensation can only adjust the average response of the entire image, making it difficult to simultaneously meet the different requirements for color difference consistency across multiple overlapping regions. In contrast, this step generates a spatially varying compensation field based on interest point neighborhood weighted statistics and thin-plate spline interpolation. This allows the tone adjustment amount in different regions to adaptively change according to local anchor point statistics, rather than forcing the entire image to use the same set of gains and biases. This allows the different color shift requirements of multiple overlapping regions to be addressed separately in the initial stage, reducing conflicts during subsequent unified optimization and avoiding the fixation of a global tone parameter unsuitable for the local structure at the initial stage. The gradient, saturation, and structure terms in the color statistical weights suppress the excessive influence of highlights, shadows, and edge pixels on the estimation of local color parameters, making the initial tone correction closer to the true color differences between stable surfaces in the scene.
[0074] In step S20, a set of high-confidence anchor points and initial image transformation parameters were obtained between the street facade image, intersection image, and traffic sign image. These parameters map the street facade image, intersection image, and traffic sign image to the same reference view coordinate system. In the overlapping area of the street facade and intersection images, the neighborhood of the shop sign interest point contains stable wall and sign surfaces. Distance weight and structure weight make these stable pixels dominate in color statistics, gradient weight suppresses high-frequency changes at the sign edges, and saturation weight avoids excessive influence of highly saturated sign colors on the statistical results. In the overlapping area of the street facade and traffic sign images, tree branch occlusion causes some pixels to be unstable, but the structure weight reduces the contribution of these pixels. After calculating the weighted color statistics of the luminance channel and two chrominance channels for each interest point neighborhood, a local linear color correspondence is established to obtain the local gain and local bias near each interest point. Using these interest point locations as interpolation nodes, spatial interpolation is performed using a thin plate spline kernel to generate an initial tone compensation field covering the entire street facade image. The compensation field moderately brightens the shadow side in the overlapping area of the image near the intersection, and locally corrects the warm color shift in the overlapping area of the image near the traffic sign, while the dark areas far from the overlapping area retain their original tonal range. After color correction of the street facade image using this compensation field, a color-corrected image is obtained, which, together with the initial color compensation field and robust anchor point evaluation data, is output as the initial color correction data for step S40 to perform geometric and color joint optimization.
[0075] Step S40: Based on the initial tone correction data, perform a joint geometry and tone optimization task using the POI-constrained Levenberg-Marquardt alternating optimization mechanism, and output the joint optimization data;
[0076] When the initial tone-corrected data enters the joint optimization stage, the tone-corrected image already contains the initial tone compensation field, and the cross-view point of interest (POI) anchor data still retains the projection position and confidence information corresponding to each POI. Before optimization begins, each POI is relocated on the tone-corrected image, with the search range limited to a preset search window near the corresponding projection position. Candidate positions simultaneously calculate image matching response and semantic matching response. The image matching response reflects the consistency of texture and structure between the candidate position's neighborhood and the original POI's neighborhood, while the semantic matching response reflects whether the candidate position still falls on the same type of land cover surface. After obtaining the relocation confidence by combining the degree of local occlusion, the relocated position, relocation confidence, and initial tone-corrected data are combined into POI constraint initialization data. The joint optimization objective consists of four terms: the POI position drift term limits the deviation between the relocated position and the geographic coordinate projection position; the overlapping area color difference term measures the brightness and chromaticity differences within the overlapping areas between images; the stitching seam residual term constrains the image intensity and gradient differences near the stitching seam; and the tone field smoothing term prevents the tone compensation field from oscillating violently in space. After the four parameters are combined with preset weights, the geometric parameters and tone parameters are updated alternately using the initial image transformation parameters and the initial tone compensation field as initial values, and the Levenberger and Marquardt methods are used until the preset convergence conditions are met. The optimized image transformation parameters, optimized tone compensation field, interest point relocation position and relocation confidence are output.
[0077] After joint optimization, the image transformation parameters and tone compensation field have been repeatedly coordinated by the same objective function, and the relocation positions and relocation confidence of interest points have also been preserved. In the subsequent step S50, when performing Laplacian pyramid fusion, the optimized image transformation parameters can be directly used to map each image to a unified reference view. Residual color differences and geometric misalignments in the overlapping areas have been significantly reduced, and the fusion stage only needs to handle minor local discontinuities. The relocation confidence of interest points, as part of the fusion weights, makes the fusion result near high-confidence interest points more stable, while the area near low-confidence interest points relies more on the structural information of the image itself. Because the tone compensation field is constrained by a smoothing term during optimization, it will not generate new color level breaks at the boundaries of the overlapping areas, making it less prone to artifacts caused by abrupt changes in the compensation field during subsequent fusion. After the geometric and tone parameters are updated alternately, the overlapping areas between images are coordinated in both color and geometry. Step S50 no longer requires additional global color adjustments to the entire image, thus avoiding local color casts or residual stitching seams in the fusion result.
[0078] Traditional methods typically perform image registration first, then adjust the colors of individual images or stitched results. Since geometric transformations are fixed before tone correction, subsequent tone adjustments cannot reverse the local color misalignment caused by geometric errors. When color differences conflict in multiple overlapping regions, global compensation for a single image can only take an average compromise, resulting in one overlapping region improving color gradation breaks while another overlapping region experiences increased stitching seam residuals. This step incorporates the drift of interest point positions, color differences in overlapping regions, stitching seam residuals, and tone field smoothing into the same objective function. During the geometric optimization stage, the tone field is fixed to avoid virtual edges introduced by tone changes interfering with image alignment. During the tone optimization stage, geometric parameters are fixed, focusing on processing color differences in overlapping regions while ensuring image alignment. Dynamic weight scheduling initially prioritizes geometric constraints to prevent the still-converged tone field from causing excessive disturbance to geometric optimization. Later, color and seam constraints are gradually strengthened to provide more refined processing of residual color differences. Interest point location constraints prevent image relocation from deviating excessively from real features during tone correction. Color difference and seam terms also promote better local consistency of geometric parameters in overlapping areas. Therefore, different color difference requirements in multiple overlapping areas can be balanced within the same optimization framework, reducing the residual local color level breaks.
[0079] After tone correction in step S30, the street facade image already has an initial tone compensation field. The shadow side of the neighborhood of the shop sign's point of interest is moderately brightened, and the warm color shift near the overlapping area of the traffic sign image is locally corrected. At the start of joint optimization, the shop sign's point of interest is relocated on the tone-corrected street facade image. The search window covers the wall area near the edge of the sign, and the candidate locations simultaneously calculate the image matching response and semantic matching response. The candidate locations at the tree branch occlusion area have lower relocation confidence due to the higher degree of local occlusion. In the joint optimization objective, the deviation between the relocated location of the shop sign's point of interest and its geographic coordinate projection location is limited. The color difference terms of the overlapping areas of the street facade and the intersection image, and the color difference terms of the overlapping areas of the street facade and the traffic sign image are retained simultaneously. The tone field smoothing term makes the compensation amount within the street facade image spatially smooth. In the geometric optimization stage, the point of interest position is kept from drifting. In the tone optimization stage, the local gain and bias are adjusted without changing the alignment relationship. After repeated iterations, the color difference of the two overlapping areas is suppressed to a certain extent. In the final output of the joint optimization data, the transformation parameters of the street facade image and the intersection image and traffic sign image are locally consistent in the overlapping area. The tone compensation field smoothly transitions between the shadow side and the warm color offset area. The relocation position and relocation confidence of the points of interest are preserved, which are used for the Laplacian pyramid fusion with confidence weighted by the points of interest in step S50.
[0080] Step S50: Based on the joint optimization data, the panoramic fusion and POI location binding tasks are performed using the POI confidence-weighted Laplacian pyramid fusion mechanism, and the panoramic stitching result data is output.
[0081] The optimized image transformation parameters, optimized tone compensation field, repositioning location of points of interest (POIs), POI repositioning confidence, and cross-view POI anchor point data extracted from the joint optimization data constitute the input for panoramic fusion. Multi-view images, POI geographic coordinates, and the set of visible POIs in each image are read from the cross-view POI anchor point data. Each multi-view image generates a final tone compensation image based on the optimized image transformation parameters and optimized tone compensation field, ensuring a consistent color reference for overlapping areas. For each pixel location in the final tone compensation image, the fusion weight is determined by three parts: the distance weight from the pixel location to the image boundary, the gradient magnitude weight of the pixel location, and the POI fusion weight enhancement term. The POI fusion weight enhancement term forms a Gaussian weighted region centered on the POI repositioning location, within a preset POI influence radius, based on the POI repositioning confidence. This makes pixels near high-confidence POIs more likely to adopt the contribution of the view containing that POI during fusion. The final tone compensation image and the final fusion weights combine to form the panoramic fusion preparation data. Subsequently, a Laplacian pyramid is constructed for the final tone-compensated image, and a Gaussian pyramid of the corresponding level is constructed for the final fusion weights. Weighted fusion and reconstruction are performed according to the corresponding levels to obtain the panoramic image. The image residuals on both sides of the stitching seam in the panoramic image are calculated, and the panoramic image and image residuals are combined to form intermediate panoramic fusion data. When the image residuals are greater than the preset stitching seam residual threshold, the final fusion weights are adjusted or a local bias correction is performed on the optimized tone-compensated field. The relocation confidence of the same point of interest in each multi-view image is obtained. Based on this confidence, the projection positions of the geographic coordinates of the point of interest in each multi-view image are weighted, and the weighted results are converted to panoramic coordinates to obtain the point of interest positions bound to the panoramic coordinates. The panoramic image and the point of interest positions bound to the panoramic coordinates are combined to form the panoramic stitching result data.
[0082] The image source entering the Laplacian pyramid has undergone joint optimization of tone field and geometric parameters, resulting in smaller color differences and more stable alignment. The pyramid processes low-frequency, large-scale color transitions and high-frequency detail information separately. Residual color differences in the low-frequency layer are smoothly transitioned, while textures and edges in the high-frequency layer are not blurred due to fusion. The point of interest (POI) fusion weight enhancement term strengthens the contribution of the corresponding view near high-confidence POIs, avoiding the misuse of low-quality views at key features and preventing information loss. Stitch seam residual detection provides a local correction mechanism, handling individual residual discontinuities without requiring global optimization of the entire panorama, ensuring smoothness in overlapping areas of the output panoramic image. POI locations are weighted by multi-view confidence projection, resulting in bound coordinates that more closely approximate the actual location of features, enhancing the traceability of POI information within the panoramic image. This panoramic stitching result data can be directly used for subsequent geographic information annotation, navigation positioning, or scene display, with the binding relationship between POI locations and panoramic coordinates ensuring spatial consistency.
[0083] While traditional Laplacian pyramid fusion can mitigate local abrupt changes near the stitching seam, if significant tonal differences still exist between input images, multi-band fusion only smooths the fusion weights and cannot eliminate existing tonal breaks in the low-frequency layers. This step reduces color and geometric residuals through joint optimization before fusion, and the final tone compensation field has been applied to each image, resulting in better color consistency in overlapping areas of the images entering the pyramid. Interest point relocation confidence is introduced into the fusion weights to prevent a decrease in visual quality of important interest point regions in the panoramic image due to a low-quality image. Stitching seam residual detection and local correction further address a small number of non-converged residuals, reducing the likelihood of visible tonal breaks in the final output. Compared to traditional methods, this step does not simply rely on fusion weights to mask input differences but achieves high color consistency between image sources through unified optimization before fusion.
[0084] After joint optimization in step S40, the overall brightness of the street facade image, intersection image, and traffic sign image is basically consistent, but a slight difference in brightness still exists at the edge of a sunshade. This location is close to a high-confidence shop sign point of interest, so the local weights of the two images containing this sign and with high confidence are strengthened in the fusion weights. The Laplacian pyramid disperses the low-frequency brightness differences near the sunshade to a wider transition area while preserving high-frequency details at the edge of the sunshade. The stitching seam residual detection found that the residual at the edge of the sunshade exceeded the preset threshold, so a Gaussian envelope bias correction was applied to this local area. In the final panoramic image, the transition between the sunshade and the surrounding walls is natural, and the location of the shop sign point of interest is accurately bound to the panoramic coordinates, while preserving the geographic coordinates, semantic category, and cross-view source information of the point of interest.
[0085] For example, such as Figure 2 As shown, the simulation sample contains 140 cross-view candidate anchor points. The horizontal axis represents the reprojection error, and the vertical axis represents the cosine consistency between the local semantic features of the image and the reference semantic features of the POI. The color of the point indicates the anchor point confidence obtained after fusing the degree of occlusion and the stability of the local structure. Samples with small reprojection errors, high semantic consistency, and complete local structures are concentrated in the upper left region of the image, with a confidence level higher than 0.7, and are classified into the high-confidence anchor point set. Samples caused by reflections, leaf swaying, or vehicle occlusion usually show both increased reprojection errors and decreased semantic consistency, and are removed when the confidence level is lower than 0.3. Anchor points between these two levels are retained as medium-confidence samples, but do not directly dominate the initial transformation solution. This distribution shows that S20 does not rely on a single distance threshold for mechanical point screening. For example, some candidate points with projection errors of about 3 pixels but significantly inconsistent semantic categories are still suppressed, while some stable wall edges with slightly larger errors can retain appropriate weights due to their high structural consistency. Multi-evidence fusion makes the initial registration constraints closer to the real object correspondence, especially suitable for situations where there are simultaneous changes in illumination, reflections, and local occlusions in multi-overlapping areas.
[0086] like Figure 3 As shown, after three street scene images are directly projected onto the same panoramic canvas according to an initial geometric transformation, a first overlapping area with a bluish and dark tone is formed between the left shadow view and the middle view, while a second overlapping area with a yellowish and bright tone is formed between the middle view and the right sunset view. The two vertical dashed lines in the image mark the main seam locations when directly stitched together. It can be observed that the continuous color gradations of the same building wall are truncated on both sides of the seam, and local brightness jumps are also observed in the awnings and road surface. If only one set of global gain and bias is used on the middle image, the first overlapping area requires increased brightness and reduced cool color components, while the second overlapping area requires reduced brightness and suppressed warm color components. Since the two sets of constraints are in opposite directions, only one side can be improved, or a compromise residual can be left on both sides. This scene directly corresponds to the tonal conflict in multiple overlapping areas that S30 needs to handle, indicating that the problem is not a simple single exposure difference, but rather the same source image facing different compensation requirements in different spatial locations, necessitating the generation of a spatially varying and continuous tonal compensation field.
[0087] like Figure 4As shown in the figure, the spatially continuous gain field of the luminance channel is presented. The dots represent the control positions of high-confidence POIs after S20 screening, and the contour lines represent the smooth change of gain with image coordinates after thin-plate spline interpolation. Control points near the left shadow overlap area need to have their brightness moderately increased, with their local gain greater than the unit value; control points near the right sunset overlap area need to suppress overbrightness response, with their local gain less than the unit value. The two types of compensation amounts with opposite directions do not form abrupt boundaries in the image, but rather transition continuously in space through the thin-plate spline kernel function and low-order polynomial terms. The control values are not directly determined by a single pixel at the center of the POI, but come from the weighted color statistics of the POI neighborhood. The distance weight reduces the influence of far-end pixels, the gradient weight suppresses edges, the saturation weight avoids highlights, and the structure weight prioritizes retaining wall and sign areas consistent with the stable descriptor. The resulting initial compensation field can respond to the color difference requirements of multiple overlap areas separately, while avoiding violent oscillations of local parameters between POIs, providing physically more reasonable initial hue values for S40 alternating optimization.
[0088] like Figure 5 As shown, the overall chromatic aberration is represented by a normalized combination of the differences between the luminance channel and the two chromaticity channels in the overlapping region. The sample distributions after direct projection, global gain correction, thin-slab spline initial correction, and S40 joint optimization are statistically analyzed. The direct projection group is affected by shadows, sunset color temperature, and local highlights, resulting in a wide distribution range and a significant long tail. Global correction can reduce the overall average difference, but it cannot simultaneously satisfy two overlapping regions in opposite directions, still retaining many high residual samples. Thin-slab spline correction utilizes local statistics in the POI neighborhood to generate spatially varying gain and bias fields, further reducing the median and upper quartile of the distribution. The joint optimization group continues to compress the tail residuals after geometric realignment. These results indicate that the role of S30 is not simply to pull all images to the same average tone, but rather to first spatially decompose conflicting chromatic aberrations, and then S40 coordinates geometric and tonal parameters. Therefore, it can reduce pseudo-chromatic aberrations caused by slight misalignment and maintain higher consistency with the input sources faced by the subsequent Laplacian pyramid.
[0089] like Figure 6As shown, the simulation starts with the initial compensation field of the thin-plate spline and the image transformation after Huber screening, recording the changes in the POI position drift term, overlapping area color difference term, seam residual term, and unified objective function in the Levenberg-Marquardt outer loop. In the first few iterations, geometric constraints are the primary focus, causing the POI position drift term to decrease rapidly, while the color difference and seam terms are only mildly adjusted to avoid the unstable hue field affecting geometric registration. As the dynamic weights gradually heat up, the overlapping area color difference and seam residual begin to converge continuously. In the later joint refinement stage, both geometric and hue parameters are updated simultaneously, and all curves enter a stable region without one term decreasing while another rebounds significantly. This convergence process corresponds to the S40 outer loop joint and inner loop block strategy, indicating that geometric alignment and local hue compensation are not two independent post-processing steps, but rather gradually achieve a balance within the unified objective function. This avoids simply pursuing color consistency leading to POI position drift, or only pursuing minimum geometric error leaving visible seams.
[0090] like Figure 7 As shown, the horizontal axis represents the pixel position across the splicing line, and the vertical axis represents the local residuals on both sides of the splicing seam, which are composed of brightness difference and gradient difference. The direct fusion curve forms a sharp peak at the center of the seam, indicating that even if the initial geometric relationship is roughly correct, the residual color difference will still form visible breaks at the wall surface, the edge of the awning, and the road texture. After completing the S40 joint optimization, the peak height and the range of influence are significantly reduced, but a small amount of local residuals are still retained at the high-frequency edges. S50 uses the Laplacian pyramid to fuse by frequency layer, and adjusts the fusion weight in local areas where the residual exceeds the threshold, and applies Gaussian envelope bias correction to the tone compensation field. The final curve remains continuous near the seam. This change shows that multi-band fusion does not use blurring to cover up all differences. The low-frequency layer is responsible for smoothing the transition between light and dark, the high-frequency layer retains the sign text and building edges, and the residual detection only makes local corrections to the non-converged positions, thus taking into account both the smoothness of the seam and the clarity of details.
[0091] Example 2: Furthermore, the panoramic image stitching system integrating tone correction and POI anchoring provided by this invention employs a panoramic image stitching method integrating tone correction and POI anchoring as described in the above embodiments, which can solve the technical problem of panoramic image stitching integrating tone correction and POI anchoring. The beneficial effects of the panoramic image stitching system integrating tone correction and POI anchoring provided by this invention are the same as those of the panoramic image stitching method integrating tone correction and POI anchoring provided in the above embodiments, and other technical features of the panoramic image stitching system integrating tone correction and POI anchoring are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0092] Example 3: This invention provides a panoramic image stitching device that integrates tone correction and POI anchoring. The device includes at least one processor and a memory communicatively connected to the processor. The memory stores instructions executable by the processor, which are then executed to enable the processor to perform the panoramic image stitching method integrating tone correction and POI anchoring described in Example 1. This panoramic image stitching device integrating tone correction and POI anchoring, as described in this embodiment, can include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. This panoramic image stitching device integrating tone correction and POI anchoring is merely an example and should not be construed as limiting the functionality or scope of this invention. A panoramic image stitching device integrating tone correction and POI anchoring may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) that can perform various appropriate actions and processes according to a program stored in a read-only memory or a program loaded from a storage device into a random access memory. The random access memory also stores various programs and data required for the operation of the panoramic image stitching device integrating tone correction and POI anchoring. The processing unit, read-only memory, and random access memory are interconnected via a bus. An I / O interface is also connected to the bus. Typically, the following systems can be connected to the I / O interface: input devices including touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices including liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices including magnetic tapes, hard disks, etc.; and communication devices. The communication device allows the panoramic image stitching device integrating tone correction and POI anchoring to communicate wirelessly or wiredly with other devices to exchange data. While a panoramic image stitching device integrating tone correction and POI anchoring with various systems has been described, it should be understood that implementation of all the systems described is not required. Alternatively, more or fewer systems may be implemented.
[0093] Example 4: This invention also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the panoramic image stitching method for fusing tone correction and POI anchoring as described above. The computer program product provided by this invention can solve the technical problem of panoramic image stitching for fusing tone correction and POI anchoring. Compared with the prior art, the beneficial effects of the computer program product provided by this invention are the same as the beneficial effects of the panoramic image stitching method for fusing tone correction and POI anchoring provided in the above embodiments, and will not be repeated here.
[0094] In particular, according to the embodiments disclosed in this invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device, or installed from a read-only memory. When the computer program is executed by a processing device, it performs the functions defined in the methods of the embodiments disclosed in this invention.
[0095] It should be understood that the various parts disclosed in this invention can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics may be combined in any suitable manner in one or more embodiments or examples.
[0096] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the present invention and its equivalents, the present invention also intends to include these modifications and variations.
Claims
1. A panoramic image stitching method integrating tone correction and POI anchoring, characterized in that, The method includes: Step S10: Obtain panoramic stitching source data, and based on the panoramic stitching source data, perform cross-view anchor point construction task using POI projection anchoring mechanism, and output cross-view POI anchor point data; Step S20: Based on the cross-view POI anchor data, the Huber robust credibility assessment mechanism is used to perform anchor filtering and initial registration tasks, and output robust anchor evaluation data; Step S30: Based on the robust anchor point evaluation data, the initial tone correction task is performed using the thin plate spline local tone correction mechanism, and the initial tone correction data is output. Step S40: Based on the initial tone correction data, perform a joint geometry and tone optimization task using the POI-constrained Levenberg-Marquardt alternating optimization mechanism, and output the joint optimization data; Step S50: Based on the joint optimization data, the panoramic fusion and POI location binding tasks are performed using the POI confidence-weighted Laplacian pyramid fusion mechanism, and the panoramic stitching result data is output.
2. The panoramic image stitching method integrating tone correction and POI anchoring as described in claim 1, characterized in that, Step S10, which involves acquiring panoramic stitching source data, and based on this data, performing a cross-view anchor point construction task using a POI projection anchoring mechanism to output cross-view POI anchor point data, specifically includes: Step S101: Obtain multi-view images, shooting poses and camera intrinsic parameter matrices corresponding to each multi-view image, and POI data. The POI data includes POI geographic coordinates, POI semantic categories, and POI reference structure descriptors. Combine the multi-view images, shooting poses, camera intrinsic parameter matrices, and POI data to form the panoramic stitching source data. Step S102: Based on the panoramic stitching source data, the geographic coordinates of the POI are projected onto the corresponding multi-view image according to the shooting pose and the camera intrinsic parameter matrix to obtain the initial projection position of the POI; the local structure descriptor of the image is extracted in the preset neighborhood of the initial projection position of the POI, and the local structure descriptor of the image is matched with the corresponding POI reference structure descriptor to obtain the initial POI anchor point data. Step S103: Based on the initial POI anchor point data, determine the visibility markers according to whether the projection positions of each POI in the initial POI anchor point data are located within the effective image plane and whether they are occluded. Combine the POI projection positions visible in at least two multi-view images, the image local structure descriptor, the POI semantic category, the visibility markers, and the panoramic stitching source data to form the cross-view POI anchor point data.
3. The panoramic image stitching method integrating tone correction and POI anchoring as described in claim 1, characterized in that, Step S20, based on the cross-view POI anchor point data, employs the Huber robust credibility assessment mechanism to perform anchor point filtering and initial registration tasks, and outputs robust anchor point assessment data. This step specifically includes: Step S201: Based on the cross-view POI anchor point data, determine the reprojection error, semantic consistency, occlusion degree and local structural stability of each POI anchor point, and combine the reprojection error, semantic consistency, occlusion degree and local structural stability to form anchor point quality feature data; Step S202: Based on the anchor point quality feature data, the reprojection error, semantic consistency, occlusion degree, and local structural stability are fused using preset weights to obtain an anchor point credibility score; POI anchor points with an anchor point credibility score greater than a preset high credibility threshold are included in a high credibility anchor point set, POI anchor points with an anchor point credibility score less than a preset rejection threshold are rejected, and the anchor point credibility score and the high credibility anchor point set are combined to form anchor point screening data; Step S203: Based on the anchor point screening data, the anchor point confidence score is used as the registration weight of the corresponding POI anchor point. The Huber robust kernel is used to minimize the weighted projection error of the high-confidence anchor point set to obtain the initial image transformation parameters. The cross-view POI anchor point data, the anchor point confidence score, the high-confidence anchor point set and the initial image transformation parameters are combined to form the robust anchor point evaluation data.
4. The panoramic image stitching method integrating tone correction and POI anchoring as described in claim 3, characterized in that, In step S202, the anchor point confidence score is calculated based on the anchor point quality feature data. The anchor point confidence score is: in, For the first The first multi-view image Anchor credibility score for each POI anchor point For multi-view image indexing, For POI indexing, For the Sigmoid function, The reprojection error is expressed in pixels. For the sake of dimensionless semantic consistency. The degree of occlusion is dimensionless. The local structural stability is dimensionless. This is a dimensionless preset bias coefficient. These are preset weighting coefficients that match the dimensions of the reprojection error. , and The preset weight coefficients are dimensionless; the registration weights and credibility categories of each POI anchor point are determined using the anchor point credibility scores to obtain the anchor point screening data.
5. The panoramic image stitching method integrating tone correction and POI anchoring as described in claim 3, characterized in that, Step S30, based on the robust anchor point evaluation data, involves performing an initial tone correction task using a thin-plate spline local tone correction mechanism and outputting the initial tone correction data. This step specifically includes: Step S301: Obtain the cross-view POI anchor data, the high-confidence anchor set, the anchor confidence score, and the initial image transformation parameters from the robust anchor evaluation data; obtain multi-view images and POI semantic categories from the cross-view POI anchor data; establish a coordinate mapping from each multi-view image to the corresponding reference view using the initial image transformation parameters; in the YCbCr color space, use robust color weights composed of distance weights, gradient weights, saturation weights, and structure weights to calculate the weighted color statistics of each POI neighborhood in each color channel within the high-confidence anchor set, and obtain POI neighborhood color statistics data. Step S302: Based on the color statistics of the POI neighborhood, determine the saturation anomaly ratio and local structural integrity of each POI neighborhood. According to the anchor confidence score, the saturation anomaly ratio, the local structural integrity, and the POI semantic category, determine the reference view from the multi-view images corresponding to anchors belonging to the same POI within the high-confidence anchor set. Establish a local linear color correspondence relationship between each POI neighborhood in the corresponding multi-view image and the reference view based on the weighted color statistics of each POI neighborhood in the POI neighborhood color statistics, and obtain local tone parameter data. The local tone parameter data includes the local gain and local bias of each color channel. Step S303: Based on the local tone parameter data, using the position of each POI anchor point in the high-confidence anchor point set in the corresponding multi-view image as the interpolation node, spatial interpolation is performed on the local gain and the local bias using a thin plate spline kernel to obtain an initial tone compensation field; the initial tone compensation field is used to perform tone correction on the multi-view image to obtain a tone-corrected image; the tone-corrected image, the initial tone compensation field, and the robust anchor point evaluation data are combined to form the initial tone correction data.
6. The panoramic image stitching method integrating tone correction and POI anchoring as described in claim 5, characterized in that, Step S40, based on the initial tone correction data, involves performing a joint geometry and tone optimization task using a Levenberg-Marquardt alternating optimization mechanism with POI constraints, and outputting the joint optimization data. This step specifically includes: Step S401: Obtain the cross-view POI anchor point data from the initial tone correction data, and obtain the corresponding POI projection position from the cross-view POI anchor point data; within the preset search window of the tone-corrected image, calculate the image matching response and semantic matching response of the candidate positions, and determine the POI repositioning position based on the image matching response and the semantic matching response; determine the local occlusion degree at the POI repositioning position, and determine the POI repositioning confidence based on the image matching response, the semantic matching response, and the local occlusion degree; combine the POI repositioning position, the POI repositioning confidence, and the initial tone correction data to form POI constraint initialization data; Step S402: Based on the POI constraint initialization data, construct a joint optimization objective including POI position drift term, overlapping area color difference term, seam residual term, and tone field smoothing term. The above four terms constrain the deviation between the repositioned POI position and the corresponding POI projection position, the overlapping area color difference, the seam continuity, and the tone compensation field smoothness in sequence. Combine the above four terms with preset weights to obtain the joint optimization objective data. Step S403: Based on the joint optimization target data, using the initial image transformation parameters and the initial tone compensation field as initial values, the Levenberg-Marquardt method is used to update one type of parameters while fixing another type of parameters until the joint optimization target satisfies the preset convergence condition, thereby obtaining the optimized image transformation parameters and the optimized tone compensation field; the optimized image transformation parameters, the optimized tone compensation field, the POI repositioning location, the POI repositioning confidence, and the initial tone correction data are combined to form the joint optimization data.
7. The panoramic image stitching method integrating tone correction and POI anchoring as described in claim 6, characterized in that, Step S50, based on the joint optimization data, employs a POI confidence-weighted Laplacian pyramid fusion mechanism to perform panoramic fusion and POI location binding tasks, and outputs panoramic stitching result data. This step specifically includes: Step S501: Obtain the optimized image transformation parameters, the optimized tone compensation field, the POI relocation location, the POI relocation confidence, and the cross-view POI anchor point data from the joint optimization data; and obtain the multi-view images, POI geographic coordinates, and the set of visible POIs in each multi-view image from the cross-view POI anchor point data; generate the final tone compensation image corresponding to each multi-view image based on the optimized image transformation parameters and the optimized tone compensation field, and determine the pixel position of the final tone compensation image. Final fusion weights at the point: in, For the first The final tone-compensated image at pixel location The dimensionless final fusion weights at the point are For multi-view image indexing, For POI indexing, The dimensionless pixel position to the first Distance weights for the boundaries of the final tone-compensated image. The gradient magnitude weights at the pixel positions are dimensionless. The dimensionless preset POI fusion weight enhancement coefficient. For the first The set of visible POIs in a multi-view image. For the first The first multi-view image Dimensionless POI relocation reliability for each POI The first one, in pixels Relocate the POI of each POI. The pixel position is expressed in pixels. The preset POI influence radius in pixels. This indicates that the summation is performed on the POIs in the set of visible POIs. It is an exponential function. The L2 norm is used; the final tone-compensated image and the final fusion weights are combined to form panoramic fusion preparation data; Step S502: Based on the panoramic fusion preparation data, construct a Laplacian pyramid for the final tone compensation image, construct a Gaussian pyramid of the corresponding level for the final fusion weights, perform weighted fusion and reconstruction according to the corresponding level to obtain a panoramic image; calculate the image residuals on both sides of the stitching seam in the panoramic image, and combine the panoramic image and the image residuals to form panoramic fusion intermediate data; Step S503: Based on the panoramic fusion intermediate data, when the image residual is greater than the preset stitching seam residual threshold, adjust the final fusion weight or perform local offset correction on the optimized tone compensation field; obtain the POI repositioning confidence in each multi-view image for the same POI, weight the projection position of the POI geographic coordinates in each multi-view image according to the POI repositioning confidence, convert the weighting result to panoramic coordinates, and obtain the POI position bound to the panoramic coordinates; combine the panoramic image and the POI position bound to the panoramic coordinates to form the panoramic stitching result data.
8. A panoramic image stitching system integrating tone correction and POI anchoring, applied to the panoramic image stitching method integrating tone correction and POI anchoring as described in any one of claims 1 to 7, characterized in that, The system includes: The POI anchor point construction module is used to acquire panoramic stitching source data, and based on the panoramic stitching source data, execute the cross-view anchor point construction task using the POI projection anchoring mechanism, and output the cross-view POI anchor point data. The robust anchor evaluation module is used to perform the anchor selection and initial registration tasks based on the cross-view POI anchor data and the Huber robust credibility evaluation mechanism, and output the robust anchor evaluation data. The initial tone correction module is used to perform the initial tone correction task based on the robust anchor point evaluation data and the thin plate spline local tone correction mechanism, and output the initial tone correction data. The joint optimization module is used to perform the geometry and tone joint optimization task based on the initial tone correction data and using the Levenberg-Marquardt alternating optimization mechanism with the POI constraint, and output the joint optimization data. The panoramic fusion module is used to perform the panoramic fusion and POI location binding task based on the joint optimized data and the POI confidence-weighted Laplacian pyramid fusion mechanism, and output the panoramic stitching result data.
9. A panoramic image stitching device integrating tone correction and POI anchoring, characterized in that, The panoramic image stitching device for fusion tone correction and POI anchoring includes: a memory, a processor, and a panoramic image stitching program for fusion tone correction and POI anchoring stored in the memory and executable on the processor. When the panoramic image stitching program for fusion tone correction and POI anchoring is executed by the processor, it implements a panoramic image stitching method for fusion tone correction and POI anchoring as described in any one of claims 1 to 7.
10. A computer program product, characterized in that, The computer program product includes a panoramic image stitching program that integrates tone correction and POI anchoring. When the panoramic image stitching program that integrates tone correction and POI anchoring is executed by the processor, it implements a panoramic image stitching method that integrates tone correction and POI anchoring as described in any one of claims 1 to 7.