Tower base video image automatic positioning method and system
By automatically calibrating the horizontal parameters and corner coordinate mapping of the tower base monitoring equipment and combining it with a registration algorithm, precise positioning of the tower base video image is achieved, solving the problems of low positioning accuracy and reliance on manual operation in existing technologies, and improving the accuracy and reliability of the monitoring system.
Patent Information
- Application Number
- CN202510918757.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-07-04
AI Technical Summary
The existing tower base video image positioning method has low accuracy, relies on manual operation and is subject to limited conditions, making it difficult to achieve fast and accurate positioning.
By acquiring the tower base video image containing calibration markers, the horizontal parameter deviation of the monitoring equipment is calculated using image processing technology for automatic calibration. Combined with corner coordinate mapping and alignment algorithms, precise positioning of the tower base video image can be achieved.
It achieves fast and precise positioning of tower base video images, overcomes the problems of low positioning accuracy and reliance on manual operation in existing technologies, and improves the accuracy and reliability of the monitoring system.
Smart Images

Figure CN120431165B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a tower base video image automatic positioning method and system, belonging to the technical field of satellite remote sensing imaging. Background Art
[0002] Tower-based remote sensing, a near-ground remote sensing observation method based on elevated platforms such as communication towers, power towers, and streetlight poles, effectively captures ground-level details thanks to its high temporal and spatial resolution and relatively low observation altitude. It is widely used in disaster monitoring, traffic management, environmental pollution monitoring, and smart city development. Compared to satellite and drone remote sensing, tower-based remote sensing provides highly accurate, dynamic, and time-series observation data, compensating for the lower temporal and spatial resolution of satellites while overcoming the limitations of drones, which struggle with sustained operation and are susceptible to weather conditions, complex terrain, and airspace regulations.
[0003] Intelligent analysis technologies for tower-based video images (such as target detection, environmental monitoring, and disaster warning) have already formed a relatively complete technical system in engineering practice. However, these technologies are based solely on the semantic information of video images, and how to quickly and accurately locate the target scene remains a key issue restricting the development of tower-based video.
[0004] Currently, methods for positioning tower base video images can be roughly divided into two categories: 2D homography-based methods and multi-point positioning-based methods. The 2D homography-based method relies on manually selecting corresponding points in surveillance video and high-definition remote sensing satellite imagery, calculating the homography matrix, and converting the surveillance video image coordinates into 2D rectangular coordinates. However, this method is not only time-consuming and labor-intensive, but also has low accuracy. The multi-point positioning method uses digital pan-tilt heads on different towers instead of angle measurement tools to determine point positions using the two-point intersection method. However, due to the relatively scattered distribution of tower base video points and the arbitrary setting of the horizontal parameter starting point and rotation direction of some tower base monitoring, the accuracy of the multi-point positioning method is hampered. Summary of the Invention
[0005] In order to solve the above problems, the present invention proposes a method and system for automatic positioning of tower base video images, which can achieve accurate positioning of tower base video images, overcoming the problems of low positioning accuracy, dependence on manual labor and conditional limitations in the existing technology.
[0006] The technical solution adopted by the present invention to solve its technical problems is:
[0007] In a first aspect, an embodiment of the present invention provides a method for automatically positioning a tower base video image, comprising the following steps:
[0008] Step S1, obtaining a video image of the tower base including calibration markers, calculating the horizontal parameter deviation of the monitoring equipment through image processing technology, and automatically calibrating the horizontal parameters of the monitoring equipment; the calibration markers are calibration markers set at multiple known positions around the tower base;
[0009] Step S2: Correct the viewing angle of the tower base video image according to the parameters of the monitoring equipment, and determine the rough position of the tower base video image in the satellite image by corner point coordinate mapping;
[0010] In step S3, a registration algorithm is used to register the corrected tower base video image with the satellite image, and a registration mapping relationship is established to achieve accurate positioning of the tower base video image.
[0011] As a possible implementation of this embodiment, step S1 includes the following steps:
[0012] Step S11, selecting two images taken from the same monitoring source at different horizontal angles, and recording the initial horizontal parameters of the two images;
[0013] Step S12, increasing the horizontal parameters of the two images by a set angle in sequence, numbering them, and performing perspective correction and rough positioning on the images using the changed horizontal parameter values;
[0014] Step S13, matching the tower base video correction image and the satellite coarse positioning image corresponding to each increased horizontal parameter using a matching module to find the optimal matching pair, and calculating the rotation angle between the matching pairs according to the affine transformation matrix;
[0015] Step S14, determine whether the rotation angles between the optimal matching pairs of the two images are equal within the allowable error range. If they are equal within the allowable error range, it is considered that the monitor is rotating clockwise, and the horizontal parameter is the angle between the true north direction and zero degrees in the increased horizontal parameter; if they are not equal, it is considered that the monitor is rotating counterclockwise, and the horizontal parameter is the other angle between the true north direction and zero degrees in the increased horizontal parameter.
[0016] As a possible implementation method of this embodiment, in step S12, the angle increased each time is a fixed value, and after each increase, the image is corrected for perspective and roughly positioned, and the image pairs with corresponding numbers are matched until the number of matching points reaches a peak and the number of matching points of two adjacent numbers changes continuously, and the horizontal parameters and numbers at this time are recorded.
[0017] As a possible implementation of this embodiment, in step S13, the process of calculating the rotation angle between the matching pairs according to the affine transformation matrix includes:
[0018] Further decomposition of the affine transformation matrix:
[0019] ,
[0020] Where, is the rotation difference, yes The scale difference in direction, yes The scale difference of the direction, is the shear coefficient, =0;
[0021] The rotation angle between the optimal matching pairs is obtained by calculation :
[0022] ,
[0023] if , the registered image is rotated counterclockwise relative to the reference image ;if , the registered image rotates clockwise relative to the reference image .
[0024] As a possible implementation of this embodiment, in step S14, the allowable error range is within 5 degrees.
[0025] As a possible implementation of this embodiment, step S1 further includes the following steps:
[0026] Step S15 , recording the parameters obtained after the adjustment of the first picture and the angle displayed in the monitoring system, as well as the parameters obtained after the adjustment of the second picture and the angle displayed in the monitoring system.
[0027] As a possible implementation of this embodiment, step S2 includes the following steps:
[0028] Step S21, using the initial data of the tower base video image and the collinearity equation, converting the four corner points of the tower base video image from image plane coordinates to spatial coordinates, and completing the coarse positioning of the tower base video data with the circumscribed rectangle of the four corner points as the boundary, the initial data includes the tower coordinates and height, horizontal parameters, vertical parameters and internal parameters of the monitoring camera;
[0029] Step S22, taking the minimum x and y coordinates of the four corner points corresponding to the ground coordinates as the starting point, setting the target spatial resolution as the sampling interval, and calculating the corresponding coordinates of each ground coordinate in the original tower base video image using the collinearity equation;
[0030] Step S23 , resampling is performed and pixel values are assigned to corresponding points of the corrected tower base video image to obtain a perspective correction result of the tower base video image.
[0031] As a possible implementation of this embodiment, the initial data also includes the width and height of the original image and the pixel size. The collinearity equation converts the scanning coordinates into image plane coordinates using affine transformation parameters, and then converts the image plane coordinates into spatial coordinates using a rotation matrix and the camera focal length.
[0032] The affine transformation parameters include a translation parameter and a scaling parameter, wherein the translation parameter is 0 and the scaling parameter is equal to the pixel size;
[0033] The rotation matrix is calculated by the following formula:
[0034] ,
[0035] The rotation matrix parameters include the horizontal parameter A of rotation around the Z axis and the complementary angle B of the vertical parameter of rotation around the Y axis, where the horizontal parameter A is the angle of clockwise rotation of the monitoring device around the Z axis, and the vertical parameter complementary angle B is the angle of counterclockwise rotation of the monitoring device around the Y axis.
[0036] As a possible implementation of this embodiment, the target spatial resolution is a preset sampling interval, and the sampling starting point is the smallest x and y coordinates among the ground coordinates of the four corner points.
[0037] As a possible implementation of this embodiment, the resampling adopts a bilinear interpolation method or a nearest neighbor interpolation method to obtain the pixel value corresponding to the coordinate position.
[0038] As a possible implementation of this embodiment, step S3 includes the following steps:
[0039] Step S311: super-resolution processing is performed on the satellite image using a pre-trained super-resolution reconstruction model to achieve resolution alignment between the satellite image and the tower base video image;
[0040] Step S312, using the DINOv2 model to extract coarse features of the satellite image and the tower base video image;
[0041] Step S313: Refine the coarse features using the fine features extracted by the VGG19 model at multiple scales to construct a feature pyramid;
[0042] Step S314: perform coarse matching based on the coarse features through a Transformer-based matching decoder, discretize the output space into evenly distributed anchor points, and predict the matching probability of each anchor point through classification;
[0043] Step S315, using the RANSAC method to eliminate incorrectly matched feature point pairs;
[0044] Step S316: Using the Delaunay triangulation constraint strategy to eliminate incorrectly matched triangle pairs, obtain matching point pairs, and improve matching robustness and accuracy;
[0045] In step S317, the obtained matching point pairs are mapped from the corrected image to the original image to achieve registration of the satellite image and the tower base video image.
[0046] As a possible implementation of this embodiment, in step 312, the DINOv2 model is used as a frozen coarse feature extractor to extract coarse features at a scale of 1 / 16.
[0047] As a possible implementation of this embodiment, in step 313, the VGG19 model is used to extract fine features at scales of 1 / 2, 1 / 4, and 1 / 8, which are combined with the coarse features extracted by the DINOv2 model to construct a feature pyramid.
[0048] As a possible implementation of this embodiment, in step 314, the Transformer-based matching decoder consists of 5 ViT blocks, each ViT block contains 8 attention heads, the hidden layer size D is set to 1024, the MLP size is 4096, the input is a vector concatenated from the DINOv2 feature and the feature output of the Gaussian process module, and the output is a matching probability vector of the classification anchor.
[0049] As a possible implementation of this embodiment, in step 314, a regression-classification loss function is used to construct a coarse matching loss function by minimizing the Kullback–Leibler divergence between the estimated matching distribution and the theoretical model distribution.
[0050] As a possible implementation of this embodiment, in step 315, the RANSAC method screens out inliers from noisy data and estimates optimal model parameters through iterative random sampling and consistency verification.
[0051] As a possible implementation of this embodiment, in step 316, the matching triangle pairs are detected using side length and angle constraints, and incorrect matching points that do not meet the constraints are eliminated.
[0052] As a possible implementation method of this embodiment, in step 317, the coordinate relationship between the original tower base video image and the corrected image in the tower base video image correction module is used to map the matching points from the corrected image to the original image, thereby achieving accurate alignment of the satellite image and the tower base video image.
[0053] As another possible implementation of this embodiment, step S3 may further include the following steps:
[0054] Step S321: Super-resolution reconstruction of the satellite image is performed using a pre-trained super-resolution reconstruction model to achieve resolution alignment between the satellite image and the tower base video image.
[0055] Step S322: Use a deep learning intensive matching model to match the satellite image with the tower base video image to obtain a preliminary matching result, wherein the deep learning intensive matching model includes:
[0056] A coarse feature encoder that uses a pre-trained DINOv2 model to extract coarse features from the input satellite imagery and tower-based video imagery;
[0057] Fine feature encoder, which uses the VGG19 model to extract fine features from the input satellite imagery and tower base video imagery;
[0058] Feature pyramid construction module, which combines the coarse features extracted by the DINOv2 model and the fine features extracted by the VGG19 model to construct a feature pyramid;
[0059] In the coarse matching stage, a Transformer-based matching decoder is used. It is designed using a regression-classification formula, discretizes the output space into a set of evenly distributed anchor points, and predicts the probability of each anchor point through classification to obtain the coarse matching result.
[0060] In the refinement stage, the rough matching results are locally adjusted to obtain the refined matching results;
[0061] And, add the coarse matching loss and the refined loss to obtain the total loss function to train the deep learning dense matching model;
[0062] Step S323: Using the RANSAC algorithm to remove incorrectly matched feature point pairs from the obtained preliminary matching results;
[0063] Step S324: Using the Delaunay triangulation constraint strategy to eliminate incorrectly matched triangle pairs, obtain matching point pairs, and improve matching robustness and accuracy;
[0064] In step S325 , the obtained matching point pairs are mapped from the corrected image to the original image to achieve registration of the satellite image and the tower base video image.
[0065] As a possible implementation of this embodiment, in the coarse matching stage, the Transformer-based matching decoder consists of 5 ViT blocks, each ViT block contains 8 attention heads, the hidden layer size D is set to 1024, the MLP size is 4096, and its input is the concatenation of the 512-dimensional projected DINOv2 features and the 512-dimensional features output by the Gaussian process module. The output is a vector, the number of vectors is equal to the number of classification anchors, and the additional "1" dimension is the matchability score.
[0066] As a possible implementation of this embodiment, the input image size of the coarse feature encoder DINOv2 model is a multiple of 14 to meet the needs of multi-scale feature extraction and is used to extract coarse features at a scale of 1 / 16.
[0067] As a possible implementation of this embodiment, the fine feature encoder VGG19 model is used to extract fine features at scales such as 1 / 2, 1 / 4, and 1 / 8.
[0068] As a possible implementation of this embodiment, when training the deep learning dense matching model, the parameters of the DINOv2 model remain frozen and are not fine-tuned; for the decoder, 10 -4 The learning rate is 5×10 for the encoder. -6 The learning rate is set to 0 and scaled linearly with the batch size; training is performed for 100 epochs on the MegaDepth dataset.
[0069] As a possible implementation of this embodiment, side length and angle constraints are used to detect mismatched triangle pairs and eliminate mismatched points. For matching triangles, the three corresponding sides and three corresponding angles meet the preset side length and angle constraint thresholds.
[0070] As a possible implementation of this embodiment, the training dataset of the deep learning dense matching model is the MegaDepth dataset, which contains image pairs with depth annotations collected from Internet photos. The number of image pairs is up to millions, and the depth map resolution varies, usually between 640×480 and 1920×1080.
[0071] In a second aspect, an embodiment of the present invention provides a tower base video image automatic positioning system, comprising:
[0072] A horizontal parameter calibration module is used to obtain a video image of the tower base containing calibration markers set at multiple known locations around the tower base, calculate the horizontal parameter deviation of the monitoring equipment through image processing technology, and automatically calibrate the horizontal parameters of the monitoring equipment;
[0073] The coarse positioning and perspective correction module is used to correct the perspective of the tower base video image according to the parameters of the monitoring equipment, and determine the rough position of the tower base video image in the satellite image by mapping the corner point coordinates;
[0074] The registration module is used to register the corrected tower base video image with the satellite image using a registration algorithm, establish a registration mapping relationship, and achieve accurate positioning of the tower base video image.
[0075] The beneficial effects of the technical solutions of the embodiments of the present invention are as follows:
[0076] The technical solution of an embodiment of the present invention provides a method for automatic positioning of tower-based video images. First, by seeking the connection between the monitoring image and the ground satellite image, the automatic calibration of the monitoring level parameters is realized; secondly, the rough positioning of the tower-based video image is realized by mapping the corner coordinates using known parameters, and the inverse digital differential correction method is improved to reduce the perspective difference between the tower-based video image and the satellite image; finally, the deep learning dense matching model with resolution alignment is used to improve the adaptability of the registration module to complex scenes, and the automatic 2D homography registration of the tower-based video image and the satellite image is realized to complete the precise positioning. The present invention combines satellite images and tower-based videos to realize the rapid and precise positioning of the target scene in the tower-based video image, overcomes the problems of low positioning accuracy, dependence on manual labor and conditional limitations in the existing technology, and provides strong support for the development of the field of tower-based remote sensing observation.
[0077] In response to the problem that the initial position and rotation direction of some tower base monitoring horizontal parameters are randomly set, the present invention proposes a calibration method for monitoring horizontal parameters. By using two images of the same monitoring to seek the connection between the monitoring image and the ground satellite image, the monitoring rotation direction and the compensation value from the true north direction to the monitoring zero degree direction are judged and calculated, thereby realizing automatic calibration of the monitoring horizontal parameters and effectively utilizing the tower base monitoring with randomly set horizontal parameters and rotation direction.
[0078] The automatic positioning process of the tower base video data of the present invention, from coarse positioning to registration and precise positioning, makes full use of the known data of tower base monitoring. The coarse position of the tower base video image is determined by using corner point coordinate mapping based on known parameters (tower coordinates and height, horizontal parameters and vertical parameters of tower base video data), thereby realizing coarse positioning of the tower base video image.
[0079] In view of the differences between tower base monitoring posture data and aerial photography posture data, the present invention adopts a rotation matrix calculation method suitable for the tower base monitoring system, and combines perspective correction, resolution alignment, deep learning dense matching model, RANSAC algorithm and Delaunay triangulation constraint strategy to achieve satellite-assisted precise positioning of tower base video data in complex scenarios.
[0080] During the registration phase, the present invention first improves the inverse digital differential correction (IDD) algorithm used for aerial and drone imagery, enabling it to be used for tower base video imagery, thereby minimizing the impact of perspective differences on tower base video and satellite imagery. The present invention also uses a resolution-aligned deep learning dense matching model to improve the registration module's adaptability to complex scenes, enabling automatic 2D homography registration of tower base video imagery. Finally, the matching points are mapped back to the original tower base video imagery, achieving precise positioning. BRIEF DESCRIPTION OF THE DRAWINGS
[0081] Figure 1 This is a flow chart of a method for automatically positioning a tower base video image according to an exemplary embodiment;
[0082] Figure 2 is a flow chart showing a first method of coarse positioning and perspective correction according to an exemplary embodiment;
[0083] Figure 3 is a flow chart of a second method for coarse positioning and viewing angle correction according to an exemplary embodiment;
[0084] Figure 4 1 is a schematic structural diagram of a tower base video image automatic positioning system according to an exemplary embodiment;
[0085] Figure 5 This is a specific implementation flow chart of tower base video image positioning using the method of the present invention;
[0086] Figure 6 is a flow chart of a horizontal angle calibration method according to an exemplary embodiment;
[0087] Figure 7 is a flow chart of a coarse positioning and perspective correction method according to an exemplary embodiment;
[0088] Figure 8 is a flow chart of a registration method according to an exemplary embodiment;
[0089] Figure 9 is a schematic diagram showing a first monitoring clockwise rotation according to an exemplary embodiment;
[0090] Figure 10 is a schematic diagram showing a second monitoring clockwise rotation according to an exemplary embodiment;
[0091] Figure 11 is a schematic diagram showing a counterclockwise rotation monitoring method according to an exemplary embodiment;
[0092] Figure 12 The figure is a schematic diagram showing a perspective correction of a tower base video image according to an exemplary embodiment. DETAILED DESCRIPTION
[0093] In order to more clearly illustrate the technical features of the present invention, the present invention is described in detail below through specific implementation methods and in conjunction with the accompanying drawings.
[0094] Satellite images not only have wide-area coverage capabilities, but also contain precise geographic coordinates and orientation information, providing a reliable geographic reference for the spatial positioning of scene targets, and making up for the shortcomings of tower-based videos in terms of geographic positioning. The present invention proposes a tower-based monitoring horizontal parameter correction method to achieve horizontal parameter correction of the tower-based monitoring system, and further proposes a tower-based video image automatic positioning method based on satellite image assistance to achieve precise positioning of tower-based video images. By fusing information from both satellite images and tower-based videos, the present invention can not only more accurately determine the actual geographic location of the target, but also obtain rich dynamic information in real time, enabling the monitoring system to more efficiently perform target tracking and situational awareness in complex environments, thereby providing more accurate decision-making support for tasks such as emergency response, disaster monitoring, and urban management.
[0095] like Figure 1 As shown, an embodiment of the present invention provides a tower base video image automatic positioning method, comprising the following steps:
[0096] Step S1, obtaining a video image of the tower base including calibration markers, calculating the horizontal parameter deviation of the monitoring equipment through image processing technology, and automatically calibrating the horizontal parameters of the monitoring equipment; the calibration markers are calibration markers set at multiple known positions around the tower base;
[0097] Step S2: Correct the viewing angle of the tower base video image according to the parameters of the monitoring equipment, and determine the rough position of the tower base video image in the satellite image by corner point coordinate mapping;
[0098] In step S3, a registration algorithm is used to register the corrected tower base video image with the satellite image, and a registration mapping relationship is established to achieve accurate positioning of the tower base video image.
[0099] As a possible implementation of this embodiment, step S1 includes the following steps:
[0100] Step S11, selecting two images taken from the same monitoring source at different horizontal angles, and recording the initial horizontal parameters of the two images;
[0101] Step S12, increasing the horizontal parameters of the two images by a set angle in sequence, numbering them, and performing perspective correction and rough positioning on the images using the changed horizontal parameter values;
[0102] Step S13, matching the tower base video correction image and the satellite coarse positioning image corresponding to each increased horizontal parameter using a matching module to find the optimal matching pair, and calculating the rotation angle between the matching pairs according to the affine transformation matrix;
[0103] Step S14, determine whether the rotation angles between the optimal matching pairs of the two images are equal within the allowable error range. If they are equal within the allowable error range, it is considered that the monitor is rotating clockwise, and the horizontal parameter is the angle between the true north direction and zero degrees in the increased horizontal parameter; if they are not equal, it is considered that the monitor is rotating counterclockwise, and the horizontal parameter is the other angle between the true north direction and zero degrees in the increased horizontal parameter.
[0104] As a possible implementation method of this embodiment, in step S12, the angle increased each time is a fixed value, and after each increase, the image is corrected for perspective and roughly positioned, and the image pairs with corresponding numbers are matched until the number of matching points reaches a peak and the number of matching points of two adjacent numbers changes continuously, and the horizontal parameters and numbers at this time are recorded.
[0105] As a possible implementation of this embodiment, in step S13, the process of calculating the rotation angle between the matching pairs according to the affine transformation matrix includes:
[0106] Further decomposition of the affine transformation matrix:
[0107] ,
[0108] Where, is the rotation difference, yes The scale difference in direction, yes The scale difference of the direction, is the shear coefficient, =0;
[0109] The rotation angle between the optimal matching pairs is obtained by calculation :
[0110] ,
[0111] if , the registered image is rotated counterclockwise relative to the reference image ;if , the registered image rotates clockwise relative to the reference image .
[0112] As a possible implementation of this embodiment, in step S14, the allowable error range is within 5 degrees.
[0113] As a possible implementation of this embodiment, step S1 further includes the following steps:
[0114] Step S15 , recording the parameters obtained after the adjustment of the first picture and the angle displayed in the monitoring system, as well as the parameters obtained after the adjustment of the second picture and the angle displayed in the monitoring system.
[0115] The present invention realizes the automatic calibration of tower base monitoring horizontal parameters through an automated calibration method, solves the problem that horizontal parameters in some monitoring systems lose their reference significance, improves the accuracy and reliability of monitoring data, and provides strong support for the optimization and upgrading of the monitoring system.
[0116] As a possible implementation of this embodiment, step S2 includes the following steps:
[0117] Step S21, using the initial data of the tower base video image and the collinearity equation, converting the four corner points of the tower base video image from image plane coordinates to spatial coordinates, and completing the coarse positioning of the tower base video data with the circumscribed rectangle of the four corner points as the boundary, the initial data includes the tower coordinates and height, horizontal parameters, vertical parameters and internal parameters of the monitoring camera;
[0118] Step S22, taking the minimum x and y coordinates of the four corner points corresponding to the ground coordinates as the starting point, setting the target spatial resolution as the sampling interval, and calculating the corresponding coordinates of each ground coordinate in the original tower base video image using the collinearity equation;
[0119] Step S23 , resampling is performed and pixel values are assigned to corresponding points of the corrected tower base video image to obtain a perspective correction result of the tower base video image.
[0120] As a possible implementation of this embodiment, the initial data also includes the width and height of the original image and the pixel size. The collinearity equation converts the scanning coordinates into image plane coordinates using affine transformation parameters, and then converts the image plane coordinates into spatial coordinates using a rotation matrix and the camera focal length.
[0121] The affine transformation parameters include a translation parameter and a scaling parameter, wherein the translation parameter is 0 and the scaling parameter is equal to the pixel size;
[0122] The rotation matrix is calculated by the following formula:
[0123] ,
[0124] The rotation matrix parameters include the horizontal parameter A of rotation around the Z axis and the complementary angle B of the vertical parameter of rotation around the Y axis, where the horizontal parameter A is the angle of clockwise rotation of the monitoring device around the Z axis, and the vertical parameter complementary angle B is the angle of counterclockwise rotation of the monitoring device around the Y axis.
[0125] As a possible implementation of this embodiment, the target spatial resolution is a preset sampling interval, and the sampling starting point is the smallest x and y coordinates among the ground coordinates of the four corner points.
[0126] As a possible implementation of this embodiment, the resampling adopts a bilinear interpolation method or a nearest neighbor interpolation method to obtain the pixel value corresponding to the coordinate position.
[0127] The present invention utilizes initial data and collinear equations for coarse positioning and view angle correction. The method is not only simple and efficient, reducing the amount of calculation, but also realizes real-time processing, improves the accuracy and practicality of monitoring video, and is suitable for practical application in tower-based monitoring systems.
[0128] As a possible implementation of this embodiment, Figure 2 As shown, step S3 includes the following steps:
[0129] Step S311: super-resolution processing is performed on the satellite image using a pre-trained super-resolution reconstruction model to achieve resolution alignment between the satellite image and the tower base video image;
[0130] Step S312, using the DINOv2 model to extract coarse features of the satellite image and the tower base video image;
[0131] Step S313: Refine the coarse features using the fine features extracted by the VGG19 model at multiple scales to construct a feature pyramid;
[0132] Step S314: perform coarse matching based on the coarse features through a Transformer-based matching decoder, discretize the output space into evenly distributed anchor points, and predict the matching probability of each anchor point through classification;
[0133] Step S315, using the RANSAC method to eliminate incorrectly matched feature point pairs;
[0134] Step S316: Using the Delaunay triangulation constraint strategy to eliminate incorrectly matched triangle pairs, obtain matching point pairs, and improve matching robustness and accuracy;
[0135] In step S317, the obtained matching point pairs are mapped from the corrected image to the original image to achieve registration of the satellite image and the tower base video image.
[0136] As a possible implementation of this embodiment, in step 312, the DINOv2 model is used as a frozen coarse feature extractor to extract coarse features at a scale of 1 / 16.
[0137] As a possible implementation of this embodiment, in step 313, the VGG19 model is used to extract fine features at scales of 1 / 2, 1 / 4, and 1 / 8, which are combined with the coarse features extracted by the DINOv2 model to construct a feature pyramid.
[0138] As a possible implementation of this embodiment, in step 314, the Transformer-based matching decoder consists of 5 ViT blocks, each ViT block contains 8 attention heads, the hidden layer size D is set to 1024, the MLP size is 4096, the input is a vector concatenated from the DINOv2 feature and the feature output of the Gaussian process module, and the output is a matching probability vector of the classification anchor.
[0139] As a possible implementation of this embodiment, in step 314, a regression-classification loss function is used to construct a coarse matching loss function by minimizing the Kullback–Leibler divergence between the estimated matching distribution and the theoretical model distribution.
[0140] As a possible implementation of this embodiment, in step 315, the RANSAC method screens out inliers from noisy data and estimates optimal model parameters through iterative random sampling and consistency verification.
[0141] As a possible implementation of this embodiment, in step 316, the matching triangle pairs are detected using side length and angle constraints, and incorrect matching points that do not meet the constraints are eliminated.
[0142] As a possible implementation method of this embodiment, in step 317, the coordinate relationship between the original tower base video image and the corrected image in the tower base video image correction module is used to map the matching points from the corrected image to the original image, thereby achieving accurate alignment of the satellite image and the tower base video image.
[0143] As another possible implementation of this embodiment, Figure 3 As shown, the step S3 may further include the following steps:
[0144] Step S321: Super-resolution reconstruction of the satellite image is performed using a pre-trained super-resolution reconstruction model to achieve resolution alignment between the satellite image and the tower base video image.
[0145] Step S322: Use a deep learning intensive matching model to match the satellite image with the tower base video image to obtain a preliminary matching result, wherein the deep learning intensive matching model includes:
[0146] A coarse feature encoder that uses a pre-trained DINOv2 model to extract coarse features from the input satellite imagery and tower-based video imagery;
[0147] Fine feature encoder, which uses the VGG19 model to extract fine features from the input satellite imagery and tower base video imagery;
[0148] Feature pyramid construction module, which combines the coarse features extracted by the DINOv2 model and the fine features extracted by the VGG19 model to construct a feature pyramid;
[0149] In the coarse matching stage, a Transformer-based matching decoder is used. It is designed using a regression-classification formula, discretizes the output space into a set of evenly distributed anchor points, and predicts the probability of each anchor point through classification to obtain the coarse matching result.
[0150] In the refinement stage, the rough matching results are locally adjusted to obtain the refined matching results;
[0151] And, add the coarse matching loss and the refined loss to obtain the total loss function to train the deep learning dense matching model;
[0152] Step S323: Using the RANSAC algorithm to remove incorrectly matched feature point pairs from the obtained preliminary matching results;
[0153] Step S324: Using the Delaunay triangulation constraint strategy to eliminate incorrectly matched triangle pairs, obtain matching point pairs, and improve matching robustness and accuracy;
[0154] In step S325 , the obtained matching point pairs are mapped from the corrected image to the original image to achieve registration of the satellite image and the tower base video image.
[0155] As a possible implementation of this embodiment, in the coarse matching stage, the Transformer-based matching decoder consists of 5 ViT blocks, each ViT block contains 8 attention heads, the hidden layer size D is set to 1024, the MLP size is 4096, and its input is the concatenation of the 512-dimensional projected DINOv2 features and the 512-dimensional features output by the Gaussian process module. The output is a vector, the number of vectors is equal to the number of classification anchors, and the additional "1" dimension is the matchability score.
[0156] As a possible implementation of this embodiment, the input image size of the coarse feature encoder DINOv2 model is a multiple of 14 to meet the needs of multi-scale feature extraction and is used to extract coarse features at a scale of 1 / 16.
[0157] As a possible implementation of this embodiment, the fine feature encoder VGG19 model is used to extract fine features at scales such as 1 / 2, 1 / 4, and 1 / 8.
[0158] As a possible implementation of this embodiment, when training the deep learning dense matching model, the parameters of the DINOv2 model remain frozen and are not fine-tuned; for the decoder, 10 -4 The learning rate is 5×10 for the encoder. -6 The learning rate is set to 0 and scaled linearly with the batch size; training is performed for 100 epochs on the MegaDepth dataset.
[0159] As a possible implementation of this embodiment, side length and angle constraints are used to detect mismatched triangle pairs and eliminate mismatched points. For matching triangles, the three corresponding sides and three corresponding angles meet the preset side length and angle constraint thresholds.
[0160] As a possible implementation of this embodiment, the training dataset of the deep learning dense matching model is the MegaDepth dataset, which contains image pairs with depth annotations collected from Internet photos. The number of image pairs is up to millions, and the depth map resolution is diverse, usually between 640×480 and 1920×1080.
[0161] The present invention realizes super-resolution processing of satellite images through a super-resolution reconstruction model, effectively reducing the resolution difference between satellite images and tower base video images, and providing a good foundation for the subsequent registration process; combining the DINOv2 model and the VGG19 model for feature extraction and refinement, constructing a feature pyramid with robustness and precise positioning capabilities, and improving the accuracy and robustness of registration; adopting the Transformer-based matching decoder and RANSAC method for coarse matching and false matching elimination, effectively improving the robustness and accuracy of matching; using the Delaunay triangulation constraint strategy to further enhance the robustness and accuracy of matching, and realizing accurate registration of satellite images with tower base video images.
[0162] like Figure 4 As shown, an embodiment of the present invention provides a tower base video image automatic positioning system, comprising:
[0163] A horizontal parameter calibration module is used to obtain a video image of the tower base containing calibration markers set at multiple known locations around the tower base, calculate the horizontal parameter deviation of the monitoring equipment through image processing technology, and automatically calibrate the horizontal parameters of the monitoring equipment;
[0164] The coarse positioning and perspective correction module is used to correct the perspective of the tower base video image according to the parameters of the monitoring equipment, and determine the rough position of the tower base video image in the satellite image by mapping the corner point coordinates;
[0165] The registration module is used to register the corrected tower base video image with the satellite image using a registration algorithm, establish a registration mapping relationship, and achieve accurate positioning of the tower base video image.
[0166] In response to the problem that the initial position and rotation direction of some tower base monitoring horizontal parameters are randomly set, the present invention proposes a calibration method for monitoring horizontal parameters, so that the problematic tower base monitoring can be effectively utilized. The present invention fully utilizes the known data of tower base monitoring to achieve efficient and accurate automatic positioning of tower base video images from coarse positioning to registration. In view of the difference between tower base monitoring posture data and aerial photography posture data, the present invention adopts a rotation matrix calculation method suitable for tower base monitoring systems, so that inverse digital differential correction can be effectively applied to tower base video images. At the same time, combined with perspective correction, resolution alignment, deep learning dense matching model, RANSAC algorithm and Delaunay triangulation constraint strategy, satellite-assisted precise positioning of tower base video data in complex scenarios is achieved.
[0167] like Figure 5 As shown, the specific implementation process of the tower base video image positioning method of the present invention is divided into three modules: a horizontal parameter calibration module, a coarse positioning and perspective correction module, and a registration module. The horizontal parameter calibration module implements automated calibration of the monitoring horizontal parameters, providing a guarantee for the implementation of subsequent modules; the coarse positioning and perspective correction module achieves rough positioning and perspective alignment of the tower base video image, reducing the impact of perspective differences on the tower base video image and satellite imagery; and the registration module registers the corrected tower base video image with satellite imagery, achieving precise positioning of the tower base video image through registration mapping.
[0168] like Figure 6 As shown in the present invention, the monitoring level parameter calibration process first selects two images of the same monitoring image with different horizontal angles. The horizontal parameters of the two images are By increasing the horizontal angle of the two images by one set angle , and number them. Then use the changed horizontal parameter values to perform coarse correction and coarse positioning using the perspective correction module. Use the matching module to match the tower base video correction image and the satellite coarse positioning image corresponding to each increased horizontal parameter to find the optimal matching pair and obtain the rotation angle between the matching pairs. , the matching pair is numbered If the two images If the error is within the allowable range (within 5 degrees), the monitor is considered to be rotating clockwise, and the horizontal parameter is , is the angle between true north and zero degrees. If they are not equal, it is considered that the monitor rotates counterclockwise, and the horizontal parameter is , , It is the angle from the true north to zero degrees. Through the automatic calibration of the monitored horizontal parameters, tower base monitoring with arbitrary horizontal parameters and rotation directions can be effectively utilized.
[0169] Error source explanation: Due to the complexity of the scene (weak texture and low-contrast repeated texture areas in many areas of the scene), the severe geometric distortion, occlusion differences and side interference information that still exist in the tower base video image after correction, the results obtained by the registration module have certain errors, which in turn causes certain errors in the affine transformation matrix obtained by registration, which ultimately affects At the same time, since the judgment requires two images from the same monitoring system, the two The value error will have an additive effect.
[0170] like Figure 7 As shown, the coarse positioning and perspective correction process of the present invention first uses the initial tower base video image data (such as tower coordinates and height, horizontal parameters, vertical parameters, and surveillance camera internal parameters) and collinearity equations to convert the four corner points of the tower base video image from image plane coordinates to spatial coordinates. The coarse positioning of the tower base video data is completed using the circumscribed rectangle of the four corner points as the bounding box. Next, starting from the minimum x and y coordinates of the four corner points corresponding to the ground coordinates, the target spatial resolution is set as the sampling interval, and the corresponding coordinates of each ground coordinate in the original tower base video image are calculated using the collinearity equation. Resampling is performed, and the pixel values are assigned to the corresponding points in the corrected tower base video image, thereby obtaining the perspective correction result for the tower base video image.
[0171] like Figure 8 As shown, the registration process of the present invention first super-resolutions the satellite imagery using a pre-trained super-resolution reconstruction model to align the two data resolutions and reduce the resolution difference between the two images. Subsequently, the DINOv2 model is used to extract coarse features from the image pair. These coarse features are then refined using fine features at multiple scales. Finally, the image pair is registered using the RANSAC method and Delaunay triangulation constraints.
[0172] 1. Tower base monitoring level parameter calibration method.
[0173] In monitoring systems, the starting point of the horizontal angle is usually set to the true north direction. However, the starting point and positive rotation direction of the horizontal angle in some monitoring systems are set arbitrarily, resulting in the horizontal parameters losing their reference significance. In order to improve the automation level of the processing flow, the present invention proposes a method for adjusting the horizontal parameters based on two images of the same monitoring source. Figure 6 As shown, the following are the detailed steps and principles:
[0174] In the first step, the image corrected according to the horizontal parameters returned by the monitoring system is matched with the satellite image of the corresponding ground range. If the matching fails, the starting point of the horizontal parameters is set arbitrarily.
[0175] The second step is to adjust based on the horizontal parameters of the monitoring system, increasing a fixed degree clockwise each time. , numbered as , according to the horizontal parameters adjusted each time, perform perspective correction and capture satellite images within the rough positioning range, and then match the image pairs with corresponding numbers until the number of matching points reaches a peak and the number of matching points between two adjacent numbers changes continuously, and record the number of matching points. and .
[0176] The third step is to obtain the affine transformation matrix of the optimal matching corresponding numbered image pair. , the affine transformation matrix can be further decomposed according to formula (1):
[0177] (1),
[0178] Where, is the rotation difference, yes The scale difference in direction, yes The scale difference of the direction, is the shear coefficient, Equal to 0. By calculation, the rotation angle between the optimal matching pairs can be obtained , as shown in formula (2):
[0179] (2),
[0180] like , the registered image is rotated counterclockwise relative to the reference image ;like , the registered image rotates clockwise relative to the reference image Since the satellite image is the reference image, , the satellite image rotates clockwise relative to the corrected image. , the satellite image is rotated counterclockwise relative to the corrected image.
[0181] In the fourth step, the parameters obtained after adjusting the first image of the same monitoring are recorded as , , The angle displayed in the monitoring system is recorded as The parameters obtained after adjusting the second picture are recorded as , , The angle displayed in the monitoring system is recorded as .
[0182] Figure 9 and Figure 10 This is a schematic diagram of the method assuming that the monitoring rotates clockwise. Figure 11 This is a schematic diagram of the method assuming that the monitoring is counterclockwise rotation. Figure 9 It can be seen that if the monitor rotates clockwise, the two images of the same monitor are equal, and are equal to the degree from the true north to the monitoring 0 degree. Figure 11 It can be seen that if the monitor rotates counterclockwise, The size of the monitored display angle The influence of image level parameters at different positions are different, so the images at different locations The size is different. Therefore, the rotation direction of the monitoring can be judged in this way. When they are equal, it is considered that the monitor is rotating clockwise, such as Figure 9 and Figure 10 As shown, otherwise it is considered to be rotating counterclockwise, such as Figure 11 If the monitor rotates clockwise, the horizontal parameter of the monitor after correction is ,in If the adjusted level parameter is not , convert it into this range, such as Figure 10 As shown. Figure 11 It can be seen that when the monitor rotates counterclockwise, the degree from the true north direction to the monitor 0 degree is equal to , then the horizontal parameter of the counterclockwise rotating monitor after correction is ,in ; If the adjusted level parameter is not , convert it to this range.
[0183] 2. Rough positioning and perspective correction of tower base video images.
[0184] Figure 12This is a schematic diagram of the perspective correction process for tower-base video images, where (a) represents the conversion process from tower-base video image scanning coordinates to image plane coordinates, (b) represents the conversion process from image plane coordinates to spatial coordinates, and (c) represents the conversion process from spatial coordinates to orthophoto coordinates. Scanning coordinates: The origin is the lower left corner of the image, with the Y axis pointing upward and the X axis pointing right. There are no units and only represent the pixel position. Image plane coordinates: The origin is the optical center of the image (which can be approximately considered the center of the photo), with the Y axis pointing upward and the X axis pointing right. The units are physical lengths. Spatial coordinates: Use the UTM coordinate system, with the X axis indicating due east, the Y axis indicating true north, and the Z axis indicating elevation. Figure 12 The asterisk in the middle indicates the change process of the corner point, and point p indicates the corresponding position of the point in the coordinates.
[0185] Assume that the four corner scanning coordinates of the original tower base video image are (0, 0), (Cols, 0), (0, Rows), and (Cols, Rows). The scanning coordinates of the four corner points on the original image are converted into image plane coordinates using formula (3):
[0186] (3),
[0187] in, and j is the coordinate value of the pixel in the scanning coordinate system, and is the coordinate value in the image plane coordinate system. and are the affine transformation parameters, and is 0. The width and height of the original image are and , the pixel size is ,but and equal , equal , equal .
[0188] The image plane coordinates of the four corner points are converted into image space coordinates through formula (4) to achieve the coarse positioning of the tower base video image:
[0189] (4),
[0190] in, and is the tower coordinate, It can be replaced by the height of the tower, is the focal length of the surveillance camera, can be approximated by the mean elevation, and are the coordinate values of the four corner points in the image plane coordinate system, are all parameters of the rotation matrix.
[0191] The expression of the exterior orientation elements of tower-based monitoring is different from that of aircraft and drones, and is usually expressed in horizontal parameters and vertical parameters. Assuming that the initial state of the monitoring device is vertical to the ground and starting from the true north direction, it needs to go through the following process in the process of rotating to the target position. First, the value of the horizontal parameter is rotated clockwise around the Z axis (usually the rotation direction of the monitoring device is clockwise). Then, the device will rotate counterclockwise around the Y axis by the complementary angle of the vertical parameter (the vertical parameter is measured from the horizontal direction). The horizontal parameter is recorded as , the complementary angle of the vertical parameter is recorded as B. By derivation, the rotation matrix can be calculated by formula (5):
[0192] (5),
[0193] The coverage of the ground is preliminarily defined by the external matrix determined by the ground coordinates of the four corner points. Coordinates and The coordinates are marked as and , as the sampling starting point. A grid is then divided according to the target spatial resolution of the corrected image, with each grid representing a pixel in the corrected image. Equations (3) and (4) are used to calculate the scanning coordinates in the tower base video image corresponding to the sampling point. Finally, the pixel value corresponding to the coordinate position is obtained through resampling and assigned to the corresponding position in the corrected image.
[0194] 3 Registration module.
[0195] The registration process first uses a pre-trained super-resolution reconstruction model to super-resolve the satellite image to achieve alignment of the two data resolutions and reduce the resolution difference between the two images. Then, a deep learning dense matching model is selected for matching. The model uses the DINOv2 model to extract the image from the two input images. and Extracting coarse features and Using the pre-trained DINOv2 model as a coarse feature encoder can not only reduce the risk of overfitting the model to the training set, but also use its powerful feature extraction ability to obtain coarse features that are highly robust to changes in perspective, illumination, etc. Since DINOv2 cannot provide fine features, the deep learning dense matching model uses VGG19 as a fine feature encoder to extract fine features from the image. and Fine features extracted from and ,Combined with the coarse features extracted by DINOv2, a feature pyramid is constructed that is both robust and capable of accurate positioning.
[0196] In the coarse matching stage, a Transformer-based matching decoder is used. This is designed using a regression-classification formula. The output space is discretized into a set of uniformly distributed anchor points, and the probability of each anchor point is predicted through classification. This approach effectively captures multimodal distributions and improves the robustness of matching. The formula is:
[0197] (6),
[0198] Where, Indicates that the model parameters are In the case of input image Pixels in , predict its Corresponding pixel point The probability of this is the prediction of the rough matching stage. K is the quantization level, and the actual value here is 64 64 classification anchors, express Corresponding to The probability of anchor points, all The sum is 1, the basic distribution Defined at the anchor point The probability distribution around Centered at The normalized image network is non-zero in the square area and zero in other areas.
[0199] In the refinement stage, the classification anchors are first decoded by performing an argmax operation:
[0200] (7),
[0201] Determine the most likely anchor point that matches pixel x , and then make local adjustments, the formula is:
[0202] (8),
[0203] in, express The set of the four closest anchor points: left, right, top, and bottom.
[0204] The decoder consists of 5 ViT blocks, each containing 8 attention heads, the hidden layer size D is set to 1024, and the MLP size is 4096. Its input is the concatenation of the 512-dimensional projected DINOv2 features and the 512-dimensional features output by the Gaussian process (GP) module, and the output is vector, is the number of classification anchors, and the additional "1" dimension is the matchability score .
[0205] The deep learning dense matching model is based on the multimodal characteristics of the coarse-scale matching distribution and adopts a regression-classification loss function. The coarse matching loss function is constructed by minimizing the Kullback–Leibler divergence between the estimated matching distribution and the theoretical model distribution. , the specific calculation formula is:
[0206] (9),
[0207] in, is a hyperparameter used to control the weight of marginal probability and conditional probability, is a discrete set of known correspondences.
[0208] In the refinement stage, the generalized Charbonnier loss ( ) as the refinement loss , and its calculation formula is:
[0209] (10),
[0210] Finally, the coarse matching loss and the refined loss are added together to get the total loss function .
[0211] It should be noted that:
[0212] 1. The input image size of the deep learning dense matching model should be a multiple of 14;
[0213] 2. DINOv2 is a self-supervised pre-trained model based on the Vision Transformer, pre-trained using mask image modeling. It is used as a frozen coarse feature extractor in this deep learning dense matching model to extract coarse features at a scale of 1 / 16. This provides robust coarse features. Its input image size is typically a multiple of 14 to accommodate multi-scale feature extraction.
[0214] 3. VGG19 is a classic convolutional neural network with a simple structure and high localization capability. In deep learning dense matching models, it is used as a specialized fine feature extractor, extracting fine features at scales such as 1 / 2, 1 / 4, and 1 / 8 for refinement.
[0215] 4. Training Dataset: The MegaDepth dataset contains image pairs with depth annotations collected from internet photos, covering landmarks and natural landscapes around the world. The dataset contains millions of images from various scenes and viewpoints, with depth map resolutions ranging from 640×480 to 1920×1080. This large dataset is effective for training the model's feature extraction capabilities and demonstrates good generalization capabilities.
[0216] 5. Training settings: At the beginning of training, the parameters of DINOv2 are kept frozen without fine-tuning, and its pre-trained feature extraction capabilities are directly used; for the decoder, 10 -4 The learning rate is 5×10 for the encoder. -6 The learning rate is set and scales linearly with the batch size; the deep learning dense matching model is trained for 100 epochs on the MegaDepth dataset using a single GPU.
[0217] After preliminary matching using the deep learning dense matching model, RANSAC is used to eliminate false matches. Through iterative random sampling and consistency verification, RANSAC filters inliers (data that conform to the model) from noisy data and estimates the optimal model parameters. Although the RANSAC algorithm can be used to identify inliers and outliers in deep learning dense matching feature sets, the dense matcher matches a large number of feature point pairs, and outlier interference still exists. This is particularly true in low-contrast, weakly textured areas of the tower base video data, where the RANSAC algorithm may not be able to completely eliminate these outliers.
[0218] This method uses the Delaunay triangulation constraint strategy to improve the robustness and accuracy of the overall matching module after RANSAC removes outliers. The triangular mesh generated by the Delaunay triangulation algorithm is unique and is processed to maximize the minimum angle, so that the triangles tend to be regular. The matching points of the left image are used to generate a triangular mesh using the Delaunay triangulation algorithm. For the right image, the triangular mesh is generated using the matching points of the left image. Since there may be mismatched points in the matching points, resulting in the existence of mismatched triangle pairs, it is necessary to detect the mismatched triangle pairs and eliminate the mismatched points. The present invention mainly uses side length and angle constraints. For matching triangles, the three corresponding sides and three corresponding angles satisfy formula (11):
[0219] (11),
[0220] Where, 、 To match corresponding sides of the triangle, 、 To match corresponding angles of triangles, 、 are the edge length and angle constraint thresholds respectively.
[0221] After obtaining the matching point pairs, the matching points are mapped from the corrected image to the original image through the coordinate relationship between the original tower base video image and the corrected image in the tower base video image correction module, thereby realizing the registration of the tower base video image and the satellite image.
[0222] To verify the practical application of the monitor horizontal angle calibration method, an experiment was designed using three monitors, as shown in Figure 1. Monitor A is the DH-TPC-PT8420A, with a tower height of 49 meters and a focal length of 6 mm; Monitor B is the DH-SD-8A2456-HNF-PA-EL, with a tower height of 35 meters and a focal length of 5.5 mm; and Monitor C is the DS-2TD6236-75C2L / V2, with a tower height of 49 meters and a focal length of 6.7 mm. The initial horizontal angles and rotation directions of Monitors A and B were randomly set, while Monitor C, starting from true north and rotating clockwise, served as a reference for evaluating the effectiveness of the proposed method.
[0223] Table 1 Horizontal angle calibration method test
[0224]
[0225] For the convenience of analysis, the clockwise rotation step is set to 10 degrees and numbered, where the number corresponding to the horizontal angle displayed by the monitoring device is set to 0, and then increases in sequence according to the step size. The results are all within 5°, which can be considered to be equal within the allowable error range. Based on this, it can be considered that the two monitoring devices are rotating clockwise, and the angle compensation values of the two monitoring devices from the true north direction to the display zero degree direction are for The average value of . The value is 84.6°, monitoring C The value is 0.35°. The two images of monitoring A correspond to The result does not fall within the 5° range, so it can be determined that the monitoring device is rotating counterclockwise. In this case, the compensation value For two images The average value of A The value is 24.25°. After calculating the compensation value of the three monitoring devices After that, we can further calculate the horizontal angle of each clockwise rotation starting from the true north direction. Among them, the horizontal angle calculation formula of monitoring A is: , the horizontal angle calculation formula of monitoring B and monitoring C is: , The horizontal angle value displayed for the monitoring device.
[0226] Monitor C is used as the reference reference. Rotating clockwise, with True North as the starting point, it achieves the same rotation direction as the horizontal angle correction method proposed in this invention, with a starting point difference of 0.35°. This fully demonstrates the practicality of the horizontal angle correction method proposed in this invention. After horizontal angle calibration of Monitors A and B, an image was selected for perspective correction and compared with the corresponding satellite image to demonstrate the effectiveness of the horizontal angle correction method proposed in this invention.
[0227] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered by the scope of protection of the claims of the present invention.
Claims
1. A tower base video image automatic positioning method, characterized in that: The steps include: Step S1, obtaining a video image of the tower base including calibration markers, calculating the horizontal parameter deviation of the monitoring equipment through image processing technology, and automatically calibrating the horizontal parameters of the monitoring equipment; the calibration markers are calibration markers set at multiple known positions around the tower base; Step S2: Correct the viewing angle of the tower base video image according to the parameters of the monitoring equipment, and determine the rough position of the tower base video image in the satellite image by corner point coordinate mapping; Step S3: Using a registration algorithm to register the corrected tower base video image with the satellite image, establish a registration mapping relationship, and achieve accurate positioning of the tower base video image; The step S1 includes the following steps: Step S11, selecting two images taken from the same monitoring source at different horizontal angles, and recording the initial horizontal parameters of the two images; Step S12, increasing the horizontal parameters of the two images by a set angle in sequence, numbering them, and performing perspective correction and rough positioning on the images using the changed horizontal parameter values; Step S13, matching the tower base video correction image and the satellite coarse positioning image corresponding to each increased horizontal parameter using a matching module to find the optimal matching pair, and calculating the rotation angle between the matching pairs according to the affine transformation matrix; Step S14, determine whether the rotation angles between the optimal matching pairs of the two images are equal within the allowable error range. If they are equal within the allowable error range, it is considered that the monitor is rotating clockwise, and the horizontal parameter is the angle between the true north direction and zero degrees in the increased horizontal parameter; if they are not equal, it is considered that the monitor is rotating counterclockwise, and the horizontal parameter is the other angle between the true north direction and zero degrees in the increased horizontal parameter.
2. The tower base video image automatic positioning method according to claim 1, characterized in that: In step S12, the angle increased each time is a fixed value, and after each increase, the image is corrected for perspective and roughly positioned, and the image pairs with corresponding numbers are matched until the number of matching points reaches a peak and the number of matching points of two adjacent numbers changes continuously, and the horizontal parameters and numbers at this time are recorded.
3. The tower base video image automatic positioning method according to claim 1, characterized in that: In step S13, the process of calculating the rotation angle between the matching pairs according to the affine transformation matrix includes: Further decomposition of the affine transformation matrix: , Where, is the rotation difference, yes The scale difference in direction, yes The scale difference of the direction, is the shear coefficient, =0; The rotation angle between the optimal matching pairs is obtained by calculation : , if , the registered image is rotated counterclockwise relative to the reference image ;if , the registered image rotates clockwise relative to the reference image .
4. The tower base video image automatic positioning method according to claim 1, characterized in that: The step S1 further includes the following steps: Step S15 , recording the parameters obtained after the adjustment of the first picture and the angle displayed in the monitoring system, as well as the parameters obtained after the adjustment of the second picture and the angle displayed in the monitoring system.
5. The tower base video image automatic positioning method according to claim 1, characterized in that: The step S2 comprises the following steps: Step S21, using the initial data of the tower base video image and the collinearity equation, converting the four corner points of the tower base video image from image plane coordinates to spatial coordinates, and completing the coarse positioning of the tower base video data with the circumscribed rectangle of the four corner points as the boundary, the initial data includes the tower coordinates and height, horizontal parameters, vertical parameters and internal parameters of the monitoring camera; Step S22, taking the minimum x and y coordinates of the four corner points corresponding to the ground coordinates as the starting point, setting the target spatial resolution as the sampling interval, and calculating the corresponding coordinates of each ground coordinate in the original tower base video image using the collinearity equation; Step S23 , resampling is performed and pixel values are assigned to corresponding points of the corrected tower base video image to obtain a perspective correction result of the tower base video image.
6. The tower base video image automatic positioning method according to claim 5, characterized in that: The initial data also includes the width and height of the original image and the pixel size. The collinearity equation converts the scanning coordinates into image plane coordinates through affine transformation parameters, and then converts the image plane coordinates into space coordinates through the rotation matrix and the camera focal length. The affine transformation parameters include a translation parameter and a scaling parameter, wherein the translation parameter is 0 and the scaling parameter is equal to the pixel size; The rotation matrix is calculated by the following formula: , The rotation matrix parameters include the horizontal parameter A of rotation around the Z axis and the complementary angle B of the vertical parameter of rotation around the Y axis, where the horizontal parameter A is the angle of clockwise rotation of the monitoring device around the Z axis, and the vertical parameter complementary angle B is the angle of counterclockwise rotation of the monitoring device around the Y axis.
7. The tower base video image automatic positioning method according to any one of claims 1 to 6, characterized in that: The step S3 comprises the following steps: Step S311: super-resolution processing is performed on the satellite image using a pre-trained super-resolution reconstruction model to achieve resolution alignment between the satellite image and the tower base video image; Step S312, using the DINOv2 model to extract coarse features of the satellite image and the tower base video image; Step S313: Refine the coarse features using the fine features extracted by the VGG19 model at multiple scales to construct a feature pyramid; Step S314: perform coarse matching based on the coarse features through a Transformer-based matching decoder, discretize the output space into evenly distributed anchor points, and predict the matching probability of each anchor point through classification; Step S315, using the RANSAC method to eliminate incorrectly matched feature point pairs; Step S316: Using the Delaunay triangulation constraint strategy to eliminate incorrectly matched triangle pairs, and obtain matching point pairs; In step S317, the obtained matching point pairs are mapped from the corrected image to the original image to achieve registration of the satellite image and the tower base video image.
8. The tower base video image automatic positioning method according to any one of claims 1 to 6, characterized in that: The step S3 comprises the following steps: Step S321: Super-resolution reconstruction of the satellite image is performed using a pre-trained super-resolution reconstruction model to achieve resolution alignment between the satellite image and the tower base video image. Step S322: Use a deep learning intensive matching model to match the satellite image with the tower base video image to obtain a preliminary matching result, wherein the deep learning intensive matching model includes: A coarse feature encoder that uses a pre-trained DINOv2 model to extract coarse features from the input satellite imagery and tower-based video imagery; Fine feature encoder, which uses the VGG19 model to extract fine features from the input satellite imagery and tower base video imagery; Feature pyramid construction module, which combines the coarse features extracted by the DINOv2 model and the fine features extracted by the VGG19 model to construct a feature pyramid; In the coarse matching stage, a Transformer-based matching decoder is used. It is designed using a regression-classification formula, discretizes the output space into a set of evenly distributed anchor points, and predicts the probability of each anchor point through classification to obtain the coarse matching result. In the refinement stage, the rough matching results are locally adjusted to obtain the refined matching results; And, add the coarse matching loss and the refined loss to obtain the total loss function to train the deep learning dense matching model; Step S323: Using the RANSAC algorithm to remove incorrectly matched feature point pairs from the obtained preliminary matching results; Step S324: using the Delaunay triangulation constraint strategy to eliminate incorrectly matched triangle pairs and obtain matching point pairs; In step S325 , the obtained matching point pairs are mapped from the corrected image to the original image to achieve registration of the satellite image and the tower base video image.
9. A tower base video image automatic positioning system, characterized in that: include: A horizontal parameter calibration module is used to obtain a video image of the tower base containing calibration markers set at multiple known locations around the tower base, calculate the horizontal parameter deviation of the monitoring equipment through image processing technology, and automatically calibrate the horizontal parameters of the monitoring equipment; The coarse positioning and perspective correction module is used to correct the perspective of the tower base video image according to the parameters of the monitoring equipment, and determine the rough position of the tower base video image in the satellite image by mapping the corner point coordinates; The registration module is used to register the corrected tower base video image with the satellite image using a registration algorithm, establish a registration mapping relationship, and achieve accurate positioning of the tower base video image; The horizontal parameter calibration module is specifically used to: Select two images taken from the same monitoring source at different horizontal angles and record the initial horizontal parameters of the two images; The horizontal parameters of the two images are sequentially increased by a set angle and numbered, and the images are corrected for viewing angle and roughly positioned using the changed horizontal parameter values; The tower base video correction image and the satellite coarse positioning image corresponding to each increased horizontal parameter are matched using the matching module to find the optimal matching pair, and the rotation angle between the matching pairs is calculated based on the affine transformation matrix; Determine whether the rotation angles between the optimal matching pairs of the two images are equal within the allowable error range. If they are equal within the allowable error range, it is considered that the monitor is rotating clockwise, and the horizontal parameter is the angle between the true north direction and zero degrees in the increased horizontal parameter; if they are not equal, it is considered that the monitor is rotating counterclockwise, and the horizontal parameter is the other angle between the true north direction and zero degrees in the increased horizontal parameter.
Citation Information
Patent Citations
Rapid north correction method for photoelectric turret
CN117570948A
Method and system for calibrating coordinate conversion precision of monitoring video image
CN117876654A