Radar vision aerial view spectrum-deep fusion cross-view geographic positioning method
Through the radar vision bird's eye view-deep fusion cross-view geolocation method, perspective projection transformation and deep learning model are used to solve the problem of alignment and matching between ground view image and bird's eye view image, and efficient and accurate cross-view geolocation fusion is achieved, improving matching accuracy and robustness.
Patent Information
- Application Number
- CN202510420403.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-04
- Publication Date
- 2025-07-22
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art is difficult to efficiently and accurately align the ground viewing image with the bird's-eye viewing image, resulting in insufficient accuracy and inefficiency of geolocation fusion across viewing angles, especially in the feature matching of images of different viewing angles.
The radar vision bird's eye view-deep fusion cross-view geolocation method is adopted, and the training is achieved through data preprocessing, perspective projection transformation, deep learning feature extraction and cross-modal matching, combined with standardized temperature scaling cross-entropy loss function, to achieve efficient alignment of ground view images and bird's eye view images.
The matching accuracy and robustness of cross-view geolocation fusion is improved, the adaptability and generalization ability of the method are enhanced, the geometric mapping error problem caused by perspective differences is solved, and the feasibility and accuracy of feature extraction and matching are improved.
Smart Images

Figure CN120355950A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and geographic information fusion, and particularly to a radar vision bird's-eye view map - deep fusion cross-view geographic positioning method. Background Art
[0002] In many fields such as geographic information processing, autonomous driving, and remote sensing monitoring, it is often necessary to fuse and match image data obtained from different perspectives to obtain more comprehensive and accurate spatial information. Ground perspective images and bird's-eye view images (such as satellite images, aerial images, etc.) are two common types of perspective images, each having its own unique information and advantages.
[0003] However, due to the perspective difference, it is quite difficult to directly match these two types of images.
[0004] Traditional methods often have problems such as insufficient accuracy, low efficiency, or poor adaptability to specific scenarios when performing perspective conversion and feature matching, and cannot meet the growing demand for high-precision cross-view geographic positioning fusion.
[0005] Therefore, the key to the technical solution of the present invention is to provide a method that can efficiently and accurately align and match ground perspective images and bird's-eye view images to achieve effective cross-view geographic positioning fusion. Summary of the Invention
[0006] In view of this, the present invention provides a radar vision bird's-eye view map - deep fusion cross-view geographic positioning method.
[0007] To solve the above technical problems, the present invention adopts the following technical solutions:
[0008] A radar vision bird's-eye view map - deep fusion cross-view geographic positioning method includes the following steps:
[0009] Step S1: Data acquisition
[0010] Collect ground perspective images and bird's-eye view images;
[0011] Step S2: Data preprocessing
[0012] Perform image cropping, image normalization processing, and data augmentation processing;
[0013] Step S3: Perform perspective projection transformation
[0014] Perform perspective projection transformation on the ground perspective image to convert it into a bird's-eye view image;
[0015] Step S4: BEV feature extraction
[0016] After the perspective conversion is completed, a deep learning model is used to extract features from the bird's-eye view image to obtain its feature representation;
[0017] Step S5: Cross-modal matching
[0018] Match the features of the extracted bird's-eye view image with the features of the bird's-eye view image;
[0019] Step S6: Loss function training
[0020] Train using the normalized temperature-scaled cross-entropy loss function.
[0021] Preferably, in step S3, the ground view image I is converted into a bird's-eye view image I using the homography matrix H: g I bev :
[0022] I bev = HI g
[0023] Among them, the homography matrix H is calculated from the camera intrinsic parameter K and the extrinsic rotation matrix R and translation vector t:
[0024]
[0025] Among them, n is the ground normal vector and d is the distance from the ground to the camera.
[0026] Preferably, in step S4, the bird's-eye view image I is subjected to feature extraction to obtain its feature representation F: bev F bev :
[0027] F bev = f θ (I bev )
[0028] Among them, f θ represents a neural network, and its parameters θ can be obtained through training and learning.
[0029] Preferably, in step S5, the features F of the extracted bird's-eye view image are matched with the features F of the bird's-eye view image, and the point-to-point cosine similarity is used to calculate the matching degree: bev F sat
[0030]
[0031]
[0032] Among them, S(i,j) represents the similarity between positions i and j.
[0032] Preferably, in step S6, the normalized temperature-scaled cross-entropy (NT-Xent) loss function is used for training:
[0033]
[0034] where τ is the temperature parameter used to adjust the similarity distribution.
[0035] Preferably, in step S2, the normalization process adjusts the pixel value range to [0, 1] or [-1, 1].
[0036] Preferably, in step S2, the data augmentation process includes: randomly rotating and flipping the ground-view image to simulate different shooting angles and directions; randomly cropping and color-adjusting the bird's-eye view image to increase the diversity of the data.
[0037] Preferably, in step S4, the deep learning model is a convolutional neural network CNN or Transformer.
[0038] The present invention has achieved the following technical effects compared with the prior art:
[0039] (1) The present invention uses an auxiliary measurement device to calculate the antenna trace vector in combination with radar design parameters, avoiding the large geometric mapping error caused by the conventional method of only measuring the coordinates at both ends of the orbit; by performing perspective projection transformation to convert the ground-view image into a bird's-eye view image, the problem of view difference is effectively solved, facilitating subsequent feature extraction and matching, and improving the feasibility and accuracy of matching;
[0040] (2) The present invention uses a deep learning model to extract the features of the bird's-eye view image, which can automatically learn the complex patterns and semantic information in the image, and has stronger expressive ability and adaptability compared with traditional manual feature extraction methods, helping to improve the matching accuracy;
[0041] (3) The present invention uses point-to-point cosine similarity for cross-modal matching, which is simple and efficient and can effectively measure the similarity between different image features, providing a reliable matching basis for cross-view geolocation fusion;
[0042] (4) The present invention's contrastive learning optimization method based on the NT-Xent loss function can further improve the model's feature extraction and matching performance, enabling the model to maintain good matching effects under different scenarios and data conditions, and enhancing the robustness and generalization ability of the method. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 It is a flowchart of the perspective projection transformation (ground → BEV) of a radar vision bird's-eye view map - deep fusion cross-view geolocation method of the present invention;
[0044] Figure 2 This is the flowchart of BEV feature extraction for a radar-vision bird's-eye view map - deep fusion cross-view geolocation method of the present invention;
[0045] Figure 3 This is the flowchart of cross-modal matching for a radar-vision bird's-eye view map - deep fusion cross-view geolocation method of the present invention;
[0046] Figure 4 This is the radar-vision combined bird's-eye view map and depth information fusion map for a radar-vision bird's-eye view map - deep fusion cross-view geolocation method of the present invention. Detailed implementation manners
[0047] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0048] The present invention discloses a radar-vision bird's-eye view map - deep fusion cross-view geolocation method, including the following steps:
[0049] Step S1: Data acquisition
[0050] Perform ground-view image acquisition and bird's-eye view image acquisition; use a ground camera to acquire ground-view images, and satellite images, aerial images, etc. can be used to obtain bird's-eye view image acquisition;
[0051] Step S2: Data preprocessing
[0052] According to actual needs, crop the acquired ground-view images and bird's-eye view images, remove the irrelevant regions in the images, and retain the target regions; perform normalization processing on the cropped images, including pixel value normalization and size normalization. Pixel value normalization adjusts the pixel value range to [0, 1] or [-1, 1] to facilitate subsequent deep learning model processing; size normalization adjusts the images to a unified size; randomly rotate and flip the ground-view images to simulate different shooting angles and directions; randomly crop and adjust the color of the bird's-eye view images to increase the diversity of the data;
[0053] Step S3: Perform perspective projection transformation, as Figure 1 shown;
[0054] Perform perspective projection transformation on the ground-view image to convert it into a bird's-eye view image. Specifically, use the homography transformation matrix H to convert the ground-view image I g into the bird's-eye view image Ibev :
[0055] I bev = HI g
[0056] Wherein, the homography matrix H is calculated from the camera internal parameters K and external parameters (rotation matrix R and translation vector t):
[0057]
[0058] In this formula, n is the ground normal vector and d is the distance from the ground to the camera. In this way, the ground-view image can be converted into an image with a similar perspective and geometric relationship to the bird's-eye view image, laying a foundation for subsequent feature extraction and matching.
[0059] Step S4: BEV feature extraction, as Figure 2 shown;
[0060] After completing the perspective transformation, use a deep learning model, such as a convolutional neural network CNN or Transformer, to extract features from the bird's-eye view image I bev to obtain its feature representation F bev :
[0061] F bev = f θ (I bev )
[0062] Wherein, f θ represents a neural network, and its parameters θ can be obtained through training and learning to extract features that can effectively represent the image content. These features will be used to match the features of bird's-eye view images (such as satellite images, aerial images, etc.) later.
[0063] Step S5: Cross-modal matching, as Figure 3 shown;
[0064] Match the extracted bird's-eye view image features F bev with the features F sat of bird's-eye view images (such as satellite images, aerial images, etc.), and use point-to-point cosine similarity to calculate the matching degree:
[0065]
[0066] Wherein, S(i,j) represents the similarity between positions i and j. By calculating the similarity between different positions, the best matching relationship between the ground-view image and the bird's-eye view image can be determined, thus realizing cross-perspective geolocation fusion.
[0067] Step S6: Loss function training
[0068] To optimize the above matching process, the Normalized Temperature Scaled Cross Entropy (NT-Xent) loss function is used for training:
[0069]
[0070] where τ is the temperature parameter used to adjust the similarity distribution. By minimizing this loss function, the model can learn more effective feature extraction and matching capabilities, thereby improving the accuracy and robustness of cross-view geolocation fusion.
[0071] As Figure 4 shown, it is a radar-visual bird's-eye view map - a deep fusion cross-view geolocation method of the present invention, which is a radar-visual combined bird's-eye view map and depth information fusion map.
[0072] Example 1:
[0073] Step S1: Data collection
[0074] Ground perspective image collection: Use a ground camera to collect ground perspective images; ensure that the camera has a high resolution and stable shooting parameters, such as focal length, aperture, etc.; the collected images should cover different scenes and perspectives of the target area to increase the diversity and representativeness of the data;
[0075] For example, in the autonomous driving scenario, images of different scenes such as roads, buildings, and traffic signs can be collected; in the geodetic surveying scenario, images of natural terrain, urban blocks, etc. can be collected;
[0076] Bird's-eye view image collection: Obtain bird's-eye view images, which can be satellite images, aerial images, etc.; satellite images can be obtained by purchasing commercial satellite data or using open-source satellite data; aerial images can be obtained by drone aerial photography; ensure that the resolution and coverage of the bird's-eye view images meet the application requirements;
[0077] For example, in urban planning, high-resolution satellite images are required to accurately identify buildings and roads; in autonomous driving, aerial images need to cover the entire road network and surrounding environment;
[0078] Step S2: Data preprocessing
[0079] Image cropping: According to actual needs, crop the collected ground perspective images and bird's-eye view images to remove irrelevant areas in the images and retain the target area;
[0080] For example, in the autonomous driving scenario, the sky part in the image can be cropped off, leaving only the road and surrounding environment;
[0081] Image Normalization: The cropped image is normalized, including pixel value normalization and size normalization. Pixel value normalization adjusts the pixel value range to [0,1] or [-1,1] for easier processing by subsequent deep learning models; size normalization resizes the image to a unified size;
[0082] For example, 256×256 or 512×512 to meet the requirements of model input;
[0083] Data Augmentation: To improve the generalization ability and robustness of the model, data augmentation operations can be performed on the preprocessed images; common data augmentation methods include random rotation, flipping, cropping, color adjustment, etc.;
[0084] For example, randomly rotate and flip the ground perspective image to simulate different shooting angles and directions; randomly crop and adjust the color of the bird's-eye view image to increase data diversity;
[0085] Step S3: Perspective Projection Transformation (Ground → BEV)
[0086] Camera Intrinsic Parameter Acquisition: The intrinsic parameter matrix K of the ground camera is obtained through camera calibration methods; camera intrinsics include parameters such as focal length and principal point coordinates, and common methods such as Zhang Zhengyou calibration method can be used to calculate camera intrinsics by shooting calibration board images;
[0087] Camera Extrinsic Parameter Acquisition: The extrinsic parameters of the ground camera are obtained, including the rotation matrix R and the translation vector t; they can be calculated by shooting images at multiple different positions in a known scene and using methods such as feature point matching and PnP algorithm;
[0088] For example, in an autonomous driving scenario, multiple marked points at known positions can be set in a parking lot or test field, and the extrinsic parameters of the camera can be calculated by shooting images of these marked points;
[0089] Homography Matrix Calculation:
[0090] According to the camera intrinsics K and the extrinsics R, t, calculate the homography matrix H:
[0091]
[0092] where n is the ground normal vector, usually assumed to be [0,0,1]; d is the distance from the ground to the camera, which can be obtained through known scene geometry information or the calibration process;
[0093] Ground Perspective Image Conversion:
[0094] Using the calculated homography matrix H, convert the ground perspective image Ig to a bird's-eye view image I bev :
[0095] I bev = HI g
[0096] In actual operation, the perspective transformation function in an image processing library (such as OpenCV) can be used to implement this process;
[0097] For example, using the cv2.warpPerspective function in OpenCV, input the ground view image and the homography matrix, and output the transformed bird's-eye view image;
[0098] Step S4: BEV feature extraction
[0099] Deep learning model selection:
[0100] Select a suitable deep learning model for feature extraction; commonly used models include convolutional neural networks (CNNs) and Transformers;
[0101] For example, pre-trained CNN models such as ResNet and VGG can be used, or Transformer models such as Vision Transformer (ViT); pre-trained models can be pre-trained on large-scale image datasets (such as ImageNet) to learn general image features;
[0102] Feature extraction:
[0103] Input the transformed bird's-eye view image I bev into the selected deep learning model to extract its feature representation F bev :
[0104] F bev = f θ (I bev )
[0105] where f θ represents a neural network, and its parameters θ can be learned through training;
[0106] In actual operation, deep learning frameworks (such as PyTorch and TensorFlow) can be used to load pre-trained models and perform forward propagation on the input image to extract features;
[0107] Feature processing:
[0108] Process the extracted features to facilitate subsequent matching operations;
[0109] For example, the features can be normalized to have unit norm for calculating cosine similarity; in addition, the features can be dimensionally reduced to reduce the computational amount and improve the matching efficiency;
[0110] Step S5: Cross-modal matching
[0111] Feature extraction of bird's-eye view images:
[0112] Extract features from bird's-eye view images (such as satellite images, aerial images, etc.) to obtain their feature representation F sat ; The same deep learning model as the one used to extract F bev can be used, or different models can be selected according to specific requirements;
[0113] For example, for high-resolution satellite images, a dedicated satellite image feature extraction model can be used;
[0114] Cosine similarity calculation:
[0115] Calculate the point-to-point cosine similarity between F bev and F sat to obtain the matching degree matrix S:
[0116]
[0117] where S(i,j) represents the similarity between positions i and j; In actual operations, matrix operation functions in deep learning frameworks can be used to efficiently calculate cosine similarity;
[0118] Optimization of matching results:
[0119] Determine the best matching relationship according to the matching degree matrix S; Methods such as the greedy algorithm and the Hungarian algorithm can be used to optimize the matching results to ensure that the matching relationship at each position is globally optimal;
[0120] For example, in an autonomous driving scenario, the road features in the ground view image can be matched with the road features in the bird's-eye view image to determine the position of the vehicle on the map;
[0121] Step S6: Loss function training
[0122] Definition of loss function:
[0123] Use the normalized temperature-scaled cross-entropy (NT-Xent) loss function for training:
[0124]
[0125] where τ is the temperature parameter used to adjust the similarity distribution; The choice of temperature parameter has an important impact on the training effect of the model and usually needs to be adjusted through experiments;
[0126] Training process:
[0127] Data Partitioning: Partition the preprocessed data into a training set, a validation set, and a test set. The training set is used for model training, the validation set is used for hyperparameter tuning and model selection, and the test set is used to evaluate the final performance of the model;
[0128] Model Initialization: Initialize the deep learning model. The weights of a pre-trained model can be used as the initial weights to accelerate the training speed and improve the model performance;
[0129] Training Loop: Perform multiple iterative trainings on the training set. Each iteration includes forward propagation, loss calculation, and backward propagation; update the model parameters through an optimizer (such as Adam, SGD) to minimize the loss function;
[0130] Hyperparameter Tuning: During the training process, adjust the hyperparameters according to the performance of the validation set, such as the learning rate, temperature parameter, regularization coefficient, etc.; methods such as a learning rate scheduler and an early stopping mechanism can be used to improve the training efficiency and prevent overfitting;
[0131] Model Evaluation: Evaluate the performance of the model on the test set. Commonly used evaluation metrics include matching accuracy, recall rate, F1 score, etc.;
[0132] According to the evaluation results, select the model with the optimal performance as the final model.
[0133] As described above, it is only a preferred embodiment of the present invention, and does not impose any limitation on the technical scope of the present invention. Therefore, any minor modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention still fall within the scope of the technical solution of the present invention.
Claims
1. A radar vision bird's-eye view atlas - deep fusion cross-perspective geolocation method, characterized in that, It includes the following steps: Step S1: Data acquisition Perform ground-view image acquisition and bird's-eye view image acquisition; Step S2: Data preprocessing Perform image cropping, image normalization processing, and data augmentation processing; Step S3: Perform perspective projection transformation Perform perspective projection transformation on the ground-view image and convert it into a bird's-eye view image; Step S4: BEV feature extraction After completing the perspective conversion, use a deep learning model to extract features from the bird's-eye view image to obtain its feature representation; Step S5: Cross-modal matching Match the features of the extracted bird's-eye view image with the features of the bird's-eye view image; Step S6: Loss function training Adopt a normalized temperature-scaled cross-entropy loss function for training.
2. A radar-vision bird's-eye view map - deep fusion cross-view geolocation method according to claim 1, characterized in that In the step S3, the ground-view image I is converted into a bird's-eye view image I g by using the homography transformation matrix H bev : I bev = HI g Among them, the homography matrix H is calculated from the camera internal parameter K and the external parameter rotation matrix R and translation vector t: Among them, n is the ground normal vector, and d is the distance from the ground to the camera.
3. A radar-vision bird's-eye view map - deep fusion cross-view geolocation method according to claim 1, characterized in that In the step S4, for the bird's-eye view image I bev feature extraction is performed to obtain its feature representation F bev : F bev = f θ (I bev ) Among them, f θ represents a neural network, and its parameters θ can be obtained through training and learning.
4. A radar vision bird's-eye view map - deep fusion cross - perspective geolocation method according to claim 1, characterized in that, In the step S5, the extracted bird's-eye view image feature F bev is matched with the feature F of the bird's-eye view image sat and the point-to-point cosine similarity is used to calculate the matching degree: Among them, S(i,j) represents the similarity between positions i and j.
5. A radar-vision bird's-eye view map - deep fusion cross-view geolocation method according to claim 1, characterized in that In the said step S6, a normalized temperature-scaled cross-entropy (NT-Xent) loss function is adopted for training: Among them, τ is the temperature parameter used to adjust the similarity distribution.
6. A radar-vision bird's-eye view map - deep fusion cross-view geolocation method according to claim 1, characterized in that In the said step S2, the normalization processing adjusts the pixel value range to [0,1] or [-1,1].
7. A radar vision bird's-eye view map - deep fusion cross-view geolocation method according to claim 1, characterized in that, In the said step S2, the data augmentation processing includes: randomly rotating and flipping the ground-view image to simulate different shooting angles and directions; randomly cropping and color-adjusting the bird's-eye view image to increase the diversity of the data.
8. A radar vision bird's-eye view map - deep fusion cross-perspective geolocation method according to claim 1, characterized in that In the said step S4, the deep learning model is a convolutional neural network CNN or Transformer.