A visual and positioning based cross-view target dual matching method and system
Patent Information
- Application Number
- CN202511400619.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-28
- Publication Date
- 2026-09-11
- Estimated Expiration
- 2045-09-28
AI Technical Summary
[0005]为了解决目前后融合方法在应用过程中,存在仅依赖单一模态信息,导致匹配精度不高的问题,本发明提供了一种基于视觉和定位的跨视角目标双重匹配方法和系统,其在目前的后融合方法上进行进一步改进,同时集成了行人重识别技术,从而能够适用于空、地间大视角差异的场景,显著提升了空-地跨视角目标匹配的准确性与鲁棒性,所述技术方案如下:
1、本发明提出了一种基于视觉和定位的跨视角目标双重匹配方法和系统,通过结合跨视角的视频片段目标视觉特征与GNSS RTK定位数据,显著提升了空-地跨视角目标匹配的准确性与鲁棒性,尤其在视角差异超过45°的复杂场景下仍能保持高匹配成功率;
Smart Images

Figure CN121330576B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a cross-view dual-matching method and system for targets based on vision and localization, belonging to the field of computer vision. Background Technology
[0002] In complex air-to-ground collaborative target reconnaissance and surveillance scenarios, joint perception using multiple UAVs and unmanned vehicles across different perspectives and modalities has become an important technical approach to enhance situational awareness capabilities. A core task in multi-view perception is to achieve accurate detection and correlation matching of the same target from multiple perspectives, both from the air and the ground. The key lies in fusing visual appearance features with geolocation information to achieve stable and robust dual matching of targets.
[0003] Currently, there are various main techniques for multi-view target matching, each with its own advantages and disadvantages. These include pre-fusion methods and ground-based image matching methods. Pre-fusion methods typically map image features from different perspectives to a unified ground plane or bird's-eye view (BEV), and perform joint detection and matching of targets in a shared feature representation space. However, this type of method has high requirements for camera deployment and image acquisition conditions, and usually relies on the overlapping coverage of the same area by multi-view images, thus limiting its applicability in scenarios with significant differences between air and ground perspectives (such as a difference of more than 45°). Ground-based image matching methods, such as achieving coarse localization through air-ground image matching in geographically denied environments, primarily aim at scene-level or region-level matching, rather than stable association for specific fine targets. Therefore, they cannot meet the requirements for high-precision target association.
[0004] Another multi-view target matching technique, the post-fusion method, first detects targets in each independent view, then projects the detection results onto a unified three-dimensional coordinate system, and associates targets based on spatial proximity and appearance similarity. Compared to pre-fusion methods and ground-based image matching methods, this type of method has stronger viewpoint adaptability and system scalability. However, when there are large cross-viewpoints and significant differences in target appearance, it is prone to mismatches and missed matches because it generally relies on only single modal information (such as visual or localization only). Therefore, there is an urgent need to provide a multi-view target matching technique to solve the above problems. Summary of the Invention
[0005] To address the issue of low matching accuracy caused by relying solely on single-modal information in current post-fusion methods, this invention provides a cross-view target dual matching method and system based on vision and localization. This method further improves upon current post-fusion methods and integrates pedestrian re-identification technology, making it applicable to scenarios with significant differences in viewpoints between air and ground. This significantly enhances the accuracy and robustness of air-to-ground cross-view target matching. The technical solution is as follows: A cross-viewpoint target dual matching method based on vision and localization includes the following steps: S1. Collect video clips from both aerial and ground perspectives, and attach corresponding timestamps and geolocation information; S2. Use an object detection network to generate target bounding results for aerial and ground-based cross-view video clips respectively; S3. The target selection results from both the air and ground perspectives are cropped out, and then the target location estimation algorithm is used to generate the target geolocation from the air perspective and the target geolocation from the ground perspective. S4. The target selection results from different perspectives are used to generate cross-view target visual features using an air-ground target re-identification network. The geographic location of the target from the air perspective and its corresponding visual features are used as a matching dependency, and the geographic location of the target from the ground perspective and its corresponding visual features are used as another matching dependency. The Hungarian matching algorithm is used to complete the first matching of the current time frame. After matching, each target across perspectives in the current time frame obtains its intra-frame ID. S5. Visual features corresponding to targets with the same intra-frame ID are used as a query set to perform cross-time frame matching with the sample library. The sample library is initially empty. After the second matching is completed, the global ID of the target in the current time frame is obtained. S6. Add the visual features corresponding to the target with the global ID to the corresponding position in the sample library. The visual features with the global ID in the sample library are saved according to the data format of global ID, timestamp, and geographic location. S7. Repeat steps S2 to S6 until every frame of the aerial and ground-based cross-view video clip has been processed.
[0006] Furthermore, in step S1, a camera mounted on a drone is used to collect video clips from an aerial perspective, and a camera mounted on an unmanned vehicle is used to collect video clips from a ground perspective. The video clips from different perspectives have the same frame rate, and each frame carries its corresponding timestamp and GNSS RTK high-precision positioning information at the time of shooting.
[0007] Furthermore, in step S2, the target detection network includes a backbone network, a multi-scale feature fusion network, and a detection head. The backbone network uses stepwise downsampling to extract features. The multi-scale feature fusion network achieves bottom-up and top-down multi-scale feature fusion through upsampling and splicing operations, enhancing the perception capability of targets of different sizes. The detection head adopts a decoupled structure to predict the target category and bounding box coordinates respectively, achieving efficient end-to-end detection.
[0008] Furthermore, in step S3, the target location estimation algorithm includes an encoder, a decoder, and a coordinate projection module. The encoder encodes the cropped target bounding box result into a high-dimensional feature map; the decoder converts it into a depth map; and the coordinate projection module combines GNSS RTK positioning data to finally output a cross-view target geolocation containing WGS84 coordinates.
[0009] Furthermore, in step S3, the coordinate projection module, in conjunction with GNSS RTK positioning data, ultimately outputs a cross-view target geolocation including WGS84 coordinates, comprising the following steps: S31. Project the 2D pixels of the depth map onto the camera coordinate system using the inverse of the camera intrinsic parameter matrix K. ; S32. Using camera extrinsics, including the rotation matrix Translation vector The camera coordinate system is converted to local world coordinates. ; S33. Combine the GNSS RTK positioning data to perform coordinate replacement, from local world coordinates to ENU coordinates. Then convert from ENU coordinates to GS84 coordinates. ,in and Let be the radius of curvature of the Earth.
[0010] Further, in step S4, the air-to-ground target re-identification network includes a feature extraction backbone network and an embedding feature head. The feature extraction backbone network is built based on ResNet-50, and the stride of the last layer is set to 1. The cross-view target bounding box result is processed by the feature extraction backbone network to output a feature map with a dimension of 2048×16×16. The embedding feature head includes a global average pooling layer and a batch normalization layer. The global average pooling layer compresses the feature map into a 2048-dimensional vector, which is then processed by the batch normalization layer to obtain... The normalized 512-dimensional vector is then processed by the network to generate visual features of cross-view targets. In the first matching, for each target, its visual feature vector is L2 normalized, and the normalized visual feature vector is concatenated with the position coordinate vector to form a multimodal joint feature vector, which serves as the matching descriptor for the target. In Hungarian matching, the matching cost between two targets is determined by the weighted cosine distance and Euclidean distance between the joint feature vectors. Finally, the accurate association of cross-view targets is achieved by minimizing the overall matching cost.
[0011] Furthermore, in S5, the cross-time frame matching is as follows: query the cosine similarity between the visual features corresponding to targets with the same intra-frame ID and all existing visual features in the sample library to form a similarity matrix; for each queried visual feature, select the maximum value of its cosine similarity with existing features in the sample library; if the maximum value exceeds a preset threshold, the queried visual feature and the most similar candidate in the sample library are considered to be the same target, and its corresponding global ID is assigned to the current queried target; if the maximum value does not reach the threshold, a new global ID is assigned to the visual feature, and it is added to the sample library as a new sample.
[0012] On the other hand, the present invention provides a cross-view target dual matching system based on vision and localization. Based on the aforementioned cross-view target dual matching method based on vision and localization, the system includes: The video acquisition module is used to simultaneously acquire video data from both aerial and ground perspectives. Each frame of the video data is accompanied by a precise timestamp and GNSS RTK high-precision positioning information. The target detection module is used to identify and select targets in images from aerial and ground perspectives. It is connected to the output of the video acquisition module, receives synchronous frame images, performs target detection, and outputs the target detection box results from each perspective. The target geolocation estimation module is used to output the absolute geographic coordinates of the target based on the location information of the detected target and the carrier. It is connected to the output end of the target detection module, receives the target detection box results and GNSS RTK positioning information, and calculates the precise geographic location of each target in the geodetic coordinate system through the target location estimation algorithm. The air-ground target re-identification module is used to extract visual features of targets across different viewpoints. It is connected to the output of the target detection module and processes target image patches under different viewpoints through the air-ground target re-identification network to extract depth visual features with viewpoint invariance. The dual matching module is used to realize the dual association of targets within and across frames. It is connected to the output of the target geolocation estimation module and the air-ground target re-identification module. The dual matching module performs cross-view target matching at the same time based on visual features and geolocation to complete the first matching and assign an intra-frame ID to the target. The dual matching module then performs target association across time frames based on visual features to complete the second matching and assign and maintain a global ID for the target. The sample library management and trajectory generation module is used to store and manage the visual features and geolocation data of the target, maintain the consistency of the global ID, and support dynamic target trajectory backtracking and static target location output. It is connected to the dual matching module, updates the sample library according to the global ID generated by matching, and records the target location information in the format of global ID, timestamp, and geolocation. Finally, it outputs the global identifier, geographical location and movement trajectory of all targets.
[0013] A computing device includes: one or more processors and a storage device; the storage device is used to store one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors perform the method described above.
[0014] The beneficial effects of this invention are: 1. This invention proposes a dual-view target matching method and system based on vision and positioning. By combining the visual features of targets in cross-view video clips with GNSS RTK positioning data, it significantly improves the accuracy and robustness of air-to-ground cross-view target matching, and can maintain a high matching success rate even in complex scenarios with view differences exceeding 45°. 2. Simultaneously, a strategy combining first-level intra-frame matching based on matching algorithms and second-level global matching based on cross-frame retrieval is adopted. This not only achieves reliable association of targets from multiple perspectives at the same time, but also has the ability to maintain IDs and integrate trajectories across time frames, effectively avoiding target mismatch and ID jump problems. 3. By constructing and dynamically updating the visual feature sample library, the system can continuously track and identify the same target, support the geolocation of static targets and the motion trajectory tracing of dynamic targets, and enhance the practicality and scalability of the system in continuous monitoring and situational awareness. 4. This invention does not rely on strict overlap or fixed deployment conditions of multiple cameras, and is applicable to various collaborative reconnaissance modes such as air-to-air, ground-to-ground, and air-to-ground. It has good system adaptability and deployment flexibility, and can be applied to complex application scenarios such as security monitoring and intelligent inspection. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is a schematic diagram of the overall process of the cross-view dual matching method according to Embodiment 1 of the present invention; Figure 2 This is a schematic diagram of the target detection network in Embodiment 1 of the present invention; Figure 3 This is a schematic diagram of the target position estimation algorithm according to Embodiment 1 of the present invention; Figure 4 This is a schematic diagram of the air-to-ground target re-identification network according to Embodiment 1 of the present invention; Figure 5 This is a schematic diagram of the cross-view target dual matching system according to Embodiment 2 of the present invention. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.
[0018] Example 1 This embodiment provides a cross-view target dual matching method based on vision and localization, see [link to documentation]. Figure 1 The method includes the following steps: S1: For complex target scenarios with multiple targets, use drones to collect aerial video clips (VID). a Using unmanned vehicles to collect video clips from a ground-level perspective (VID) g The captured cross-view video clips from both aerial and ground perspectives have the same frame rate, and each frame carries its corresponding timestamp and GNSS RTK high-precision positioning information at the time of capture.
[0019] S2: Take aerial view images (IMG) with the same timestamp. a and ground-view images IMG g The target detection network was used to process the target boxes to obtain the target bounding box result R from the aerial view. a Target selection results from ground view R g .
[0020] The specific structure of the object detection network is shown below. Figure 2 It consists of three parts: a backbone network, a multi-scale feature fusion network (Neck), and a detection head. The backbone network receives an input image (an aerial view image, IMG) with a size of H×W×3. a Or ground-view image IMG g This network employs a series of convolutional (Conv) modules, C2f modules, spatial-channel decoupled downsampling (SCDown), compact inverted block structure (C2fCIB), spatial pyramid fast pooling (SPPF), and partial self-attention modules (PSA) to progressively downsample and extract multi-scale features. The multi-scale feature fusion network (Neck) achieves bottom-up and top-down multi-scale feature fusion through upsampling and concatenation operations, enhancing the perception of targets of different sizes. The detection head adopts a decoupled structure, predicting the target category and bounding box coordinates separately, and introduces a consistent dual-label assignment strategy without non-maximum suppression (NMS), completely abandoning NMS post-processing during the inference stage to achieve efficient end-to-end detection. Through lightweight classification heads, deep separable operations in large-kernel convolutions, and rank-guided module design, this network significantly reduces computational redundancy and latency while maintaining high accuracy, making it particularly suitable for real-time detection tasks involving multiple views and multiple targets in air-ground collaborative scenarios.
[0021] Specifically, the object detection network is trained using supervised learning, where the average loss of the statistical image sequence is used as the error for backpropagation during training. First, the aerial view image (IMG) is analyzed. a and ground-view images IMG g The object detection task features in the dataset are modeled as an optimization model consisting of a mixture of classification error, regression error, and matching assignment error. A training dataset is constructed from labeled image sequences, and the loss function is designed as follows:
[0022] in, The classification loss is represented by the cross-entropy error between the predicted class score and the true label, calculated using a task-aligned classifier, which characterizes the semantic similarity between the output prediction and the true class. The regression loss is represented by the combination of Distribution Focal Loss (DFL) and Complete IoU Loss (CIoU) to calculate the predicted bounding box. With the true bounding box The positional deviation and overlap error between them characterize the accuracy of bounding box regression; The consistent matching metric loss is represented by the consistent dual-label assignment strategy, which is used to calculate the matching metric. ,in For spatial priors, For categorized scores, and To balance hyperparameters, predict bounding boxes True bounding box ), ensuring supervisory consistency error between one-to-many and one-to-one heads; The attention module loss is represented by the consistency error of high-dimensional features extracted by the Partial Self-Attention (PSA) module, which characterizes the model's ability to capture global context. and It is an adjustable parameter used to balance the weights of different loss terms.
[0023] S3: Aerial View Image (IMG) a and ground-view images IMG g The target selection results from the aerial perspective are shown in the image below. a Target selection results from ground view R g Cropping out the target selection result from the aerial view R a Inputting the UAV GNSS RTK geolocation information with the current timestamp into the target position estimation algorithm, the target geolocation P from the aerial perspective is obtained. a ; The target selection result from the ground view is R g The target location estimation algorithm is input with the GNSS RTK geolocation information of the unmanned vehicle at the current timestamp to obtain the target geolocation P from the ground view. g .
[0024] The specific structure of the target location estimation algorithm is shown below. Figure 3 It consists of an encoder, a decoder, and a coordinate projection module. The encoder crops the target bounding box result R from the aerial view of the image. a Or the target selection result from the ground view R g(The image is H×W×3, where H and W represent the height and width of the cropped image, respectively) encoded into a high-dimensional feature map; then, the decoder uses upsampling and feature fusion to generate a disparity map, which is then converted into a depth map; the coordinate projection module combines the platform's GNSS RTK data to project the decoder's depth map (containing the target's depth information) onto the geographic coordinate system, outputting the target (P) a or P g WGS84 coordinates (including latitude) (longitude λ, altitude h).
[0025] Specifically, the encoder consists of an initial convolutional layer, a max-pooling layer, and four residual block groups. The initial convolutional layer has a 7×7 kernel, a stride of 2, padding of 3, 64 output channels, a ReLU activation function, and an output size of [missing information]. The max pooling layer has a 3×3 pooling kernel, a stride of 2, a fill size of 1, and further downsampling to... After processing through four residual block groups, the final feature map is output. The size is This is used for subsequent depth estimation.
[0026] Specifically, the decoder receives the feature map. Its specific structure includes an upsampling and fusion layer group, a disparity regression layer group, and a depth transformation. The upsampling and fusion layer group contains five upsampling and fusion layers, which sequentially perform upsampling and feature fusion using 3×3 convolutional kernels (stride 1, padding 1, activation function ELU). Combined with the encoder's skip connection features, the output channels are gradually reduced (256, 128, 64, 32, 16), and the size is restored sequentially. , , , H×W. The disparity regression layer group is followed by a 3×3 convolutional kernel (stride 1, padding 1, output channel 1, activation function Sigmoid) after each of the fusion layers from the 2nd to the 5th group, generating a multi-scale disparity map. to The final fusion layer is followed by a 3×3 convolutional kernel (stride 1, padding 1, output channel 1, activation function sigmoid) to generate a high-resolution disparity map σ with dimensions H×W×1. The disparity map σ is obtained using the formula... Convert to a depth map, where a is the disparity scaling factor and b is the disparity offset, both of which are preset parameters.
[0027] Specifically, the coordinate projection module measures the depth of the 2D pixel (u, v) at the target center of the depth map D. Perform 3D projection to obtain WGS84 coordinates. First, project the 2D pixels onto the camera coordinate system using the inverse of the camera intrinsic matrix K. Then, use the camera extrinsic parameters, i.e., the rotation matrix. Translation vector Convert to local world coordinates Finally, combining GNSS RTK high-precision positioning information (latitude) ,longitude ,high Perform coordinate replacement from local world coordinates to East-North-Sky (ENU) coordinates. For rotation matrices of drone or unmanned vehicle platforms, (This is the translation vector of the platform; here, the platform is the origin, so it is the zero vector); ENU to WGS84 coordinates. ( and (where is the radius of curvature of the Earth).
[0028] S4: Select the target outline from the aerial view. a Target selection results from ground view R g Inputting the air-ground target re-identification network, visual features VIS_F of the target across different viewpoints are obtained. a and VIS_F g At this point, the aerial perspective is set to visual feature VIS_F a and target geographic location P a To match dependencies, the ground view uses visual feature VIS_F g and target geographic location P g To match dependencies, a Hungarian matching algorithm is used for the current time frame. This is the first level of matching. After matching, each target across the viewpoint in the current time frame obtains its intra-frame ID.
[0029] The specific structure of the air-to-ground target re-identification network is as follows: Figure 4 As shown, it is built on the FastReID framework. Its core structure consists of a feature extraction backbone network and an embedding head, which are used to extract discriminative and viewpoint-invariant visual features from target images from different perspectives.
[0030] Specifically, the feature extraction backbone is built on ResNet-50, with the stride of the last layer set to 1 to maintain high spatial resolution, thus preserving more detail in scenes with significant cross-viewpoint differences. The aerial viewpoint target bounding box selection result is shown in Figure R. a Target selection results from an aerial perspective (R) g After processing by this feature extraction backbone network, feature maps with dimensions of 2048×16×16 are output respectively.
[0031] Specifically, the embedding head includes a global average pooling layer and a batch normalization layer, and employs a BNNeck structure to optimize feature distribution, enhancing training stability and feature discriminative power. The global average pooling layer compresses the aforementioned feature map into a 2048-dimensional vector, which is then processed by the BNNeck layer to obtain a normalized 512-dimensional feature representation.
[0032] Finally, the target selection result from the aerial perspective is R. a The network outputs the visual features VIS_F after L2 normalization. a Target selection result from ground view R g The network outputs the visual features VIS_F after L2 normalization. g The two represent the target's appearance features from aerial and ground perspectives, respectively.
[0033] Specifically, during the training phase, the network jointly employs Cross Entropy Loss and Triplet Loss. Triplet Loss introduces a hard sample mining strategy to narrow the feature distance between similar targets from different viewpoints and push away features of dissimilar targets, thereby enhancing the model's robustness to changes in viewpoint, lighting, and scale. For data augmentation, strategies such as random erasure, horizontal flipping, and random padding are used to further improve the model's generalization ability. This network is pre-trained on large-scale pedestrian re-identification data and can be fine-tuned using air-to-ground collaborative target data. Ultimately, it can effectively handle the significant appearance differences and geometric deformations between air-to-ground and ground-to-ground targets, providing highly discriminative visual feature representations for subsequent dual matching.
[0034] Specifically, the matching dependencies in the first level of matching are constructed by fusing visual features and geolocation information. For each target, its visual feature vector (VIS_F) is... a or VIS_F g L2 normalization is performed on the visual feature vector, and the normalized visual feature vector is compared with the position coordinate vector (P). a or P g The features are concatenated to form a multimodal joint feature vector, which serves as the matching descriptor for the target. In the Hungarian matching algorithm, the matching cost between two targets is determined by the weighted cosine distance and Euclidean distance between the joint feature vectors. The visual part weight is used to measure appearance reliability, and the localization part weight is used to measure spatial consistency. Finally, the accurate association of cross-view targets is achieved by minimizing the overall matching cost.
[0035] S5: Label the visual features corresponding to targets with the same intra-frame ID as VIS_F.ID_intra VIS_F ID_intra The query set is used for cross-time frame matching against the sample database. This is the second level of matching. Targets in the current time frame obtain their global IDs, and targets with the same intra-frame ID obtain the same global ID. The sample database is initially empty.
[0036] The cross-time frame matching specifically involves: calculating the cosine similarity between the query feature and all existing features in the sample library to form a similarity matrix; for each query feature, selecting the maximum cosine similarity between it and features in the sample library; if this maximum value exceeds a preset threshold, the query target and the most similar candidate in the sample library are considered to be the same target, and its corresponding global ID is assigned to the current query target; if the maximum value does not reach the threshold, a new global ID is assigned to the target, and the feature is added to the sample library as a new sample. This strategy draws on the Rank-1 identification criterion in target re-identification, controlling the strictness of matching through a threshold, effectively ensuring the accuracy and robustness of cross-time frame target association, while also exhibiting good scalability and adaptability to dynamically added targets.
[0037] S6: Obtain the visual feature VIS_F corresponding to the target with the same global ID in the current time frame. ID_intra Each target is added to the corresponding position in the sample library according to its obtained global ID. The target of each specific global ID in the sample library maintains its independent visual feature VIS_F. ID_inter List; saves the target geographic location of the current time frame according to the data format of "global ID-timestamp-geographic location".
[0038] S7: Repeat steps S2 to S6 until every frame in the cross-view video clip has been processed. At this point, the sample library can provide all sensitive targets with global IDs in the cross-view and cross-time frame video clips without repetition or omission. At the same time, it can provide the geographic location of static targets with specific global IDs in the scene, and can trace the motion trajectory of dynamic targets according to the timestamp and provide the geographic location information of their final moment.
[0039] Example 2 This embodiment provides a cross-view target dual matching system based on vision and localization. See [link to documentation]. Figure 5 The system includes: The video acquisition module is used to simultaneously acquire video data from both aerial and ground perspectives. This module acquires aerial video data through a camera mounted on a drone and ground-view video data through a camera mounted on an unmanned vehicle. Each frame of the video data is accompanied by a precise timestamp and GNSS RTK high-precision positioning information. The target detection module is used to identify and select targets in images from aerial and ground perspectives. This module receives synchronous frame images from the video acquisition module, performs target detection on each, and outputs the target detection bounding box results for each perspective. The target geolocation estimation module is used to output the absolute geographic coordinates of the target based on the location information of the detected target and the carrier. This module receives the target detection frame results from the target detection module and the video acquisition module, as well as the GNSS RTK positioning data of the corresponding platform, and calculates the precise geographic location of each target in the geodetic coordinate system through the target location estimation algorithm. The air-ground target re-identification module is used to extract visual features of targets across different viewpoints. This module is connected to the output of the target detection module and uses the air-ground target re-identification network to process target image patches under different viewpoints and extract depth visual features with viewpoint invariance. The dual matching module is used to achieve dual association of targets within and across frames. This module is connected to the output of the target geolocation estimation module and the air-to-ground target re-identification module, respectively. First, it performs cross-view target matching at the same time based on visual features and geolocation to complete the first matching and assign an intra-frame ID to the target. Then, it performs target association across time frames based on visual features to complete the second matching and assign and maintain a global ID for the target. The sample library management and trajectory generation module is used to store and manage the visual features and geolocation data of the target, maintain the consistency of the global ID, and support dynamic target trajectory backtracking and static target location output. This module is connected to the dual matching module, updates the sample library according to the global ID generated by matching, and records the target location information in the format of "global ID-timestamp-geolocation". Finally, it outputs the global identifier, geographical location and movement trajectory of all targets.
[0040] As an example, a computing device includes: one or more processors and a storage device; the storage device is used to store one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the above-described vision- and localization-based cross-view target dual matching method.
[0041] In summary, the present invention provides a vision- and location-based cross-view target dual matching method and system. By fusing visual appearance features with high-precision geolocation information, combining intra-frame cross-view matching and cross-time frame global matching mechanisms, and introducing a dynamically updated sample library management strategy, it effectively solves the matching and association problem caused by large viewpoint differences and significant target shape changes in air-to-ground collaborative scenarios. This method not only significantly improves the accuracy and robustness of target matching, but also has the ability to trace the trajectory of dynamic targets and accurately locate static targets. It can be widely applied in multiple fields such as intelligent monitoring and collaborative perception of unmanned systems, and has significant practical value and promising prospects for promotion.
[0042] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A visual and positioning based cross-view target dual matching method, characterized in that, The method includes: S1. Collect video clips from both aerial and ground perspectives, and attach corresponding timestamps and geolocation information; S2. Use an object detection network to generate target bounding results for aerial and ground-based cross-view video clips respectively; S3. The target selection results from both the air and ground perspectives are cropped out, and then the target location estimation algorithm is used to generate the target geolocation from the air perspective and the target geolocation from the ground perspective. S4. The target selection results from different perspectives are used to generate cross-view target visual features using an air-ground target re-identification network. The geographic location of the target from the air perspective and its corresponding visual features are used as a matching dependency, and the geographic location of the target from the ground perspective and its corresponding visual features are used as another matching dependency. The Hungarian matching algorithm is used to complete the first matching of the current time frame. After matching, each target across perspectives in the current time frame obtains its intra-frame ID. S5. Visual features corresponding to targets with the same intra-frame ID are used as a query set to perform cross-time frame matching with the sample library. The sample library is initially empty. After the second matching is completed, the global ID of the target in the current time frame is obtained. S6. Add the visual features corresponding to the target with the global ID to the corresponding position in the sample library. The visual features with the global ID in the sample library are saved according to the data format of global ID, timestamp, and geographic location. S7. Repeat steps S2 to S6 until every frame of the aerial and ground-based cross-view video clip has been processed.
2. The method of claim 1, wherein, S1 includes: using a camera mounted on a drone to capture video clips from an aerial perspective, and using a camera mounted on an unmanned vehicle to capture video clips from a ground perspective. The video clips from different perspectives have the same frame rate, and each frame carries its corresponding timestamp and GNSS RTK high-precision positioning information at the time of capture.
3. The method of claim 1, wherein, In step S2, the target detection network includes a backbone network, a multi-scale feature fusion network, and a detection head. The backbone network uses stepwise downsampling to extract features. The multi-scale feature fusion network achieves bottom-up and top-down multi-scale feature fusion through upsampling and splicing operations, enhancing the perception capability of targets of different sizes. The detection head adopts a decoupled structure to predict the target category and bounding box coordinates respectively, achieving efficient end-to-end detection.
4. The method of claim 1, wherein, In step S3, the target location estimation algorithm includes an encoder, a decoder, and a coordinate projection module. The encoder encodes the cropped target bounding box result into a high-dimensional feature map; the decoder converts it into a depth map; and the coordinate projection module combines GNSS RTK positioning data to finally output a cross-view target geolocation containing WGS84 coordinates.
5. The method of claim 4, wherein, The coordinate projection module, combined with GNSS RTK positioning data, ultimately outputs cross-view target geographic positioning including WGS84 coordinates, including: S31. Project the 2D pixels of the depth map onto the camera coordinate system (X) using the inverse of the camera intrinsic parameter matrix K. C ,Y C Z C )=d center ·K-1·[u,v,1] T ; S32. Use camera extrinsics, including rotation matrix R cam and translation vector t cam to convert the camera coordinate system to local world coordinates (x, y, z) = R cam · [X C , Y C , Z C ] T + t cam ; S33. Coordinate replacement in combination with the GNSS RTK positioning data, from local world coordinates to ENU coordinates (E, N, U) = R platform • [x, y, z] T + t platform and from ENU coordinates to GS84 coordinates where R N and R E are the radii of curvature of the earth.
6. The cross-view target dual matching method based on vision and localization according to claim 1, characterized in that, In step S4, the air-to-ground target re-identification network includes a feature extraction backbone network and an embedding feature head. The feature extraction backbone network is built based on ResNet-50, and the stride of the last layer is set to 1. The cross-view target bounding box result is processed by the feature extraction backbone network to output a feature map with a dimension of 2048×16×16. The embedding feature head includes a global average pooling layer and a batch normalization layer. The global average pooling layer compresses the feature map into a 2048-dimensional vector, which is then processed by the batch normalization layer to obtain a normalized 512-dimensional vector. Finally, the 512-dimensional vector is normalized by the network to generate the visual features of the cross-view target. In the first level of matching, for each target, its visual feature vector is L2 normalized, and the normalized visual feature vector is concatenated with the position coordinate vector to form a multimodal joint feature vector, which serves as the matching descriptor for the target. In Hungarian matching, the matching cost between two targets is determined by the weighted cosine distance and Euclidean distance between the joint feature vectors. Finally, the accurate association of cross-view targets is achieved by minimizing the overall matching cost.
7. The cross-view target dual matching method based on vision and localization according to claim 1, characterized in that, In step S5, the cross-time frame matching is as follows: query the cosine similarity between the visual features corresponding to targets with the same intra-frame ID and all existing visual features in the sample library to form a similarity matrix; for each queried visual feature, select the maximum value of its cosine similarity with existing features in the sample library; if the maximum value exceeds a preset threshold, the queried visual feature and the most similar candidate in the sample library are considered to be the same target, and its corresponding global ID is assigned to the current queried target; if the maximum value does not reach the threshold, a new global ID is assigned to the visual feature, and it is added to the sample library as a new sample.
8. A vision- and location-based cross-view target dual matching system, used in any one of claims 1 to 7, characterized in that, The system includes: The video acquisition module is used to simultaneously acquire video data from both aerial and ground perspectives. Each frame of the video data is accompanied by a precise timestamp and GNSS RTK high-precision positioning information. The target detection module is used to identify and select targets in images from aerial and ground perspectives. It is connected to the output of the video acquisition module, receives synchronous frame images, performs target detection, and outputs the target detection box results from each perspective. The target geolocation estimation module is used to output the absolute geographic coordinates of the target based on the location information of the detected target and the carrier. It is connected to the output end of the target detection module, receives the target detection box results and GNSS RTK positioning information, and calculates the precise geographic location of each target in the geodetic coordinate system through the target location estimation algorithm. The air-ground target re-identification module is used to extract visual features of targets across different viewpoints. It is connected to the output of the target detection module and processes target image patches under different viewpoints through the air-ground target re-identification network to extract depth visual features with viewpoint invariance. The dual matching module is used to realize the dual association of targets within and across frames. It is connected to the output of the target geolocation estimation module and the air-ground target re-identification module. The dual matching module performs cross-view target matching at the same time based on visual features and geolocation to complete the first matching and assign an intra-frame ID to the target. The dual matching module then performs target association across time frames based on visual features to complete the second matching and assign and maintain a global ID for the target. The sample library management and trajectory generation module is used to store and manage the visual features and geolocation data of the target, maintain the consistency of the global ID, and support dynamic target trajectory backtracking and static target location output. It is connected to the dual matching module, updates the sample library according to the global ID generated by matching, and records the target location information in the format of global ID, timestamp, and geolocation. Finally, it outputs the global identifier, geographical location and movement trajectory of all targets.
9. A computing device, comprising: One or more processors or storage devices; Storage device for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the vision- and localization-based cross-view target dual matching method as described in any one of claims 1-7.
Citation Information
Patent Citations
Multi-target cooperative tracking method based on high-altitude view angle and ground view angle
CN114648557A
Focus target detection method based on air-ground collaborative view angle
CN114648716A