UAV Ground Target Localization Method and System Based on Multimodal Information Alignment
Through the multimodal information alignment method, a satellite feature base map is constructed and super voxel segmented. Combined with three-dimensional area and scene feature maps, the accuracy problem of drone target positioning is solved, and more accurate target detection and matching is achieved.
Patent Information
- Application Number
- CN202410071163.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-18
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2044-01-18
AI Technical Summary
In the prior art, when multiple feature maps representing different features are segmented separately, the relationship between features is not obtained, resulting in the drone's inaccurate matching and detection of ground target positioning.
The multimodal information alignment method is used to construct a satellite feature base map through satellite reference map and aerial images, perform super voxel segmentation, and arrange the feature map according to the loss value to form a three-dimensional area, and combine the scene feature map to match the ground scene and target positioning.
The accuracy of the drone's target positioning on the ground is improved, and by fusing similar feature areas, unrelated areas are eliminated, and more accurate matching and target detection are achieved.
Smart Images

Figure CN119478722B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and more particularly, to a method and system for positioning ground targets by an unmanned aerial vehicle based on multimodal information alignment. Background Art
[0002] Currently, super-voxel segmentation is performed on two-dimensional images or three-dimensional point clouds and then they are matched. The segmentation mainly uses similar points under the positional relationship between images or point clouds. At the same time, after super-voxel segmentation is performed on the feature maps for extracting features, positional matching can also be carried out. However, when super-voxel segmentation is performed on multiple feature maps representing different features respectively, the relationships between features are not obtained, and the influence of features on target detection and positioning is not detected, which will lead to inaccurate matching and target detection. Summary of the Invention
[0003] The purpose of the present invention is to provide a method and system for positioning ground targets by an unmanned aerial vehicle based on multimodal information alignment to solve the above problems existing in the prior art.
[0004] In a first aspect, an embodiment of the present invention provides a method for positioning ground targets by an unmanned aerial vehicle based on multimodal information alignment, including:
[0005] Obtaining multimodal information; the multimodal information includes a satellite reference map, an aerial image, the altitude of the unmanned aerial vehicle, the attitude of the unmanned aerial vehicle, and aerial elevation information; the aerial elevation information is the elevation information of the area corresponding to the aerial image;
[0006] Constructing a satellite feature base map according to the satellite reference map and the aerial elevation information; the satellite feature base map contains multiple two-dimensional feature maps;
[0007] Performing scene feature extraction on the aerial image based on the altitude of the unmanned aerial vehicle, the attitude of the unmanned aerial vehicle, and the aerial image to obtain a scene feature map;
[0008] Performing super-voxel segmentation on the multiple two-dimensional feature maps in the satellite feature base map respectively to obtain super-voxel segmentation feature maps;
[0009] Based on the super-voxel segmentation feature maps, selectively forming different three-dimensional regions according to the two-dimensional feature maps to obtain a segmented satellite feature map;
[0010] Performing ground scene matching and target positioning based on the segmented satellite feature map and the scene feature map.
[0011] Optionally, the step of selectively forming different three-dimensional regions according to the two-dimensional feature maps based on the super-voxel segmentation feature maps to obtain a segmented satellite feature map includes:
[0012] Adjust the arrangement of the two-dimensional feature maps in the supervoxel segmentation feature map to obtain a supervoxel feature map; the supervoxel feature map is a three-dimensional feature map;
[0013] The supervoxel feature map contains multiple segmented two-dimensional regions; the segmented two-dimensional regions are the regions after segmentation of the two-dimensional feature maps in the supervoxel segmentation feature map;
[0014] If the area of the segmented two-dimensional region is greater than the segmentation threshold, mark the segmented two-dimensional region as a two-dimensional region to be fused;
[0015] Perform three-dimensional fusion on multiple two-dimensional regions to be fused to obtain multiple fused three-dimensional regions;
[0016] Replace the corresponding two-dimensional regions to be fused with multiple fused three-dimensional regions to form a segmented satellite feature map.
[0017] Optionally, the adjusting the arrangement of the two-dimensional feature maps in the supervoxel segmentation feature map to obtain a supervoxel feature map includes:
[0018] Input the two-dimensional feature maps in the supervoxel segmentation feature map into a trained first object detection network respectively to obtain multiple loss values; one loss value corresponds to one two-dimensional feature map;
[0019] Arrange the multiple loss values from small to large to obtain a supervoxel feature map;
[0020] Obtain important two-dimensional feature maps; the important two-dimensional feature maps are the two-dimensional feature maps with loss values less than those of other two-dimensional feature maps.
[0021] Optionally, the performing three-dimensional fusion on multiple two-dimensional regions to be fused to obtain multiple fused three-dimensional regions includes:
[0022] Based on the two-dimensional regions to be fused, number them according to their positions to obtain multiple numbered regions; the two-dimensional regions to be fused with the same number coincide in position; the supervoxel segmentation feature maps corresponding to the two-dimensional regions to be fused with the same number are different;
[0023] Obtain multiple first numbered regions; the first numbered regions are the regions with the same number and the number less than other numbers in the supervoxel segmentation feature map; one supervoxel segmentation feature map corresponds to one first numbered region;
[0024] Taking the arranged supervoxel segmentation feature map as the abscissa and the area of the first numbered region as the ordinate to establish a histogram to obtain a first numbered region histogram;
[0025] Based on the first numbered region histogram, judge whether three-dimensional regions are formed to obtain multiple first three-dimensional fusion regions;
[0026] Repeat the above method. By repeatedly judging multiple labeled regions with the same label, multiple fused three-dimensional regions are obtained; the multiple fused three-dimensional regions include multiple first three-dimensional fused regions.
[0027] Optionally, judging whether the region constitutes a three-dimensional region based on the first labeled region histogram to obtain multiple first three-dimensional fused regions includes:
[0028] Obtain a fusion set; the fusion set is an empty set;
[0029] Add the values corresponding to the important two-dimensional feature maps in the first labeled region histogram to the fusion set;
[0030] Obtain adjacent values; the adjacent values are the values in the first labeled region histogram that have not been added to the fusion set and are horizontally adjacent to the value newly added to the fusion set;
[0031] Judge whether the absolute value of the value newly added to the fusion set and the adjacent value is less than the fusion threshold;
[0032] If the absolute value of the value newly added to the fusion set and the adjacent value is less than or equal to the fusion threshold, add the adjacent unadded fusion set ordinates to the fusion set;
[0033] Construct the two-dimensional regions to be fused corresponding to the fusion set into three-dimensional regions to obtain the first three-dimensional fused regions;
[0034] If the absolute value of the value newly added to the fusion set and the adjacent value is greater than the fusion threshold, obtain a new fusion set, add the adjacent unadded fusion set ordinates to the new fusion set, and repeat the above method to obtain multiple first three-dimensional fused regions; the new fusion set is an empty set.
[0035] Optionally, numbering the regions to be fused according to their positions based on the two-dimensional regions to be fused to obtain multiple labeled regions includes:
[0036] Number the two-dimensional regions to be fused in the important two-dimensional feature maps to obtain multiple important two-dimensional regions;
[0037] Find the overlapping area between the important two-dimensional regions and the two-dimensional regions to be fused other than the important two-dimensional feature maps to obtain the overlapping area;
[0038] Divide the overlapping area by the area of the important two-dimensional regions to obtain the overlapping value;
[0039] If the overlapping value is greater than the overlapping threshold, mark the two-dimensional regions to be fused corresponding to the overlapping value with the same label as the important two-dimensional regions.
[0040] Optionally, the labels of the two-dimensional regions to be fused at different positions in each important two-dimensional feature map are different.
[0041] Optionally, the extracting of scene features from the aerial image based on the UAV altitude, UAV attitude and aerial image to obtain a scene feature map includes:
[0042] Performing orthorectification on the aerial image according to the UAV altitude and UAV attitude to obtain a rectified aerial image;
[0043] Passing the rectified aerial image through a histogram of oriented gradients operator to obtain a histogram of oriented gradients;
[0044] Passing the rectified aerial image through an LBP operator to obtain an LBP feature map;
[0045] The histogram of oriented gradients and the LBP feature map have the same size;
[0046] Superimposing the histogram of oriented gradients and the LBP feature map to obtain a superimposed feature map;
[0047] Performing supervoxel segmentation on the histogram of oriented gradients and the LBP feature map in the superimposed feature map respectively and constructing different three-dimensional regions to obtain a scene feature map.
[0048] Optionally, the performing of ground scene matching and target positioning based on the segmented satellite feature map and the scene feature map includes:
[0049] Matching the segmented satellite feature map and the scene feature map to obtain a matching position;
[0050] Passing the segmented satellite feature map and the scene feature map corresponding to the matching position through a second target detection network for target positioning to obtain a target position.
[0051] In a second aspect, an embodiment of the present invention provides a UAV ground target positioning system based on multi-modal information alignment, including:
[0052] An acquisition module: acquiring multi-modal information; the multi-modal information includes a satellite reference map, an aerial image, a UAV altitude, a UAV attitude and aerial elevation information; the aerial elevation information is the elevation information of the area corresponding to the aerial image;
[0053] A satellite feature base map construction module: constructing a satellite feature base map according to the satellite reference map and the aerial elevation information; the satellite feature base map includes multiple two-dimensional feature maps;
[0054] A scene feature map acquisition module: extracting scene features from the aerial image based on the UAV altitude, UAV attitude and aerial image to obtain a scene feature map;
[0055] 2D Feature Map Segmentation Module: Perform supervoxel segmentation on multiple 2D feature maps in the satellite feature base map respectively to obtain supervoxel segmentation feature maps;
[0056] 3D Region Segmentation Module: Based on the supervoxel segmentation feature maps, selectively form different 3D regions according to the 2D feature maps to obtain segmented satellite feature maps;
[0057] Ground Scene Matching and Target Location Module: Based on the segmented satellite feature maps and scene feature maps, perform ground scene matching and target location
[0058] Compared with the prior art, the embodiments of the present invention achieve the following beneficial effects:
[0059] The embodiments of the present invention also provide a UAV ground target location method and system based on multi-modal information alignment. The method includes: obtaining multi-modal information; the multi-modal information includes a satellite reference map, an aerial image, the UAV altitude, the UAV attitude, and aerial elevation information; the aerial elevation information is the elevation information of the area corresponding to the aerial image; constructing a satellite feature base map according to the satellite reference map and the aerial elevation information; the satellite feature base map contains multiple 2D feature maps; performing scene feature extraction on the aerial image based on the UAV altitude, the UAV attitude, and the aerial image to obtain a scene feature map; performing supervoxel segmentation on multiple 2D feature maps in the satellite feature base map respectively to obtain supervoxel segmentation feature maps; based on the supervoxel segmentation feature maps, selectively form different 3D regions according to the 2D feature maps to obtain segmented satellite feature maps; based on the segmented satellite feature maps and the scene feature maps, perform ground scene matching and target location.
[0060] The present invention uses multi-modal information for supervoxel segmentation, matches and aligns the satellite reference map and the aerial image, and thus locates the UAV ground target. Perform supervoxel segmentation on multiple 2D feature maps in the satellite feature base map constituted by the satellite reference map. After arranging the multiple loss values from small to large, combine the 2D regions of the supervoxel segmentation in the feature direction to obtain 3D regions. It can fuse regions with similar segmentation situations and ignore other regions. Different features are segmented into regions at similar positions, indicating that these features have the same role in detecting this position and can be detected as a 3D region. And arranging them in descending order of influence to judge the same role of these features can construct the relationship of the regions that can best represent the importance. The combination of 3D regions and 2D regions can make the subsequent matching more accurate, and thus the target location more accurate. Description of the Drawings
[0061] Figure 1 is a flowchart of a UAV ground target location method based on multi-modal information alignment provided by the embodiments of the present invention.
[0062] Figure 2 It is a schematic block diagram of an electronic device provided by an embodiment of the present invention.
[0063] Markings in the figure: bus 500; receiver 501; processor 502; transmitter 503; memory 504; bus interface 505. Specific implementation manner
[0064] The present invention will be described in detail below with reference to the accompanying drawings.
[0065] Embodiment 1
[0066] As Figure 1 shown, an embodiment of the present invention provides a method for positioning a UAV ground target based on multi-modal information alignment. The method includes:
[0067] S101: Obtain multi-modal information; the multi-modal information includes a satellite reference map, an aerial image, the UAV altitude, the UAV attitude, and aerial elevation information; the aerial elevation information is the elevation information of the area corresponding to the aerial image.
[0068] Among them, the aerial elevation information is the elevation information of the area corresponding to the aerial image taken by the UAV at the current position, the current UAV altitude, and the current UAV attitude.
[0069] S102: Construct a satellite feature base map according to the satellite reference map and the aerial elevation information; the satellite feature base map includes multiple two-dimensional feature maps.
[0070] Among them, the satellite feature base map is a three-dimensional image, and each layer of two-dimensional feature map represents different features.
[0071] Among them, the satellite feature base map refers to a map marked and drawn using feature points in satellite images and aerial elevation information, and contains multiple feature information.
[0072] S103: Based on the UAV altitude, the UAV attitude, and the aerial image, extract scene features from the aerial image to obtain a scene feature map
[0073] S104: Perform supervoxel segmentation on each of the multiple two-dimensional feature maps in the satellite feature base map to obtain a supervoxel segmentation feature map.
[0074] Among them, the supervoxel segmentation feature map is a three-dimensional map.
[0075] S105: Based on the supervoxel segmentation feature map, selectively form different three-dimensional regions according to the two-dimensional feature maps to obtain a segmented satellite feature map.
[0076] S106: Based on the segmented satellite feature map and the scene feature map, perform ground scene matching and target positioning;
[0077] Among them, match the positions of the segmented satellite feature map and the scene feature map.
[0078] Optionally, the method for obtaining the segmented satellite feature map by selectively forming different three-dimensional regions according to the two-dimensional feature map based on the supervoxel segmentation feature map includes:
[0079] Adjust the arrangement of the two-dimensional feature maps in the supervoxel segmentation feature map to obtain a supervoxel feature map; the supervoxel feature map is a three-dimensional feature map.
[0080] Among them, because the arrangements of the feature maps are different, the three-dimensional regions formed by adding and fusing are different. Therefore, it is necessary to arrange according to the importance of the two-dimensional feature maps, so that the formed three-dimensional regions can be based on the feature with the highest importance, and the features with similar importance are fused together for subsequent determination.
[0081] The supervoxel feature map contains multiple segmented two-dimensional regions; the segmented two-dimensional regions are the regions after the two-dimensional feature maps in the supervoxel segmentation feature map are segmented;
[0082] If the area of the segmented two-dimensional region is greater than the segmentation threshold, mark the segmented two-dimensional region as a two-dimensional region to be fused.
[0083] Among them, in this embodiment, the segmentation threshold is 1 / 400 of the area of the satellite feature base map.
[0084] Among them, the number of the segmented two-dimensional regions is greater than or equal to the number of the two-dimensional regions to be fused. One two-dimensional region to be fused corresponds to one segmented two-dimensional region, but one segmented two-dimensional region does not necessarily correspond to one two-dimensional region to be fused.
[0085] Perform three-dimensional fusion on the multiple two-dimensional regions to be fused to obtain multiple fused three-dimensional regions.
[0086] Among them, the number of the two-dimensional regions to be fused is greater than or equal to the number of the fused three-dimensional regions. One fused three-dimensional region corresponds to one or more two-dimensional regions to be fused, and one two-dimensional region to be fused corresponds to one fused three-dimensional region.
[0087] Replace the corresponding two-dimensional regions to be fused with multiple fused three-dimensional regions to form a segmented satellite feature map.
[0088] Among them, the segmented satellite feature map is a three-dimensional map. The two-dimensional regions to be fused corresponding to the fused three-dimensional regions form the fused three-dimensional regions, and the other segmented two-dimensional regions remain unchanged to obtain the segmented satellite feature map.
[0089] Optionally, arranging the two-dimensional feature maps in the super-voxel segmentation feature map to obtain a super-voxel feature map includes:
[0090] Inputting the two-dimensional feature maps in the super-voxel segmentation feature map into a trained first object detection network respectively to obtain a plurality of loss values; one loss value corresponds to one two-dimensional feature map.
[0091] Wherein, the trained first object detection network is the YOLOv5 model.
[0092] Arranging the plurality of loss values from small to large to obtain a super-voxel feature map.
[0093] Optionally, performing three-dimensional fusion on the plurality of two-dimensional regions to be fused to obtain a plurality of fused three-dimensional regions includes:
[0094] Numbering the two-dimensional regions to be fused based on their positions to obtain a plurality of numbered regions; the two-dimensional regions to be fused with the same number overlap in position; the two-dimensional feature maps corresponding to the two-dimensional regions to be fused with the same number are different;
[0095] Wherein, according to the positional relationship of the two-dimensional regions to be fused in the super-voxel segmentation feature map, numbering is performed based on the feature with the smallest loss value, that is, the feature with the largest influence value and the most important feature. There are no repeated numbers in one two-dimensional feature map, and the numbers in each two-dimensional feature map start from 0.
[0096] Wherein, the numbering reflects the positional relationship of the two-dimensional regions to be fused, and it is judged whether they can be fused based on the positional relationship. In this embodiment, different features are segmented into regions at similar positions, indicating that these features have the same role in detecting this position and can be detected as a three-dimensional region. And arranging these features in descending order of influence to judge their same role can construct the relationship of the most important regions.
[0097] Obtaining a plurality of first numbered regions; the first numbered regions are the regions in the super-voxel segmentation feature map with the same number and the number is less than other numbers; one super-voxel segmentation feature map corresponds to one first numbered region;
[0098] Taking the arranged super-voxel segmentation feature map as the abscissa and the area of the first numbered region as the ordinate to establish a histogram to obtain a first numbered region histogram.
[0099] Wherein, in this embodiment, the number of feature points of the first numbered region is used to represent the area.
[0100] Based on the first numbered region histogram, judging whether three-dimensional regions are formed to obtain a plurality of first three-dimensional fusion regions.
[0101] Repeat the above method. By repeatedly judging multiple labeled regions with the same label, multiple fused three-dimensional regions are obtained; the multiple fused three-dimensional regions include multiple first three-dimensional fused regions.
[0102] Optionally, determining whether multiple first labeled regions are fused based on the first labeled region histogram to obtain multiple first three-dimensional fused regions includes:
[0103] Obtain a fusion set; the fusion set is an empty set;
[0104] Add the values corresponding to the important two-dimensional feature maps in the first labeled region histogram to the fusion set.
[0105] Among them, the value corresponding to the important two-dimensional feature map is the area of the first labeled region corresponding to the important two-dimensional feature map.
[0106] Obtain adjacent values; the adjacent values are the values whose abscissas are adjacent to the value newly added to the fusion set and that are not added to the fusion set in the first labeled region histogram;
[0107] Judge whether the absolute value of the value newly added to the fusion set and the adjacent value is greater than the fusion threshold.
[0108] Among them, in this embodiment, the fusion threshold is 1 / 1600 of the area of the satellite feature base map.
[0109] If the absolute value of the value newly added to the fusion set and the adjacent value is less than or equal to the fusion threshold, add the unadded adjacent ordinate to the fusion set.
[0110] Among them, if the absolute value of the value newly added to the fusion set and the adjacent value is less than or equal to the fusion threshold, it means that the area change between the first labeled regions corresponding to two adjacent two-dimensional feature maps is small, so the influence of the two features on the segmentation at this position is not much different.
[0111] Construct the corresponding two-dimensional regions to be fused in the fusion set into three-dimensional regions to obtain the first three-dimensional fused regions.
[0112] Among them, the method of constructing three-dimensional regions is to label the corresponding two-dimensional regions to be fused in the fusion set with the label of a three-dimensional region.
[0113] If the absolute value of the value newly added to the fusion set and the adjacent value is greater than the fusion threshold, obtain a new fusion set, add the unadded adjacent ordinate to the new fusion set, and repeat the above method to obtain multiple first three-dimensional fused regions; the new fusion set is an empty set.
[0114] Optionally, numbering the regions to be fused according to positions to obtain multiple numbered regions includes:
[0115] Label the two-dimensional regions to be fused in the important two-dimensional regions to obtain multiple important two-dimensional regions.
[0116] Among them, the labels of the two-dimensional regions to be fused in the important supervoxel segmentation feature map are not repeated.
[0117] Calculate the overlapping area between the important two-dimensional regions and the two-dimensional regions to be fused other than the important two-dimensional feature map to obtain the overlapping area.
[0118] Among them, in this embodiment, the number of feature points of the important two-dimensional region and the two-dimensional region to be fused is used to represent the area. The overlapping area is the feature points representing the overlap between the important two-dimensional region and the two-dimensional region to be fused.
[0119] Divide the overlapping area by the area of the important two-dimensional region to obtain the overlapping value.
[0120] Among them, based on the important two-dimensional region, it is necessary to divide by the important two-dimensional region to obtain the overlapping ratio.
[0121] If the overlapping value is greater than the overlapping threshold, mark the two-dimensional region to be fused corresponding to the overlapping value with the same label as the important two-dimensional region.
[0122] Among them, in this embodiment, the overlapping threshold is 0.5.
[0123] Optionally, the labels of the two-dimensional regions to be fused at different positions in each important two-dimensional feature map are different.
[0124] Optionally, the extracting the scene features from the aerial image based on the drone altitude, the drone attitude, and the aerial image to obtain a scene feature map includes:
[0125] Orthorectify the aerial image according to the drone altitude and the drone attitude to obtain a rectified aerial image.
[0126] Among them, in this embodiment, an orthorectification is performed using a collinearity model. In this embodiment, since the aerial image is obtained through a sensor on the drone, the orthorectification is performed through the drone altitude and the drone attitude. If a camera is used for shooting, the orthorectification is performed through the camera internal parameters and the camera external parameters.
[0127] Apply the Histogram of Oriented Gradient (HOG) operator to the rectified aerial image to obtain a histogram of oriented gradients.
[0128] Among them, the Histogram of Oriented Gradient (HOG) feature is a feature descriptor used for object detection in computer vision and image processing. The HOG feature is constructed by calculating and statistically analyzing the gradient direction histogram of local regions of the image.
[0129] The corrected aerial image is processed by the LBP operator to obtain an LBP feature map.
[0130] Among them, the Local Binary Pattern (LBP), and the LBP operator is an operator used to describe the local texture features of an image.
[0131] The size of the histogram of oriented gradients is the same as that of the LBP feature map;
[0132] The histogram of oriented gradients and the LBP feature map are superimposed to obtain a superimposed feature map;
[0133] Among them, the superimposing direction is the feature direction, and the superimposed feature map is three-dimensional.
[0134] The histogram of oriented gradients and the LBP feature map in the superimposed feature map are respectively subjected to supervoxel segmentation and form different three-dimensional regions to obtain a scene feature map.
[0135] Among them, the method of performing supervoxel segmentation and forming different three-dimensional regions is the same as the method of performing supervoxel segmentation on the satellite feature base map and forming different three-dimensional regions to obtain a segmented satellite feature map.
[0136] Optionally, the ground scene matching and target positioning based on the second segmented satellite feature map and the scene feature map include:
[0137] The segmented satellite feature map and the scene feature map are matched to obtain a matching position.
[0138] Among them, the similarity is judged based on the regions obtained by supervoxel segmentation, and thus the matching is performed to obtain a matching position.
[0139] The segmented satellite feature map and the scene feature map corresponding to the matching position are subjected to target positioning through a second target detection network to obtain a target position.
[0140] Among them, in this embodiment, the second target detection network is the YOLOV5 model trained by multiple historical segmented satellite feature maps and corresponding scene feature maps.
[0141] Optionally, multi-modal information: (1) satellite reference map + elevation information (2) UAV aerial scene + UAV height / point cloud information / UAV attitude.
[0142] High-precision positioning information is obtained through scene matching of the satellite reference map. Aiming at the high-precision positioning requirements of UAVs for ground targets under GNSS denial conditions, technologies such as the construction of multi-scale weighted satellite feature base maps, real-time extraction of aerial scene features, ground scene matching and target positioning are studied to achieve the high-precision positioning ability of ground targets under denial conditions and provide reliable target indication information for our equipment.
[0143] Step 1: Construction of multi-scale weighted satellite feature base map
[0144] Combine elevation information
[0145] Create a satellite feature base map (grayscale conversion, Gaussian filtering, Sobel convolution, multi-scale filtering, normalized weighting, feature direction).
[0146] Organize and manage the satellite feature base map.
[0147] Step 2: Real-time extraction of aerial scene features
[0148] Orthorectification of aerial images: UAV altitude, UAV attitude angle, payload attitude angle, camera internal parameters, camera external parameters.
[0149] Step 3: Perform supervoxel segmentation on the two types of images.
[0150] Step 4: Ground scene matching and target positioning
[0151] Cosine-based similarity measurement, search for homologous points in the satellite feature base map, registration of the aerial image and the satellite feature base map, calculation of the longitude and latitude of the ground target.
[0152] Embodiment 2
[0153] Based on the above UAV ground target positioning method based on multi-modal information alignment, an embodiment of the present invention also provides a UAV ground target positioning system based on multi-modal information alignment. The system includes an acquisition module, a satellite feature base map construction module, a scene feature map acquisition module, a two-dimensional feature map segmentation module, a three-dimensional region segmentation module, and a ground scene matching and target positioning module.
[0154] The acquisition module is used to obtain multi-modal information; the multi-modal information includes a satellite reference map, an aerial image, UAV altitude, UAV attitude, and aerial elevation information; the aerial elevation information is the elevation information of the area corresponding to the aerial image;
[0155] The satellite feature base map construction module is used to construct a satellite feature base map according to the satellite reference map and the aerial elevation information; the satellite feature base map contains multiple two-dimensional feature maps;
[0156] The scene feature map acquisition module is used to perform scene feature extraction on the aerial image based on the UAV altitude, UAV attitude, and aerial image to obtain a scene feature map;
[0157] The two-dimensional feature map segmentation module is used to perform supervoxel segmentation on the multiple two-dimensional feature maps in the satellite feature base map respectively to obtain a supervoxel segmentation feature map;
[0158] The three-dimensional region segmentation module is used to selectively form different three-dimensional regions according to the two-dimensional feature map based on the supervoxel segmentation feature map, and obtain the segmented satellite feature map;
[0159] The ground scene matching and target positioning module is used to perform ground scene matching and target positioning based on the segmented satellite feature map and the scene feature map.
[0160] An embodiment of the present invention also provides an electronic device, as Figure 2 shown, including a memory 504, a processor 502, and a computer program stored on the memory 504 and executable on the processor 502. When the processor 502 executes the program, it implements the steps of any one of the foregoing methods for positioning a ground target by an unmanned aerial vehicle based on multi-modal information alignment.
[0161] Among them, in Figure 2 , the bus architecture (represented by bus 500), bus 500 can include any number of interconnected buses and bridges. Bus 500 links together various circuits including one or more processors represented by processor 502 and a memory represented by memory 504. Bus 500 can also link together various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art, and therefore, will not be further described herein. Bus interface 505 provides an interface between bus 500 and receiver 501 and transmitter 503. Receiver 501 and transmitter 503 can be the same element, i.e., a transceiver, providing a unit for communicating with various other devices on the transmission medium. Processor 502 is responsible for managing bus 500 and general processing, while memory 504 can be used to store data used by processor 502 when performing operations.
[0162] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps of any one of the foregoing methods for positioning a ground target by an unmanned aerial vehicle based on multi-modal information alignment and the data involved above.
[0163] The algorithms and displays provided herein are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems can also be used in conjunction with the teachings herein. The structure required to construct such a system will be apparent from the above description. In addition, the present invention is not directed to any particular programming language. It should be understood that the content of the present invention described herein can be implemented using various programming languages, and the description of a particular language above is for the purpose of disclosing the best mode of the present invention.
[0164] In the specification provided herein, a number of specific details are set forth. However, it will be understood that embodiments of the invention may be practiced without these specific details. In some instances, well-known methods, structures and techniques have not been shown in detail in order not to obscure an understanding of this description.
[0165] Similarly, it should be understood that in order to streamline this disclosure and assist in understanding one or more of the various inventive aspects, in the foregoing description of exemplary embodiments of the invention, various features of the invention are sometimes grouped together in a single embodiment, figure, or description thereof. However, the disclosed method should not be construed as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as the claims reflect, inventive aspects lie in less than all the features of the single foregoing disclosed embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into this detailed description, with each claim standing on its own as a separate embodiment of the invention.
[0166] Those skilled in the art will appreciate that the modules in the devices in the embodiments can be adaptively changed and disposed in one or more devices different from the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and in addition, they can be divided into multiple sub-modules or sub-units or sub-components. Except that at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all the features disclosed in this specification (including the accompanying claims, abstract and drawings) and all the processes or units of any method or device so disclosed. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstract and drawings) can be replaced by an alternative feature that provides the same, equivalent or similar purpose.
[0167] In addition, those skilled in the art will be able to understand that although some of the embodiments herein include certain features included in other embodiments but not other features, the combination of features of different embodiments means that it is within the scope of the invention and forms different embodiments. For example, in the following claims, any one of the claimed embodiments can be used in any combination.
[0168] Each component embodiment of the present invention may be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all of the components in the device according to the embodiments of the present invention. The present invention can also be implemented as a device or device program (for example, a computer program and a computer program product) for executing part or all of the methods described herein. Such a program for implementing the present invention can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.
[0169] It should be noted that the above embodiments illustrate the present invention rather than limit the present invention, and those skilled in the art can design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in the claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present invention can be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In the unit claims listing several devices, several of these devices can be embodied by the same item of hardware. The use of the words first, second, and third, etc. does not denote any order. These words can be interpreted as names.
Claims
1. A method for positioning an unmanned aerial vehicle (UAV) to a ground target based on multi-modal information alignment, characterized in that, Including: Obtaining multimodal information; The multimodal information includes a satellite reference map, an aerial image, the altitude of the drone, the attitude of the drone, and aerial elevation information; The aerial elevation information is the elevation information of the area corresponding to the aerial image; Based on the satellite reference map and the aerial elevation information, constructing a satellite feature base map; the satellite feature base map contains multiple two-dimensional feature maps; Based on the altitude of the drone, the attitude of the drone, and the aerial image, extracting scene features from the aerial image to obtain a scene feature map; Performing supervoxel segmentation on the multiple two-dimensional feature maps in the satellite feature base map respectively to obtain supervoxel segmentation feature maps; Based on the supervoxel segmentation feature maps, selectively forming different three-dimensional regions according to the two-dimensional feature maps to obtain a segmented satellite feature map; including: Adjusting the arrangement of the two-dimensional feature maps in the supervoxel segmentation feature map to obtain a supervoxel feature map; the supervoxel feature map is a three-dimensional feature map; The supervoxel feature map contains multiple segmented two-dimensional regions; the segmented two-dimensional regions are the regions after segmentation of the two-dimensional feature maps in the supervoxel segmentation feature map; If the area of the segmented two-dimensional region is greater than the segmentation threshold, marking the segmented two-dimensional region as a two-dimensional region to be fused; Performing three-dimensional fusion on multiple two-dimensional regions to be fused to obtain multiple fused three-dimensional regions; specifically including: Based on the two-dimensional regions to be fused, numbering them according to their positions to obtain multiple numbered regions; the two-dimensional regions to be fused with the same number overlap in position; the supervoxel segmentation feature maps corresponding to the two-dimensional regions to be fused with the same number are different; Obtaining multiple first numbered regions; the first numbered region is the region with the same number and a number smaller than other numbers in the supervoxel segmentation feature map; one supervoxel segmentation feature map corresponds to one first numbered region; Taking the arranged supervoxel segmentation feature maps as the abscissa and the area of the first numbered region as the ordinate to establish a histogram to obtain a first numbered region histogram; Based on the first numbered region histogram, judging whether a three-dimensional region is formed to obtain multiple first three-dimensional fusion regions; repeating the above method, and judging multiple numbered regions with the same number multiple times to obtain multiple fused three-dimensional regions; the multiple fused three-dimensional regions include multiple first three-dimensional fusion regions Replacing the corresponding two-dimensional regions to be fused with multiple fused three-dimensional regions to form a segmented satellite feature map; Based on the segmented satellite feature map and the scene feature map, performing ground scene matching and target positioning.
2. The method for positioning a UAV ground target based on multimodal information alignment according to claim 1, wherein The adjusting the arrangement of the two-dimensional feature maps in the supervoxel segmentation feature map to obtain a supervoxel feature map includes: Inputting the two-dimensional feature maps in the supervoxel segmentation feature map into a trained first object detection network respectively to obtain multiple loss values; one loss value corresponds to one two-dimensional feature map; Arranging the multiple loss values from small to large to obtain a supervoxel feature map; Obtaining important two-dimensional feature maps; the important two-dimensional feature maps are the two-dimensional feature maps with loss values smaller than the loss values of other two-dimensional feature maps.
3. The method for positioning an unmanned aerial vehicle (UAV) against a ground target based on multimodal information alignment according to claim 1, wherein The based on the first numbered region histogram, judging whether a three-dimensional region is formed to obtain multiple first three-dimensional fusion regions includes: Obtaining a fusion set; the fusion set is an empty set; Add the values corresponding to the important two-dimensional feature maps in the first labeled region histogram to the fusion set; Obtain adjacent values; the adjacent values are the values in the first labeled region histogram that have not been added to the fusion set and whose abscissas are adjacent to the value of the latest added value in the fusion set; Determine whether the absolute value of the latest added value in the fusion set and the adjacent value is less than the fusion threshold; If the absolute value of the latest added value in the fusion set and the adjacent value is less than or equal to the fusion threshold, add the adjacent ordinate that has not been added to the fusion set to the fusion set; Construct the corresponding two-dimensional region to be fused in the fusion set into a three-dimensional region to obtain the first three-dimensional fusion region; If the absolute value of the latest added value in the fusion set and the adjacent value is greater than the fusion threshold, obtain a new fusion set, add the adjacent ordinate that has not been added to the fusion set to the new fusion set, and repeat the above method to obtain multiple first three-dimensional fusion regions; the new fusion set is an empty set.
4. The method for positioning a ground target by an unmanned aerial vehicle based on multimodal information alignment according to claim 1, wherein Based on the two-dimensional region to be fused, label it by position to obtain multiple labeled regions, including: Label the two-dimensional region to be fused in the important two-dimensional feature map to obtain multiple important two-dimensional regions; Find the overlapping area between the important two-dimensional region and the two-dimensional region to be fused other than the important two-dimensional feature map to obtain the overlapping area; Divide the overlapping area by the area of the important two-dimensional region to obtain the overlapping value; If the overlapping value is greater than the overlapping threshold, label the two-dimensional region to be fused corresponding to the overlapping value with the same label as the important two-dimensional region.
5. The method for positioning an unmanned aerial vehicle (UAV) against a ground target based on multi-modal information alignment according to claim 4, wherein The labels of the two-dimensional regions to be fused at different positions in each important two-dimensional feature map are different.
6. The method for positioning a UAV ground target based on multi-modal information alignment according to claim 1, characterized in that Based on the drone height, drone attitude, and aerial image, perform scene feature extraction on the aerial image to obtain a scene feature map, including: Perform orthorectification on the aerial image according to the drone height and drone attitude to obtain a rectified aerial image; Pass the rectified aerial image through the histogram of oriented gradients operator to obtain the histogram of oriented gradients; Pass the rectified aerial image through the LBP operator to obtain the LBP feature map; The histogram of oriented gradients and the LBP feature map have the same size; Overlay the histogram of oriented gradients and the LBP feature map to obtain an overlaid feature map; Perform supervoxel segmentation and construct different three-dimensional regions on the histogram of oriented gradients and the LBP feature map in the overlaid feature map respectively to obtain the scene feature map.
7. The method for positioning a UAV ground target based on multi-modal information alignment according to claim 1, wherein Based on the segmented satellite feature map and the scene feature map, perform ground scene matching and target positioning, including: Match the segmented satellite feature map and the scene feature map to obtain the matching position; Perform target positioning on the segmented satellite feature map and the scene feature map corresponding to the matching position through the second target detection network to obtain the target position.
8. An unmanned aerial vehicle ground target positioning system based on multi-modal information alignment, characterized in that, Including: Acquisition module: Obtain multimodal information; The multimodal information includes a satellite reference map, an aerial image, drone height, drone attitude, and aerial elevation information; The aerial elevation information is the elevation information of the area corresponding to the aerial image; Satellite feature base map construction module: Construct a satellite feature base map according to the satellite reference map and the aerial elevation information; the satellite feature base map contains multiple two-dimensional feature maps; Scene Feature Map Acquisition Module: Based on the UAV altitude, UAV attitude, and aerial image, perform scene feature extraction on the aerial image to obtain a scene feature map; Two-Dimensional Feature Map Segmentation Module: Perform supervoxel segmentation on multiple two-dimensional feature maps in the satellite feature base map respectively to obtain supervoxel segmentation feature maps; Three-Dimensional Region Segmentation Module: Based on the supervoxel segmentation feature map, selectively form different three-dimensional regions according to the two-dimensional feature maps to obtain a segmented satellite feature map; including: Adjust the arrangement of the two-dimensional feature maps in the supervoxel segmentation feature map to obtain a supervoxel feature map; the supervoxel feature map is a three-dimensional feature map; The supervoxel feature map contains multiple segmented two-dimensional regions; the segmented two-dimensional region is the region after segmentation of the two-dimensional feature map in the supervoxel segmentation feature map; If the area of the segmented two-dimensional region is greater than the segmentation threshold, mark the segmented two-dimensional region as a two-dimensional region to be fused; Perform three-dimensional fusion on multiple two-dimensional regions to be fused to obtain multiple fused three-dimensional regions; specifically including: Based on the two-dimensional regions to be fused, number them according to their positions to obtain multiple numbered regions; the two-dimensional regions to be fused with the same number coincide in position; the supervoxel segmentation feature maps corresponding to the two-dimensional regions to be fused with the same number are different; Obtain multiple first numbered regions; the first numbered region is the region in the supervoxel segmentation feature map with the same number and a number smaller than other numbers; one supervoxel segmentation feature map corresponds to one first numbered region; Take the arranged supervoxel segmentation feature map as the abscissa and the area of the first numbered region as the ordinate to establish a histogram to obtain the first numbered region histogram; Based on the first numbered region histogram, judge whether a three-dimensional region is formed to obtain multiple first three-dimensional fusion regions; repeat the above method, and judge multiple numbered regions with the same number multiple times to obtain multiple fused three-dimensional regions; the multiple fused three-dimensional regions include multiple first three-dimensional fusion regions Replace multiple fused three-dimensional regions with the corresponding two-dimensional regions to be fused to form a segmented satellite feature map; Ground Scene Matching and Target Location Module: Based on the segmented satellite feature map and the scene feature map, perform ground scene matching and target location.
Citation Information
Patent Citations
Image feature extraction method and device, storage medium and electronic equipment
CN112580641A
Multi-unmanned aerial vehicle high-precision matching positioning method
CN115187798A