A cross-modal terrain matching and positioning method
The cross-modal terrain matching method effectively addresses the precision limitations in matching digital elevation models and Doppler delay maps by using neural networks for feature extraction and cosine similarity, enhancing flight platform positioning accuracy.
Patent Information
- Application Number
- CN202510257840.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-03-05
AI Technical Summary
In the existing aircraft positioning methods, two-dimensional time-lapse Doppler diagram (DDM) and digital elevation model (DEM) belong to different modes and cannot be matched directly, resulting in limited matching accuracy, especially in non-flat terrain and complex scattering characteristics.
The cross-modal terrain matching positioning method is adopted, and the target digital elevation model and two-dimensional delay Doppler diagram are obtained through the on-board synthetic aperture radar altimeter. The encoder and decoder are used for feature encoding and decoding. Combined with the feature extraction module, the cross-modal fusion vector is calculated, and the positioning data of the flight platform is determined using cosine similarity.
It effectively solves the problem of insufficient generalization capabilities of aircraft positioning-dependent global positioning system (GPS) scenarios and the heterogeneous gap in cross-modal terrain matching, and improves the accuracy of terrain matching and positioning accuracy in complex scenarios.
Smart Images

Figure CN119762928B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of aircraft positioning, and in particular, to a cross-modal terrain matching and positioning method. Background Art
[0002] Aircraft positioning is crucial for flight safety. An airborne synthetic aperture radar altimeter (SARAL) illuminates the ground in a near-vertical manner. When the platform flies directly above a ground scatter point, the radar emits a broadband signal, and after a certain delay, the echo is received. The echo is processed to obtain a two-dimensional delay Doppler map (DDM). The characteristics of the two-dimensional delay Doppler map (DDM) reflect the unexpected undulations of the lowest point of the aircraft and can be used to detect terrain changes.
[0003] Since the two-dimensional delay Doppler map (DDM) projects the three-dimensional terrain onto a two-dimensional plane by the synthetic aperture radar altimeter (SARAL), people cannot directly interpret the terrain changes intuitively from the two-dimensional delay Doppler map (DDM). When matching the two-dimensional delay Doppler map (DDM) formed by radar echo data with terrain data, since the two belong to two heterogeneous modalities and there are differences in quantization scales, there is a semantic gap between modalities, which easily causes the problem of limited matching accuracy; especially in the case of non-flat terrain, complex changes in scattering characteristics, platform errors, etc., this problem is more serious. Summary of the Invention
[0004] In view of this, embodiments of this application provide a cross-modal terrain matching and positioning method to solve the problem that the existing digital elevation model (DEM) and two-dimensional delay Doppler map (DDM) belong to different modalities and there is a heterogeneous gap that is difficult to directly match. The following specific technical solutions are adopted.
[0005] A cross-modal terrain matching and positioning method for a flight platform, on which an airborne synthetic aperture radar altimeter is carried, the method includes:
[0006] Obtain the target digital elevation model and the target two-dimensional delay Doppler map of the flight platform in the current terrain, where the two-dimensional delay Doppler map is obtained by the airborne synthetic aperture radar altimeter scanning the current terrain by radar;
[0007] Encode the target digital elevation model and the target two-dimensional delay Doppler map into a first coding space through a first encoder and a second encoder respectively, to obtain a first coding feature corresponding to the target digital elevation model and a second coding feature corresponding to the target two-dimensional delay Doppler map;
[0008] In the first coding space, calculate first feature data corresponding to the first coding feature and calculate second feature data corresponding to the second coding feature;
[0009] Based on the second feature data, a first cross-modal fusion vector corresponding to the first encoded feature is calculated, and based on the first feature data, a second cross-modal fusion vector corresponding to the second encoded feature is calculated;
[0010] The first cross-modal fusion vector and the second cross-modal fusion vector are respectively decoded by a first decoder and a second decoder to obtain a first decoded feature corresponding to the first cross-modal fusion vector and a second decoded feature corresponding to the second cross-modal fusion vector;
[0011] The first decoded feature and the second decoded feature are respectively subjected to feature extraction by different feature extraction modules to obtain a first feature vector of the target digital elevation model and a second feature vector of the target two-dimensional delayed Doppler map;
[0012] Based on the first feature vector and the second feature vector, the positioning data of the flight platform is determined.
[0013] In the above solution, the cross-modal terrain matching and aircraft positioning network fuses the features of the cross-modal digital elevation model (DEM) and the two-dimensional delayed Doppler map (DDM), then performs feature extraction, and based on the extracted features, obtains the positioning result, effectively solving the problems such as the dependence of aircraft positioning on the global positioning system (GPS), insufficient scene generalization ability, and the existence of heterogeneous gaps in cross-modal terrain matching.
[0014] In some embodiments, the step of respectively encoding the target digital elevation model and the target two-dimensional delayed Doppler map into a first encoding space by a first encoder and a second encoder to obtain a first encoded feature corresponding to the target digital elevation model and a second encoded feature corresponding to the target two-dimensional delayed Doppler map specifically includes:
[0015] The target digital elevation model is encoded by a first encoder to obtain a first encoded feature map;
[0016] And, the target two-dimensional delayed Doppler map is encoded by a second encoder to obtain a second encoded feature map;
[0017] The first encoded feature map and the second encoded feature map are respectively subjected to matrix multiplication calculations by unit matrices of different dimensions to obtain a first encoded feature corresponding to the target digital elevation model and a second encoded feature corresponding to the target two-dimensional delayed Doppler map, and the first encoded feature and the second encoded feature have the same dimension.
[0018] In some embodiments, the steps of calculating the first feature data corresponding to the first encoded feature and calculating the second feature data corresponding to the second encoded feature in the first encoding space specifically include:
[0019] For the first encoded feature, perform a linear mapping through a first weight matrix to obtain a first mapping matrix corresponding to the first encoded feature. For the second encoded feature, perform a linear mapping through a second weight matrix and a third weight matrix to obtain a second mapping matrix and a third mapping matrix corresponding to the second encoded feature;
[0020] For the second encoded feature, perform a linear mapping through the first weight matrix to obtain a fourth mapping matrix corresponding to the second encoded feature. For the first encoded feature, perform a linear mapping through a second weight matrix and a third weight matrix to obtain a sixth mapping matrix and a third mapping matrix corresponding to a fifth encoded feature.
[0021] In some embodiments, the steps of calculating a first cross-modal fusion vector corresponding to the first encoded feature based on the second feature data and calculating a second cross-modal fusion weight matrix corresponding to the second encoded feature based on the first feature data specifically include:
[0022] Based on the first mapping matrix and the second mapping matrix, obtain a second cross-modal fusion weight matrix corresponding to the first encoded feature;
[0023] Based on the fourth mapping matrix and the fifth mapping matrix, obtain a first cross-modal fusion weight matrix corresponding to the second encoded feature;
[0024] Perform matrix multiplication on the second cross-modal fusion weight matrix and the third mapping matrix, and perform a normalization process on the matrix multiplication result to obtain a first cross-modal fusion vector;
[0025] And perform matrix multiplication on the first cross-modal fusion weight matrix and the sixth mapping matrix, and perform a normalization process on the matrix multiplication result to calculate a second cross-modal fusion vector.
[0026] In some embodiments, further, the steps of respectively performing feature extraction on the first decoded feature and the second decoded feature through different feature extraction modules to obtain a first feature vector of the target digital elevation model and a second feature vector of the target two-dimensional delayed Doppler map specifically include:
[0027] Perform feature extraction processing on the first decoded data through a first feature extraction module to obtain a first feature vector of the target digital elevation model, and the first feature vector corresponds to the first decoded data;
[0028] The second feature extraction module performs feature extraction processing on the second decoded data to obtain a second feature vector of the target two-dimensional delayed Doppler map, and the second feature vector corresponds to the second decoded data.
[0029] In some embodiments, the step of determining the positioning data of the flying platform based on the first feature vector and the second feature vector specifically includes:
[0030] Calculate the cosine similarity between the first feature vector and the second feature vector;
[0031] Based on the cosine similarity, determine the positioning data of the flying platform.
[0032] In some embodiments, the step of calculating the cosine similarity between the first feature vector and the second feature vector specifically includes:
[0033] The cosine similarity satisfies the following formula:
[0034]
[0035] where, represents the cosine similarity between the first feature vector and the second feature vector, represents the first feature vector, represents the second feature vector, represents the coordinates in the target two-dimensional delayed Doppler map, represents the coordinates of the target digital elevation model.
[0036] In some embodiments, the step of determining the positioning data of the flying platform based on the cosine similarity specifically includes:
[0037] For each position coordinate in the target two-dimensional delayed Doppler map, based on the second feature vector, select multiple second feature vectors with the largest similarity;
[0038] Based on multiple second feature vectors with the largest similarity, inversely retrieve the candidate coordinates mapped in the target digital elevation model, and each candidate coordinate corresponds to one second feature vector;
[0039] Based on multiple candidate coordinates, determine the positioning data of the flying platform.
[0040] In some embodiments, the step of determining the positioning data of the flying platform based on multiple candidate coordinates specifically includes:
[0041] Perform weighted calculation on multiple candidate coordinates to obtain a weighted coordinate value;
[0042] Based on the weighted coordinate values, the positioning data of the flight platform is determined.
[0043] The cross-modal terrain matching positioning method of the present application obtains the target digital elevation model and the target two-dimensional delayed Doppler map of the flight platform in the current terrain; performs cross-modal fusion on the target digital elevation model and the target two-dimensional delayed Doppler map to obtain fusion matching data; respectively performs feature extraction on the fusion matching data through different feature extraction modules to obtain the first feature vector of the target digital elevation model and the second feature vector of the target two-dimensional delayed Doppler map; based on the first feature vector and the second feature vector, the positioning data of the flight platform is determined. The present application fuses the features of the cross-modal digital elevation model and the two-dimensional delayed Doppler map, and obtains the positioning result based on the features extracted after feature fusion. This method effectively solves the problems such as the dependence of aircraft positioning on the Global Positioning System (GPS), insufficient scene generalization ability, and the existence of heterogeneous gaps in cross-modal terrain matching. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0045] Figure 1 is a schematic flowchart of a cross-modal terrain matching positioning method provided by an embodiment of the present application;
[0046] Figure 2 is a schematic diagram of a model architecture for a cross-modal terrain matching positioning method provided by an embodiment of the present application;
[0047] Figure 3 is a schematic diagram of a simulated two-dimensional delayed Doppler map provided by an embodiment of the present application;
[0048] Figure 4 is a schematic diagram of a measured two-dimensional delayed Doppler map provided by an embodiment of the present application;
[0049] Figure 5 is a schematic diagram of a regional digital elevation model (DEM) and a certain flight track provided by an embodiment of the present application;
[0050] Figure 6 is a result diagram of terrain matching and real-time aircraft positioning using this embodiment provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without making creative efforts shall fall within the protection scope of the present application.
[0052] Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present application. It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof. In the description of the present application, it should be understood that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present application. The terms "first", "second", "third", and "fourth" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. In the description of the present application, unless otherwise stated, the meaning of "a plurality" is two or more, unless otherwise clearly defined.
[0053] The Synthetic Aperture Radar Altimeter (SARAL) is not restricted by illumination and climate conditions and can generate a two-dimensional Delay-Doppler Map (DDM). The airborne Synthetic Aperture Radar Altimeter (SARAL) illuminates the ground in a near-vertical manner. When the platform flies directly above the ground scatter point, the radar emits a broadband signal, receives the echo after a certain delay, and processes the echo to obtain a two-dimensional Delay-Doppler Map (DDM). The characteristics of the two-dimensional Delay-Doppler Map (DDM) reflect the unexpected undulations of the lowest point of the aircraft and can be used to detect terrain changes.
[0054] A Digital Elevation Model (DEM) is a digital representation of the elevation of the Earth's surface, which records elevation information at different locations in a discrete form. At the same time, DEM data is available globally. Since the DEM reflects the ground elevation, in the case of not relying on the Global Positioning System (GPS), this application uses the DEM as a reference to match with the real-time two-dimensional delayed Doppler map (DDM) to achieve aircraft positioning.
[0055] Deep learning has strong non-linear fitting and massive data processing capabilities, attracting researchers to use deep learning to solve image matching problems in different scenarios. These methods can be generally divided into two categories: single-modal matching and positioning, and cross-modal matching and positioning methods. Specifically, the single-modal matching and positioning method makes predictions by analyzing data from a single information source or form, while the cross-modal matching and positioning method needs to combine information from two or more modalities for prediction. In addition, existing image matching methods all rely on visual images that can be recognized by the human eye and understood by the human brain. However, the two-dimensional delayed Doppler map (DDM) is formed by the Synthetic Aperture Radar Altimeter (SARAL) projecting the three-dimensional terrain onto a two-dimensional plane, and people cannot directly interpret the terrain changes from the DDM intuitively.
[0056] Due to the following two reasons, it is difficult to match the DDM and the DEM in the existing solutions. First, the DDM and the DEM belong to different modalities and cannot be directly matched. Second, deep learning lacks formula guidance and has limited interpretability, and it is difficult to achieve the mapping between the DEM and the DDM with the help of this. Third, there is no publicly available dataset and dataset generation method on the market, and it is difficult to obtain sufficient training datasets.
[0057] To meet the practical needs of cross-modal terrain matching and real-time aircraft positioning, this application directly performs cross-modal matching between the two-dimensional delayed Doppler map (DDM) formed by the Synthetic Aperture Radar Altimeter (SARAL) echo and the Digital Elevation Model (DEM), deeply studies the neural network representation learning method, studies the elevation features that improve the data representation ability, establishes a neural network model for explicit matching and positioning, establishes a loss function introducing quantization features, realizes the explicit derivation between the optimal matching position coordinates and the terrain data, ensures that the objects of terrain matching have the same physical meaning, the parameters can be quantitatively evaluated, effectively improves the cross-modal two-dimensional delayed Doppler map (DDM) terrain data matching accuracy, complex scene generalization ability and network interpretability, and provides reliable theoretical and technical support for practical scene matching and positioning.
[0058] As Figure 1 shown, Figure 1It is a schematic flowchart of a cross-modal terrain matching and positioning method provided by an embodiment of the present application. The cross-modal terrain matching and positioning method is used for a flight platform, and an airborne synthetic aperture radar altimeter is carried on the flight platform. The method includes the following steps 101 to 107.
[0059] Step 101, obtain the target digital elevation model and the target two-dimensional delay Doppler map of the flight platform in the current terrain.
[0060] In the embodiment of the present application, the digital elevation model (DEM) is a digital representation of the elevation of the earth's surface, which records the elevation information at different positions in a discrete form. At the same time, digital elevation model (DEM) data can be obtained globally. Therefore, the target digital elevation model of the flight platform in the current terrain can be directly obtained.
[0061] The two-dimensional delay Doppler map is obtained by the airborne synthetic aperture radar altimeter scanning the current terrain by radar. The synthetic aperture radar altimeter emits a linear frequency modulated wide beam pulse to the current terrain range in a nearly vertical observation mode, and receives the echo signal after a certain time delay, and obtains the corresponding target two-dimensional delay Doppler map according to the echo signal. The target two-dimensional delay Doppler map is the measured two-dimensional delay Doppler map, and its specific calculation method is the same as that of the measured two-dimensional delay Doppler map.
[0062] Step 102, respectively encode the target digital elevation model and the target two-dimensional delay Doppler map into the first coding space through the first encoder and the second encoder, to obtain the first coding feature corresponding to the target digital elevation model and the second coding feature corresponding to the target two-dimensional delay Doppler map.
[0063] In the embodiment of the present application, the encoder 1 (the first encoder) is used to encode the target digital elevation model (DEM), and it is gradually compressed into a 50-dimensional low-dimensional representation through multiple linear transformations and activation functions in the encoder 1 to obtain the first coding feature of the target digital elevation model. The encoder 2 (the second encoder) is used to encode the target two-dimensional delay Doppler map (DDM), and the second coding feature of the target two-dimensional delay Doppler map is obtained through multiple linear transformations and activation functions in the encoder 2.
[0064] For the output coding data of the encoder 1 and the encoder 2, a linear transformation is performed to obtain the first coding feature and the second coding feature with the same dimension. The first coding feature and the second coding feature with the same dimension can be used for cross-modal fusion.
[0065] Step 103, in the first coding space, calculate the first feature data corresponding to the first coding feature and calculate the second feature data corresponding to the second coding feature.
[0066] In the embodiment of the present application, the first feature data includes the feature data of the target Digital Elevation Model (DEM), and the second feature data includes the feature data of the target two-dimensional Delay-Doppler Map (DDM). Specifically, the first feature data may include a first mapping matrix, a second mapping matrix, and a third mapping matrix. The second feature data may include a fourth mapping matrix, a fifth mapping matrix, and a sixth mapping matrix.
[0067] Step 104: Based on the second feature data, calculate the first cross-modal fusion vector corresponding to the first encoded feature, and based on the first feature data, calculate the second cross-modal fusion vector corresponding to the second encoded feature.
[0068] In the embodiment of the present application, based on the second feature data, calculate the first cross-modal fusion vector corresponding to the first encoded feature, so that the first cross-modal fusion vector fuses the feature data of the target two-dimensional Delay-Doppler Map (DDM) on the basis of the first encoded feature. Based on the first feature data, calculate the second cross-modal fusion weight matrix corresponding to the second encoded feature, so that the second cross-modal fusion vector fuses the feature data of the target Digital Elevation Model (DEM) on the basis of the second encoded feature.
[0069] In this embodiment, since the two-dimensional Delay-Doppler Map (DDM) has relatively rich low-level features, the feature fusion of the Digital Elevation Model (DEM) and the two-dimensional Delay-Doppler Map (DDM) adopts the Figure 2 feature fusion algorithm shown. Figure 2 It is a schematic diagram of the model architecture of a cross-modal terrain matching and positioning method provided by the embodiment of the present application, showing a schematic diagram of feature fusion. This feature fusion utilizes the connection and interaction between the low-level features of each modality, has a low algorithm complexity, and only requires single-model training. In the embodiment of the present application, an encoder-decoder structure can be used to establish a connection between the features of the two-dimensional Delay-Doppler Map (DDM) and the Digital Elevation Model (DEM), and obtain two cross-modal fusion vectors that fuse the feature information of the two modalities. Specifically, a first encoder structure is used to fuse the features of the Digital Elevation Model (DEM) and the two-dimensional Delay-Doppler Map (DDM) to obtain a first cross-modal fusion vector containing the features of both, and a second encoder structure is used to fuse the features of the Digital Elevation Model (DEM) and the two-dimensional Delay-Doppler Map (DDM) to obtain a second cross-modal fusion vector containing the features of both. Then, a decoder structure is used to decode the fused cross-modal fusion vector, and finally a new feature vector containing the information of the other party after fusion is obtained. Therefore, the terrain matching and aircraft positioning algorithm proposed in this embodiment is applicable to non-visual and different-modal data matching, and can specifically implement the matching between the Digital Elevation Model (DEM) and the two-dimensional Delay-Doppler Map (DDM).
[0070] Subsequently, inspired by the dot product operation of the attention mechanism, the cross-modal fusion weight matrix is calculated through the following formula:
[0071]
[0072] where represents the dimension of the weight matrix. And the calculated cross-modal fusion weight matrix is used to and perform weighted summation, and finally the first cross-modal fusion vector and the second cross-modal fusion vector after fusion are obtained.
[0073] Step 105: Respectively perform decoding processing on the first cross-modal fusion vector and the second cross-modal fusion vector through the first decoder and the second decoder to obtain the first decoded feature corresponding to the first cross-modal fusion vector and the second decoded feature corresponding to the second cross-modal fusion vector.
[0074] In the embodiment of the present application, the first decoder 1 and the second decoder 2 are respectively used to decode the first cross-modal fusion vector and the second cross-modal fusion vector to obtain the first decoded data fused with the target two-dimensional delayed Doppler map and the target digital elevation model, and the second decoded data fused with the target digital elevation model and the target two-dimensional delayed Doppler map. Specifically, both the first decoder 1 and the second decoder 2 are composed of a linear layer and an activation layer, and each linear layer is followed by an activation layer. The first decoder 1 obtains the first decoded feature with a size of 256*512 through 2 linear layers and 2 activation layers. The second decoder 2 obtains the second decoded feature with a size of 512*2048 through 2 linear layers and 2 activation layers.
[0075] Step 106: Respectively perform feature extraction on the first decoded feature and the second decoded feature through different feature extraction modules to obtain the first feature vector of the target digital elevation model and the second feature vector of the target two-dimensional delayed Doppler map.
[0076] In the embodiment of the present application, the network architecture 1 can be used to perform feature extraction on the first decoded feature fused with the target two-dimensional delayed Doppler map and the target digital elevation model to obtain the first feature vector, and the network architecture 2 can be used to perform feature extraction on the second decoded feature fused with the target digital elevation model and the target two-dimensional delayed Doppler map to obtain the second feature vector.
[0077] Step 107: Determine the positioning data of the flight platform based on the first feature vector and the second feature vector.
[0078] In the embodiments of the present application, after feature extraction, the feature vectors of the reference digital elevation model (DEM) (the first feature vectors) and the real-time two-dimensional delayed Doppler map (DDM) (the second feature vectors) can be obtained respectively.
[0079] In order to measure the similarity between the feature vectors of the reference digital elevation model (DEM) and the real-time two-dimensional delayed Doppler map (DDM) belonging to different modalities after feature extraction, for obtaining the terrain matching result, as Figure 2 shown, the cosine similarity is used to measure the similarity between the feature vectors of the digital elevation model (DEM) and the two-dimensional delayed Doppler map (DDM). Different from the similarity measurement algorithm that directly solves the linear distance between two data in the n-dimensional space, the cosine similarity measures the difference between the two by solving the cosine value of the included angle between the two vectors in the vector space, which reflects the difference in direction in the vector space and does not consider the magnitude of the vectors. This characteristic is adapted to the lower attention degree of cross-modal to the magnitudes of heterogeneous data.
[0080] In order to measure the similarity between the first feature vector and the second feature vector, the cosine similarity is used for similarity measurement. The cosine similarity satisfies the following formula:
[0081]
[0082] where represents the cosine similarity between the first feature vector and the second feature vector, represents the first feature vector, represents the second feature vector, represents the coordinates in the target two-dimensional delayed Doppler map, represents the coordinates of the target digital elevation model.
[0083] During the training process, the mean square error (MSE) loss and the contrast loss function are weighted as the global loss function. The MSE loss is a commonly used regression loss measurement method in machine learning. It can represent the square of the absolute difference between the predicted value and the true value. The smaller the loss value, the more accurate the model prediction result. The MSE loss satisfies the following formula:
[0084]
[0085] where N is the total number of features of the first feature vector, the second feature vector, is the first feature vector. The MSE loss can also be used to evaluate the deviation between the positioning coordinates and the reference coordinates, and the expression form is as follows:
[0086]
[0087] Where M is the number of terrain coordinates, denotes retrieving candidate coordinates mapped in the target digital elevation model in reverse based on multiple second feature vectors with the maximum similarity, denotes the positioning coordinates.
[0088] The matching degree of the two-modal mapping matrix is evaluated using a contrast loss function, and the expression form of the contrast loss is as follows:
[0089]
[0090] where the label denotes is a similar sample, that is, the first and second feature vectors with the same label, denotes is a dissimilar sample. denotes the Euclidean distance between the two samples, is a hyperparameter that defines the minimum distance of the sample pair.
[0091] Based on the established MSE loss and contrast loss, the global loss function can be modeled as:
[0092]
[0093] where, is the MSE loss coefficient of the first and second feature vectors, is the contrast loss coefficient of the first and second feature vectors, is the MSE loss coefficient of the candidate coordinates and the positioning coordinates.
[0094] After forward propagation through the above steps to obtain the first and second feature vectors and the candidate coordinates and calculate the global loss, the gradients of the parameters are calculated layer by layer using the backpropagation algorithm. The optimizer adjusts the parameters of each layer according to the gradients, and the global loss is backpropagated to the weights of each layer of the encoder, decoder, feature fusion module, and feature extraction module to obtain the optimized first feature vector and the second feature vector , as well as the candidate coordinates mapped in the target digital elevation model retrieved in reverse through cosine similarity , for a new round of global loss function calculation. The above steps are continuously repeated until the loss function converges.
[0095] In the embodiments of the present application, a target digital elevation model and a target two-dimensional delayed Doppler map of the flight platform in the current terrain are obtained. The two-dimensional delayed Doppler map is obtained by performing radar scanning on the current terrain using an airborne synthetic aperture radar altimeter. The target digital elevation model and the target two-dimensional delayed Doppler map are fused across modalities to obtain fused matching data. Different feature extraction modules are used to extract features from the fused matching data respectively, to obtain a first feature vector of the target digital elevation model and a second feature vector of the target two-dimensional delayed Doppler map. Based on the first feature vector and the second feature vector, the positioning data of the flight platform is determined. The present application fuses the features of the cross-modal digital elevation model (DEM) and the two-dimensional delayed Doppler map (DDM), then performs feature extraction, and measures the similarity of the feature vectors after feature extraction. According to the similarity score, the positioning result is obtained, effectively solving the problems such as the dependence of aircraft positioning on the Global Positioning System (GPS), insufficient scene generalization ability, and the existence of heterogeneous gaps in cross-modal terrain matching.
[0096] It can be understood that in the specific implementation of the present application, relevant data such as terrain data, flight data, and digital elevation model data are involved. When the embodiments of the present application are applied to specific products or technologies, user permission or consent needs to be obtained, and the collection, use, and processing of relevant data, as well as the training and use of various models, need to comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0097] Optionally, the step of encoding the target digital elevation model and the target two-dimensional delayed Doppler map into the first encoding space respectively through the first encoder and the second encoder to obtain the first encoding feature corresponding to the target digital elevation model and the second encoding feature corresponding to the target two-dimensional delayed Doppler map specifically includes: encoding the target digital elevation model through the first encoder to obtain a first encoding feature map; and encoding the target two-dimensional delayed Doppler map through the second encoder to obtain a second encoding feature map; performing matrix multiplication calculations on the first encoding feature map and the second encoding feature map respectively through unit matrices of different dimensions to obtain the first encoding feature corresponding to the target digital elevation model and the second encoding feature corresponding to the target two-dimensional delayed Doppler map, and the first encoding feature and the second encoding feature have the same dimension.
[0098] In the embodiments of the present application, as Figure 2As shown in the figure, the encoder 1 (the first encoder) is used to encode the target digital elevation model (DEM). Through the multi-layer linear transformation and activation function in the encoder 1, it is gradually compressed into a 50-dimensional low-dimensional representation to obtain the first encoded feature of the target digital elevation model. The encoder 2 (the second encoder) is used to encode the target two-dimensional delayed Doppler map (DDM). Through the multi-layer linear transformation and activation function in the encoder 2, the second encoded feature of the target two-dimensional delayed Doppler map is obtained. Specifically, both the encoder 1 and the encoder 2 are composed of a linear layer and an activation layer, and each linear layer is followed by an activation layer. The encoder 1 reduces the first encoded data from 512*256 to 256*50 through 2 layers of linear layers and 2 layers of activation layers. The encoder 2 reduces the second encoded feature from 2048*512 to 512*50 through 2 layers of linear layers and 2 layers of activation layers.
[0099] For the output encoded data of the encoder 1 and the encoder 2, assuming that the feature map output by the encoder 1 is , and the feature map output by the encoder 2 is . The output of the encoder 1 is right-multiplied by an identity matrix of size to obtain the first encoded feature of size . The output of the encoder 2 is left-multiplied by an identity matrix of size and right-multiplied by an identity matrix of size to obtain the second encoded feature of size . Among them, the parameter is manually adjustable.
[0100] Optionally, in the first encoded space, the steps of calculating the first feature data corresponding to the first encoded feature and the second feature data corresponding to the second encoded feature specifically include: for the first encoded feature, performing a linear mapping through the first weight matrix to obtain the first mapping matrix corresponding to the first encoded feature; for the second encoded feature, performing a linear mapping through the second weight matrix and the third weight matrix to obtain the second mapping matrix and the third mapping matrix corresponding to the second encoded feature; for the second encoded feature, performing a linear mapping through the first weight matrix to obtain the fourth mapping matrix corresponding to the second encoded feature; for the first encoded feature, performing a linear mapping through the second weight matrix and the third weight matrix to obtain the sixth mapping matrix and the third mapping matrix corresponding to the fifth encoded feature.
[0101] In the embodiment of the present application, the first feature data may include the first mapping matrix, the second mapping matrix, and the third mapping matrix. The second feature data may include the fourth mapping matrix, the fifth mapping matrix, and the sixth mapping matrix.
[0102] For the first encoded feature , through a linear mapping generate the first mapping matrix , for the second encoded feature , through a linear mapping generate the second mapping matrix and the third mapping matrix ; conversely, for the second encoded feature , through a linear mapping generate the fourth mapping matrix , for the first encoded feature , through a linear mapping generate the second mapping matrix and the sixth mapping matrix , where is the weight matrix.
[0103] Optionally, based on the second feature data, calculating the first cross-modal fusion vector corresponding to the first encoded feature, and based on the first feature data, calculating the steps of the second cross-modal fusion weight matrix corresponding to the second encoded feature specifically include: obtaining the second cross-modal fusion weight matrix corresponding to the first encoded feature based on the first mapping matrix and the second mapping matrix; obtaining the first cross-modal fusion weight matrix corresponding to the second encoded feature based on the fourth mapping matrix and the fifth mapping matrix; performing matrix multiplication on the second cross-modal fusion weight matrix and the third mapping matrix, and performing normalization processing on the matrix multiplication result to obtain the first cross-modal fusion vector; and performing matrix multiplication on the first cross-modal fusion weight matrix and the sixth mapping matrix, and performing normalization processing on the matrix multiplication result to calculate the second cross-modal fusion vector.
[0104] In the embodiments of the present application, after obtaining the first mapping matrix, the second mapping matrix, and the third mapping matrix, the fourth mapping matrix, the fifth mapping matrix, and the sixth mapping matrix, calculate the cross-modal fusion weight matrix through the following formula,
[0105]
[0106] where represents the dimension of the weight matrix. And use the calculated cross-modal fusion weight matrix to and perform weighted summation, and finally obtain the fused first cross-modal fusion vector and the second cross-modal fusion vector .
[0107] For the fused matrix and Perform standardization processing to make it have zero mean and unit standard deviation, and obtain the second standardized coding feature of the target two-dimensional delayed Doppler map. The expression is as follows:
[0108]
[0109] Among them, is the second cross-modal fusion vector obtained by fusing encoder 1 and encoder 2, is the mean value of, is the variance of.
[0110] Perform a mean adjustment operation on the second standardized coding feature of the target two-dimensional delayed Doppler map to make it have the same distribution characteristics as the first coding feature, and obtain the cross-modal fusion vector. The expression is as follows:
[0111]
[0112] Among them, is the second coding data represents the second standardized coding feature of the target two-dimensional delayed Doppler map, is the finally obtained cross-modal fusion vector, is the first coding data obtained by fusing encoder 1 and encoder 2 the mean value of, is the variance of. After standardization and mean adjustment operations, the first coding feature of the target digital elevation model obtained by encoder 1 and the second coding feature of the target two-dimensional delayed Doppler map obtained by encoder 2 are mapped to the same feature space and aligned in the feature space to obtain the cross-modal fusion vector of the first coding feature and the second standardized coding feature.
[0113] Optionally, the steps of respectively extracting features from the first decoded feature and the second decoded feature through different feature extraction modules to obtain the first feature vector of the target digital elevation model and the second feature vector of the target two-dimensional delayed Doppler map specifically include: extracting feature processing on the first decoded data through the first feature extraction module to obtain the first feature vector of the target digital elevation model, and the first feature vector corresponds to the first decoded data; extracting feature processing on the second decoded data through the second feature extraction module to obtain the second feature vector of the target two-dimensional delayed Doppler map, and the second feature vector corresponds to the second decoded data.
[0114] In the embodiments of the present application, the near-vertical observation mode of the synthetic aperture radar altimeter causes the coupling of the nadir echo and the echoes of surrounding scatterers, i.e., near-vertical coupling, resulting in a high similarity and low image quality of the two-dimensional delay-Doppler maps (DDMs) of adjacent apertures, making it difficult to effectively distinguish the two-dimensional delay-Doppler maps (DDMs) of adjacent apertures and posing challenges to correct terrain matching and high-precision aircraft positioning. Therefore, a suitable network architecture and loss function are selected to effectively extract features from the fused data.
[0115] As Figure 2 shown in the middle part, it can be seen from Figure 2 that different network architectures are respectively used to extract features from the first cross-modal fusion vector of the digital elevation model (DEM) that fuses the information of the other party and the second cross-modal fusion vector of the two-dimensional delay-Doppler map (DDM), so as to efficiently extract the global, local, and fine features of the digital elevation model (DEM) and the two-dimensional delay-Doppler map (DDM), realize the effective distinction of the digital elevation model (DEM) at different positions, realize the effective distinction of the two-dimensional delay-Doppler maps (DDMs) of different apertures, and lay a solid foundation for subsequent terrain matching.
[0116] The first feature vector can be obtained by using network architecture 1 to extract features from the first decoded feature that fuses the target two-dimensional delay-Doppler map and the target digital elevation model, and the second feature vector can be obtained by using network architecture 2 to extract features from the second decoded feature that fuses the target digital elevation model and the target two-dimensional delay-Doppler map.
[0117] Specifically, the network architecture 1 has 7 convolutional layers, followed by a max pooling layer for downsampling the feature map, and finally the feature map is converted into a one-dimensional feature vector through a flattening operation. Among them, for the 7 convolutional layers, the kernel size of the first and second layers is 3×3, the number of output channels is 32, and the stride is 1; the kernel size of the third layer is 3×3, the number of output channels is 64, and the stride is 2; the kernel size of the fourth layer is 3×3, the number of output channels is 64, and the stride is 1; the kernel size of the fifth layer is 3×3, the number of output channels is 128, and the stride is 2; the kernel size of the sixth layer is 3×3, the number of output channels is 128, and the stride is 1; the kernel size of the seventh layer is 8×8, the number of output channels is 128, and the stride is 1. The network structure 2 is ResNet18, which consists of 18 weighted layers including convolutional layers and fully connected layers, including an initial convolutional layer, followed by 16 convolutional layers, and then a fully connected layer. Among them, the kernel size of the initial convolutional layer is 7×7, the number of output channels is 64, and the stride is 2. For the 16 convolutional layers, the kernel size of the first to fourth layers is 3×3, the number of output channels is 64, and the stride is 1; the kernel size of the fifth layer is 3×3, the number of output channels is 128, and the stride is 2; the kernel size of the sixth to eighth layers is 3×3, the number of output channels is 128, and the stride is 1; the kernel size of the ninth layer is 3×3, the number of output channels is 256, and the stride is 2; the kernel size of the tenth to twelfth layers is 3×3, the number of output channels is 256, and the stride is 1; the kernel size of the thirteenth layer is 3×3, the number of output channels is 512, and the stride is 2; the kernel size of the fourteenth to sixteenth layers is 3×3, the number of output channels is 128, and the stride is 1.
[0118] Optionally, the steps of determining the positioning data of the flight platform based on the first feature vector and the second feature vector specifically include: calculating the cosine similarity between the first feature vector and the second feature vector; determining the positioning data of the flight platform based on the cosine similarity.
[0119] In the embodiments of the present application, the feature vectors of the reference digital elevation model (DEM) and the real-time two-dimensional delay Doppler map (DDM) can be obtained respectively. In order to measure the similarity between the feature vectors after feature extraction of the reference digital elevation model (DEM) and the real-time two-dimensional delay Doppler map (DDM) belonging to different modalities, for obtaining the terrain matching result, as Figure 4 shown in the similarity measurement schematic diagram, the cosine similarity is used to measure the similarity between the feature vectors of the digital elevation model (DEM) and the two-dimensional delay Doppler map (DDM). Different from the similarity measurement algorithm that directly solves the linear distance between two data in the n-dimensional space, the cosine similarity measures the difference between the two by solving the cosine value of the included angle between the two vectors in the vector space, reflecting the difference in direction in the vector space, rather than considering the magnitude of the vectors. This feature is adapted to the lower degree of attention to the magnitude of heterogeneous data in cross-modal.
[0120] Optionally, the steps of calculating the cosine similarity between the first eigenvector and the second eigenvector specifically include:
[0121] Assume is the eigenvector of the real-time two-dimensional delayed Doppler map (DDM), is the eigenvector of the reference digital elevation model (DEM), then the specific calculation process of the cosine similarity between the two is shown in the following expression (1):
[0122] (1)
[0123] where, represents the cosine similarity between the first eigenvector and the second eigenvector, represents the first eigenvector, represents the second eigenvector, represents the coordinates in the target two-dimensional delayed Doppler map, represents the coordinates of the target digital elevation model. The × operator is the inner product operator. It can be seen from expression (1) that the smaller the angle between the two vectors, the larger the cosine value and the higher the vector similarity. According to the cosine similarity, the similarity scores between each pair of digital elevation models (DEM) and two-dimensional delayed Doppler maps (DDM) are obtained for subsequent aircraft positioning.
[0124] Optionally, the steps of determining the positioning data of the flight platform based on the cosine similarity specifically include: for each position coordinate in the target two-dimensional delayed Doppler map, based on the second eigenvector, select multiple second eigenvectors with the largest similarity; based on the multiple second eigenvectors with the largest similarity, reverse search for the candidate coordinates mapped in the target digital elevation model, and each candidate coordinate corresponds to a second eigenvector; based on the multiple candidate coordinates, determine the positioning data of the flight platform.
[0125] In the embodiments of the present application, after obtaining the similarity scores between each pair of digital elevation models (DEM) and two-dimensional delayed Doppler maps (DDM), they are sorted in descending order of similarity scores to obtain a descending index. Select the top three reference digital elevation models (DEM) with the highest similarity scores to the real-time two-dimensional delayed Doppler map (DDM) according to the similarity scores, and reverse search for the coordinates mapped by these digital elevation models (DEM).
[0126] Optionally, the steps of determining the positioning data of the flight platform based on the multiple candidate coordinates specifically include: performing weighted calculation on the multiple candidate coordinates to obtain a weighted coordinate value; based on the weighted coordinate value, determine the positioning data of the flight platform.
[0127] In the embodiments of the present application, the coordinates mapped by these digital elevation models (DEMs) are obtained through reverse retrieval. To improve the credibility of the coordinate positioning results, in engineering applications, the three positioning coordinates with the highest similarity scores are weighted and averaged to determine the positioning results of the aircraft.
[0128] Optionally, the method further includes: obtaining a simulation data set and a measured data set, where the simulation data set includes a simulated two-dimensional delay Doppler map calculated based on a digital elevation model, and the measured data set includes a measured two-dimensional delay Doppler map obtained based on an airborne synthetic aperture radar altimeter during actual measurement; based on the simulation data set and the measured data set, performing a simulation test on the positioning data of the flight platform.
[0129] In the embodiments of the present application, the simulation conditions are set as follows:
[0130]
[0131] Figure 3 The result of the simulated two-dimensional delay Doppler map (DDM) is obtained using Matlab software.
[0132] Optionally, the steps of obtaining the simulation data set specifically include: obtaining simulation data, where the simulation data includes the current position of the platform of the flight platform, the platform flight speed, and the coordinates of each scatter point in the digital elevation model; based on the current position of the platform and the coordinates of the scatter points, obtaining the position vector of each scatter point; based on the platform flight speed, the position vector of each scatter point, and the radar wavelength of the airborne synthetic aperture radar altimeter, calculating the Doppler frequency of each scatter point; based on the current position of the platform and the coordinates of the scatter points, determining the relative distance between the coordinates of each scatter point and the current position of the platform; based on the Doppler frequency of each scatter point, the relative distance between the coordinates of each scatter point and the current position of the platform, and the backscatter coefficient, determining the reflection of each scatter point; based on the Doppler frequency of each scatter point and the reflection of each scatter point, obtaining the simulated two-dimensional delay Doppler map; constructing a simulation data set based on the simulated two-dimensional delay Doppler map.
[0133] In the embodiments of the present application, the simulation data set is specifically constructed as follows: The high-precision two-dimensional delay Doppler map (DDM) training data is the basis for terrain matching and aircraft positioning. According to the coordinates of each scatter point in the digital elevation model (DEM), it can be obtained as , the current position of the platform is , the flight speed is . Although the synthetic aperture radar altimeter (SARAL) observes the scatter points at the nadir point, its energy is mainly based on backscattering, and the position vector can be obtained as the following expression (2):
[0134] (2)
[0135] Among them, is the current position of the platform, are the coordinates of each scattering point, and is the set of scattering points.
[0136] The Doppler frequency of each scattering point is given by the following expression (3):
[0137] (3)
[0138] Among them, is the real-time velocity of the platform, is the radar wavelength. The relative distance of each scattering point is given by expression (3), and the backscattering coefficient is given by the following expression (4):
[0139] (4)
[0140] (5)
[0141] Among them, , and depends on the type of land cover medium. For expressions (4) and (5), the reflection of a single scattering point can be obtained as the following expression (6):
[0142] (6)
[0143] Accumulating each element of the two-dimensional delay-Doppler map (DDM) matrix according to the distance and Doppler index sets gives the following expression (7):
[0144] (7)
[0145] Among them, is the th distance gate index set, is the th Doppler channel index set. After the above processing, the two-dimensional delay-Doppler map (DDM) power matrix can be obtained as , where is the number of pulses, is the number of distance gates. Therefore, a sufficient reference two-dimensional delay-Doppler map (DDM) based on the digital elevation model (DEM) can be obtained, and the simulation two-dimensional delay-Doppler map (DDM) imaging results are as Figure 3 shown.
[0146] Optionally, the steps of the measured data set specifically include: controlling the flight platform to carry the airborne synthetic aperture radar altimeter to fly along the X-axis direction at a preset altitude and a fixed speed, and obtaining the ground scatter point coordinates; when the airborne synthetic aperture radar altimeter passes through the target scatter point, emitting a linear frequency modulation wide beam pulse, and receiving the echo signal after a certain time delay; after downsampling and range compression of the echo signal, obtaining the baseband echo signal, and the baseband echo signal includes the azimuth slow time; based on the azimuth chirp rate, giving the linear relationship between the Doppler frequency and the azimuth slow time, performing Fourier transform in the azimuth direction, and transforming the baseband echo signal to the delay-Doppler domain; performing Fourier transform on the delay-Doppler domain again to obtain the echo signal in the two-dimensional frequency domain; multiplying with the phase function for range migration correction and the echo signal in the two-dimensional frequency domain to obtain the measured two-dimensional delay-Doppler map.
[0147] In the embodiment of the present application, the construction of the measured two-dimensional delay-Doppler map (DDM) data set is specifically as follows: The airborne synthetic aperture radar altimeter (SARAL) flies along the X-axis direction at an altitude and a fixed speed . The ground scatter point coordinates are , where . The coordinates of the airborne radar position B are , which is above the L point. The synthetic aperture radar altimeter (SARAL) can obtain the high-resolution two-dimensional delay-Doppler map (DDM) at the L position. When the airborne synthetic aperture radar altimeter (SARAL) passes through the L point, it emits a linear frequency modulation wide beam pulse and receives the echo after a certain time delay. The received echo can be expressed as the following expression (8):
[0148] (8)
[0149] Where is the range envelope after range compression, usually expressed in the form of a Sinc function. is the center frequency. is the fast time in the range direction, is the slow time in the azimuth direction, is the chirp rate, is the speed of light, is a constant representing the backscattering coefficient, represents the phase change of the radar signal caused by the surface scattering process, is the instantaneous distance between the synthetic aperture radar altimeter (SARAL) and the ground scatter point. After downsampling and range compression, the baseband echo signal is as the following expression (9):
[0150] (9)
[0151] Where is a constant in the complex number field and is ignored in the subsequent derivation; is the range envelope after range compression, often expressed in the form of a Sinc function. The instantaneous range of the scatterer is given by the following expression (10):
[0152] (10)
[0153] where is the reference range. Substituting expression (9) into expression (8), the baseband signal of the target can be obtained as the following expression (11): (11)
[0155] where is the imaginary unit, satisfying . The azimuth phase modulation can be seen from the second exponential term in expression (10). Since the phase is a function of
[0156] (12)
[0157] where the azimuth modulation frequency gives the Doppler frequency and the slow time The linear relationship between them is given by the following expression (13):
[0158] (13)
[0159] Next, performing a Fourier transform in the azimuth direction, expression (10) is transformed into the delay-Doppler domain as the following expression (14):
[0160] (14)
[0161] where is the range migration, is the frequency-domain form of the azimuth antenna pattern .
[0162] Performing a Fourier transform on expression (14), the echo in the two-dimensional frequency domain can be obtained as the following expression (15):
[0163] (15)
[0164] where , is the range frequency domain, is 's frequency-domain form.
[0165] To achieve range migration correction, a phase function can be obtained as the following expression (16):
[0166] (16)
[0167] Multiplying expression (16) by expression (14), the range-compressed signal can be obtained as the following expression (17):
[0168] (17)
[0169] According to the processing of expressions (8)-(17), a two-dimensional delay Doppler map (DDM) can be obtained from the echo of the Synthetic Aperture Radar Altimeter (SARAL) to construct a measured data set to meet the training, verification, and testing requirements. The imaging result of the measured two-dimensional delay Doppler map (DDM) is as Figure 4 shown.
[0170] Using the method of this embodiment can perform cross-modal terrain matching and real-time aircraft positioning. To verify the feasibility of the proposed cross-modal network for terrain matching and the real-time aircraft positioning effect, simulation and measured data sets are constructed; then, through feature fusion, feature extraction, and similarity measurement using the method of this embodiment, the aircraft positioning result is finally obtained. As Figures 5 - 6 shown. Figure 5 FIG. is a regional digital elevation model (DEM) and a schematic diagram of a certain flight track, where the red line represents the flight direction of the aircraft, and the letter v represents the velocity vector. Figure 6 FIG. is a result diagram of terrain matching and real-time aircraft positioning using this application, where the black dashed line represents the preset flight track, the blue circle represents the point to be matched, the red triangle represents the true coordinates of the aircraft, and the red cross represents the positioning result obtained by this application. It can be seen that the method of this application can achieve cross-modal digital elevation model (DEM) and two-dimensional delay Doppler map (DDM) terrain matching, and thus can achieve real-time aircraft positioning, and the positioning result is reliable.
[0171] Finally, it should be noted that the above specific embodiments are only used to illustrate the technical solutions of the present application rather than to limit them. Although the present application has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present application can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present application, and they should all be covered within the scope of the claims of the present application.
Claims
1. A cross-modal terrain matching positioning method, characterized in that For a flight platform on which an airborne synthetic aperture radar altimeter is carried, the method includes: Obtaining a target digital elevation model and a target two-dimensional delay Doppler map of the current terrain of the flight platform, where the two-dimensional delay Doppler map is obtained by the airborne synthetic aperture radar altimeter scanning the current terrain by radar; Encoding the target digital elevation model and the target two-dimensional delay Doppler map into a first encoding space through a first encoder and a second encoder respectively, to obtain a first encoded feature corresponding to the target digital elevation model and a second encoded feature corresponding to the target two-dimensional delay Doppler map; In the first encoding space, calculating first feature data corresponding to the first encoded feature and calculating second feature data corresponding to the second encoded feature; specifically, for the first encoded feature, performing a linear mapping through a first weight matrix to obtain a first mapping matrix corresponding to the first encoded feature, for the second encoded feature, performing a linear mapping through a second weight matrix and a third weight matrix to obtain a second mapping matrix and a third mapping matrix corresponding to the second encoded feature; for the second encoded feature, performing a linear mapping through the first weight matrix to obtain a fourth mapping matrix corresponding to the second encoded feature, for the first encoded feature, performing a linear mapping through the second weight matrix and the third weight matrix to obtain a fifth mapping matrix and a sixth mapping matrix corresponding to the first encoded feature; Based on the second feature data, calculating a first cross-modal fusion vector corresponding to the first encoded feature, and based on the first feature data, calculating a second cross-modal fusion vector corresponding to the second encoded feature; specifically, based on the first mapping matrix and the second mapping matrix, obtaining a second cross-modal fusion weight matrix corresponding to the second encoded feature; based on the fourth mapping matrix and the fifth mapping matrix, obtaining a first cross-modal fusion weight matrix corresponding to the first encoded feature; multiplying the second cross-modal fusion weight matrix by the third mapping matrix and performing a normalization process on the matrix multiplication result to obtain a second cross-modal fusion vector; and multiplying the first cross-modal fusion weight matrix by the sixth mapping matrix and performing a normalization process on the matrix multiplication result to calculate a first cross-modal fusion vector; Performing decoding processing on the first cross-modal fusion vector and the second cross-modal fusion vector through a first decoder and a second decoder respectively, to obtain a first decoded feature corresponding to the first cross-modal fusion vector and a second decoded feature corresponding to the second cross-modal fusion vector; Performing feature extraction on the first decoded feature and the second decoded feature respectively through different feature extraction modules, to obtain a first feature vector of the target digital elevation model and a second feature vector of the target two-dimensional delay Doppler map; Calculating the cosine similarity between the first feature vector and the second feature vector; Based on the cosine similarity, determine the positioning data of the flying platform; specifically, for each position coordinate in the target two-dimensional delayed Doppler map, select multiple first eigenvectors with the largest similarity based on the second eigenvector; Based on multiple first eigenvectors with the largest similarity, inversely retrieve the candidate coordinates mapped in the target digital elevation model, and each candidate coordinate corresponds to one first eigenvector; Based on multiple candidate coordinates, determine the positioning data of the flying platform.
2. The cross-modal terrain matching and positioning method according to claim 1, wherein The step of encoding the target digital elevation model and the target two-dimensional delayed Doppler map into the first encoding space through the first encoder and the second encoder respectively to obtain the first encoding feature corresponding to the target digital elevation model and the second encoding feature corresponding to the target two-dimensional delayed Doppler map specifically includes: Perform encoding processing on the target digital elevation model through the first encoder to obtain a first encoded feature map; And perform encoding processing on the target two-dimensional delayed Doppler map through the second encoder to obtain a second encoded feature map; Perform matrix multiplication calculations on the first encoded feature map and the second encoded feature map respectively through unit matrices of different dimensions to obtain the first encoding feature corresponding to the target digital elevation model and the second encoding feature corresponding to the target two-dimensional delayed Doppler map, and the first encoding feature and the second encoding feature have the same dimension.
3. The cross-modal terrain matching positioning method according to claim 1, wherein The step of respectively performing feature extraction on the first decoded feature and the second decoded feature through different feature extraction modules to obtain the first eigenvector of the target digital elevation model and the second eigenvector of the target two-dimensional delayed Doppler map specifically includes: Perform feature extraction processing on the first decoded feature through the first feature extraction module to obtain the first eigenvector of the target digital elevation model, and the first eigenvector corresponds to the first decoded feature; Perform feature extraction processing on the second decoded feature through the second feature extraction module to obtain the second eigenvector of the target two-dimensional delayed Doppler map, and the second eigenvector corresponds to the second decoded feature.
4. The cross-modal terrain matching and positioning method according to claim 1, wherein The step of calculating the cosine similarity between the first eigenvector and the second eigenvector specifically includes: The cosine similarity satisfies the following formula: Among them, represents the cosine similarity between the first eigenvector and the second eigenvector, represents the first eigenvector, represents the second eigenvector, represents the coordinates in the target two-dimensional delayed Doppler map, represents the coordinates of the target digital elevation model.
5. The cross-modal terrain matching positioning method according to claim 4, wherein The step of determining the positioning data of the flying platform based on multiple candidate coordinates specifically includes: Perform weighted calculation on multiple candidate coordinates to obtain a weighted coordinate value; Based on the weighted coordinate value, determine the positioning data of the flying platform.
Citation Information
Patent Citations
Satellite-borne GNSS-R important feature selection and forest above-ground biomass and canopy height estimation method combining tree model ensemble learning algorithm and SHAP interpreter
CN119312680A
Sensing device for providing three dimensional information
US20230377181A1