A Remote Sensing Image Elevation Prediction Method Combining Semantic Information
The integration of semantic information in single-view remote sensing imagery using a shared weight encoder-decoder enhances elevation prediction accuracy by leveraging prior relationships between semantic and geometric features, addressing the limitations of existing methods.
Patent Information
- Application Number
- CN202210557539.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-19
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-05-19
AI Technical Summary
In the existing remote sensing image elevation prediction methods, there is a problem of low accuracy when using single-view remote sensing images, especially due to occlusion and complex calculation processes, which lead to slow speed and low accuracy.
Using an elevation prediction method combining semantic information, we learn the prior relationship between the geographic classification task and the elevation prediction task through a shared weight encoding-decoder, use the elevation extraction network model to classify elevation and geographic objects, use a rational function imaging model to obtain the training set, and improve the prediction accuracy through multi-task strategies and improved decoder structure.
The accuracy and speed of elevation prediction of single-view remote sensing images is improved, and elevation prediction is assisted by geographic classification tasks to achieve higher prediction accuracy and faster computing speed.
Smart Images

Figure CN114821192B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of geographic information technology, and particularly relates to a method for predicting elevation of remote sensing images combined with semantic information. Background Art
[0002] In existing methods for predicting elevation of remote sensing images, most of them simultaneously mount multiple camera lenses on a flying platform to obtain richer image information from vertical and other inclined directions at the same time. Then, using the obtained multi-view remote sensing images, through a series of calculations such as preprocessing the remote sensing images, extracting feature points from multiple remote sensing images, and performing adjustment calculations, the elevation geometric information is extracted using this calculation process. Since there are many remote sensing images to be recognized, the calculation process is complex and slow, resulting in a slow speed of obtaining the elevation prediction result. When using a single-view remote sensing image to identify elevation geometric information, there is only an image from one angle. Due to occlusion in the real world, the remote sensing image is occluded, which in turn leads to the problem of low accuracy in subsequent extraction of elevation information. Summary of the Invention
[0003] The purpose of the present invention is to provide a method for predicting elevation of remote sensing images combined with semantic information to solve the problem of low accuracy when using a single-view remote sensing image for elevation prediction.
[0004] To solve the above technical problems, the technical solutions provided by the present invention and the corresponding beneficial effects of the technical solutions are as follows:
[0005] A method for predicting elevation of remote sensing images combined with semantic information according to the present invention includes the following steps:
[0006] Obtain a single-view remote sensing image and input it into an elevation extraction network model to obtain an elevation prediction result and a ground object classification result of the single-view remote sensing image; the elevation extraction network model is trained using the single-view remote sensing image, the corresponding digital elevation model, and the ground object classification result as a training set, and the elevation extraction network model includes a shared-weight encoder-decoder, a semantic prediction branch, and an elevation prediction branch; the shared-weight encoder-decoder is used to perform a ground object classification task and an elevation prediction task on the single-view remote sensing image, learn the prior relationship between the semantic information extracted from the ground object classification task and the geometric information extracted from the elevation prediction task, and obtain a feature vector using the prior relationship; input the feature data of the feature vector except the last channel into the semantic prediction branch to extract semantic features and obtain a ground object classification result; input the feature data of the last channel of the feature vector into the elevation prediction branch to extract elevation features and obtain an elevation prediction result.
[0007] The beneficial effects of the above technical solution are as follows: In the present invention, a shared-weight encoder-decoder is used for ground object classification tasks and elevation prediction tasks. The ground object classification tasks and elevation prediction tasks share weights and supervise each other, and learn the prior relationship existing between the semantic information extracted from the ground object classification tasks and the geometric information extracted from the elevation prediction tasks. For example, the average elevation of "water area" should be lower than that of other ground objects, and the elevation of "houses" (in flat areas) should be higher than that of the surrounding flat terrain. Similarly, the places where the elevation changes are often the boundaries of ground object classification. The purpose of the present invention is to enable the shared-weight encoder-decoder to learn this prior relationship during the process of identifying the ground object distribution and predicting the elevation, and make the above two tasks supervise each other. The ground object classification task assists the elevation prediction task to more accurately extract geometric information. Then, on the one hand, elevation features are extracted through the elevation prediction branch to obtain a more accurate elevation prediction result. On the other hand, semantic features are extracted through the semantic prediction branch to obtain a more accurate ground object classification result, thereby ultimately enhancing the cognitive ability of the elevation extraction network model for real ground objects themselves. Thereby, a remote sensing image elevation prediction method combining semantic information with high prediction accuracy when using single-view remote sensing images for elevation prediction and ground object classification is provided.
[0008] Further, the decoder in the shared-weight encoder-decoder includes a four-layer structure connected in sequence, and each layer structure includes in sequence: a transposed convolution module and two first convolution modules;
[0009] The transposed convolution module includes a transposed convolution calculation, an activation function, and a padding operation after the transposed convolution arranged in sequence.
[0010] The beneficial effects of the above technical solution are as follows: The decoder with shared weights in the present invention further improves on the U-net decoder structure. Different from the U-Net decoder directly using the upsample layer with expanded scale, the present invention uses step-by-step convolution, and adds a transposed convolution layer at the front of each layer structure of the decoder, so as to better depict details. Specifically, the elevation decoder in the present invention performs convolutions with corresponding convolution kernel sizes in the x direction and the y direction respectively, so as to help the network perceive the changes in terrain gradients and thus enrich the texture details of terrain prediction, and ordinary convolution is added after the transposed convolution to eliminate checkerboard shadows. Experimental results prove that the improved decoder with shared weights has improved the elevation prediction accuracy by 0.3m. In addition, the present invention adopts the U-net decoder structure and realizes feature fusion through splicing, and the structure is concise and stable.
[0011] Further, the method for constructing the training set includes the following steps:
[0012] The method for constructing the training set includes the following steps:
[0013] Obtain a single-view remote sensing image I and a digital elevation model D of the corresponding geographic range; divide the single-view remote sensing image I into multiple small images, traverse each pixel point of each small image pixel by pixel I(i,j), and use the least squares method to obtain the longitude X, latitude Y, and elevation H corresponding to the digital elevation model D for each current pixel point according to the following method to obtain a training set; the method for obtaining the longitude X, latitude Y, and elevation H using the least squares method is as follows:
[0014] a. Select the longitude X0, latitude Y0, and elevation offset H0 in the rational function imaging model RFM corresponding to the single-view remote sensing image I, and record them as the iteration initial values;
[0015] b. Calculate the pixel coordinates (r p ,c p ) projected by the iteration initial values and the rational function imaging model RFM, and calculate the difference between the pixel coordinates and the current pixel point, which is recorded as the projection error;
[0016] c. Obtain the partial derivatives of the projection error with respect to the longitude and latitude, and construct a partial derivative arrangement matrix;
[0017] d. Solve the corrections in the longitude and latitude directions based on the partial derivative arrangement matrix and the projection error;
[0018] f. Update the iteration initial values according to the corrections, record them as the current iteration initial values, and obtain the elevation corresponding to the longitude and latitude in the current iteration initial values from the digital elevation model by interpolation, and update the elevation value in the current iteration initial values;
[0019] g. Repeat steps b-f until convergence, so as to obtain the elevation value corresponding to the current pixel point.
[0020] The beneficial effects of the above technical solution are as follows: In the prior art, digital elevation model (DEM) is used to obtain training data. When using the DEM model to obtain training data, on the one hand, remote sensing images are not corrected for geographical information; on the other hand, the triple data coordinates composed of longitude, latitude, and elevation obtained are interpolated, rather than the coordinate values on the actual DEM model. Therefore, the regular triple data coordinates obtained after grid sampling do not correspond one-to-one with the pixels of the remote sensing image, but there are certain deviations, resulting in inaccurate training data and elevation prediction results. In the present invention, a rational function imaging model is used to obtain the training set. By traversing each pixel point by pixel, and then using the least squares method to obtain the training set by setting the initial iteration value, constructing the error, linearly approximating, correcting the error of the triple data coordinates, and updating the initial iteration value, etc., the coordinate data composed of longitude, latitude, and elevation on the rational function imaging model strictly corresponds to each pixel on the remote sensing image, and the obtained triple data coordinates are more accurate, so as to facilitate obtaining accurate elevation prediction results later.
[0021] Further, the encoder in the shared weight encoder-decoder is ResNet.
[0022] Further, the structure of the semantic prediction branch includes: a second convolution module and a softmax layer; the second convolution module includes a convolutional layer, a batch normalization layer, and an activation function.
[0023] Further, when training the elevation extraction network model, the loss function used by the semantic prediction branch is the cross-entropy loss function, and the cross-entropy loss function is:
[0024]
[0025] where, for each pixel in the sample to be predicted, y i represents the classification label in one-hot form. If the current pixel belongs to the i-th class, then y i =1, otherwise y i =0, p is the output after the softmax layer of the semantic prediction branch, which is an n×1 vector, and p i indicates the probability that the current pixel belongs to the i-th class, i∈1,2,3...n.
[0026] Further, the structure of the elevation prediction branch includes a third convolution module and a tanh() activation function; the third convolution module includes: a convolutional layer and a batch normalization layer.
[0027] Further, when training the elevation extraction network model, the loss function used by the elevation prediction branch includes the elevation error loss function L g with scale invariance, and its formula is as follows:
[0028]
[0029] d = log(h p ) - log(h t )
[0030] where M is the number of sample pixels, (r, c) is the pixel coordinate of the remote sensing image, hp represents the predicted elevation, and ht represents the true elevation value.
[0031] Furthermore, when training the elevation extraction network model, the loss function used by the elevation prediction branch includes the reprojection loss function L r , and the formula is as follows:
[0032]
[0033] grid = RFM(X′, Y′, h p )
[0034]
[0035] where M is the number of sample pixels, SSIM is the image pattern consistency descriptor, μ is the mean operator, σ is the variance operator, C1 and C2 are constants to prevent the denominator from being zero, α is the weight control parameter, X' is the matrix composed of longitudes in the sample, Y' is the matrix composed of latitudes in the sample, h p is the matrix composed of the predicted elevations, I is the original remote sensing image, RFM represents the rational function imaging model, grid is the pixel coordinate grid obtained by reprojection of X', Y', h p , and is the image generated by sampling on the original image using grid. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is the data flow schematic diagram of a method for predicting the elevation of remote sensing images by combining semantic information according to the present invention;
[0037] Figure 2-1 is the first original remote sensing image in the embodiment of the method of the present invention;
[0038] Figure 2-2 is in the embodiment of the method of the present invention Figure 2-1 corresponding true elevation distribution map;
[0039] Figure 2-3 is in the embodiment of the method of the present invention Figure 2-1 corresponding model predicted elevation;
[0040] Figure 2-4 is in the embodiment of the method of the present inventionFigure 2-1 The ground object classification result map output by the corresponding model;
[0041] Figure 3-1 is the second original remote sensing image in the embodiment of the method of the present invention;
[0042] Figure 3-2 is in the embodiment of the method of the present invention Figure 3-1 The true elevation distribution map corresponding thereto;
[0043] Figure 3-3 is in the embodiment of the method of the present invention Figure 3-1 The predicted elevation of the corresponding model;
[0044] Figure 3-4 is in the embodiment of the method of the present invention Figure 3-1 The ground object classification result map output by the corresponding model;
[0045] Figure 4-1 is the third original remote sensing image in the embodiment of the method of the present invention;
[0046] Figure 4-2 is in the embodiment of the method of the present invention Figure 4-1 The true elevation distribution map corresponding thereto;
[0047] Figure 4-3 is in the embodiment of the method of the present invention Figure 4-1 The predicted elevation of the corresponding model;
[0048] Figure 4-4 is in the embodiment of the method of the present invention Figure 4-1 The ground object classification result map output by the corresponding model. Detailed implementation manners
[0049] The present invention designs a multi-task architecture for ground object segmentation and elevation prediction, effectively utilizing the semantic information of ground objects. As Figure 1 shown, adopting a multi-task strategy, first, the input original remote sensing image first passes through an encoder-decoder with shared weights, so as to learn the texture features of the remote sensing image and the relative relationship with the potential elevation. Then, the collected feature maps are respectively input into the semantic prediction module and the elevation prediction module, so as to obtain the ground object classification result and the elevation prediction result. By adding the ground object semantic information in the present invention, it can provide strong prior guidance for elevation inversion. For example, the average elevation corresponding to the water area is generally lower than that of other ground objects.
[0050] In the above, the core idea of the multi-task strategy is that by enabling the network to simultaneously complete multiple related prediction tasks T1, T2,....T n , so as to be constrained in different aspects L1, L2,....L nUnder learning, the most essential features of the current task are obtained, enhancing the generalization ability of the model, which is widely applied to various tasks in computer vision processing. For example, if the network needs to simultaneously complete related classification task T1 and regression task T2, the corresponding L1 can be the cross-entropy loss function, and L2 can be the mean square error loss function or a combination of numerical loss functions. The model under the multi-task strategy usually has a basic module B with shared weights and prediction branches C1, C2,... C n In the optimization process, L1, L2,....L n calculate the losses for the results predicted by C1, C2,... C n The generated gradients will all update the parameters of B during backpropagation optimization, so that B encodes the essential features serving T1, T2,....T n simultaneously, improving the situation where the model may fall into local optimality in a single task.
[0051] In addition, since most real objects in the present invention are continuous, the boundaries of semantic judgments are usually also the boundaries of geometric shapes (and vice versa), so the restoration and estimation of real-world objects can also use the multi-task strategy of mutual constraint between semantics and geometry. In the present invention, the distribution of ground objects is semantic information, and the elevation of ground objects is geometric information, and there is a prior relationship between the two. For example, the average elevation of "water area" should be lower than the average elevation of other ground objects, and the elevation of "houses" (in flat areas) should be higher than the elevation of the surrounding flat terrain. Similarly, the places where the elevation changes are often also the boundaries of ground object classification. The present invention aims to enable the model to learn this prior relationship during the process of identifying ground objects and predicting elevation, and make these two tasks supervise each other, ultimately enhancing the model's cognitive ability for real ground objects themselves.
[0052] Next, in combination with the accompanying drawings and embodiments, a remote sensing image elevation prediction method combining semantic information according to the present invention will be described in detail.
[0053] Method embodiment:
[0054] A method embodiment of a remote sensing image elevation prediction method combining semantic information according to the present invention will be described in detail below in combination with the accompanying drawings for the prediction method of the present invention.
[0055] Step 1, dataset construction.
[0056] Limited by the size of remote sensing images, device computing power, and geographical attributes, the present invention uses the Rational Function Imaging Model (RFM) to re-extract training samples. Given the original large-format remote sensing image I (single-view remote sensing image) and the Digital Elevation Model D of the corresponding geographical range. First, the original remote sensing image I is sliced into m×n small images I(i, j), where 0 ≤ i ≤ m and 0 ≤ j ≤ n. Then, start traversing I(i, j) pixel by pixel, and obtain the longitude, latitude, and elevation coordinate tuple (X, Y, H) corresponding to the current pixel (r, c) from the Digital Elevation Model D according to the following method:
[0057] Step 01: Select the longitude, latitude, and elevation offsets X0, Y0, H0 in the RFM corresponding to I as the iterative initial values;
[0058] Step 02: Calculate the pixel coordinates (r p , c p ) projected using X0, Y0, H0 and the RFM, and construct the projection error L = (r p , c p ) - (r, c);
[0059] Step 03: Obtain the partial derivatives of the projection error L with respect to X and Y and construct the Jacobian matrix A (partial derivative permutation matrix), where
[0060] Step 04: Solve for the corrections △X and △Y in the X and Y directions, △X, △Y = (A T A) -1 L;
[0061] Step 05: Update (X0, Y0) ← (X0, Y0) + (△X, △Y), and obtain the elevation H corresponding to the current (X0, Y0) from the Digital Elevation Model D by bilinear interpolation, and update H0 ← H.
[0062] Step 06: Repeat Steps 02 - 05 until convergence. Dense elevation samples that satisfy the RFM projection conditions for each pixel of the sample I(i, j) can be obtained. Produce all small images within the range of 0 ≤ i ≤ m and 0 ≤ j ≤ n to construct a training set for DEM elevation prediction.
[0063] Step Two: Construct an elevation extraction network model, and use the dataset constructed in Step One to train the elevation extraction network model to obtain a trained elevation extraction network model.
[0064] As Figure 1As shown in the figure, the elevation extraction network model includes a shared-weight encoder-decoder, a semantic prediction branch, and an elevation prediction branch. The shared-weight encoder-decoder is used to perform a ground object classification task and an elevation prediction task on a single-view remote sensing image, learn the prior relationship between the semantic information extracted from the ground object classification task and the geometric information extracted from the elevation prediction task, and obtain a feature vector using the prior relationship. The feature data of all channels except the last one in the feature vector is input into the semantic prediction branch to extract semantic features and obtain a ground object classification result. The feature data of the last channel in the feature vector is input into the elevation prediction branch to extract elevation features and obtain an elevation prediction result.
[0065] The following will introduce these parts of the structure and the loss function used during training in detail.
[0066] 1) Feature encoder-decoder with shared weights.
[0067] The decoder in the shared-weight encoder-decoder includes four layers connected in sequence. Each layer structure includes, in sequence: a transposed convolution module and two first convolution modules. The transposed convolution module includes a transposed convolution layer, an activation function, and a padding operation layer after transposed convolution arranged in sequence. The encoder in the shared-weight encoder-decoder is ResNet. The structure of the feature encoder-decoder with shared weights obtained by combining and improving based on the ResNet and U-Net structures is shown in Table 1:
[0068] Table 1 Structure of the feature encoder-decoder with shared weights
[0069]
[0070]
[0071]
[0072] Among them, Conv represents a convolution layer, and the parameter list is (number of input channels, number of output channels, convolution kernel size, edge padding size). BatchNorm represents a batch normalization layer, and the parameter list is (number of feature channels). ReLU is an activation function.
[0073] Pad represents an operation of edge padding according to the step combination, and the purpose is to smoothly complete the subsequent convolution. ConvTr represents a transposed convolution, and the parameter list is (number of input channels, number of output channels, convolution kernel size, step). Pad_Tr represents the padding operation after transposed convolution, and the parameter is the step. conv is as shown before, which is a normal convolution operation.
[0074] During specific application, assume the dimension of the input image is [B, C, H, W], where B is the number of input images per time, C is the number of channels, H is the image height (in pixels), and W is the image width (in pixels). The number of classes during the learning process is n_class, then the dimension of the final output of the feature encoder-decoder is [B, n_class + 1, H, W]. The first n_class channels will enter the semantic prediction branch as the class vectors to be processed, and the last channel will enter the elevation prediction branch.
[0075] The feature encoder-decoder with shared weights is obtained by combining and improving based on the ResNet and U-Net structures, denoted as the feature encoder-decoder with shared weights. The encoding part (layer identifier starting with Enc) follows the pattern of the basic layers of ResNet. The decoding part (the first convolutional module, layer identifier starting with Dec) receives the feature maps of different scales output by the encoder in the U-Net format. Different from the U-Net which directly uses the upsample layer with an extended scale, the present invention uses stepwise convolution and transposed convolution to better depict details. Specifically, (1) the decoder performs convolutions with convolutional kernels of (3, 1) and (1, 3) in the x and y directions respectively, so as to help the network perceive the changes in terrain gradients; (2) a transposed convolutional layer with learnable parameters is used, and a normal convolution is added after the transposed convolution to eliminate checkerboard shadows, thereby enriching the texture details of terrain prediction. Experimental results show that the improvement of the model contributes a 0.3m accuracy improvement in the elevation direction.
[0076] 2) Semantic prediction branch and training.
[0077] The semantic prediction branch is used to extract semantic features to output the final land cover classification map, where n is the number of classes. The structure of the semantic prediction branch is shown in Table 2, including: a second convolutional module and a softmax layer; the second convolutional module includes a convolutional layer, a batch normalization layer, and an activation function.
[0078] Table 2 Structure of the semantic prediction branch
[0079]
[0080] Land cover classification is trained according to the cross-entropy loss function L s :
[0081]
[0082] where, for each pixel in the sample to be predicted, y i represents the classification label in one-hot form. If the current pixel belongs to the i-th class, then y i = 1, otherwise y i = 0. p is the output of the semantic prediction branch after passing through the softmax layer, which is an n×1 vector, pi The probability that the current pixel belongs to the $i$-th category, where $i \in \{1, 2, 3, \ldots, n\}$.
[0083] 3) Elevation prediction branch and training.
[0084] The elevation prediction branch is mainly used to calculate the elevation by combining the rational function model of the remote sensing image. The structure of the elevation prediction branch is shown in Table 3, including a third convolutional module and a $\tanh()$ activation function; the third convolutional module includes a convolutional layer and a batch normalization layer.
[0085] Table 3 Structure of the elevation prediction branch
[0086]
[0087] The elevation prediction branch adopts the following two loss functions.
[0088] First, the elevation error loss $L$ with scale invariance g :
[0089]
[0090] $d = \log(h^{(\text{pred})}) - \log(h^{(\text{gt})})$ p ) - \log(h^{(\text{gt})})$ t )
[0091] where $M$ is the number of sample pixels, $(r, c)$ is the pixel coordinate of the remote sensing image, $h^{(\text{pred})}$ represents the predicted elevation, and $h^{(\text{gt})}$ represents the ground truth elevation. p represents the predicted elevation, $h^{(\text{gt})}$ t represents the ground truth elevation.
[0092] Second, the reprojection loss function $L$ r :
[0093]
[0094] $\text{grid} = \text{RFM}(X', Y', h^{(\text{pred})})$ p )
[0095]
[0096] where $M$ is the number of sample pixels, $\text{SSIM}$ is the image pattern consistency descriptor, $\mu$ is the mean operator, $\sigma$ is the variance operator, $C_1$ and $C_2$ are constants to prevent the denominator from being zero, usually set to $0.0001$ and $0.0003$ respectively. $\alpha$ is the weight control parameter, usually $0.85$, $X'$ is the matrix composed of longitudes in the sample, $Y'$ is the matrix composed of latitudes in the sample, $h^{(\text{pred})}$ p is the matrix composed of the predicted elevation, $I$ is the original remote sensing image, $\text{RFM}$ represents the rational function imaging model, and $\text{grid}$ is the result of applying $\text{RFM}$ to $X'$, $Y'$, $h^{(\text{pred})}$ pThe pixel coordinate grid obtained after reprojection is an image generated by sampling on the original image using grid. Theoretically, when the predicted elevation is consistent with the ground truth, it will be exactly the same as I. Comparing it with the original remote sensing image I, the loss can be used as the error to drive the training.
[0097] Step 3: Obtain the single-view remote sensing image to be predicted and input it into the elevation extraction network model to obtain the elevation prediction result and the ground object classification result.
[0098] Next, combined with specific examples, the effectiveness of the method of the present invention is verified. The test results on the plain and hilly test sets are shown in Table 4:
[0099] Table 4 Elevation prediction results
[0100]
[0101] Partial visualization results are compared and shown through the original remote sensing image, the ground truth map of elevation distribution, the elevation predicted by the model, and the ground object classification result output by the model. For example Figure 2-1 is the first original remote sensing image, Figure 2-2 is the corresponding ground truth map of elevation distribution, Figure 2-3 is the corresponding elevation predicted by the model, Figure 2-4 is the corresponding ground object classification result output by the model; Figure 3-1 is the second original remote sensing image, Figure 3-2 is the corresponding ground truth map of elevation distribution, Figure 3-3 is the corresponding elevation predicted by the model, Figure 3-4 is the corresponding ground object classification result output by the model; For example Figure 4-1 is the third original remote sensing image, Figure 4-2 is the corresponding ground truth map of elevation distribution, Figure 4-3 is the corresponding elevation predicted by the model, Figure 4-4 is the corresponding ground object classification result output by the model.
[0102] The present invention aims to enable the shared-weight encoder-decoder to learn the prior relationship in the process of identifying the ground object distribution and predicting the elevation, and to make the elevation prediction task and the ground object classification task supervise each other. The ground object classification task assists the elevation prediction task to more accurately extract geometric information. Then, on the one hand, the elevation features are extracted through the elevation prediction branch to obtain a more accurate elevation prediction result. On the other hand, the semantic features are extracted through the semantic prediction branch to obtain a more accurate ground object classification result, so as to ultimately improve the cognitive ability of the elevation extraction network model for the real ground objects themselves.
Claims
1. A method for predicting the elevation of remote sensing images combined with semantic information, characterized in that: The method includes the following steps: Obtain a single-view remote sensing image and input it into an elevation extraction network model to obtain an elevation prediction result and a ground object classification result of the single-view remote sensing image; The elevation extraction network model is trained using the single-view remote sensing image, the corresponding digital elevation model, and the ground object classification result as a training set, and the elevation extraction network model includes a shared-weight encoder-decoder, a semantic prediction branch, and an elevation prediction branch; The shared-weight encoder-decoder is used to perform a ground object classification task and an elevation prediction task on the single-view remote sensing image, learn the prior relationship between the semantic information extracted by the ground object classification task and the geometric information extracted by the elevation prediction task, and obtain a feature vector using the prior relationship; the decoder of the shared-weight encoder-decoder is used to receive the different-scale features output by the encoder in imitation of U-Net and includes four layers of structures. Each layer of structure includes a transposed convolution, an activation function, a padding operation after the transposed convolution, an edge padding operation according to the step combination, a first step convolution, an activation function, an edge padding operation according to the step combination, a second step convolution, and an activation function in sequence. The first step convolution is a convolution with a convolution kernel of (3,1) in the x direction, and the second step convolution is a convolution with a convolution kernel of (1,3) in the y direction; Input the feature data of the channels except the last one in the feature vector into the semantic prediction branch to extract semantic features and obtain a ground object classification result; input the feature data of the last channel in the feature vector into the elevation prediction branch to extract elevation features and obtain an elevation prediction result.
2. The remote sensing image elevation prediction method combining semantic information according to claim 1, wherein: The encoder of the shared-weight encoder-decoder includes four layers of structures. The first layer of structure includes a normal convolution, a batch normalization, an activation function, a normal convolution, a batch normalization, and an activation function in sequence. The last three layers of structures all include a max pooling, a normal convolution, a batch normalization, an activation function, a normal convolution, a batch normalization, and an activation function in sequence.
3. The method for predicting the elevation of a remote sensing image combining semantic information according to claim 1, wherein: The method for constructing the training set includes the following steps: Obtain a single-view remote sensing image I and a digital elevation model D of the corresponding geographical range; divide the single-view remote sensing image I into multiple small images, traverse each pixel point of each small image pixel by pixel I(i,j), and use the least squares method to obtain the longitude X, latitude Y, and elevation H corresponding to the digital elevation model D for each current pixel point (r,c) in the following way to obtain a training set; the method for obtaining the longitude X, latitude Y, and elevation H using the least squares method is: a. Select the longitude X0, latitude Y0, and elevation offset H0 in the rational function imaging model RFM corresponding to the single-view remote sensing image I as the iterative initial values; b. Calculate the pixel coordinates (r p , c p ) obtained by projecting the initial iteration value and the rational function imaging model RFM, and calculate the difference between the pixel coordinates and the current pixel point, which is denoted as the projection error; c. Obtain the partial derivatives of the projection error with respect to the longitude and latitude, and construct a partial derivative permutation matrix; d. Solve the corrections in the longitude and latitude directions based on the partial derivative permutation matrix and the projection error; f. Update the iteration initial value according to the correction, denoted as the current iteration initial value, and obtain the elevation corresponding to the longitude and latitude in the current iteration initial value from the digital elevation model by interpolation, and update the elevation value in the current iteration initial value; g. Repeat steps b - f until convergence, so as to obtain the elevation value corresponding to the current pixel point.
4. The remote sensing image elevation prediction method combining semantic information according to claim 1, characterized in that: The encoder in the shared weight encoder - decoder is ResNet.
5. The remote sensing image elevation prediction method combining semantic information according to claim 1, characterized in that: The structure of the semantic prediction branch includes: a second convolution module and a softmax layer; the second convolution module includes a convolutional layer, a batch normalization layer and an activation function.
6. The remote sensing image elevation prediction method combining semantic information according to claim 1 or 5, characterized in that: When training the elevation extraction network model, the loss function used by the semantic prediction branch is the cross-entropy loss function, and the cross-entropy loss function L s is as follows: Among them, for each pixel in the sample to be predicted, y i represents the classification label in one-hot form. If the current pixel belongs to the i-th class, then y i = 1; otherwise, y i = 0. p is the output after the semantic prediction branch passes through the softmax layer, which is an n×1 vector. p i indicates the probability that the current pixel belongs to the i-th class, where i ∈ 1, 2, 3... n.
7. The remote sensing image elevation prediction method combining semantic information according to claim 1, characterized in that: The structure of the elevation prediction branch includes a third convolution module and an activation function; the third convolution module includes: a convolutional layer and a batch normalization layer.
8. The remote sensing image elevation prediction method combining semantic information according to claim 1 or 7, characterized in that: When training the elevation extraction network model, the loss function used by the elevation prediction branch includes the elevation error loss function L with scale invariance g , and its formula is as follows: d = log(h p ) - log(h t ) where M is the number of sample pixels, (r, c) is the pixel coordinate of the remote sensing image, and h p represents the predicted elevation, and h t represents the true elevation value.
9. The remote sensing image elevation prediction method combining semantic information according to claim 1 or 7, characterized in that: When training the elevation extraction network model, the loss function used by the elevation prediction branch includes a reprojection loss function \(L\) r , and the formula is as follows: grid = RFM(X′, Y′, h p ) Among them, M is the number of sample pixels, SSIM is the image pattern consistency descriptor, μ is the mean calculation operator, σ is the variance calculation operator, C1 and C2 are constants to prevent the denominator from being zero, α is the weight control parameter, X' is the matrix composed of longitudes in the sample, Y' is the matrix composed of latitudes in the sample, and h p is the matrix composed of predicted elevations, I is the original remote sensing image, RFM represents the rational function imaging model, and grid is the pixel coordinate grid obtained by reprojection of X', Y', and h p after reprojection, and is the image generated by sampling on the original image using grid.
Citation Information
Patent Citations
Multi-stereoscopic-image fusion drawing method considering different illumination imaging conditions
CN108305237A