Satellite image ground object height calculation method and system based on monocular parallax estimation

By using a satellite imagery feature height calculation method based on monocular parallax estimation, azimuth and parallax maps are determined using ground control points and RPC models, and height point clouds are generated. This solves the unambiguity problem in height restoration of single satellite images, and achieves accurate restoration and improved stability of height information.

CN120997275AActive Publication Date: 2025-11-21WUHAN UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202511123973.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-11-21
Estimated Expiration
2045-08-12

AI Technical Summary

Technical Problem

In existing technologies, the height restoration of a single satellite image suffers from ambiguity, making it difficult to accurately infer the depth information of objects. This is especially true when objects change significantly at high altitudes, resulting in a lack of height information.

Method used

The method for calculating the height of ground features in satellite imagery based on monocular parallax estimation determines the azimuth angle through ground control points and RPC models. Combining the parallax map and scale relationship, a trained monocular parallax estimation network model is used to generate a height point cloud. After pre-processing, a height map consistent with the geographic range of a single oblique satellite image is generated.

Benefits of technology

It improves the accuracy and stability of height recovery, is suitable for time-sensitive monitoring of changes in ground elevation, solves the problem of geometric position distortion of objects caused by differences in viewing angle, and significantly improves the accuracy and stability of height recovery.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120997275A_ABST
    Figure CN120997275A_ABST
Patent Text Reader

Abstract

The invention provides a satellite image ground object height calculation method based on monocular parallax estimation. The satellite image ground object height calculation method comprises the following steps: accurately calculating an azimuth angle and a height parallax proportion during image shooting by using an RPC model of a satellite image; utilizing a trained monocular parallax estimation network model to extract parallax information in the image along the plumb line direction, wherein the monocular parallax estimation network is composed of an encoder, a decoder layer-by-layer optimization fusion module, a ground classification module and a parallax optimization module; and recovering the height information of each position point in the image by combining the parallax estimation value, the azimuth angle and the height angle. According to the method, the consistency of the image and the height information in the same time phase is ensured, the method is particularly suitable for time sensitivity monitoring of surface height change, a monocular parallax estimation method is adopted, and the problem of object geometric position distortion caused by view angle difference is solved, so that the height recovery precision and stability are remarkably improved; therefore, the technology shows a wide application prospect in the field of monocular tilt satellite image height recovery.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer vision, and particularly relates to a satellite image ground object height calculation method and system based on monocular disparity estimation. BACKGROUND

[0002] Traditional three-dimensional reconstruction methods mainly rely on multiple view stereo (MVS), laser radar (LiDAR) and other technologies. Although these methods can provide high-precision reconstruction results, they often require a large number of multi-view images or rely on expensive equipment and complex sensor systems. In the past three decades, optical satellite remote sensing observation technology has experienced a leap-forward development, and significant progress has been made in key indicators such as spatial resolution, imaging quality and geometric accuracy. High-resolution satellite images not only capture the details of the earth's surface, but also provide accurate geometric information. By introducing advanced geometric correction models such as the rational polynomial coefficient (RPC) model, the positioning accuracy and geometric consistency of the image are further improved, which makes it possible to apply monocular satellite images to three-dimensional reconstruction.

[0003] In near-range applications such as unmanned aerial vehicles and indoor environments, monocular depth estimation has become an important method for restoring three-dimensional scenes using single images. Satellite images are usually taken from high altitudes above 500 kilometers from the ground. Due to the high viewing angle, the scale of the object changes significantly, making it difficult for the depth estimation model to accurately infer the depth information of the object. In addition, the complex imaging mechanism of satellite images (three-line array push-broom imaging model) further increases the difficulty of depth calculation. Therefore, in three-dimensional reconstruction based on a single satellite image, height estimation is usually chosen to avoid the uncertainty brought by long-distance depth estimation. Monocular height estimation refers to the process of using images and digital surface models (DSM) or above-ground level (AGL) to construct a monocular height estimation model to realize the conversion from image information to three-dimensional information. Although depth estimation and height estimation are different in task, the rich depth estimation methods and theories are still valuable knowledge for three-dimensional reconstruction using a single satellite image. Conventional monocular height estimation methods usually face the problem of solution ambiguity, which is caused by the fact that the same two-dimensional image can be generated from different three-dimensional scenes. For example, when a relatively low building is photographed from a relatively low height angle, its image performance may be similar to that of a relatively high building photographed from a relatively high height angle. This loss of height information due to the non-unique mapping relationship between imaging angle and target height is one of the main challenges of monocular height estimation. SUMMARY

[0004] The application provides a satellite image ground object height calculation method and system based on monocular disparity estimation to solve the defects in height recovery of a single oblique satellite image in the prior art.

[0005] In a first aspect, the application provides a satellite image ground object height calculation method based on monocular disparity estimation, comprising: Based on ground control points or known ground object height data, the pixel positions of the top and bottom of the ground object corresponding to the ground control points in the image are determined in combination with the RPC model of the DTM and the satellite image, and the azimuth angle is calculated. The proportional relationship between the disparity and the actual height of the ground object in the image along the vertical line direction is determined through the ground control points or known ground object height data. The disparity value of the ground object top to the ground object bottom along the vertical line direction is extracted from the oblique satellite image using the trained monocular disparity estimation network model to generate a disparity map. Based on the azimuth angle, the proportional relationship, and the disparity map, the height point cloud is obtained. The height point cloud is preprocessed to generate a height map consistent with the geographical range of a single oblique satellite image.

[0006] According to the satellite image ground object height calculation method based on monocular disparity estimation provided by the application, based on ground control points or known ground object height data, the pixel positions of the top and bottom of the ground object corresponding to the ground control points in the image are determined in combination with the RPC model of the DTM and the satellite image, and the azimuth angle is calculated, comprising:

[0007]

[0008]

[0009] wherein, , , , are polynomial functions of the numerator and denominator in the RPC model, , are proportional coefficients in the normalization parameters provided by the RPC model, , are image coordinate offsets in the normalization parameters provided by the RPC model, is the ground control point coordinate, is the ground surface height at the location of the ground object, , , and are the calculated azimuth angles a process variable.

[0010] According to the satellite image ground object height calculation method based on monocular disparity estimation provided by the application, the proportional relationship between the disparity and the actual height of the ground object in the image along the vertical line direction is determined through the ground control points or the known ground object height data, and the method comprises the following steps:

[0011] wherein, is the ellipsoidal height of the position of the ground object, is the surface height of the position of the ground object, which is obtained by DTM interpolation, is the proportional relationship.

[0012] According to the satellite image ground object height calculation method based on monocular disparity estimation provided by the application, the proportional relationship between the disparity and the actual height of the ground object in the image along the vertical line direction is determined through the ground control points or the known ground object height data, and the method comprises the following steps: An original data set composed of multiple oblique satellite images is obtained, and the original data set is divided into a training set, a verification set and a test set, wherein the training set and the verification set contain oblique satellite image blocks, image azimuth angle information and label data, and the test set contains oblique satellite image blocks and image azimuth angle information; A monocular disparity estimation network model is constructed based on a large depth estimation model Depth Anything as a basic framework; A loss function is constructed based on a disparity prediction loss component, a ground loss component and a direction gradient loss component; A preset error evaluation index is established for the network prediction result; The training set is input into the monocular disparity estimation network model, a gradient descent optimization algorithm is used to update the model parameters, the learning rate is adjusted according to the convergence of the loss function in the training process, the model performance is monitored in real time by using the verification set to prevent overfitting, and the trained monocular disparity estimation network model is obtained through multiple iterations; The performance of the trained network model is evaluated on the verification set, the disparity prediction ability of the model is quantitatively analyzed through the preset error evaluation index, and the consistency of the network prediction result and the true result is verified through a visual means; The test set data is input into the trained network model to generate a disparity prediction result, and the disparity prediction result is post-processed to obtain an optimized disparity map.

[0013] According to the satellite image ground object height calculation method based on monocular disparity estimation provided by the application, the proportional relationship between the disparity and the actual height of the ground object in the image along the vertical line direction is determined through the ground control points or the known ground object height data, and the method comprises the following steps: The input image is encoded by using the encoder of the large depth estimation model Depth Anything, the last four layers of the DPT Head module and the DINO-v2 pre-training model of the depth anything are unfreezed to perform feature reconstruction and prediction; The decoding layer layer-by-layer optimization fusion module is constructed, which includes five layers, wherein the channel numbers of the first four layers are the same, and the sizes are sequentially increased to 1 / 16, 1 / 8, 1 / 4 and 1 / 2, the channel number of the fifth layer is 32, and the size is consistent with the original image, for the first four layers, each layer is through ResNet module to the unfreezed decoding layer Feature refinement, through step-by-step upsampling, convolution, connection and convolution operation, the fusion feature layer is obtained:

[0014] Among them, is the decoding feature map of the i-th layer, is the fusion feature layer, , , and are residual, upsampling and convolution in turn; The ground classification module is constructed, and the ground classification module is added after the fifth layer feature of the decoding layer The ground classification module is added after the fifth layer feature of the decoding layer is processed through two 1x1 convolution layers, the first layer convolution reduces the input channel number by half, and the batch normalization operation is used to improve the training stability, the second layer convolution compresses the channel number to 1, and outputs a single channel feature map , wherein each pixel value represents the probability that the corresponding position is a ground point, the Sigmoid activation function is used to enhance the feature and map the classification probability to the range of [0, 1], realize ground point classification, specifically:

[0015] Among them, is the first layer convolution, is the second layer convolution, is the batch normalization operation, is the activation function; The disparity optimization module is constructed based on the fusion feature layer , the position encoding feature is constructed:

[0016] Among them, and represent the arrays of W and H elements uniformly distributed from 0 to 1 respectively; This means stacking the x and y coordinates along a new dimension to obtain a shape like... tensor; This indicates that the tensor will be copied. This yields the final normalized position code, with a size of [value missing]. The first channel represents the x-coordinate, and the second channel represents the y-coordinate. The feature map and the ground classification map are resized using bilinear interpolation to ensure consistency. Then, location features and the ground classification map are encoded separately using convolutional layers to obtain the location-encoded features. and classification coding features ; Location-encoded features are obtained through addition. With fusion feature layer Addition, combining classification coding features through multiplication operations. Adjusting parallax prediction :

[0017] in, This represents the element-wise addition operation between two feature vectors at corresponding positions. The operation of multiplying two feature vectors element by element at corresponding positions; The optimized disparity map is output through the convolutional layer.

[0018] According to the present invention, a method for calculating the height of ground features in satellite imagery based on monocular disparity estimation is provided. A loss function is constructed based on disparity prediction loss components, ground loss components, and orientation gradient loss components, including: The disparity prediction loss components are:

[0019] in, Let be the disparity prediction loss, representing the mean squared error between the predicted disparity map and the true disparity map. To predict disparity maps The disparity value corresponding to the location, For true disparity maps The disparity value corresponding to the location, These are the height and width of the image, respectively; The ground loss component is:

[0020] in, For ground classification loss, represents the binary cross-entropy loss between the predicted classification probability and the classification label. For classification diagram Category value of location, Predicted classification probability map The probability value of the location. These are the height and width of the image, respectively; The directional gradient loss components are:

[0021]

[0022] in, Let be the disparity directional gradient loss, representing the gradient loss function between the predicted disparity map and the true disparity map along the azimuth direction. For the direction vector directional derivative, direction vector Expressed in azimuth angle as , To predict disparity maps The disparity value corresponding to the location, For true disparity maps The disparity value corresponding to the location, It is expressed as mean squared error and is used to measure the difference between the predicted gradient and the true gradient. Loss component predicted by disparity Ground loss component and directional gradient loss components Construct the overall loss function. :

[0023] in, , , These are the corresponding loss weights.

[0024] According to the present invention, a method for calculating the height of ground features in satellite imagery based on monocular parallax estimation is provided. For network prediction results, a preset error evaluation index is established, including: Establish root mean square error :

[0025] in, This represents the number of samples in the validation set. For the first The true value of each sample For the first Predicted values ​​for each sample; Establish the root mean square error of classification :

[0026] in, For category index, the number of samples belonging to the class in the validation set, the true value of the th sample belonging to the class , the predicted value of the th sample belonging to the class ; establish a correlation coefficient R:

[0027] wherein, is the number of samples in the validation set, is the true value of the th sample, is the predicted value of the th sample, is the mean of the true values of the samples, is the mean of the predicted values of the samples.

[0028] According to the satellite image ground object height calculation method based on monocular disparity estimation provided by the application, the height point cloud is obtained based on the azimuth angle, the proportional relationship and the disparity map, and the method comprises the following steps:

[0029]

[0030] wherein, is the proportional relationship, is the azimuth angle, is the original image coordinate, is the correct image coordinate after disparity compensation, is the disparity value of the corresponding position of the coordinate , and is the height value corresponding to the disparity compensation position .

[0031] In a second aspect, the application further provides a satellite image ground object height calculation system based on monocular disparity estimation, comprising: a first calculation module, configured to determine the pixel positions of the top and bottom of the ground control point corresponding ground object in the image based on the ground control point or known ground object height data, combine the DTM and the RPC model of the satellite image, and calculate the azimuth angle; a second calculation module, configured to determine the proportional relationship between the disparity and the actual height of the ground object in the image along the vertical line direction through the ground control point or known ground object height data; The third calculation module is configured to extract a parallax value of a ground object from a top of the ground object to a bottom of the ground object along a vertical line direction from the oblique satellite image by using the trained monocular parallax estimation network model, and generate a parallax map; The fourth calculation module is configured to obtain a height point cloud based on the azimuth angle, the scale relationship and the parallax map. The fifth calculation module is configured to perform a preset processing on the height point cloud, and generate a height map consistent with a geographic range of a single oblique satellite image.

[0032] In a third aspect, the present application further provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method for calculating the height of a ground object in a satellite image based on monocular parallax estimation according to any one of the above aspects.

[0033] The method and system for calculating the height of a ground object in a satellite image based on monocular parallax estimation provided by the present application ensure the consistency of the image and the height information in the same phase, are suitable for time-sensitive monitoring of the height change of the ground surface, solve the problem of geometric position distortion of an object caused by the difference in viewing angle by using the monocular parallax estimation method, thereby significantly improving the accuracy and stability of height recovery, and making the technology have a wide application prospect in the field of height recovery of monocular oblique satellite images. BRIEF DESCRIPTION OF DRAWINGS

[0034] In order to more clearly illustrate the technical solutions in the present application or prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0035] Figure 1 is one of the flowcharts of the method for calculating the height of a ground object in a satellite image based on monocular parallax estimation provided by the present application; Figure 2 is another flowchart of the method for calculating the height of a ground object in a satellite image based on monocular parallax estimation provided by the present application; Figure 3 is a schematic diagram of the monocular depth estimation network structure provided by the present application; Figure 4 is a comparison diagram of the results of the monocular parallax estimation algorithm provided by the present application; Figure 5 is a result diagram of the height map restored from the monocular parallax provided by the present application Figure 6 is a structural diagram of the system for calculating the height of a ground object in a satellite image based on monocular parallax estimation provided by the present application; Figure 7 is a structural schematic diagram of an electronic device provided by the present application. DETAILED DESCRIPTION

[0036] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below with reference to the drawings in the present application. Obviously, the described embodiments are some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work belong to the protection scope of the present application.

[0037] Figure 1 is one of flowcharts of a satellite image ground object height calculation method based on monocular disparity estimation provided by an embodiment of the present application, as shown in Figure 1 , comprising: Step 100: Based on ground control points or known ground object height data, the pixel positions of the top and bottom of the ground object corresponding to the ground control points in the image are determined in combination with the RPC model of the DTM and the satellite image, and the azimuth angle is calculated. Step 200: Through the ground control points or known ground object height data, the proportional relationship between the disparity and the actual height of the ground object in the image along the vertical line direction is determined. Step 300: The trained monocular disparity estimation network model is used to extract the disparity value of the ground object top to the ground object bottom along the vertical line direction from the oblique satellite image, and a disparity map is generated. Step 400: Based on the azimuth angle, the proportional relationship and the disparity map, a height point cloud is obtained. Step 500: The height point cloud is preprocessed to generate a height map consistent with the geographical range of a single oblique satellite image.

[0038] It can be understood that in the past monocular height estimation research, attention has been paid to the concept of “geocentric attitude”. The geocentric attitude describes the direction and height of the ground object relative to the center of gravity of the ground. Through in-depth analysis, it is found that the position offset of the ground object along the center of gravity direction in the image (i.e. the vertical offset of the ground object top to the ground object bottom) is called “center of gravity disparity”, which is more suitable for monocular three-dimensional reconstruction tasks of oblique satellite images than directly estimating the height of the ground object. The main reasons for this adaptability include the following points: (1) One-to-one stable relationship: there is a one-to-one correspondence between the center of gravity disparity and the image, which fundamentally avoids the common many-to-one ill-conditioned problem in traditional height estimation, and improves the stability of the solution space of the task.

[0039] (2) Simple 3D reconstruction: Under the imaging model of approximate parallel projection of satellite images, one image corresponds to a fixed set of azimuth and elevation parameters, so the 3D point cloud information of the ground object can be directly reconstructed by the center parallax without complex nonlinear optimization process.

[0040] (3) Elevation deviation compensation: In a single oblique satellite image, the geographical position distortion of the objects above the ground will cause the predicted elevation to deviate. The center parallax effectively compensates for the error caused by the imaging tilt by estimating the relative relationship between the top and bottom positions of the ground object, thereby improving the accuracy of the elevation estimation.

[0041] (4) Method transferability: Compared with traditional monocular height estimation, the monocular estimation method based on center parallax is more likely to use the technical framework of monocular depth estimation. The difference is that depth estimation relies on perspective geometry, and the light direction of each pixel in the image is affected by different azimuth and elevation angles; while the center parallax estimation is based on parallel projection principle, and the light direction of all pixels in the image shares the same azimuth and elevation angle, the model complexity is significantly reduced.

[0042] Based on the above advantages, the embodiment of the present application proposes a deep learning oblique satellite image height recovery method based on monocular parallax estimation. The method estimates the center parallax in the image to replace the direct elevation estimation, and combines the RPC model of satellite image to calculate the azimuth and elevation information, and accurately restores the height information of each position point in the image.

[0043] Specifically, as shown in Figure 2 , the following steps are included: Step 1, azimuth calculation, based on ground control points or known ground object height data, combined with digital terrain model (Digital Terrain Model, DTM) and RPC model of satellite image, to determine the pixel positions of the top and bottom of the ground control point corresponding to the ground object in the image. Calculate the azimuth , that is, the direction from the top of the ground object to the bottom (also the direction of the plumb line in the image), which provides a spatial reference for subsequent parallax analysis and height recovery. The specific calculation formula of the azimuth is as follows: (1) (2) (3) In formula (1), (2), (3), , , , is a polynomial function of the numerator and denominator in the RPC model, , a scale factor in the normalization parameters provided for the RPC model, , an image coordinate offset in the normalization parameters provided for the RPC model, all of which are provided by the RPC model; is a ground control point coordinate, is a surface height at a location of the object, , , and are process variables for calculating an azimuth angle , respectively.

[0044] Step 2, parallax and height scale calculation, determine the proportionality relationship k between the parallax D and the actual height H of the object in the image along the vertical line through the ground control point or known object height data. The specific calculation formula of the ratio k is: (4) In formula (4), is the ellipsoidal height at the location of the object; is the surface height at the location of the object, which can be obtained by DTM interpolation; and are calculated by formula (2), formula (3), respectively.

[0045] Step 3, parallax estimation information extraction, use the trained monocular parallax estimation network model to automatically extract the parallax value from the top of the object to the bottom of the object along the vertical line direction from a single oblique satellite image, and generate a parallax map D. The specific steps include: Step 3.1. Data set preprocessing. Divide the original data set into training set, validation set and test set, both the training set and the validation set include oblique satellite image small blocks and image azimuth angle information as input data and parallax map along the vertical line direction of the image as label data, and the test set only contains oblique satellite image small blocks and image azimuth angle information as input.

[0046] The embodiment of the application adopts the DFC19 dataset as the basic data source for model training and verification. The original data includes: a small piece of oblique satellite image with a size of 2048x2048, a parallax true value map consistent with the size of the image along the vertical direction, azimuth information corresponding to the image, and the ratio information of parallax and height value for calculating the terrain height from the parallax. In addition, the DFC19 dataset also contains high-density three-dimensional point cloud data, and the above-ground height map aligned with the image is generated through point cloud filtering and interpolation operation. In order to adapt to the model training requirements, the original data is processed such as cropping, resampling, training set and test set division, etc. The size of the cropped image adapts to the input requirements of the model, the resampling ensures the consistency of the data, and the data is strictly divided into training set and test set according to the random division principle. Finally, the DFC19 dataset contains 6450 groups of training data and 1500 groups of test data, and the size of the preprocessed dataset is 512x512.

[0047] Step 3.2. Design the network model. The network model takes the large depth estimation model Depth Anything as the basic framework, uses the encoder enhancement network to improve the parallax estimation ability of the oblique satellite image. The DPT Head in the decoding module is unfrozen, and a layer-by-layer optimization fusion module is designed in the decoding layer to extract high-resolution feature representation. In addition, a ground classification module is integrated at the end of the decoding module to predict the ground point probability, and a parallax optimization module is designed combining the position encoding and ground classification information to further improve the accuracy of parallax prediction. The overall network structure is as shown in Figure 3 Module 1. Encoder module. The method adopts the large depth estimation model Depth Anything as the encoder to enhance the generalization ability of the parallax estimation task, and unfreezes the DPT Head module in the decoder, so that the depth estimation large model feature extraction network can realize efficient feature reconstruction and prediction in the monocular parallax estimation task of oblique satellite image.

[0048] Module 2. Layer-by-layer optimization fusion module in decoding layer. The decoding module mentioned in module 1 contains five layers. The channel numbers of the first four layers are the same, and the sizes are 1 / 16, 1 / 8, 1 / 4 and 1 / 2 of the original image size in turn. The channel number of the fifth layer is 32, and the size is consistent with the original image. For the first four layers, each layer refines the unfreezed decoding layer through the ResNet module, and the feature map represents the decoding feature of the first layer. After the following operations, the fusion feature is obtained: (5) ​​​decoded feature map of the layer, for the fusion feature layer, 、 and Residual, up-sampling and convolution are sequentially performed; deep-level feature fusion and nonlinear mapping further enhance the network's expression ability for different scale details and improve its adaptability to complex terrain.

[0049] Module 3. Ground classification module. After obtaining the fifth layer feature of the decoding layer described in module 2 , a ground classification module is added. This module processes the input through two 1x1 convolution layers. The first layer of convolution reduces the number of input channels by half and improves the training stability through batch normalization operation. The second layer of convolution compresses the channel number to 1, outputting a single-channel feature map , where each pixel value represents the probability of the position being a ground point. Finally, using the Sigmoid activation function, the feature is enhanced and the classification probability is mapped to the range of [0, 1], realizing the classification of ground points. The specific implementation steps are as follows: (6) where, is the first layer of convolution, is the second layer of convolution, is the batch normalization operation, and Sigmoid activation function.

[0050] Module 4. Disparity optimization module. This module is used to assist disparity estimation by combining position encoding and ground classification information to optimize disparity prediction.

[0051] First, based on the fusion feature described in module 2, the position encoding feature is constructed as follows: (7) where, and represent arrays of W and H elements respectively generated from a uniform distribution from 0 to 1; represents stacking the x and y coordinates along the new dimension to obtain a tensor with shape ; represents copying the tensor times to obtain the final normalized position encoding with size , where the first channel is the x coordinate and the second channel is the y coordinate; Then, the size of the feature map and the ground classification map is adjusted to be consistent through bilinear interpolation, and the position feature and the ground classification map are encoded through convolution layers respectively to obtain the position encoding feature and the classification encoding feature​ .

[0052] Then, the position encoding feature is added to the fusion feature obtained by the module 2 through an addition operation, and the ground classification feature is combined through a multiplication operation to further adjust the disparity prediction. : (8) wherein, represents an element-wise addition operation of two feature vectors at corresponding positions, an element-wise multiplication operation of two feature vectors at corresponding positions; Finally, an optimized disparity map is output through a convolution layer.

[0053] Step 3.3, constructing a loss function. The loss function is composed of multiple components, including: disparity prediction loss, calculating the difference between the predicted disparity map and the true disparity map through mean square error (MSE); ground classification loss, using a binary cross-entropy loss function to constrain the ground classification result; direction gradient loss: according to the direction information of the input image, the predicted disparity map is constrained along the feature direction, and the structural consistency of the disparity map is optimized. In the embodiment of the present application, the weights of the disparity prediction loss, the ground classification loss and the direction gradient loss function are usually set to 1, 10 and 1 in turn.

[0054] The disparity prediction loss component is: (9) wherein, is the disparity prediction loss, representing the mean square error between the predicted disparity map and the true disparity map, is the predicted disparity map the disparity value corresponding to the position, is the true disparity map the disparity value corresponding to the position, respectively the height and width of the image; The ground loss component is: (10) wherein, is the ground classification loss, representing the binary cross-entropy loss between the predicted classification probability and the classification label, is the classification map the classification value of the position, is the predicted classification probability map the probability value of the position, respectively the height and width of the image; The direction gradient loss component is: (11) ​ (12) in, Let be the disparity directional gradient loss, representing the gradient loss function between the predicted disparity map and the true disparity map along the azimuth direction. For the direction vector directional derivative, direction vector Expressed in azimuth angle as , To predict disparity maps The disparity value corresponding to the location, For true disparity maps The disparity value corresponding to the location, It is expressed as mean squared error and is used to measure the difference between the predicted gradient and the true gradient. The overall loss function consists of the three loss components mentioned above, namely: (13) in, , , These are the corresponding loss weights.

[0055] Step 3.4: Establish error evaluation metrics. Based on the network prediction results, the established error evaluation metrics include: Root Mean Square Error (RMSE), which measures the overall deviation between the predicted disparity map and the true disparity map; Correlation Coefficient (R), which measures the linear correlation between the predicted disparity map and the true disparity map; and Classification Disparity Error, which categorizes disparities into three classes: low (disparity value less than 5), medium (disparity value between 5 and 15), and high (true disparity value greater than 15). The RMSE is calculated for each class of disparity to measure the model's prediction deviation across different disparity categories, comprehensively evaluating the model's adaptability and performance. The specific formula is: Establish root mean square error : (14) in, This represents the number of samples in the validation set. For the first The true value of each sample For the first Predicted values ​​for each sample; Establish the root mean square error of classification : (15) in, For category index, To verify that the set belongs to the category The number of samples, To belong to category The The true value of each sample To belong to category The Predicted values ​​for each sample; Establish the correlation coefficient R: (16) in, This represents the number of samples in the validation set. For the first The true value of each sample For the first The predicted value for each sample, for The mean of the true values ​​of each sample for The mean of the predicted values ​​for each sample.

[0056] Step 3.5. Network Model Training. Input the training data into the network model and update the model parameters using the gradient descent optimization algorithm. Adjust the learning rate based on the convergence of the loss function during training, and monitor model performance in real time using the validation set to prevent overfitting. Through multiple iterations, the final model is obtained.

[0057] Step 3.6. Network Model Validation. Evaluate the performance of the trained network model on the validation set. Quantitatively analyze the model's disparity prediction ability using evaluation metrics, and verify the consistency between the network prediction results and the actual results using visualization methods.

[0058] Step 3.7. Network Model Testing. Input the test set data into the trained network model to generate disparity prediction results. Post-process the prediction results (such as denoising, interpolation, etc.) to optimize the final disparity map.

[0059] After completing the network design, the preprocessed training set is input, and the loss function is calculated using the backpropagation algorithm. Parameter learning is performed by iteratively reducing the error to obtain the optimal weight model. In actual training, the PyTorch Lightening structure is used, with a maximum iteration count of 200, a batch processing parameter of 4, and the Adam optimizer with a learning rate of 0.0001. The pre-trained weight model is loaded, and the test set is input into the network model to obtain a monocular disparity map. The obtained disparity map is compared with the ground truth disparity map to evaluate the weight model. It is also compared with five other methods (img2height, MQTransformer, geocentric pose, zoedepth, and depthanything). The comparison results are shown in Table 1. The disparity map comparison results are displayed as follows: Figure 4 As shown.

[0060] Table 1 Comparison of monocular disparity estimation results of DFC19 dataset

[0061] As can be seen from Table 1, in the monocular disparity estimation task, the present application is superior to the other five methods in performance, especially in complex scenes where the true value disparity is greater than 10. When using a large model (vitl) and direction gradient enhancement (grad), the accuracy can be further improved, showing better accuracy and robustness. Figure 4 The results of five comparison methods and the present application using different models are shown, which also shows the advantages of the present application.

[0062] Step 4, height point cloud extraction, based on the azimuth angle obtained in step 1.1 and the scale factor k obtained in step 1.2, combined with the disparity map D estimated by the network model in step 1.3, to obtain the height point cloud. The specific calculation formula is: (17) (18) wherein, is the proportional relationship, is the azimuth angle, is the original image coordinate, is the correct image coordinate after disparity compensation, is the coordinate corresponding to the position of the disparity value, is the disparity compensation position corresponding to the height value.

[0063] Step 5, height point cloud post-processing, filtering, constructing an irregular triangular net for interpolation, etc. on the obtained height point cloud, to finally generate a height map consistent with the geographical range of the image.

[0064] Figure 5 The process of recovering from the disparity map to the height map is shown: first, combine the azimuth angle and the height scale of the disparity to recover the disparity map into a height point cloud; then filter the height point cloud to remove the point cloud with a density of less than 5 points in a spatial range of 2 to obtain a filtered height point cloud; finally, interpolate to obtain a height map. The height map recovered from the disparity map is compared with the results of training using four height map true values, and the comparison results are shown in Table 2. In Table 2, "the present application" refers to the result of converting the disparity result using the vitl model and using the direction gradient loss (grad) into height.

[0065] Table 2 Comparison of height map results of DFC19 dataset

[0066] As can be seen from Table 2, the application performs more outstandingly in the scene where the true disparity is greater than 10, and the comprehensive performance leads other methods. This shows that the application has not only higher accuracy, but also better adaptability and application potential when dealing with complex scenes with larger disparity.

[0067] The satellite image ground object height calculation system based on monocular disparity estimation provided by the application is described below, and the satellite image ground object height calculation system based on monocular disparity estimation described below can be correspondingly referred to the satellite image ground object height calculation method based on monocular disparity estimation described above.

[0068] Figure 6 The satellite image ground object height calculation system based on monocular disparity estimation provided by the embodiment of the application is a structural schematic diagram, as shown in Figure 6 The satellite image ground object height calculation system based on monocular disparity estimation provided by the embodiment of the application is a structural schematic diagram, as shown in The first calculation module 61 is used to determine the pixel positions of the top and bottom of the ground control point corresponding ground object in the image based on the ground control point or known ground object height data, combined with the RPC model of the DTM and the satellite image, and calculate the azimuth angle; the second calculation module 62 is used to determine the proportional relationship between the disparity and the actual height of the ground object in the image along the vertical line direction through the ground control point or known ground object height data; the third calculation module 63 is used to extract the disparity value of the ground object top to the ground object bottom along the vertical line direction from the oblique satellite image by using the trained monocular disparity estimation network model, and generate a disparity map; the fourth calculation module 64 is used to obtain the height point cloud based on the azimuth angle, the proportional relationship and the disparity map; the fifth calculation module 65 is used to perform a preset processing on the height point cloud, and generate a height map consistent with the geographical range of a single oblique satellite image.

[0069] Figure 7 An example of an entity structure schematic diagram of an electronic device is shown in Figure 7As shown, the electronic device can include a processor 710, a communications interface 720, a memory 730, and a communications bus 740, wherein the processor 710, the communications interface 720, and the memory 730 complete mutual communication through the communications bus 740. The processor 710 can invoke a logic instruction in the memory 730 to execute a satellite image ground object height calculation method based on monocular disparity estimation, which includes: based on ground control points or known ground object height data, combining a DTM and an RPC model of a satellite image, determining pixel positions of a top and a bottom of a ground object corresponding to the ground control points in the image, and calculating an azimuth angle; determining a proportional relationship between a disparity and an actual height of a ground object in the image along a vertical line direction through the ground control points or the known ground object height data; extracting a disparity value from the top of the ground object to the bottom of the ground object along the vertical line direction from the inclined satellite image using a trained monocular disparity estimation network model to generate a disparity map; obtaining a height point cloud based on the azimuth angle, the proportional relationship, and the disparity map; and performing a preset processing on the height point cloud to generate a height map consistent with a geographical range of a single inclined satellite image.

[0070] In addition, the logic instruction in the memory 730 described above can be implemented in the form of a software function unit and sold or used as an independent product, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, includes a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the method described in various embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various program code storage media.

[0071] The device embodiments described above are only schematic, wherein the units shown as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the present embodiment scheme according to actual needs. Those skilled in the art can understand and implement without creative labor.

[0072] Those skilled in the art can clearly understand the technical solutions of the various embodiments from the above description of the embodiments, and the various embodiments can be implemented by means of software with the necessary general hardware platforms, and of course, can also be implemented by hardware. Based on such understanding, the above technical solutions, essentially or in other words, the part of the prior art that makes a contribution, can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, and the like, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0073] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for some technical features therein; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for calculating the height of ground objects in satellite images based on monocular disparity estimation, characterized in that, The application comprises the following steps: Based on ground control points or known height data of ground objects, a rational polynomial coefficient RPC model is combined with a digital terrain model DTM and satellite images to determine the pixel positions of the top and bottom of the ground objects corresponding to the ground control points in the images, and the azimuth is calculated; The proportion relationship between the parallax and the actual height of the ground objects in the images along the vertical line direction is determined through the ground control points or known height data of the ground objects; A trained monocular disparity estimation network model is used to extract the parallax value from the top of the ground objects to the bottom of the ground objects along the vertical line direction from the oblique satellite images to generate a parallax map; Based on the azimuth, the proportion relationship and the parallax map, the height point cloud is obtained; The height point cloud is subjected to a preset treatment to generate a height map consistent with the geographical range of a single oblique satellite image.

2. The method of claim 1, wherein, Based on ground control points or known height data of ground objects, a rational polynomial coefficient RPC model is combined with a digital terrain model DTM and satellite images to determine the pixel positions of the top and bottom of the ground objects corresponding to the ground control points in the images, and the azimuth is calculated, comprising: wherein, , , , are polynomial functions of the numerator and denominator in the RPC model, respectively, , are scale factors in the normalization parameters provided for the RPC model, , are image coordinate offsets in the normalization parameters provided for the RPC model, are ground control point coordinates, is the surface height at the location of the object on the ground, , , and are process variables for calculating the azimuth , respectively.

3. The method of claim 2, wherein the method further comprises: The proportion relationship between the parallax and the actual height of the ground objects in the images along the vertical line direction is determined through the ground control points or known height data of the ground objects, comprising: wherein, is the ellipsoidal height of the location of the ground object, is the ground surface height of the location of the ground object, obtained by DTM interpolation, is the proportionality relationship. 4.The method of claim 1, wherein, A trained monocular disparity estimation network model is used to extract the parallax value from the top of the ground objects to the bottom of the ground objects along the vertical line direction from the oblique satellite images to generate a parallax map, comprising: An original data set composed of multiple oblique satellite images is obtained, and the original data set is divided into a training set, a validation set and a test set, wherein the training set and the validation set contain small blocks of oblique satellite images, image azimuth information and label data, and the test set contains small blocks of oblique satellite images and image azimuth information; A monocular disparity estimation network model is constructed based on a large depth estimation model Depth Anything as a basic framework; A loss function is constructed based on a parallax prediction loss component, a ground loss component and a direction gradient loss component; A preset error evaluation index is established for the network prediction result; The training set is input into the monocular disparity estimation network model, a gradient descent optimization algorithm is used to update the model parameters, the learning rate is adjusted according to the convergence of the loss function in the training process, the model performance is monitored in real time by using the validation set to prevent overfitting, and through multiple rounds of iteration, a trained monocular disparity estimation network model is obtained; The performance of the trained network model is evaluated on the validation set, the parallax prediction ability of the model is quantitatively analyzed through the preset error evaluation index, and the consistency of the network prediction result and the true result is verified through a visual means; The test set data is input into the trained network model to generate a parallax prediction result, and the parallax prediction result is post-processed to obtain an optimized parallax map.

5. The method for calculating the height of ground objects from satellite images based on monocular parallax estimation according to claim 4, characterized in that, A monocular disparity estimation network model is constructed based on a large depth estimation model Depth Anything as a basic framework, comprising: The input image is encoded by using the encoder of the large depth estimation model Depth Anything, the DPT Head module of Depth Anything and the last four layers of the DINO-v2 pre-training model are unfrozen, and feature reconstruction and prediction are performed. A decoding layer layer-by-layer optimization fusion module is constructed, and the decoding layer layer-by-layer optimization fusion module includes five layers, wherein the first four layers have the same number of channels, and the sizes are sequentially increased to 1 / 16, 1 / 8, 1 / 4 and 1 / 2 of the original image size, the fifth layer has 32 channels, and the size is consistent with the original image, for the first four layers, each layer refines the features of the unfrozen decoding layer through a ResNet module, and through step-by-step upsampling, convolution, connection and convolution operations, a fusion feature layer is obtained: The optimized disparity map is output through a convolution layer. wherein, is the first layer of decoded feature maps, is the fusion feature layer, , and are residual, up-sampling and convolution, respectively. A ground classification module is constructed, and the fifth layer feature of the decoding layer is input into the ground classification module Then, the ground classification module is added, and the input is processed through two 1x1 convolution layers The first layer of convolution reduces the number of input channels by half, and the batch normalization operation is used to improve the training stability, and the second layer of convolution compresses the number of channels to 1 to output a single-channel feature map where each pixel value represents the probability that the corresponding position is a ground point, a Sigmoid activation function is used to enhance the feature and map the classification probability to the range of [0, 1], and the ground point classification is realized, specifically: wherein, is a first layer convolution, is a second layer convolution, is a batch normalization operation, is a Sigmoid activation function; A disparity optimization module is constructed based on a fused feature layer , a position encoding feature is constructed : where, and denotes an array of W and H elements uniformly distributed from 0 to 1 respectively; denotes stacking the x and y coordinates along a new dimension, resulting in a tensor of shape ; denotes copying the tensor times, resulting in the final normalized position encoding of size , where the 1st channel is the x coordinate and the 2nd channel is the y coordinate; The sizes of the feature map and the ground classification map are adjusted through bilinear interpolation to be consistent, and the position feature and the ground classification map are respectively processed through a convolution layer to obtain position encoding features and classification encoding features and classification encoding features ; Position encoding features are added by an addition operation to the fusion feature layer Classification encoding features are combined by a multiplication operation , adjusting disparity prediction : wherein, represents an element-wise addition operation of two feature vectors at corresponding positions, represents an element-wise multiplication operation of two feature vectors at corresponding positions; A loss function is constructed based on a disparity prediction loss component, a ground loss component and a direction gradient loss component, including:

6. The method of claim 4, wherein, The disparity prediction loss component is: The ground loss component is: wherein, is the disparity prediction loss, representing the mean square error between the predicted disparity map and the ground truth disparity map, is the predicted disparity map is the disparity value corresponding to the position, is the ground truth disparity map is the disparity value corresponding to the position, are the height and width of the image, respectively; The direction gradient loss component is: wherein, is a ground classification loss, representing a binary cross-entropy loss between the predicted classification probability and the classification label, is a classification map a classification value for a location, is a predicted classification probability map a probability value for a location, are respectively height and width of the image; A preset error evaluation index is established for the network prediction result, including: wherein, is the disparity directional gradient loss, representing the loss function of the gradient along the azimuthal direction between the predicted disparity map and the ground truth disparity map, is the directional derivative along the direction vector , the direction vector is represented by the azimuthal angle as , is the disparity value of the predicted disparity map at the corresponding position, is the disparity value of the ground truth disparity map at the corresponding position, is represented as the mean square error, used to measure the difference between the predicted gradient and the ground truth gradient; from the parallax prediction loss component , the ground loss component and the orientation gradient loss component , to build a total loss function : wherein, , , are the respective loss weights.

7. The method of claim 4, wherein the method further comprises: The correlation coefficient R is established: establishing root mean square error : in, The number of samples in the validation set. For the first The true value of each sample For the first Predicted values ​​for each sample; establishing a classification root mean square error : wherein, is the number of samples in the validation set belonging to class is the number of samples in the validation set belonging to class is the number of samples in the validation set belonging to class is the true value of the th sample belonging to class is the predicted value of the th sample belonging to class is the predicted value of the th sample belonging to class Based on the azimuth angle, the proportional relationship and the disparity map, a height point cloud is obtained, including: wherein, the number of samples of the validation set, the true value of the th sample, the predicted value of the th sample, the mean of the true values of the th samples, the mean of the predicted values of the th samples.

8. The method of claim 1, wherein, Including: wherein, is a proportional relationship, is an azimuth angle, is an original image coordinate, is a correct image coordinate after parallax compensation, is a coordinate a parallax value of a corresponding position, is a parallax compensation position a corresponding height value. 9.A system for calculating the height of ground objects in satellite images based on monocular disparity estimation, characterized in that, A first calculation module is configured to determine the pixel positions of the top and bottom of the ground control point corresponding ground object in the image based on the ground control point or known ground object height data, combine the DTM and the RPC model of the satellite image, and calculate the azimuth angle; A second calculation module is configured to determine the proportional relationship between the disparity and the actual height of the ground object in the image along the vertical line direction based on the ground control point or known ground object height data; A third calculation module is configured to extract the disparity value of the ground object from the top to the bottom of the ground object along the vertical line direction from the oblique satellite image using the trained monocular disparity estimation network model, and generate a disparity map; A fourth calculation module is configured to obtain a height point cloud based on the azimuth angle, the proportional relationship and the disparity map; A fifth calculation module is configured to perform a preset processing on the height point cloud to generate a height map consistent with the geographical range of a single oblique satellite image. The processor executes the program to implement the satellite image ground object height calculation method based on monocular disparity estimation according to any one of claims 1 to 8.

10. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, ​

Citation Information

Patent Citations

  • Height-level-guided single-view-angle remote sensing image building height estimation method and device

    CN116503744A

  • Satellite image parallax estimation method and system based on multi-scale geometric coding and texture decoding

    CN120298469A

  • Satellite image cascade matching method and system based on double-branch context awareness

    CN120298730A

  • Water Tank with Horizontal Waveform Panel

    KR1020250006586A

  • Satellite image processing system and method for estimating point spread function based on satellite image

    KR102465292B1