Remote sensing image generation method and device based on air-to-ground perspective geometric transformation
Patent Information
- Application Number
- CN202311084673.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-25
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-08-25
AI Technical Summary
[0004]本申请提供一种基于地空视角几何变换的遥感图像生成方法及装置,以解决相关技术中的通过插值算法生成遥感图像,无法对多变的地貌进行准确的建模,造成遥感图像存在失真,导致生成的遥感图像的质量较差,无法满足实际场景的应用需求的问题
[0017]本申请实施例可以根据原始圆柱体投影获得第一地面全景图像,将第一地面全景图像进行立方体投影获得第二地面全景图像,对第一地面全景图像进行深度估计,并获得第一深度图和初始地对空注意力掩膜,进而对第二地面全景图像进行特征提取,得到目标方向的特征图并聚合,得到特征向量,从而构建经度和纬度的环路超图,进而构建地对空注意力系数,将地对空注意力系数和初始地对空注意力掩膜进行加权处理,获得地对空注意力系数加权后的地对空注意力掩膜,将第一深度图、第一地面全景图像以及地对空注意力掩膜进行几何转换,获得遥感视角下的目标视角投影图像,并修复目标视角投影图像的纹理和细节,生成最终的遥感图像,进而有效的提升遥感图像的质量,满足实际场景的应用需求。由此,解决了相关技术中的通过插值算法生成遥感图像,无法对多变的地貌进行准确的建模,造成遥感图像存在失真,导致生成的遥感图像的质量较差,无法满足实际场景的应用需求的问题。
Smart Images

Figure CN117274514B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of remote sensing image generation technology, and in particular to a remote sensing image generation method and apparatus based on ground-to-air perspective geometric transformation. Background Technology
[0002] In related technologies, pixel-level interpolation can be performed on the original image using interpolation algorithms, or semantic matching and extraction can be performed based on datasets to reconstruct the relationship between ground features in detail. By modeling the terrain, remote sensing images can be generated, thus enabling the application of remote sensing technology in real-world scenarios.
[0003] However, the interpolation algorithms used in related technologies to generate remote sensing images cannot accurately model the varied terrain, resulting in distortion of the remote sensing images and poor quality of the generated images, which cannot meet the application needs of real-world scenarios and urgently needs to be solved. Summary of the Invention
[0004] This application provides a remote sensing image generation method and apparatus based on ground-to-air perspective geometric transformation, in order to solve the problem in related technologies that remote sensing images generated by interpolation algorithms cannot accurately model the varied terrain, resulting in distortion of the remote sensing images and poor quality of the generated remote sensing images, which cannot meet the application needs of real-world scenarios.
[0005] The first aspect of this application provides a remote sensing image generation method based on ground-to-air perspective geometric transformation, comprising the following steps: acquiring an original cylindrical projection and obtaining a first ground panoramic image based on the original cylindrical projection; performing a cube projection on the first ground panoramic image and obtaining a second ground panoramic image based on the cube projection; performing depth estimation on the first ground panoramic image to obtain a depth estimation result of the first ground panoramic image, and obtaining a first depth map and an initial ground-to-air attention mask based on the depth estimation result of the first ground panoramic image; performing feature extraction on the second ground panoramic image to obtain a feature map of the target direction, and then... The feature maps are aggregated to obtain feature vectors, and a loop hypergraph of longitude and latitude is constructed using the feature vectors. Ground-to-air attention coefficients are constructed based on the loop hypergraphs, and the ground-to-air attention coefficients and the initial ground-to-air attention mask are weighted to obtain a weighted ground-to-air attention mask. The first depth map, the first ground panoramic image, and the ground-to-air attention mask are geometrically transformed to obtain a target view projection image from a remote sensing perspective. The target view projection image is input into a target remote sensing image generation module to repair the texture and details of the target view projection image, generating the final remote sensing image.
[0006] Optionally, in one embodiment of this application, the step of performing depth estimation on the first ground panoramic image to obtain a depth estimation result of the first ground panoramic image, and obtaining a first depth map and an initial ground-to-air attention mask based on the depth estimation result, includes: using a target image generation network to obtain a first depth map and an initial ground-to-air attention mask prediction of the first ground panoramic image, wherein the target image generation network has 3 input channels and 2 output channels, wherein the first output channel of the target image generation network is the first depth map, and the second output channel is the initial ground-to-air attention mask prediction.
[0007] Optionally, in one embodiment of this application, the step of extracting features from the second ground panoramic image to obtain a feature map in the target direction, aggregating the feature maps to obtain feature vectors, and constructing a loop hypergraph of longitude and latitude using the feature vectors includes: inputting the panoramic image in the target direction into a feature extraction network, using a ResNet (Residual Network) structure to obtain a feature map in the target direction; inputting the feature map into a feature aggregation network for aggregation to obtain at least one feature vector; and using the at least one feature vector as a node to construct the loop hypergraph of longitude and latitude.
[0008] Optionally, in one embodiment of this application, the step of constructing ground-to-air attention coefficients based on the loop hypergraph, and performing weighted processing based on the ground-to-air attention coefficients and the initial ground-to-air attention mask to obtain a weighted ground-to-air attention mask includes: performing a preset convolution on the features of the loop hypergraph nodes to obtain new features of the target nodes of the loop hypergraph; combining the new features into two sets of feature matrices; concatenating the target features in the two sets of feature matrices to obtain two new feature vectors; combining the two new feature vectors into a binary array; and performing matrix multiplication on the binary array to obtain ground-to-air attention coefficients; and performing weighted processing on the predicted initial ground-to-air attention mask and the ground-to-air attention coefficients to obtain a weighted ground-to-air attention mask.
[0009] Optionally, in one embodiment of this application, the step of geometrically transforming the first depth map, the first ground panoramic image, and the ground-to-air attention mask to obtain a target view projection image from a remote sensing perspective includes: weighting the first depth map and the ground-to-air attention mask to obtain a weighted second depth map; using the second depth map, converting the homogeneous panoramic image coordinates of the first ground panoramic image into three-dimensional coordinates in the camera coordinate system to obtain converted non-homogeneous panoramic image coordinates; based on the non-homogeneous panoramic image coordinates, converting the RGB pixel values of the first ground panoramic image into homogeneous remote sensing image coordinates from a remote sensing perspective; based on the homogeneous remote sensing image coordinates, converting each pixel in the first ground panoramic image to obtain the RGB value of each pixel in the final remote sensing image, and obtaining the target view projection image from the remote sensing perspective based on the RGB values.
[0010] A second aspect of this application provides a remote sensing image generation device based on ground-to-air perspective geometric transformation, comprising: a first acquisition module for acquiring an original cylindrical projection and obtaining a first ground panoramic image based on the original cylindrical projection; a second acquisition module for performing a cube projection on the first ground panoramic image and obtaining a second ground panoramic image based on the cube projection; a third acquisition module for performing depth estimation on the first ground panoramic image to obtain a depth estimation result of the first ground panoramic image, and obtaining a first depth map and an initial ground-to-air attention mask based on the depth estimation result; and a first determination module for performing feature extraction on the second ground panoramic image to obtain features of the target direction. The system comprises: a feature map, a feature vector, and a loop hypermap of longitude and latitude; a construction module, which constructs ground-to-air attention coefficients based on the loop hypermap, and weights the ground-to-air attention coefficients and the initial ground-to-air attention mask to obtain a weighted ground-to-air attention mask; a second determination module, which performs geometric transformation on the first depth map, the first ground panoramic image, and the ground-to-air attention mask to obtain a target view projection image from a remote sensing perspective; and a generation module, which inputs the target view projection image into a target remote sensing image generation module to repair the texture and details of the target view projection image and generate the final remote sensing image.
[0011] Optionally, in one embodiment of this application, the third acquisition module includes: a first acquisition unit, configured to use a target image generation network to obtain a first depth map and an initial ground-to-air attention mask prediction of the first ground panoramic image, wherein the target image generation network has 3 input channels and 2 output channels, wherein the first output channel of the target image generation network is the first depth map and the second output channel is the initial ground-to-air attention mask prediction.
[0012] Optionally, in one embodiment of this application, the first determining module includes: a first determining unit, configured to input the panoramic image of the target direction into a feature extraction network and obtain a feature map in the target direction using a ResNet structure; a second determining unit, configured to input the feature map into a feature aggregation network for aggregation to obtain at least one feature vector; and a construction unit, configured to use the at least one feature vector as nodes to construct the loop hypergraph of the longitude and the latitude.
[0013] Optionally, in one embodiment of this application, the construction module includes: a second acquisition unit, configured to perform a preset convolution on the features of the loop hypergraph nodes to obtain new features of the target nodes of the loop hypergraph, and combine the new features into two sets of feature matrices; a third acquisition unit, configured to concatenate the target features in the two sets of feature matrices to obtain two new feature vectors, combine the two new feature vectors into a binary array, and perform matrix multiplication on the binary array to obtain ground-to-air attention coefficients; and a processing unit, configured to perform weighted processing on the predicted initial ground-to-air attention mask and the ground-to-air attention coefficients to obtain a weighted ground-to-air attention mask.
[0014] Optionally, in one embodiment of this application, the second determining module includes: a third determining unit, configured to perform weighted processing on the first depth map and the ground-to-air attention mask to obtain a weighted second depth map; a fourth determining unit, configured to use the second depth map to convert the homogeneous panoramic image coordinates of the first ground panoramic image into three-dimensional coordinates in the camera coordinate system to obtain converted non-homogeneous panoramic image coordinates; a conversion unit, configured to convert the RGB pixel values of the first ground panoramic image into homogeneous remote sensing image coordinates under the remote sensing perspective based on the non-homogeneous panoramic image coordinates; and a fifth determining unit, configured to convert each pixel in the first ground panoramic image based on the homogeneous remote sensing image coordinates to obtain the RGB value of each pixel in the final remote sensing image, and obtain the target view projection image under the remote sensing perspective based on the RGB values.
[0015] A third aspect of this application provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the remote sensing image generation method based on ground-to-space perspective geometric transformation as described in the above embodiments.
[0016] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the remote sensing image generation method based on ground-to-air perspective geometric transformation as described above.
[0017] This application embodiment can obtain a first ground panoramic image based on the original cylindrical projection, and then obtain a second ground panoramic image by performing cube projection on the first ground panoramic image. Depth estimation is performed on the first ground panoramic image to obtain a first depth map and an initial ground-to-air attention mask. Feature extraction is then performed on the second ground panoramic image to obtain a feature map of the target direction, which is then aggregated to obtain a feature vector, thereby constructing a longitude and latitude loop hypergraph. Ground-to-air attention coefficients are then constructed, and the ground-to-air attention coefficients and the initial ground-to-air attention mask are weighted to obtain a weighted ground-to-air attention mask. Geometric transformation is then performed on the first depth map, the first ground panoramic image, and the ground-to-air attention mask to obtain a target-view projection image from a remote sensing perspective. The texture and details of the target-view projection image are then repaired to generate the final remote sensing image, effectively improving the quality of the remote sensing image and meeting the application requirements of practical scenarios. This solves the problem in related technologies where remote sensing images generated by interpolation algorithms cannot accurately model varied terrain, resulting in distortion and poor quality of the generated remote sensing images, failing to meet the application requirements of practical scenarios.
[0018] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0019] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:
[0020] Figure 1 This is a flowchart of a remote sensing image generation method based on ground-to-air perspective geometric transformation provided in an embodiment of this application;
[0021] Figure 2 This is a schematic diagram illustrating the construction process of a hyperedge in a loop hypergraph according to a specific embodiment of this application;
[0022] Figure 3This is a schematic diagram of a remote sensing image generation algorithm framework based on ground-to-air perspective geometric transformation, which is a specific embodiment of this application.
[0023] Figure 4 This is a schematic diagram of a ground-to-air projection method according to a specific embodiment of this application;
[0024] Figure 5 This is a schematic diagram of the remote sensing image generation device based on ground-to-air perspective geometric transformation according to an embodiment of this application.
[0025] Figure 6 This is a schematic diagram of the structure of an electronic device provided according to an embodiment of this application. Detailed Implementation
[0026] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.
[0027] The following describes a remote sensing image generation method and apparatus based on ground-to-air perspective geometric transformation according to embodiments of this application, with reference to the accompanying drawings. Addressing the problem mentioned in the background section that remote sensing images generated using interpolation algorithms cannot accurately model varied terrain, resulting in image distortion and poor image quality that fails to meet the application requirements of real-world scenarios, this application provides a remote sensing image generation method based on ground-to-air perspective geometric transformation. In this method, a first panoramic ground image is obtained based on the original cylindrical projection; a second panoramic ground image is obtained by performing cube projection on the first panoramic ground image; depth estimation is performed on the first panoramic ground image to obtain a first depth map and an initial ground-to-air attention mask; and then the second panoramic ground image is generated... Feature extraction is performed on the ground panoramic image to obtain feature maps of the target direction, which are then aggregated to obtain feature vectors. This constructs a longitude and latitude loop hypergraph, which in turn generates ground-to-air attention coefficients. These coefficients are then weighted with an initial ground-to-air attention mask to obtain a weighted ground-to-air attention mask. A geometric transformation is performed on the first depth map, the first ground panoramic image, and the ground-to-air attention mask to obtain a target-view projection image from a remote sensing perspective. The texture and details of this projection image are then restored to generate the final remote sensing image, effectively improving its quality and meeting the application requirements of real-world scenarios. This solves the problem in related technologies where interpolation algorithms fail to accurately model varied terrain, resulting in distorted and low-quality remote sensing images that do not meet the needs of practical applications.
[0028] Specifically, Figure 1 This is a flowchart illustrating a remote sensing image generation method based on ground-to-air perspective geometric transformation, provided in an embodiment of this application.
[0029] like Figure 1 As shown, the remote sensing image generation method based on ground-to-air perspective geometric transformation includes the following steps:
[0030] In step S101, the original cylindrical projection is acquired, and a first panoramic ground image is obtained based on the original cylindrical projection.
[0031] It is understood that the embodiments of this application can acquire the original cylindrical projection. For example, a ground panoramic camera can be used to acquire the original cylindrical projection of the camera at a selected ground location, and a first ground panoramic image P can be obtained based on the original cylindrical projection. e This effectively improves the feasibility of remote sensing image generation.
[0032] In step S102, the first ground panoramic image is subjected to cube projection, and the second ground panoramic image is obtained based on the cube projection.
[0033] It is understood that, in the embodiments of this application, a first ground panoramic image can be cube-projected to obtain a second ground panoramic image, for example, as shown below. Figure 2 As shown, the second ground panoramic image P c It can contain six directions, where the six directions are P c1 ,P c2 ,P c3 ,P c4 ,P c5 ,P c6 These represent the six directions: left, front, right, back, top, and bottom, respectively, which effectively improves the accuracy of modeling the relationship between ground features.
[0034] In step S103, depth estimation is performed on the first ground panoramic image to obtain the depth estimation result of the first ground panoramic image, and the first depth map and the initial ground-to-air attention mask of the first ground panoramic image are obtained based on the depth estimation result.
[0035] It is understood that the embodiments of this application can perform depth estimation on the first ground panoramic image in the following steps to obtain the depth estimation result of the first ground panoramic image, and obtain the first depth map and the initial ground-to-air attention mask of the first ground panoramic image based on the depth estimation result, thereby effectively improving the quality of the remote sensing image.
[0036] In one embodiment of this application, depth estimation is performed on a first ground panoramic image to obtain a depth estimation result of the first ground panoramic image. Based on the depth estimation result, a first depth map and an initial ground-to-air attention mask of the first ground panoramic image are obtained. This includes: using a target image generation network to obtain a first depth map and an initial ground-to-air attention mask prediction of the first ground panoramic image. The target image generation network has 3 input channels and 2 output channels. The first output channel of the target image generation network is the first depth map, and the second output channel is the initial ground-to-air attention mask prediction.
[0037] In actual implementation, the embodiments of this application can utilize an image generation network to process the first ground panoramic image P. e The first depth map and initial ground-to-air attention mask prediction are performed. The target image generation network input has 3 channels and a size of 3×H×W, where H is the number of pixels along the vertical axis of the first panoramic image and W is the number of pixels along the horizontal axis. The output has 2 channels and a size of 2×H×W. The first channel output by the target image generation network is the first depth map D. e The size is 1×H×W, and the second output channel is the initial ground-to-space attention mask prediction M. e With a size of 1×H×W, the quality of the generated remote sensing images is effectively improved.
[0038] In step S104, feature extraction is performed on the second ground panoramic image to obtain a feature map of the target direction. The feature maps are aggregated to obtain a feature vector. The feature vector is then used to construct a loop hypermap of longitude and latitude.
[0039] It is understood that the embodiments of this application can extract features from the second ground panoramic image in the following steps to obtain a feature map of the target direction, aggregate the feature maps to obtain a feature vector, and use the feature vector to construct a loop hypergraph of longitude and latitude, thereby effectively reducing the distortion rate of the remote sensing image and improving the quality of the remote sensing image.
[0040] In one embodiment of this application, feature extraction is performed on a second ground panoramic image to obtain a feature map in the target direction. The feature maps are then aggregated to obtain feature vectors. A loop hypergraph of longitude and latitude is constructed using the feature vectors. This includes: inputting the panoramic image in the target direction into a feature extraction network and using a ResNet structure to obtain a feature map in the target direction; inputting the feature map into a feature aggregation network for aggregation to obtain at least one feature vector; and using the at least one feature vector as a node to construct a loop hypergraph of longitude and latitude.
[0041] As one possible implementation, this application embodiment can perform feature extraction on a second ground panoramic image. The panoramic images in six directions of the second ground panoramic image can be input into a feature extraction network. First, a ResNet structure is used to obtain feature maps in the six directions. The images in the six directions after feature extraction are input into a feature aggregation network to aggregate the feature maps and obtain feature vectors. The six feature vectors are V1, V2, V3, V4, V5, and V6, and the size of each feature vector is 1×C, that is, the number of elements on the vertical axis of each feature vector is 1, and the number of elements on the horizontal axis is C.
[0042] In addition, in this embodiment, feature vectors in six directions of the second ground panoramic image can be used as nodes to construct two loop hypergraphs for longitude and latitude. Each loop hypergraph has 4 nodes, and the latitude loop hypergraph has H nodes. lat The nodes are V2, V4, V5, and V6, which correspond to the image feature vectors in the six directions of the second ground panoramic image: front, right, top, and bottom, respectively. The longitude loop hypergraph H... lon The nodes are V1, V2, V3, and V4, which correspond to the image feature vectors of the left, front, right, and rear directions of the second ground panoramic image, respectively. The hyperedges of two loop hypergraphs are constructed using the K-Hop nearest neighbor method. The specific construction method is as follows:
[0043] ε D ={N hopk (V i )|V i ∈V},
[0044] Among them, V i Let V represent the feature vector of a surface image, and let N represent all nodes in the hypergraph. hopk The k nearest neighbor is calculated, which represents the set of the k other nodes that are closest to (Vi) in Euclidean distance.
[0045] In step S105, ground-to-air attention coefficients are constructed based on the loop hypergraph, and the ground-to-air attention coefficients and the initial ground-to-air attention mask are weighted to obtain the ground-to-air attention mask after weighting the ground-to-air attention coefficients.
[0046] It is understood that, in the embodiments of this application, ground-to-air attention coefficients can be constructed based on the loop hypergraph in the following steps, and the ground-to-air attention coefficients and the initial ground-to-air attention mask can be weighted to obtain a ground-to-air attention mask with weighted ground-to-air attention coefficients, thereby effectively improving the quality of remote sensing images.
[0047] In one embodiment of this application, ground-to-air attention coefficients are constructed based on the loop hypergraph, and a weighted ground-to-air attention mask is obtained by weighting the ground-to-air attention coefficients and the initial ground-to-air attention mask. This includes: performing a preset convolution on the features of the loop hypergraph nodes to obtain new features of the target nodes in the loop hypergraph; combining the new features into two sets of feature matrices; concatenating the target features in the two sets of feature matrices to obtain two new feature vectors; combining the two new feature vectors into a binary array; and performing matrix multiplication on the binary array to obtain the ground-to-air attention coefficients; and weighting the predicted initial ground-to-air attention mask and the ground-to-air attention coefficients to obtain the weighted ground-to-air attention mask.
[0048] For example, in the embodiments of this application, the two loop hypergraphs constructed in the above steps can be subjected to HCNNConv+ convolution on the hypergraph node features under the guidance of all hyperedges to obtain new features of all 4 nodes of each loop hypergraph, and combined into two sets of feature matrices.
[0049] Next, the two sets of feature matrices obtained after hypergraph convolution are 4×C. lat and 4×C lon , where C lat C is the latitudinal characteristic dimension. lon To obtain the longitude feature size, the features of the two feature matrices are concatenated to obtain a size of 1×4C. lat and 1×4C lon Two eigenvectors, where C lat =W / 4, C lon =H / 4, combine the two eigenvectors into a binary array, and perform matrix multiplication to obtain a result of size 1×4C. lon ×4C lat The ground-to-air attention coefficient A is used to predict the initial ground-to-air attention mask M obtained in the above steps. e Multiplying the ground-to-air attention coefficient A by the ground-to-air attention coefficient weighted ground-to-air attention mask M yields the ground-to-air attention mask M. a This effectively improves the accuracy of remote sensing images.
[0050] In step S106, the first depth map, the first ground panoramic image, and the ground-to-air attention mask are geometrically transformed to obtain a target view projection image from a remote sensing perspective.
[0051] It is understood that the embodiments of this application can perform geometric transformation on the first depth map, the first ground panoramic image, and the ground-to-air attention mask in the following steps to obtain the target view projection image under the remote sensing perspective. By performing geometric transformation on the original remote sensing image, the quality of the remote sensing image is effectively improved to meet the application needs of actual scenarios.
[0052] In one embodiment of this application, a geometric transformation is performed on a first depth map, a first ground panoramic image, and a ground-to-air attention mask to obtain a target view projection image from a remote sensing perspective. This includes: weighting the first depth map and the ground-to-air attention mask to obtain a weighted second depth map; using the second depth map, converting the homogeneous panoramic image coordinates of the first ground panoramic image into three-dimensional coordinates in the camera coordinate system to obtain converted non-homogeneous panoramic image coordinates; based on the non-homogeneous panoramic image coordinates, converting the RGB pixel values of the first ground panoramic image into homogeneous remote sensing image coordinates from a remote sensing perspective; based on the homogeneous remote sensing image coordinates, transforming each pixel in the first ground panoramic image to obtain the RGB value of each pixel in the final remote sensing image; and obtaining the target view projection image from a remote sensing perspective based on the RGB values.
[0053] In actual implementation, the embodiments of this application can use the first depth map D obtained in the above steps. e and ground-to-air attention mask M a Perform multiplication to obtain the weighted second depth map D. a =D e ×M a Using the second depth map D a The homogeneous coordinates (u) of the first ground panoramic image p ,v p 1) Convert to three-dimensional coordinates in the ground panoramic camera coordinate system The conversion method is as follows:
[0054]
[0055] Among them, K p r is the intrinsic parameter of the ground panoramic camera. p This represents the depth value of the ground-to-air attention mask.
[0056] Next, the converted non-homogeneous panoramic image coordinates can be obtained. Based on the non-homogeneous panoramic image coordinates, the RGB pixel values of the first ground panoramic image are converted into homogeneous remote sensing image coordinates from the remote sensing perspective. The specific conversion relationship is as follows:
[0057]
[0058] Where F is the transformation function for three-dimensional points in rectangular coordinates and spherical coordinates, and K...s H represents the camera intrinsic parameters of the remote sensing camera, and H represents the orbital height of the joystick camera.
[0059] Therefore, based on homogeneous remote sensing image coordinates, each pixel in the first ground panoramic image can be transformed to obtain a final size H. s ×W s The RGB values of each pixel in the remote sensing image are used to obtain the target view projection image from the remote sensing perspective.
[0060] In step S107, the target view projection image is input into the target remote sensing image generation module to repair the texture and details of the target view projection image and generate the final remote sensing image.
[0061] It is understood that the embodiments of this application can be based on U-Net structure for image generation. The target view projection image is input into the target remote sensing image generation module to repair the texture and details of the target view projection image and generate the final remote sensing image. The image generation network has 3 input channels and 3 output channels. By performing geometric transformation on the original remote sensing image, a remote sensing image with higher quality and richer details is generated.
[0062] For example, such as Figure 3 As shown, the embodiments of this application can be composed of three main modules: ground-to-air projection, aerial embedded attention, and satellite image generation. It can accept different types of ground panoramic images as input, including cube projection P. c That is, the second ground panoramic image, equidistant projection P e The input is a second ground panoramic image and its corresponding edge map. The input can be used to generate a synthetic satellite image of a given location through three modules. The ground-to-air projection module uses an encoder-decoder structure, starting from the second ground panoramic image P. e Its edge map predicts the first depth map D e Utilizing geometric and distance information from RGB and depth panoramic images from a ground perspective, the ground-to-air projection module reconstructs the geometric distribution of the input location from a satellite perspective using a geometry-based projection method. The satellite image generation module uses a generator and discriminator network to synthesize missing textures and generates high-quality remote sensing images based on the reconstructed geometric distribution.
[0063] Furthermore, due to the differences in visibility and occlusion relationships between the two viewpoints, each pixel captured by the terrestrial panoramic camera has different visibility in the satellite camera. Therefore, an airborne embedded attention module can be introduced, first using an encoder-decoder structure based on the second terrestrial panoramic image P. e Its edge map prediction is a simple mask, i.e., an initial empty attention mask prediction M. eThen, features are extracted from the images in each direction of the second ground panoramic image to construct two cyclic hypergraphs for latitude and longitude. Cyclic attention, i.e. ground-to-air attention coefficients, is obtained in these two directions through hypergraph learning. By combining simple masking and cyclic attention, an attention mask, i.e. ground-to-air attention mask, can be obtained to guide the cross-view geometric transformation during ground-to-air projection.
[0064] In summary, the embodiments of this application employ a discriminator to establish a conditional generative adversarial network structure, utilize a cyclic hypergraph constructed from cubic image features, introduce aerial attention, effectively model the visibility and occlusion relationships between two viewpoints, and focus on ground-to-aerial view synthesis, enabling the generation of satellite images across viewpoints and scales from ground panoramic images.
[0065] For example, such as Figure 4 The image shows the geometric correspondence between ground panoramic images and satellite images when capturing the same target. Figure 4 On the left, the alignment of the ground image and the satellite image is shown. Figure 4 The right side shows the pixel distribution on the image plane. Assuming the X-axis of the world coordinate system is parallel to the v-axis of the satellite image coordinate system, the Y-axis is parallel to the U-axis of the satellite image coordinate system, and the z-axis is perpendicular to the ground plane, the coordinates of the target object in the world coordinate system can be represented as:
[0066] P W =(x w ,y w ,z w )
[0067] Additionally, assume the extrinsic parameters of the ground panoramic camera are R. P and T P In the coordinate system of a ground-based panoramic camera, the coordinates of an object are represented as (x... g ,y g ,z g ), then (x g ,y g ,z g ) = R P (x w ,y w ,z w )+T P Assuming the ground panoramic camera is located at the origin of the world coordinate system, let R... P =I (identity matrix) and T P =0, therefore, (x g ,y g ,z g )=(x w ,y w ,z wThe first panoramic ground image was captured using a spherical coordinate system, and the coordinates of points in the first panoramic ground image are represented as follows: The corresponding relationship can be represented as:
[0068]
[0069]
[0070]
[0071] Next, the intrinsic parameters of the ground panoramic camera can be expressed as K. g The relationship between the coordinates of a point in the first ground panoramic image and its spherical coordinates is expressed as: Assume the external parameters of the remote sensing camera are R. s and T s The intrinsic parameters of the remote sensing camera are represented by K. s In the coordinate system of the remote sensing camera, the coordinates of the target point are (x... s ,y s ,z s Therefore, (x) s ,y s ,z s ) = R s (x w ,y w ,z w )+T s Furthermore, the image coordinates of the target point in the remote sensing camera are determined by... get.
[0072] Furthermore, in the geometric restoration process of recovering satellite images from terrestrial panoramic images, it is possible to determine any pixel (u) in the terrestrial panoramic image. g ,v g The coordinates in the ground panoramic camera coordinate system are: Next, based on the extrinsic parameters of the ground panoramic camera, the spherical space coordinates of the target object in the world coordinate system can be obtained. For any pixel (u) in the satellite image plane s ,v s The coordinates of the satellite camera in the coordinate system can be represented as: Therefore, by combining the extrinsic parameters of the remote sensing camera, the Cartesian coordinates of the target object in the world coordinate system can be obtained as follows: Assume the extrinsic parameters of the ground panoramic camera are defined as R. p =diag(1,1,1) and T p = [0,0,0], while the extrinsic parameters of the remote sensing camera are determined by R. s =diag(1,1,-1) and T s= [0,0,H] is obtained, where H represents the height difference between the ground panoramic camera and the remote sensing camera, since (x w ,y w ,z w ) and (θ w ,φ w ,r w The coordinates corresponding to the same target in both Cartesian and spherical coordinate systems can be represented by introducing the extrinsic parameters of the ground panoramic camera, i.e.:
[0073]
[0074] Where F is the transformation function for three-dimensional points in rectangular coordinates and spherical coordinates, and K... s Here, H represents the camera intrinsic parameters of the remote sensing camera, and H represents the orbital height of the joystick camera. s ,v s ) and (u p ,v p ) represent the pixels of the satellite image and the ground panoramic image, respectively, and F is the transformation function from the Cartesian coordinate system to the spherical coordinate system in the equation. -1 Let be the inverse transformation function from Cartesian coordinates to spherical coordinates in the equation.
[0075] In the ground-to-air projection module, the U-Net model is used to predict the depth map from the isometric projection. The predicted depth map is then used as the intrinsic parameter estimate for the ground panoramic camera, ensuring that for a given pixel (u) in the first panoramic image... p ,v p The spherical coordinates (θ) of the target are inferred. p ,φ p ,r p Therefore, the geometric distribution of the ground can be reconstructed against the sky background. The reconstructed geometric results are input into the satellite image generation module, which estimates the intrinsic parameters of the satellite camera, thereby synthesizing realistic satellite images and restoring missing textures to solve the artifact problem caused by resolution differences.
[0076] In some embodiments, the preliminary ground-to-air projection conversion method developed based on depth projection technology in this application does not accurately capture the complexity of satellite image formation. Specifically, firstly, due to the height occlusion relationship introduced by the aerial-to-ground shooting perspective, the satellite image only captures images with the same ground position (x). s ,y s The highest point among the pixels (with the largest z-axis) s The target point (value) makes it difficult to accurately represent the vertical distribution of a specific point using only a ground-view depth map. Secondly, directly using the distance r p Convert to z sChoosing from the z-buffer near the visibility boundary leads to the network's non-differentiability. Small changes in the estimated distance / height can cause abrupt changes in the output projection, resulting in gradient sparsity and zero gradients for invisible points. This is suboptimal for points that should be visible but are invisible due to slight inaccuracies in height estimation. Conversely, accurately predicting depth values is difficult because the discontinuous depths of extremely distant regions such as the sky and horizon differ significantly from other regions. Furthermore, unlike the air-to-ground projection transformation, not all pixels are visible from the ground viewpoint; therefore, including redundant sky data in the ground-to-air projection transformation is unnecessary.
[0077] Therefore, the factors in the above steps lead to different contributions of each pixel in the ground image to the satellite image. Since the contribution is affected by spatial location, an attention mask related to spatial location can be established. The attention mask guides the utilization of depth data in a pixel-by-pixel manner in the ground-to-air projection. The second ground panoramic image consists of images taken from different directions in a spherical space centered on the focal point. Due to the rich directional information in the second ground panoramic image, spatial distribution relationships can be analyzed and constructed by extracting diverse input features. Due to the effectiveness of the hypergraph structure in modeling multi-point relationships and capturing higher-order relationships, the features of the second ground panoramic image can be used to construct a hypergraph, where nodes represent image features.
[0078] For example, embodiments of this application can use ResNet and SAFA to extract features and aggregate the six-directional images obtained from the cube map. Each feature vector extracted from the cube map represents information in a specific direction. Feature vectors corresponding to the same position (direction) have a certain similarity. Therefore, by considering the similarity of features, the association between directions can be established, such as... Figure 2 The diagram illustrates how a panoramic image is divided into two loops based on longitude (vertical) and latitude (horizontal). The longitude loop includes four directions: up, front, down, and back, while the latitude loop includes four directions: left, front, right, and back. Two hypergraphs are constructed based on these loops. Each hypergraph's nodes represent feature vectors from its corresponding four directions. Each hypergraph is constructed using the K-Hop nearest neighbor method, specifically as follows:
[0079] ε D ={N hopk (V i )|V i ∈V},
[0080] Among them, V i Let V be the feature vector of a surface image, and let V be all the nodes in the hypergraph.
[0081] Therefore, the embodiments of this application can effectively improve the quality of remote sensing images and meet the application needs of real-world scenarios.
[0082] The remote sensing image generation method based on ground-to-air perspective geometric transformation proposed in this application can obtain a first ground panoramic image by projecting the original cylindrical image, obtain a second ground panoramic image by performing cube projection on the first ground panoramic image, estimate the depth of the first ground panoramic image to obtain a first depth map and an initial ground-to-air attention mask, extract features from the second ground panoramic image to obtain a feature map of the target direction and aggregate it to obtain a feature vector, thereby constructing a loop hypergraph of longitude and latitude, and constructing ground-to-air attention coefficients. The ground-to-air attention coefficients and the initial ground-to-air attention mask are weighted to obtain a weighted ground-to-air attention mask. The first depth map, the first ground panoramic image and the ground-to-air attention mask are geometrically transformed to obtain a target perspective projection image under the remote sensing perspective, and the texture and details of the target perspective projection image are repaired to generate the final remote sensing image, thereby effectively improving the quality of remote sensing images and meeting the application needs of practical scenarios. This solves the problem in related technologies where remote sensing images generated by interpolation algorithms cannot accurately model varied terrains, resulting in distorted remote sensing images and poor image quality that fails to meet the application requirements of real-world scenarios.
[0083] Next, referring to the accompanying drawings, a remote sensing image generation apparatus based on ground-to-air perspective geometric transformation proposed in accordance with the embodiments of this application is described.
[0084] Figure 5 This is a block diagram of a remote sensing image generation device based on ground-to-air perspective geometric transformation according to an embodiment of this application.
[0085] like Figure 5 As shown, the remote sensing image generation device 10 based on ground-to-air perspective geometric transformation includes: a first acquisition module 100, a second acquisition module 200, a third acquisition module 300, a first determination module 400, a construction module 500, a second determination module 600, and a generation module 700.
[0086] Specifically, the first acquisition module 100 is used to acquire the original cylindrical projection and obtain a first ground panoramic image based on the original cylindrical projection.
[0087] The second acquisition module 200 is used to perform cube projection on the first ground panoramic image and obtain a second ground panoramic image based on the cube projection.
[0088] The third acquisition module 300 is used to perform depth estimation on the first ground panoramic image, obtain the depth estimation result of the first ground panoramic image, and obtain the first depth map and the initial ground-to-air attention mask based on the depth estimation result.
[0089] The first determining module 400 is used to extract features from the second ground panoramic image to obtain a feature map of the target direction, aggregate the feature maps to obtain a feature vector, and use the feature vector to construct a loop hypermap of longitude and latitude.
[0090] Module 500 is used to construct ground-to-air attention coefficients based on the loop hypergraph, and to perform weighted processing on the ground-to-air attention coefficients and the initial ground-to-air attention mask to obtain the weighted ground-to-air attention mask.
[0091] The second determining module 600 is used to perform geometric transformation on the first depth map, the first ground panoramic image, and the ground-to-air attention mask to obtain a target perspective projection image from a remote sensing perspective.
[0092] The generation module 700 is used to input the target view projection image into the target remote sensing image generation module to repair the texture and details of the target view projection image and generate the final remote sensing image.
[0093] Optionally, in one embodiment of this application, the third acquisition module 300 includes: a first acquisition unit.
[0094] The first acquisition unit is used to obtain a first depth map and an initial ground-to-air attention mask prediction of a first ground panoramic image using a target image generation network. The target image generation network has 3 input channels and 2 output channels. The first output channel of the target image generation network is the first depth map, and the second output channel is the initial ground-to-air attention mask prediction.
[0095] Optionally, in one embodiment of this application, the first determining module 400 includes: a first determining unit, a second determining unit, and a constructing unit.
[0096] The first determining unit is used to input the panoramic image in the target direction into the feature extraction network and use the ResNet structure to obtain the feature map in the target direction.
[0097] The second determining unit is used to input the feature map into the feature aggregation network for aggregation to obtain at least one feature vector.
[0098] A building unit is used to construct a loop hypergraph of longitude and latitude by using at least one feature vector as a node.
[0099] Optionally, in one embodiment of this application, the construction module 500 includes: a second acquisition unit, a third acquisition unit, and a processing unit.
[0100] The second acquisition unit is used to perform a pre-defined convolution on the features of the nodes of the loop hypergraph to obtain new features of the target nodes of the loop hypergraph, and combine the new features into two sets of feature matrices.
[0101] The third acquisition unit is used to concatenate the target features in the two sets of feature matrices to obtain two new feature vectors, combine the two new feature vectors into a binary array, and perform matrix multiplication on the binary array to obtain the ground-to-air attention coefficient.
[0102] The processing unit is used to weight the initial ground-to-air attention mask and ground-to-air attention coefficients obtained from the prediction to obtain a ground-to-air attention mask after weighting the ground-to-air attention coefficients.
[0103] Optionally, in one embodiment of this application, the second determining module 600 includes: a third determining unit, a fourth determining unit, a conversion unit, and a fifth determining unit.
[0104] The third determining unit is used to perform weighted processing on the first depth map and the ground-to-air attention mask to obtain a weighted second depth map.
[0105] The fourth determining unit is used to convert the homogeneous panoramic image coordinates of the first ground panoramic image into three-dimensional coordinates in the camera coordinate system using the second depth map, thereby obtaining the converted non-homogeneous panoramic image coordinates.
[0106] The conversion unit is used to convert the RGB pixel values of the first ground panoramic image into homogeneous remote sensing image coordinates from the remote sensing perspective, based on the non-homogeneous panoramic image coordinates.
[0107] The fifth determining unit is used to transform each pixel in the first ground panoramic image based on the homogeneous remote sensing image coordinates to obtain the RGB value of each pixel in the final remote sensing image, and obtain the target view projection image under the remote sensing view based on the RGB value.
[0108] It should be noted that the foregoing explanation of the remote sensing image generation method based on ground-to-air perspective geometric transformation also applies to the remote sensing image generation device based on ground-to-air perspective geometric transformation of this embodiment, and will not be repeated here.
[0109] The remote sensing image generation device based on ground-to-air perspective geometric transformation proposed in this application can obtain a first ground panoramic image based on the original cylindrical projection, obtain a second ground panoramic image by performing cube projection on the first ground panoramic image, perform depth estimation on the first ground panoramic image to obtain a first depth map and an initial ground-to-air attention mask, then extract features from the second ground panoramic image to obtain a feature map of the target direction and aggregate it to obtain a feature vector, thereby constructing a loop hypergraph of longitude and latitude, and then constructing ground-to-air attention coefficients. The ground-to-air attention coefficients and the initial ground-to-air attention mask are weighted to obtain a ground-to-air attention mask with weighted ground-to-air attention coefficients. The first depth map, the first ground panoramic image and the ground-to-air attention mask are geometrically transformed to obtain a target perspective projection image under the remote sensing perspective, and the texture and details of the target perspective projection image are repaired to generate the final remote sensing image, thereby effectively improving the quality of remote sensing images and meeting the application needs of practical scenarios. This solves the problem in related technologies where remote sensing images generated by interpolation algorithms cannot accurately model varied terrains, resulting in distorted remote sensing images and poor image quality that fails to meet the application requirements of real-world scenarios.
[0110] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device may include:
[0111] The memory 601, the processor 602, and the computer program stored on the memory 601 and capable of running on the processor 602.
[0112] When the processor 602 executes the program, it implements the remote sensing image generation method based on ground-to-air perspective geometric transformation provided in the above embodiments.
[0113] Furthermore, electronic devices also include:
[0114] Communication interface 603 is used for communication between memory 601 and processor 602.
[0115] The memory 601 is used to store computer programs that can run on the processor 602.
[0116] The memory 601 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.
[0117] If the memory 601, processor 602, and communication interface 603 are implemented independently, then the communication interface 603, memory 601, and processor 602 can be interconnected via a bus to complete communication between them. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 6 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0118] Optionally, in a specific implementation, if the memory 601, processor 602, and communication interface 603 are integrated on a single chip, then the memory 601, processor 602, and communication interface 603 can communicate with each other through an internal interface.
[0119] The processor 602 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application.
[0120] This embodiment also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described remote sensing image generation method based on ground-to-air perspective geometric transformation.
[0121] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0122] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0123] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or N executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.
[0124] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.
[0125] It should be understood that the various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0126] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.
[0127] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0128] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.
Claims
1. A remote sensing image generation method based on ground-to-air perspective geometric transformation, characterized in that, Includes the following steps: Acquire the original cylindrical projection and obtain a first panoramic ground image based on the original cylindrical projection; The first ground panoramic image is cube-projected, and a second ground panoramic image is obtained based on the cube projection. Depth estimation is performed on the first ground panoramic image to obtain the depth estimation result of the first ground panoramic image. Based on the depth estimation result, a first depth map of the first ground panoramic image and an initial ground-to-air attention mask are obtained. Feature extraction is performed on the second ground panoramic image to obtain a feature map of the target direction. The feature maps are aggregated to obtain a feature vector. The feature vector is then used to construct a loop hypergraph of longitude and latitude. Based on the loop hypergraph, a ground-to-air attention coefficient is constructed. The ground-to-air attention coefficient and the initial ground-to-air attention mask are weighted to obtain a ground-to-air attention mask with weighted ground-to-air attention coefficients. Geometric transformation is performed on the first depth map, the first ground panoramic image, and the ground-to-air attention mask to obtain a target view projection image from a remote sensing perspective. as well as The target view projection image is input into the target remote sensing image generation module to repair the texture and details of the target view projection image and generate the final remote sensing image.
2. The method according to claim 1, characterized in that, The step of performing depth estimation on the first ground panoramic image to obtain a depth estimation result of the first ground panoramic image, and obtaining a first depth map and an initial ground-to-air attention mask based on the depth estimation result, includes: Using a target image generation network, a first depth map and an initial ground-to-air attention mask prediction of the first ground panoramic image are obtained. The target image generation network has 3 input channels and 2 output channels. The first output channel of the target image generation network is the first depth map, and the second output channel is the initial ground-to-air attention mask prediction.
3. The method according to claim 1, characterized in that, The step of extracting features from the second ground panoramic image to obtain a feature map of the target direction, aggregating the feature maps to obtain a feature vector, and constructing a loop hypergraph of longitude and latitude using the feature vector includes: The panoramic image in the target direction is input into the feature extraction network, and the feature map in the target direction is obtained using the ResNet residual network structure; The feature map is input into a feature aggregation network for aggregation to obtain at least one feature vector; Using the at least one feature vector as a node, construct the loop hypergraph of the longitude and the latitude.
4. The method according to claim 1, characterized in that, The step of constructing ground-to-air attention coefficients based on the loop hypergraph, and then performing weighted processing on the ground-to-air attention coefficients and the initial ground-to-air attention mask to obtain a weighted ground-to-air attention mask, includes: The loop hypergraph is subjected to a pre-defined convolution on the hypergraph node features to obtain new features of the target node of the loop hypergraph, and the new features are combined into two sets of feature matrices. The target features in the two sets of feature matrices are concatenated to obtain two new feature vectors. The two new feature vectors are combined into a binary array, and the binary array is multiplied by matrix to obtain the ground-to-air attention coefficient. The initial ground-to-air attention mask and the ground-to-air attention coefficients obtained from the prediction are weighted to obtain the ground-to-air attention mask after weighting the ground-to-air attention coefficients.
5. The method according to claim 1, characterized in that, The step of geometrically transforming the first depth map, the first ground panoramic image, and the ground-to-air attention mask to obtain a target view projection image from a remote sensing perspective includes: The first depth map and the ground-to-air attention mask are weighted to obtain a weighted second depth map. Using the second depth map, the homogeneous panoramic image coordinates of the first ground panoramic image are converted into three-dimensional coordinates in the camera coordinate system to obtain the converted non-homogeneous panoramic image coordinates. Based on the non-homogeneous panoramic image coordinates, the RGB pixel values of the first ground panoramic image are converted into homogeneous remote sensing image coordinates from the remote sensing perspective. Based on the homogeneous remote sensing image coordinates, each pixel in the first ground panoramic image is transformed to obtain the RGB value of each pixel in the final remote sensing image, and the target view projection image under the remote sensing view is obtained based on the RGB values.
6. A remote sensing image generation device based on ground-to-air perspective geometric transformation, characterized in that, include: The first acquisition module is used to acquire the original cylindrical projection and obtain a first ground panoramic image based on the original cylindrical projection. The second acquisition module is used to perform cube projection on the first ground panoramic image and obtain a second ground panoramic image based on the cube projection. The third acquisition module is used to perform depth estimation on the first ground panoramic image to obtain the depth estimation result of the first ground panoramic image, and obtain a first depth map and an initial ground-to-air attention mask based on the depth estimation result. The first determining module is used to extract features from the second ground panoramic image to obtain a feature map of the target direction, aggregate the feature maps to obtain a feature vector, and use the feature vector to construct a loop hypermap of longitude and latitude. The construction module is used to construct ground-to-air attention coefficients based on the loop hypergraph, and to perform weighted processing on the ground-to-air attention coefficients and the initial ground-to-air attention mask to obtain the weighted ground-to-air attention mask. The second determining module is used to perform geometric transformation on the first depth map, the first ground panoramic image and the ground-to-air attention mask to obtain a target perspective projection image under the remote sensing perspective. as well as The generation module is used to input the target view projection image into the target remote sensing image generation module to repair the texture and details of the target view projection image and generate the final remote sensing image.
7. The apparatus according to claim 6, characterized in that, The third acquisition module includes: The first acquisition unit is used to obtain a first depth map and an initial ground-to-air attention mask prediction of the first ground panoramic image using a target image generation network. The target image generation network has 3 input channels and 2 output channels. The first output channel of the target image generation network is the first depth map, and the second output channel is the initial ground-to-air attention mask prediction.
8. The apparatus according to claim 6, characterized in that, The first determining module includes: The first determining unit is used to input the panoramic image in the target direction into the feature extraction network and use the ResNet residual network structure to obtain the feature map in the target direction; The second determining unit is used to input the feature map into a feature aggregation network for aggregation to obtain at least one feature vector. A construction unit is used to construct the loop hypergraph of the longitude and the latitude by using the at least one feature vector as nodes.
9. An electronic device, characterized in that, include: The memory, the processor, and the computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the remote sensing image generation method based on ground-to-air perspective geometric transformation as described in any one of claims 1-5.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement the remote sensing image generation method based on ground-to-air perspective geometric transformation as described in any one of claims 1-5.
Citation Information
Patent Citations
Scene image classification method and system based on double hypergraph neural network
CN116206158A
Panoramic image generation method and device, equipment and storage medium
CN116245734A