Remote sensing scene new visual angle synthesis method based on depth supervision and texture discriminator
By combining neural radiation fields with point cloud depth supervision and generative adversarial network texture discriminators, the problem of poor quality of new perspective synthesis under sparse perspectives of remote sensing images is solved, especially the difficulty of texture synthesis in occluded areas, and high-quality new perspective image synthesis is achieved.
Patent Information
- Application Number
- CN202511654679.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-12
- Publication Date
- 2026-03-27
AI Technical Summary
The sparse input viewpoints in remote sensing images result in poor quality of new viewpoint synthesis, especially in the difficulty of texture synthesis in occluded areas, and existing methods are difficult to effectively transfer to remote sensing scenes.
We employ a method based on depth supervision and texture discriminator, which implicitly models scene geometry through neural radiation fields, combines differentiable volume rendering technology, introduces point cloud depth supervision optimization and a generative adversarial network texture discriminator, and improves the quality of texture synthesis in occluded areas.
It achieves high-quality new perspective image synthesis under sparse perspective, avoids model overfitting, and improves the synthesis quality and realism of textures in occluded areas.
Smart Images

Figure CN121746757A_ABST
Abstract
Description
Technical Field
[0001] This invention is a method for synthesizing novel perspectives in remote sensing scenes based on depth supervision and texture discriminators. This method relies on implicit scene representation using neural radiation fields and differentiable rendering technology, and is suitable for synthesizing novel perspectives in remote sensing images under sparse input conditions. It focuses on solving the problem of poor quality of novel perspective synthesis caused by sparse input perspectives in remote sensing images, and technically falls under the category of remote sensing image 3D reconstruction and perspective synthesis. Background Technology
[0002] Novel perspective synthesis is an important research direction in the field of 3D vision. Its goal is to recover the 3D structure of a given scene from a set of source images and corresponding camera poses, and then render images from any new perspective. Novel perspective synthesis is closely related to 3D reconstruction, and the two often complement each other. However, 3D reconstruction focuses on recovering the geometric model of the scene, while novel perspective synthesis focuses more on generating highly realistic rendered images. Using a limited number of input images, novel perspective synthesis can synthesize scene images from different perspectives, simulating the human observation process of 3D objects, thus providing rich 3D visual information and enhancing the observer's perception and understanding of the scene. This technology has broad application prospects in many fields such as virtual reality, augmented reality, autonomous driving, and film and animation production.
[0003] Novel perspective synthesis of remote sensing scenes is another important application area of this technology. In recent years, my country's remote sensing field has developed rapidly, and research content has gradually shifted from two-dimensional to three-dimensional. Utilizing high-resolution visible light images, combined with advanced technologies such as 3D reconstruction and depth estimation, stereo imaging of remote sensing images can be achieved, obtaining precise three-dimensional information of targets, including terrain, building structures, and target textures. This has significant implications for both military and civilian applications. However, since remote sensing images are usually taken by drones or satellites, their acquisition angles are limited by flight trajectories and orbital planning, making it difficult to obtain sufficient multi-view data. This results in a sparse perspective problem in the input images used for training, thus affecting the quality of novel perspective synthesis. Currently, methods to solve the problem of sparse perspective synthesis are mainly divided into two categories: pre-training-based methods and regularization-based methods.
[0004] Pre-trained methods compensate for the lack of supervision information in sparse viewpoints by learning prior knowledge across scenes. PixelNeRF (Pixel Neural Radiation Field) adjusts the neural radiation field architecture by extracting image input features through convolutional networks, and synthesizes new viewpoints from a sparse view set in a feedforward manner. MVSNeRF (Multi-View Stereo Fast General Radiation Field) uses cost volume in multi-view stereo geometry for geometric perception scene inference, and then synthesizes new viewpoints through volume rendering. Image-Based Rendering Network (IBRNet) integrates an image-based rendering framework, learning general inter-view difference relationships through pre-training, thus improving generalization and sparse viewpoint synthesis performance. The advantage of pre-trained methods is that they significantly improve the generalization performance of the model; after large-scale pre-training, only fine-tuning on a single scene is needed to achieve high rendering quality. The disadvantages are that they require a large amount of 3D scene data and consume significant computational resources and time.
[0005] Regularization-based methods constrain the model's optimization process by introducing scene-related geometric inductive biases, ensuring the quality of synthesized images from new perspectives. Regularized Neural Radiation Field (RegNeRF) introduces geometric consistency and color distribution priors, improving the quality of synthesized views at minimal cost. Semantic Consistent Neural Radiation Field (DietNeRF) leverages the property that the semantic information of objects does not change with the viewing perspective. It extracts semantic features of images through a pre-trained model, calculates the semantic consistency loss between images from different perspectives, and improves the quality of view synthesis using two-dimensional prior knowledge. The advantage of regularization-based methods lies in their flexibility; regularization constraints can be seamlessly integrated into other models, improving rendering quality under sparse perspectives without increasing computational burden. However, the disadvantage is that such regularization designs are often scene-specific and lack generality, requiring manual design for different scenarios.
[0006] While previous methods achieved novel perspective synthesis tasks under sparse perspectives, they are difficult to transfer to remote sensing scenes, mainly due to the following two challenges: First, remote sensing scenes have diverse features and significant differences in foreground and background depth, making it easy for models to overfit to the input perspective under sparse perspectives. Second, sparse perspectives result in numerous occluded areas or areas not covered by the input perspective, making texture synthesis difficult in these locations.
[0007] To address the aforementioned challenges, this invention proposes a novel perspective synthesis method for remote sensing scenes based on depth supervision and a texture discriminator. It aims to implicitly model scene geometry using Neural Radiance Field (NeRF) and combine this with differentiable volume rendering technology to achieve high-fidelity novel perspective synthesis from any viewpoint. To address the overfitting problem in sparse viewpoints, this invention introduces a point cloud-based depth supervision optimization method. This method obtains coarse 3D point clouds through motion structure estimation and stereo matching, and with sparse depth supervision, guides the model to recover the correct scene geometry. To address the difficulty of texture synthesis in occluded regions, this invention introduces a texture discriminator based on a Generative Adversarial Network (GAN). This invention trains the discriminator adversarially, constraining the texture of the rendered image. Image patches are collected from the input reference viewpoint as positive samples and from the rendered image as negative samples. The discriminator is trained to recognize different textures, and the discriminator loss is used as the basis for model optimization. Simultaneously, differentiable amplification of real and fake samples improves the discriminator's performance, thereby enhancing the quality of synthesized textures in occluded regions. Summary of the Invention
[0008] 1. Objective: The objective of this invention is to provide a method for synthesizing novel perspectives in remote sensing scenes based on depth supervision and a texture discriminator. This method addresses the difficulties in synthesizing novel perspectives in remote sensing scenes with sparse perspectives and in synthesizing textures in occluded areas. It implicitly models the scene geometry based on neural radiation fields, synthesizes arbitrary novel perspective images through differentiable rendering, and avoids overfitting during optimization through depth supervision based on point cloud depth. Furthermore, this invention addresses occluded areas in 3D scenes by introducing a discriminator based on generative adversarial networks to improve the quality of synthesized textures, ultimately achieving high-quality novel perspective rendering results. This invention improves and promotes the development of novel perspective synthesis technology for remote sensing scenes, providing strong support for fields such as urban planning and environmental monitoring.
[0009] 2. Technical Solution: The present invention is achieved through the following technical solution: The present invention is a new perspective synthesis method for remote sensing scenes based on depth supervision and texture discriminator, which consists of three steps: constructing an implicit reconstruction network based on neural radiation field, constructing a depth supervision optimization framework based on sparse point cloud, and constructing a texture discriminator based on generative adversarial network.
[0010] Step 1: Constructing an implicit reconstruction network based on neural radiation fields This invention proposes an implicit reconstruction network based on neural radiation fields, which enhances the model's ability to perceive scene structure by fusing features from adjacent reference images. Based on this, it uses differentiable volume rendering to synthesize corresponding new perspective images.
[0011] The network input consists of the 3D coordinates of the sampling points and the direction vector of the ray, and cosine position encoding is performed:
[0012] in, Representative to Perform cosine position encoding. These are the spatial coordinates of the sampling point. w It is the scaling factor, a hyperparameter that is adjusted according to the experimental scenario. To maximize the coding frequency, positional coding is not used for the view direction vector.
[0013] The basic network unit employs a multi-layer perceptron (MLP) with residual connections, referred to as ResMLP. Specifically, each ResMLP block consists of a bilinear layer and residual connections, using the ReLU activation function. The network is divided into two parts: the first part is a multi-branch structure with shared weights, processing the input coordinate vector and features from different reference images. This invention encodes the feature vectors from the spatial reference images and inputs them to each layer in residual form. The second part performs average pooling on the features output from different branches, and then passes them through several ResMLPs to predict the volume density and color attributes of the corresponding spatial sampling points. The model's input and output can be represented as:
[0014] in, Represents a neural radiation field network. It is the volume density of the sampling points. It's the color. It's the perspective direction. It comes from an image. Reference features.
[0015] After obtaining the volume density and color information, this invention utilizes differentiable volume rendering to aggregate the information from all sampling points along the ray to obtain the final color information on the rendered image:
[0016]
[0017] in, Volume density and color at each sampling point, The distance between adjacent sampling points. Cumulative transmittance The number of sampling points. This represents the final RGB color value of the corresponding pixel.
[0018] Step 2: Construct a deep supervised optimization framework based on sparse point clouds This invention uses point clouds obtained through Structure from Motion (SfM) estimation as depth guidance to ensure that the model reconstructs correct geometry without overfitting at the training viewpoint. Since the point clouds obtained through SfM estimation are very sparse and difficult to effectively guide model optimization, this invention further performs multi-view stereo matching and point cloud fusion to generate denser 3D point clouds.
[0019] Given a 3D point cloud returned by SfM Projecting these onto a 2D view yields key points, and simultaneously calculating the depth of these key points. :
[0020] in, Let these be the homogeneous coordinates of the point cloud in the camera coordinate system. Let be the homogeneous coordinates of the point cloud in the world coordinate system, where These correspond to the coordinate positions on the X, Y, and Z axes, respectively, with the last dimension (1) representing the scaling factor. and Let be the rotation and translation matrix of the camera. The depth of the point cloud at the current viewpoint can be represented as:
[0021] in, That is, the depth corresponding to the projection point, denoted as , Here are the pixel coordinates of the projection point, where u represents the x-coordinate and v represents the y-coordinate. This is a camera intrinsic parameter. Additional depth information is rendered in each iteration.
[0022] in, It represents the position of each sampling point on the ray.
[0023] Because the estimated 3D point cloud has depth errors, this invention is based on multi-view... Figure 1Consistency is achieved by adaptively assigning weights to each point cloud, reducing the impact of point cloud depth error on the model. Depth error mainly stems from two sources: incorrect spatial location estimation of the point cloud and ambiguity arising during projection of points located at edges. For point clouds with incorrect spatial locations, the projected points exhibit significant color differences across different views. For point clouds at edge structures, even small positional shifts can correspond to substantial depth and color variations. Therefore, by ignoring color changes from different viewpoints, the color information of the point cloud and projected points can be used to measure depth error and apply it as weights for depth supervision, preventing errors from excessively influencing model optimization.
[0024] Step 3: Construct a texture discriminator based on generative adversarial networks Because invisible regions caused by occlusion often produce blurred texture shapes in the synthesized image, and unobserved regions due to viewpoint limitations appear as simple copies of edge pixels in the synthesized new viewpoint, this invention employs a discriminator based on generative adversarial networks to guide texture synthesis of the new viewpoint image, avoiding the aforementioned degradation problems. This invention trains the discriminator in a generative adversarial manner, with positive samples derived from image patches acquired from the input reference image and negative samples from rendered synthesized new viewpoint image patches. Here, the generator is a neural radiation field, and the discriminator uses a cascaded convolutional neural network, employing spectral normalization to satisfy the Lipschitz continuity condition. Due to the limited number of usable samples, sample augmentation is performed to avoid overfitting and affecting model performance. This invention applies differentiable augmentation to both positive and negative samples; the differentiability property allows gradients to be passed to the generator through the augmentation operation, optimizing without altering the target distribution.
[0025] Generative adversarial networks use the hinge loss function, and the loss functions for the generator and discriminator are expressed as follows:
[0026] in, This represents a generator network, also known as a neural radiation field. Represents the discriminator network. This represents a differentiable amplification operation on positive and negative samples. This invention addresses the generator loss... It is added to the model's total loss function, and Used for training the discriminator separately. The term "sampled image" refers to image patches acquired from different perspectives. This invention treats image patches acquired from the training perspective as positive samples and synthesizes image patches from new perspectives as negative samples. By optimizing the radiation field through adversarial loss, it improves the quality of texture synthesis in occluded regions.
[0027] 3. Advantages and effects: This invention implements a novel perspective synthesis method for remote sensing scenes based on deep supervision and a texture discriminator. Compared with existing novel perspective synthesis methods, this method can achieve high-quality novel perspective image synthesis even under conditions of sparse input perspectives, insufficient supervision information, and difficult model optimization in remote sensing scenes. Based on the classic neural radiation field framework, this invention implicitly models the scene geometry by aggregating features from adjacent reference images and uses differentiable volume rendering to synthesize novel perspective images. This invention designs a deep supervision optimization framework based on sparse point clouds, effectively alleviating the model overfitting problem during sparse perspective reconstruction and improving training efficiency. This invention proposes using generative adversarial methods to train the texture discriminator, utilizing random sample amplification to improve discriminator performance, and designing a joint optimization strategy for the discriminator and neural radiation field to guide the synthesis of textures and colors, especially in invisible areas, in novel perspective images, enhancing the accuracy and realism of the synthesized novel perspective images. Attached Figure Description
[0028] Figure 1 This is an implicit reconstruction framework based on neural radiation fields.
[0029] Figure 2 This is a schematic diagram of a deep supervised optimization framework based on sparse point clouds.
[0030] Figure 3 This is a schematic diagram of a texture discriminator based on a generative adversarial network.
[0031] Figure 4 A composite image providing a new perspective on remote sensing scenes. Detailed Implementation
[0032] To better understand the technical solution of the present invention, the following is in conjunction with the appendix. Figure 1-4 The embodiments of the present invention will be further described as follows: This invention is based on the PyTorch framework and implemented using the Python language. First, the implicit reconstruction network for neural radiation fields is built. Then, a deep supervised optimization framework based on sparse point clouds and a texture discriminator based on generative adversarial networks are constructed, and the corresponding hyperparameters are determined. Next, scene-by-scene training is performed, followed by the synthesis and rendering of new perspective images. The model is then validated on a test set to evaluate its performance.
[0033] The computer's hardware configuration includes an Intel(R) Xeon(R) Silver 4210R CPU @ 2.40GHz processor, 512GB of memory, and an NVIDIA GeForce GTX 4080 graphics card with 16.0GB of video memory.
[0034] The new perspective synthesis method for remote sensing scenes includes the following steps: Step 1: Constructing an implicit reconstruction network based on neural radiation fields Figure 1 The implicit reconstruction network based on neural radiation fields proposed in this invention mainly consists of three parts: a feature extraction and fusion module, a neural radiation field reconstruction module, and a differentiable volume rendering module. The feature extraction module uses a ResNet34 model with pre-trained weights on the ImageNet dataset. During neural radiation field optimization, the network parameters are frozen. Feature maps from the first pooling input and the outputs of the first three layers of ResNet34 are extracted to form a feature pyramid. Low-resolution feature maps are upsampled and merged to obtain a pixel-wise aligned feature map with 512 channels and a size half that of the reference image. The feature pyramid effectively fuses high-level semantic information and local detail features, providing richer prior information for subsequent scene representation and new perspective synthesis.
[0035] NeRF projects a virtual ray from the camera origin to each image pixel. The origin coordinates are denoted as o, and the ray direction vector is denoted as d. The sampling ray... It can be represented as:
[0036] in, It is a manually set sampling interval. This represents the specific sampling location. The specific method is to first determine the number of sampling points N, divide the sampling interval into N equally spaced intervals, and then perform uniform random sampling within each interval to ensure the uniformity of the distribution of sampling points in space and the randomness within the interval.
[0037]
[0038] in, This indicates the position of each sampling point on the ray. Represents uniform distribution sampling. The maximum value at the sampling location. This represents the minimum value at the sampling location.
[0039] After obtaining the spatial location information of the sampling points, this invention uses projection transformation to determine the projection points of the spatial points on the reference image, thereby obtaining the corresponding feature vectors. Given a spatial point... In its first The pixel coordinates on the reference image are:
[0040] in This is the corresponding camera intrinsic and extrinsic parameter matrix. This represents the depth of a point in space from the current viewpoint. Since the coordinates are not always integers, bilinear interpolation is used to obtain the corresponding feature vector.
[0041] Next, according to step one of the invention's description, the spatial position is cosine-encoded, with the maximum encoding frequency... , scale factor The encoded position coordinates, viewpoint direction, and reference features are fed into the neural radiation field for training.
[0042] The basic unit of the neural radiation field is a residual-connected MLP, called ResMLP, with a hidden layer width of 512 and a depth of 2, using the ReLU activation function. In the reference feature fusion stage, three ResMLPs are cascaded and share weights to fuse reference features in residual form. Then, average pooling is used to process the intermediate layer features from different reference views. The result of average pooling is then passed through two more ResMLP cascades to predict the final spatial point color and volume density.
[0043] This invention employs secondary sampling to improve training efficiency. By sampling with weight distribution along the ray direction, sampling points are more densely distributed on the object's surface and less densely distributed in blank areas, avoiding waste of computational resources. Specifically, the neural radiation field parameters are frozen during the first uniform sampling, and the volume density at each spatial point is calculated. and cumulative transmittance The method for calculating the sampling point weights is as follows:
[0044] in, The weight of each sampling point during volume rendering. Given the distance between adjacent sampling points, normalize all weights to obtain the probability density function:
[0045] Then, secondary sampling is performed based on the probability density function. The higher the probability interval, the more sampling points are included, thus improving the model's training efficiency. The first stage uses 64 sampling points, and the second stage uses 128.
[0046] Finally, the final color is obtained by applying the differentiable volume rendering formula described in step one of the invention's description. This invention calculates the pixels of the rendered image. With predicted pixels Photometric loss:
[0047] In implementation, all image pixels are shuffled, and 1024 rays are randomly selected for training in each round. The input image is 256x256, and a total of 30,000 training rounds are conducted. The Adam optimizer is used, and the learning rate is [missing information]. .
[0048] Step 2: Construct a deep supervised optimization framework based on sparse point clouds To avoid overfitting during sparse viewpoint training, this invention designs a depth-supervised optimization framework based on sparse point clouds. Overfitting is a common failure mode in few-shot novel viewpoint synthesis. Neural radiation fields implicitly learn the 3D structure of a scene through cross-evaluation of multi-view consistency among 2D images. However, when the number of input views is very small, neural radiation fields tend to overfit these input images, failing to accurately reconstruct the geometry and resulting in blurry synthesized images from novel views. Overfitting specifically manifests as excessively high volume density at the origin of the light rays, causing the rendered depth map to be too close to the camera plane. Therefore, this invention uses point cloud positions to anchor the depth map, guiding the model to learn the correct geometry.
[0049] This invention supports reconstructing scene geometry from at least three input images. For each input image and its corresponding camera pose, sparse reconstruction of the known pose is first performed using colmap. The invention employs a motion structure estimation process, which mainly includes feature extraction and matching on the input images to find the one-to-one correspondence between keypoints, and triangulation based on the known camera pose to reconstruct the spatial positions of the keypoints. Due to the limited number of training images, the resulting point cloud is extremely sparse, containing only a small number of images. Left and right angles alone cannot effectively guide the depth of the model. Therefore, based on sparse point clouds, this invention further performs stereo matching and point cloud fusion based on a binocular stereo matching algorithm, generating a dense 3D point cloud, increasing the number of point clouds to [number missing]. The above process is all performed on colmap and takes less than 5 seconds.
[0050] By reprojecting the reconstructed sparse point cloud onto the reference image, corresponding key points can be obtained. This invention utilizes the depth information of each key point to guide model optimization. First, a set of key points is established. In each round of iteration optimization, in addition to the conventional 1024 rays, this invention additionally collects 64 rays from these key points and calculates their predicted depths. The error between the predicted depth and the ground truth is calculated as a supervision signal to ensure that the model converges to the correct scene geometry.
[0051] For sparse point clouds, spatial location errors can lead to incorrect geometric relationships learned by the model, impacting model performance. This error can manifest as inconsistencies between multiple views. For point clouds with incorrect spatial locations, the projected points exhibit significant color differences across different views. Furthermore, point clouds at edge structures are also unreliable, as small spatial shifts can correspond to large changes in depth and color. Therefore, this invention utilizes these color differences to measure the spatial error of point clouds.
[0052] First, define the color error between different pixels as:
[0053] in, These represent different pixel colors. The error of the projection point of the i-th point cloud onto the k-th reference image is denoted as... ,So This can be expressed as the sum of the color errors between the standard and the point cloud and the projected points at different projected points:
[0054] in, The number of reference images, This represents the color of the projection point of the i-th point cloud onto the j-th reference image. The average color of the i-projected points in the point cloud. It is the color of the i-th point cloud. It is the error between the color of each projected point and the average projected color. It is the error between the color of the projected point and the color of the point cloud.
[0055] Then, adaptive weights are calculated for each keypoint:
[0056]
[0057] here, These are the weights corresponding to the projection points of the i-th point cloud onto the k-th reference image. As an intermediate variable, This is the truncation function, indicating that the weight values are between 0 and 1. Finally, the depth-weighted mean squared error of each keypoint is calculated as the depth consistency loss:
[0058] in, The number of keypoints in each round of sampling. These are the weights corresponding to the key points. It is a prediction of depth. It is the true depth value calculated by projection transformation. The calculation is described in step two of the invention's description, using the formula for determining the depth corresponding to the projection point to predict the depth. The calculation is described in the formula for additionally rendering depth information in step two of the invention description.
[0059] Because of spatial location errors in point clouds, their depth information can only be considered a guiding signal rather than the ground truth. Later in the optimization process, this error can damage the scene's geometry and reduce training efficiency. This invention introduces depth consistency loss only in the early stages of training, continuing until a rough scene geometry is reconstructed, after which it is combined with photometric consistency loss and other regularization strategies to further optimize the model. Specifically, this invention uses depth supervision only in the first 10,000 iterations, optimizing the neural radiation field together with photometric consistency loss. Depth supervision is discontinued after the neural radiation field reconstructs a rough scene geometry.
[0060] Step 3: Construct a texture discriminator based on generative adversarial networks To address the challenge of texture synthesis in occluded or invisible regions under sparse viewpoints, this invention proposes training a texture discriminator based on a generative adversarial network (GAN) to improve the quality of synthesized images. Since training performance is poor due to limited samples in sparse viewpoints, differentiable augmentation is employed to increase the sample count. If only positive samples are augmented, the generator will merely match the distribution of augmented samples; if only the discriminator's samples are augmented, the discriminator will be unable to identify unaugmented fake samples. Therefore, differentiable augmentation is applied to both positive and negative samples. The differentiability property allows gradients to be passed to the generator through the augmentation operation, optimizing without altering the target distribution. Specifically, in each training round, a sample of size [missing information] is randomly collected from the input image. Image patches are used as positive samples, while new viewpoint image patches of the same size are randomly rendered as negative samples. The samples are then randomly amplified with a 1 / 2 probability. The amplification operations include adjusting brightness, saturation, and contrast, as well as randomly masking regions in the image.
[0061] During training, the neural radiation field acts as a generator to render new perspective image patches, while the discriminator network is responsible for identifying whether a given sample is real or fake. During training, the generator loss is added to the total loss function to improve the realism of the texture synthesized by the neural radiation field, while the discriminator loss is used to train the discriminator network separately. This invention uses a cascaded convolutional neural network as the discriminator. Specifically, the discriminator network has three layers, each employing spectral normalization, with a convolutional kernel size of [missing information]. The step size is 2, the number of channels is doubled, and the LeakyReLU activation function is used. The discriminator also uses the Adam optimization algorithm, with a learning rate of 1 / 5 of the overall model learning rate. At the same time, in order to stabilize the training process, the discriminator network parameter update frequency is 1 / 5 of the entire model.
[0062] The final total loss function of the network It can be represented as:
[0063] in, This represents a loss of photometric uniformity. For deep consistency loss, For generator loss, To adjust the weighting coefficients of each component of the loss function. In this experiment, for the first 10,000 rounds, , After 10,000 rounds , .
[0064] Experimental results: Figure 4 This paper demonstrates the effectiveness of the sparse viewpoint synthesis method of this invention. The invention utilizes the LEVIR-NVS dataset, which contains 16 scenes, each with 21 multi-view images and camera poses. In the experiments, training images from three viewpoints were input, and after 30,000 iterations, 17 new viewpoint images were synthesized. Evaluation metrics included Peak Signal-to-Noise Ratio (PSNR), Structure Similarity Index Measure (SSIM), and Learned Perceptual Image Patch Similarity (LPIPS). Higher PSNR and SSIM indicate more realistic synthesis results, while lower LPIPS indicate better synthesized image quality. The quantitative experimental results are shown in the table below.
[0065]
[0066] Experimental results show that, compared with previous novel perspective synthesis algorithms, the quality of synthesized images from novel perspectives in remote sensing scenes is significantly improved, verifying the effectiveness of the design of each part of the invention. The results demonstrate that in remote sensing scenes, due to the complex distribution of ground features and sparse perspectives, previous novel perspective synthesis methods struggled to accurately reconstruct the scene. This invention, by constructing a depth-guided optimization framework, effectively avoids the overfitting problem of the model. Simultaneously, by introducing a texture discriminator based on generative adversarial networks, it effectively improves the model's texture synthesis capability in occluded areas, demonstrating high application value.
Claims
1. A novel perspective synthesis method for remote sensing scenes based on depth supervision and texture discriminator, characterized in that, The steps include the following: Step 1: Constructing an implicit reconstruction network based on neural radiation fields The implicit reconstruction network based on neural radiation fields consists of three parts: a feature extraction and fusion module, a neural radiation field reconstruction module, and a differentiable volume rendering module. The feature extraction module uses a ResNet34 model with pre-trained weights on the ImageNet dataset. During neural radiation field optimization, network parameters are frozen. Feature maps from the first pooling input and the outputs of the first three layers of ResNet34 are extracted to form a feature pyramid. Low-resolution feature maps are upsampled and merged to obtain a pixel-wise aligned feature map with 512 channels and a size half that of the reference image. Differentiable volume rendering is then used to synthesize the corresponding new perspective image. Step 2: Construct a deep supervised optimization framework based on sparse point clouds Point clouds obtained by structure-of-motion (SfM) estimation are used as depth guides to ensure that the model reconstructs the correct geometry without overfitting at the training viewpoint; multi-view stereo matching and point cloud fusion are employed to generate denser 3D point clouds. Step 3: Construct a texture discriminator based on generative adversarial networks A discriminator based on generative adversarial networks is used to guide the texture synthesis of new perspective images. The discriminator is trained in a generative adversarial manner. Positive samples are taken from image patches acquired from the input reference image, and negative samples are taken from the rendered synthesized new perspective image patches. The generator is a neural radiation field. The discriminator adopts a cascaded convolutional neural network and uses spectral normalization to satisfy the Lipshitz continuity condition. Differentiable enhancement is applied to both positive and negative samples. Differentiable enhancement allows the gradient to be passed to the generator through the amplification operation, and optimization is performed without changing the target distribution.
2. The method for synthesizing novel perspectives of remote sensing scenes based on depth supervision and texture discriminator according to claim 1, characterized in that: In step one, the network input consists of the three-dimensional coordinates of the sampling points and the direction vector of the ray, and cosine position encoding is performed: ; in, Representative to Perform cosine position encoding. These are the spatial coordinates of the sampling point. w It is the scaling factor, a hyperparameter that is adjusted according to the experimental scenario; To maximize the coding frequency, positional coding is not used for the view direction vector.
3. A method for synthesizing novel perspectives of remote sensing scenes based on depth supervision and texture discriminator according to claim 1 or 2, characterized in that: The network employs a multilayer perceptron (MLP) with residual connections, referred to as ResMLP. Each ResMLP block consists of a bilinear layer and residual connections, using the ReLU activation function. The network is divided into two parts: the first part is a multi-branch structure with shared weights, processing the input coordinate vector and features from different reference images; the feature vectors from the spatial reference images are encoded and input to each layer as residuals; the second part performs average pooling on the features output from different branches, and then passes them through several ResMLPs to predict the volume density and color attributes of the corresponding spatial sampling points. The input and output are represented as follows: ; in, Represents a neural radiation field network. It is the volume density of the sampling points. It's the color. It's the perspective direction. It comes from an image. Reference features.
4. The method for synthesizing novel perspectives of remote sensing scenes based on depth supervision and texture discriminator according to claim 3, characterized in that: After obtaining the volume density and color information, differentiable volume rendering is used to aggregate the information from all sampling points along the ray to obtain the final color information on the rendered image. ; ; in, Volume density and color at each sampling point, The distance between adjacent sampling points. Cumulative transmittance The number of sampling points. This represents the final RGB color value of the corresponding pixel.
5. The method for synthesizing novel perspectives of remote sensing scenes based on depth supervision and texture discriminator according to claim 1, characterized in that: In step two, given the 3D point cloud returned by SfM Projecting these onto a 2D view yields key points, and simultaneously calculating the depth of these key points. : ; in, Let these be the homogeneous coordinates of the point cloud in the camera coordinate system. Let be the homogeneous coordinates of the point cloud in the world coordinate system, where These correspond to the coordinate positions on the X, Y, and Z axes, respectively, with the last dimension (1) representing the scaling factor. and Let be the rotation and translation matrix of the camera; the depth of the point cloud at the current viewpoint is represented as: ; in, That is, the depth corresponding to the projection point, denoted as , Here are the pixel coordinates of the projection point, where u represents the x-coordinate and v represents the y-coordinate. For camera intrinsics; additional depth information is rendered in each iteration: ; in, It represents the position of each sampling point on the ray.
6. The method for synthesizing novel perspectives of remote sensing scenes based on depth supervision and texture discriminator according to claim 5, characterized in that: Color difference is used to measure the spatial error of point clouds; the color error between different pixels is defined as: ; in, Representing different pixel colors; the error of the projection point of the i-th point cloud onto the k-th reference image is denoted as . ,So This is expressed as the sum of the color errors between the standard and the point cloud and the projected points, representing the differences between the colors of different projected points: ; in, The number of reference images, This represents the color of the projection point of the i-th point cloud onto the j-th reference image. The average color of the i-projected points in the point cloud. It is the color of the i-th point cloud. It is the error between the color of each projected point and the average projected color. It is the error between the color of the projected point and the color of the point cloud.
7. The method for synthesizing novel perspectives of remote sensing scenes based on depth supervision and texture discriminator according to claim 6, characterized in that: Calculate adaptive weights for each keypoint: ; ; here, These are the weights corresponding to the projection points of the i-th point cloud onto the k-th reference image. As an intermediate variable, The truncation function indicates that the weight values are between 0 and 1; finally, the depth-weighted mean square error of each keypoint is calculated as the depth consistency loss. ; in, The number of keypoints in each round of sampling. These are the weights corresponding to the key points. It is a prediction of depth. It is the true depth value calculated by projection transformation.
8. The method for synthesizing novel perspectives of remote sensing scenes based on depth supervision and texture discriminator according to claim 1, characterized in that: In step three, the generative adversarial network uses the hinge loss function. The loss functions for the generator and discriminator are expressed as follows: ; in, This represents a generator network, also known as a neural radiation field. Represents the discriminator network. This represents a differentiable amplification operation on positive and negative samples; the generator loss... It is added to the model's total loss function, and Used for training the discriminator separately; This represents the sampled data.
9. The method for synthesizing novel perspectives of remote sensing scenes based on depth supervision and texture discriminator according to claim 8, characterized in that: The total loss function of the network Represented as: ; in, This represents a loss of photometric uniformity. For deep consistency loss, For generator loss, To adjust the weighting coefficients of each component of the loss function.