A method for rendering sparse remote sensing images based on season specificity
By introducing seasonally specific parameters and solar azimuth angles, and combining dense depth supervision and scaling correction techniques, the rendering model of sparse remote sensing images is optimized, solving the distortion problem caused by illumination and seasonal changes in remote sensing image rendering, and improving rendering accuracy and speed.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-03
- Publication Date
- 2026-03-27
AI Technical Summary
Existing remote sensing image rendering technologies struggle to handle changes in scene appearance caused by variations in lighting and seasonality, and they produce poor reconstruction results under sparse remote sensing image conditions, especially prone to geometric distortion under limited viewpoints.
By introducing seasonally specific parameters and solar azimuth angle, combined with volume density, albedo, solar visibility, and sky color, the seasonal characteristics of the land surface are dynamically captured. Furthermore, dense depth supervision and scaling correction techniques are used to optimize the rendering model of sparse remote sensing images, enabling shadow masking and seasonally specific rendering.
It improves the physical consistency and accuracy of remote sensing image rendering, reduces the computational overhead of non-critical areas, enhances rendering speed, and effectively alleviates distortion problems in sparse remote sensing rendering.
Smart Images

Figure CN121074227B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of remote sensing image rendering, and particularly relates to a sparse remote sensing image rendering method based on season specificity. BACKGROUND
[0002] Remote sensing image rendering refers to a process of using multi-view image data obtained by remote sensing technology to model and render the earth's surface through a series of algorithms and steps. Remote sensing image three-dimensional reconstruction technology can not only intuitively and accurately evaluate geographic spatial information, but also accurately model the three-dimensional information of the earth's surface scene. Although a large number of excellent remote sensing image rendering technical solutions have emerged based on neural radiance field (NeRF) technology or using sparse remote sensing images for remote sensing image rendering or three-dimensional reconstruction, the current remote sensing image rendering technical solutions still have many technical problems: first, the existing remote sensing image three-dimensional reconstruction method can handle illumination changes, but it does not take into account the appearance changes of the scene due to seasonal changes, nor does it exclude the influence of transient objects on the final rendering result. When processing images taken at different time points (such as different seasons or different times), the reconstruction result of these objects will be distorted. Second, in remote sensing image three-dimensional reconstruction, most methods still require a large number of input images to obtain good reconstruction results. For the case of only a few views, the model may produce incorrect geometric structures. In the shooting orbit environment of remote sensing devices, it is difficult to obtain multiple view images of the same scene within a specific time window, which makes it difficult for the existing remote sensing image rendering method to directly perform three-dimensional reconstruction. SUMMARY
[0003] In view of the above problems, the present application provides a sparse remote sensing image rendering method based on season specificity, which is used to at least solve one of the above technical problems.
[0004] According to a first aspect of the present application, a sparse remote sensing image rendering method based on season specificity is provided, comprising: performing ray sampling on a target sparse remote sensing image to obtain a sampling ray; inputting the three-dimensional coordinates of the sampling points on the sampling ray, the time of generating the target sparse remote sensing image, and the solar direction angle associated with the target sparse remote sensing image into a trained sparse remote sensing image rendering model for processing to obtain a volume density, a seasonally adjusted albedo, a solar visibility, and a sky color; determining the volume rendering contribution weight of the sampling points on the sampling ray to the target sparse remote sensing image using the volume density; and performing shadow mask rendering and season-specific rendering on the target sparse remote sensing image using the volume rendering contribution weight, the seasonally adjusted albedo, the solar visibility, and the sky color to obtain the rendered target sparse remote sensing image.
[0005] According to the embodiment of the present application, the above-mentioned performing shadow mask rendering and season-specific rendering on the target sparse remote sensing image by using the volume rendering contribution weight, the seasonally adjusted albedo, the solar visibility and the sky color, to obtain the rendered target sparse remote sensing image, comprises: performing operation on the volume rendering contribution weight and the seasonally adjusted albedo to obtain the seasonally adjusted rendering color, wherein the seasonally adjusted rendering color represents the seasonal difference of the pixel color in the target sparse remote sensing image; performing operation on the volume rendering contribution weight and the solar visibility to obtain the shadow mask, wherein the shadow mask represents the probability that the pixel in the target sparse remote sensing image is in the shadow; performing the shadow mask rendering and the season-specific rendering on the target sparse remote sensing image by using the seasonally adjusted rendering color, the shadow mask and the sky color, to obtain the rendered target sparse remote sensing image.
[0006] According to the embodiment of the present application, the above-mentioned training completed sparse remote sensing image rendering model is obtained by: performing preprocessing on the remote sensing image sample to obtain a low-resolution depth map sample, wherein the resolution of the low-resolution depth map sample is lower than the resolution of the remote sensing image sample; performing scale adjustment and offset correction on the low-resolution depth map sample by using a dense depth supervision module to obtain a multi-view Figure 1 consistency accumulated depth map, and performing depth supervision on the training process of the sparse remote sensing image rendering model by using the multi-view Figure 1 consistency accumulated depth map; performing processing on the remote sensing image sample by using the sparse remote sensing image rendering model to obtain rendering prediction parameters, and performing processing on the rendering prediction parameters by using a seasonal change module to obtain shadow mask prediction information and seasonal adjustment prediction information; performing RGB supervision on the training process of the sparse remote sensing image rendering model by using the shadow mask prediction information and the seasonal adjustment prediction information; and performing parameter update on the sparse remote sensing image rendering model by using the loss value in the depth supervision process and the loss value in the RGB supervision process, to obtain the training completed sparse remote sensing image rendering model.
[0007] According to the embodiment of the present application, the above-mentioned performing preprocessing on the remote sensing image sample to obtain a low-resolution depth map sample comprises: performing beam adjustment on the remote sensing image sample to optimize the rational polynomial coefficients of the remote sensing image sample, to obtain an optimized remote sensing image sample; and performing multiple independent semi-global matching on the optimized remote sensing image sample to obtain the low-resolution depth map sample.
[0008] According to the embodiment of the present application, the above-mentioned performing scale adjustment and offset correction on the low-resolution depth map sample by using the dense depth supervision module comprises: performing scale adjustment on the low-resolution depth map sample to obtain a scale-adjusted low-resolution depth map sample; performing offset correction on the scale-adjusted low-resolution depth map sample to obtain a multi-view Figure 1The consistency accumulated depth map comprises: taking a line connecting a camera view angle and a low-resolution depth map sample as a given light ray, and calculating accumulated depth prediction information of the low-resolution depth map sample under the given light ray; performing scale adjustment and offset correction on the low-resolution depth map sample by using a dense depth supervision module to obtain a scale factor and an offset factor; and performing operation on the scale factor, the offset factor and the accumulated depth prediction information based on multi-view Figure 1 consistency constraints to obtain a multi-view Figure 1 consistency accumulated depth map.
[0009] According to an embodiment of the present application, the above-mentioned training process of the sparse remote sensing image rendering model by using the multi-view Figure 1 consistency accumulated depth map comprises: taking a line connecting a camera view angle and a low-resolution depth map sample as a given light ray, and calculating accumulated depth prediction information of the low-resolution depth map sample under the given light ray; performing scale adjustment and offset correction on the low-resolution depth map sample by using a dense depth supervision module to obtain a scale factor and an offset factor; and performing operation on the scale factor, the offset factor and the accumulated depth prediction information based on multi-view Figure 1 consistency constraints to obtain a multi-view Figure 1 consistency accumulated depth map. Figure 1 consistency accumulated depth map, and the transmittance of the low-resolution depth map sample and the given light ray sampling point corresponding to the multi-view Figure 1 consistency accumulated depth map to obtain depth prior loss information; and performing depth supervision on a light ray sampling process of the remote sensing image sample and on a training process of the sparse remote sensing image rendering model by using the depth prior loss information.
[0010] According to an embodiment of the present application, the above-mentioned processing of the rendering prediction parameter by using the seasonal change module to obtain shadow mask prediction information and seasonal adjustment prediction information comprises: determining a volume rendering prediction contribution weight by using a predicted volume density in the rendering prediction parameter, wherein the volume rendering prediction contribution weight represents a color contribution degree of a sampling point on a sampling light ray corresponding to the remote sensing image sample to the remote sensing image sample; and performing activation processing on the volume rendering prediction contribution weight and a predicted solar visibility in the rendering prediction parameter by using the seasonal change module to obtain the shadow mask prediction information.
[0011] According to an embodiment of the present application, the above processing of the rendering prediction parameter by the seasonal change module to obtain the shadow mask prediction information and the seasonal adjustment prediction information further comprises: obtaining a seasonal category prediction probability corresponding to the remote sensing image sample, and obtaining an expected seasonal adjustment prediction probability representing that the remote sensing image sample belongs to different seasonal categories; performing operation on the seasonal category prediction probability and the expected seasonal adjustment prediction probability to obtain a seasonal adjustment prediction term; performing nonlinear activation on the seasonal adjustment prediction term and the albedo color corresponding to the remote sensing image sample to obtain a seasonally adjusted predicted albedo; and performing operation on the seasonally adjusted predicted albedo and the volume rendering prediction contribution weight by the seasonal change module to obtain the seasonal adjustment prediction information.
[0012] According to an embodiment of the present application, the above RGB supervision of the training process of the sparse remote sensing image rendering model by the shadow mask prediction information and the seasonal adjustment prediction information comprises: performing operation on the shadow mask prediction information, the seasonal adjustment prediction information, and the sky prediction color in the rendering prediction parameter to obtain a predicted rendering color; and performing RGB supervision of the training process of the sparse remote sensing image rendering model by the predicted rendering color.
[0013] According to an embodiment of the present application, the loss value in the above RGB supervision process comprises a loss value of a sampling light corresponding to the remote sensing image sample and a loss value of a sun light.
[0014] The present application provides a sparse remote sensing image rendering method based on season specificity, which can dynamically capture the seasonal characteristics of the ground surface by introducing seasonal time parameters and solar direction angle parameters to simulate the lighting and seasonal conditions in the real world, and realize the time-varying accurate expression of the ground surface properties. At the same time, the solar visibility parameter and the sky color parameter combined with the shadow mask rendering can simulate the lighting effect caused by the change of the solar elevation angle in different seasons, and enhance the physical consistency of the image. In addition, the rendering resources are dynamically allocated by the contribution weight, which reduces the calculation overhead of non-critical areas, improves the rendering speed while maintaining the accuracy. Moreover, the volume density prediction based on the neural radiation field (NeRF) of the present application can reconstruct the details of the three-dimensional scene from the limited sampling points, effectively alleviating the distortion problem in the sparse remote sensing rendering process. BRIEF DESCRIPTION OF DRAWINGS
[0015] The above content and other purposes, features and advantages of the present application will be more clearly understood through the following description of the embodiments of the present application with reference to the accompanying drawings, in which:
[0016] Figure 1 is an application scenario diagram of the sparse remote sensing image rendering method based on season specificity according to an embodiment of the present application.
[0017] Figure 2is a flowchart of a season-specific based sparse remote sensing image rendering method according to an embodiment of the present application.
[0018] Figure 3 is a training architecture schematic diagram of a sparse remote sensing image rendering model according to an embodiment of the present application.
[0019] Figure 4 is a block diagram of an electronic device suitable for implementing a season-specific based sparse remote sensing image rendering method according to an embodiment of the present application. DETAILED DESCRIPTION
[0020] Hereinafter, embodiments of the present application will be described with reference to the accompanying drawings. It is to be understood, however, that these descriptions are merely exemplary of the application and are intended to provide an overall understanding of the application. Many modifications and variations will be apparent to those of ordinary skill in the art from this description, which is to be construed as illustrative only. Furthermore, where known, equivalent alternative implementations can be substituted for elements described without departing from the scope of the present application. Thus, it is intended that the application not be limited to the implementations presented herein but that all such modifications and variations are intended to be included herein.
[0021] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. As used herein, the term "includes" and tautological equivalents thereof, means that the claimed features, steps, operations, and / or components are present, but does not exclude the presence or addition of one or more other features, steps, operations, or components.
[0022] All terms used herein including technical and scientific terms have the same meanings as commonly understood by one of ordinary skill in the art unless otherwise defined herein. It should be noted that the terms used herein are not intended to have ideal or excessively formal meanings, but are interpreted based on the context of the present specification.
[0023] In the case of using expressions similar to "at least one of A, B, and C, etc.", it is generally intended to include any and all combinations of one or more of the associated listed items, as well as the item alone (e.g., "a system having at least one of A, B, and C" is intended to cover a system with A alone, a system with B alone, a system with C alone, a system with both A and B, a system with both A and C, a system with both B and C, and a system with all of A, B, and C, etc.).
[0024] Existing NeRF-based remote sensing image rendering technical solutions often rely on a large number of remote sensing image samples. For a small number of remote sensing images, the NeRF-based remote sensing image rendering technical solution often cannot fit the correct geometry because it does not know that most of the scene is composed of empty space and opaque surfaces. In the orbit environment of a remote sensing satellite, there is little opportunity to obtain a large number of images of a given scene from multiple perspectives within a custom time window. Therefore, using sparse input remote sensing images for three-dimensional reconstruction becomes an important choice to solve these problems. Common methods for solving sparse input three-dimensional reconstruction include using semantic learning prior information, using sparse depth supervision, and using dense depth supervision.
[0025] Three-dimensional reconstruction using semantic learning priors. By training a model to learn information about objects, textures, structures, and other information in a scene, additional prior knowledge can be provided during the reconstruction process. This prior knowledge can help the model fill in missing data, especially in terms of scene understanding. In this way, even if the input data is very sparse, the model can use existing semantic information to infer the shape and texture of the missing parts. PixelNeRF (Neural Radiance Fields from One or Few Images) has achieved outstanding results on the task of synthesizing new views of an unknown scene using only a single view. To achieve this, PixelNeRF extends the standard NeRF model by introducing a depth feature and pre-training the entire architecture to enhance its generalization ability to new scenes. Similarly, DietNeRF utilizes a pre-trained visual Transformer to ensure semantic consistency across all views, including new views. SinNeRF (Training Neural Radiance Fields on Complex Scenes from a Single Image) further develops this concept by combining a self-supervised Dino ViT (Self-DIstillation with NO labels Vision Transformer) and adopting a classifier token representation instead of traditional image feature embeddings, reducing the impact of pixel misalignment between views. In addition, SinNeRF also achieves local texture regularization and depth supervision in new perspectives through depth mapping. On the other hand, MVSNeRF (Fast Generalizable Radiance Field Reconstruction from Multi-View Stereo) draws on multi-view stereo matching technology by projecting 2D Convolutional Neural Network (CNN) features onto planes that pass through the scene, and finally using a 3D CNN to extract a neural encoding volume that is converted into RGB images and density information after regression.
[0026] Sparse depth supervision refers to the use of a small number of depth measurements as supervisory signals to train the model. These depth measurements can be discontinuous or distributed in certain specific areas of the image. During the training process, the model will try to predict the complete depth map, and then adjust its prediction by comparing the predicted result with the actual depth measurements. A key challenge of this method is how to effectively use these sparse depth information to guide the model to learn more accurate depth estimation. DS-NeRF (Depth-supervised NeRF: Fewer Views and Faster Training for Free) is the first method to use 3D points obtained from Structure from Motion (SfM) for sparse depth supervision. This research introduces an adaptive ray sampling strategy and designs a depth termination loss based on the weighted 3D point reprojection error. Sat-NeRF (Learning Multi-View Satellite Photogrammetry With Transient Objects and Shadow Modeling Using RPC Cameras) also applies a similar sparse depth supervision mechanism to process multi-date satellite images, and successfully reduces the number of required training images to about 15. It is worth noting that the model architecture of Sat-NeRF also contains physical characteristic parameters for earth observation satellite images, such as albedo and solar angle correction, in order to better handle data collected at different time points.
[0027] Compared to sparse depth supervision, dense depth supervision utilizes more comprehensive depth information, typically obtained through binocular stereo vision, structured light sensors, or other techniques that can provide depth data over a larger range. This approach provides high-density depth maps during the training phase, enabling the model to learn more detailed spatial structure information. As the input information is more comprehensive, the model can better understand the geometric details in the scene and generate more accurate three-dimensional reconstruction results during the testing phase. NerfingMVS (Guided Optimization of Neural Radiance Fields for Indoor Multi-view Stereo) combines learning-based multi-view stereo techniques with NeRF and applies it to indoor three-dimensional reconstruction. This method starts with a set of sparse 3D points output by Structure from Motion (SfM), first trains a monocular dense depth prediction network. Consistency checks between the predicted depths of each view provide an error map to guide the sampling of rays during the final NeRF optimization process. In the most sparse view scene they handled, a total of 35 images were used for training.
[0028] In view of the distortion problem of the reconstruction result caused by the use of sparse input views when using multi-view satellite images for three-dimensional reconstruction and the appearance change problem of the scene caused by seasonal change, the present application proposes a sparse remote sensing image rendering method based on season specificity. The present application uses the low-resolution dense depth map generated by traditional Multiple View Stereo (MVS) technology as supervision, so that even if there are only two to three input views, new views and 3D surfaces can be generated, and the dense depth can be optimized and corrected to improve the accuracy of depth guidance. At the same time, the present application uses the metadata of satellite images to more accurately simulate the real-world lighting and seasonal conditions, which can calculate the albedo that changes with the season, which helps to capture the changes of natural features such as vegetation in different seasons. And by limiting the range of solar angles and time of year, it is ensured that the change of time only affects the seasonal characteristics, and the change of solar angle only affects the position and intensity of the shadow. This can handle changes in lighting and adapt to changes in seasonal features.
[0029] Figure 1 is an application scenario diagram of the sparse remote sensing image rendering method based on season specificity according to an embodiment of the present application.
[0030] As Figure 1As shown, the application scenario 100 according to this embodiment can include scenarios such as remote sensing image rendering or three-dimensional reconstruction. The network 104 is a medium for providing a communication link between the first terminal device 101, the second terminal device 102, the third terminal device 103, and the server 105. The network 104 can include various connection types, such as wired, wireless communication links, or optical fiber cables, and the like.
[0031] The user can use the first terminal device 101, the second terminal device 102, and the third terminal device 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications can be installed on the first terminal device 101, the second terminal device 102, and the third terminal device 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples).
[0032] The first terminal device 101, the second terminal device 102, and the third terminal device 103 can be various electronic devices with display screens and supporting web browsing, including but not limited to smartphones, tablet computers, laptop computers, desktop computers, and the like.
[0033] The server 105 can be a server that provides various services, such as a background management server that provides support for websites browsed by users using the first terminal device 101, the second terminal device 102, and the third terminal device 103 (only as an example). The background management server can analyze and process received user requests and the like, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.
[0034] It should be noted that the method for rendering sparse remote sensing images based on season specificity provided by the embodiment of the present application can generally be executed by the server 105. Accordingly, the device for rendering sparse remote sensing images based on season specificity provided by the embodiment of the present application can generally be provided in the server 105. The method for rendering sparse remote sensing images based on season specificity provided by the embodiment of the present application can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105. Accordingly, the device for rendering sparse remote sensing images based on season specificity provided by the embodiment of the present application can also be provided in a server or a server cluster different from the server 105 and capable of communicating with the first terminal device 101, the second terminal device 102, the third terminal device 103, and / or the server 105.
[0035] It should be understood that, Figure 1The number of terminal devices, networks and servers in the figure is only illustrative. According to the implementation needs, there can be any number of terminal devices, networks and servers.
[0036] The following will be based on Figure 1 The described scenario, by Figures 2-3 The season-specific sparse remote sensing image rendering method of the disclosed embodiment is described in detail.
[0037] Figure 2 The flowchart of the season-specific sparse remote sensing image rendering method according to the embodiment of the application.
[0038] As Figure 2 The season-specific sparse remote sensing image rendering method of the embodiment includes operations S210-S240.
[0039] In operation S210, ray sampling is performed on the target sparse remote sensing image to obtain a sampling ray.
[0040] Ray sampling on the target sparse remote sensing image can extract sampling rays with specific spatiotemporal information and locate the geometric profile of the ground target in the remote sensing image; in addition, the sampling ray gives each pixel multi-dimensional spectral information, which constitutes the color basis of the remote sensing image.
[0041] In operation S220, the three-dimensional coordinates of the sampling points on the sampling ray, the time of generating the target sparse remote sensing image, and the solar direction angle associated with the target sparse remote sensing image are input into the trained sparse remote sensing image rendering model for processing to obtain the volume density, the seasonally adjusted albedo, the solar visibility and the sky color.
[0042] The three-dimensional coordinates (or spatial coordinates, three-dimensional spatial coordinates) of the sampling points, the solar direction angle (including the solar azimuth angle and the solar elevation angle), and the time of generating the target sparse remote sensing image are input into the trained sparse remote sensing image rendering model as inputs, and after processing by the trained sparse remote sensing image rendering model, the rendering parameters for rendering (or three-dimensional reconstruction) of the target sparse remote sensing image are obtained, wherein the rendering parameters can include the volume density, the seasonally adjusted albedo, the solar visibility and the sky color.
[0043] In operation S230, the volume density is used to determine the volume rendering contribution weight of the sampling points on the sampling ray to the target sparse remote sensing image.
[0044] The volume density represents the absorption ability of a sampling point in space to light; the volume rendering contribution weight refers to the degree of influence of a sampling point on the final pixel color; the volume density quantifies this contribution through a volume rendering equation, which is specifically represented as: the higher the value, the greater the probability of light being absorbed or scattered, resulting in the more significant color contribution of the point on the light path; the lower the value (close to 0), the almost unobstructed light passing through (similar to air), and the contribution is negligible.
[0045] In operation S240, the shadow mask rendering and the season-specific rendering of the target sparse remote sensing image are performed by using the volume rendering contribution weight, the seasonally adjusted albedo, the solar visibility and the sky color, to obtain the rendered target sparse remote sensing image.
[0046] The seasonally adjusted albedo is used to represent the degree of influence of the albedo color of the target sparse remote sensing image by the season.
[0047] The season-specific sparse remote sensing image rendering method provided by the present application can dynamically capture the seasonal characteristics of the ground surface by introducing the season time parameter and the solar direction angle parameter to simulate the real-world lighting and seasonal conditions, and can realize the time-varying accurate expression of the ground surface properties; at the same time, the solar visibility parameter and the sky color parameter are combined with the shadow mask rendering to simulate the lighting effect caused by the change of the solar altitude angle in different seasons, and to enhance the physical consistency of the image; in addition, the rendering resources are dynamically allocated by the contribution weight to reduce the calculation overhead of non-critical areas, and the rendering speed is improved while the accuracy is maintained. Moreover, the volume density prediction based on the neural radiation field (NeRF) can reconstruct the details of the three-dimensional scene from the limited sampling points, and effectively alleviate the distortion problem in the sparse remote sensing rendering process.
[0048] According to the embodiment of the present application, the above-mentioned execution of the shadow mask rendering and the season-specific rendering of the target sparse remote sensing image by using the volume rendering contribution weight, the seasonally adjusted albedo, the solar visibility and the sky color to obtain the rendered target sparse remote sensing image comprises: performing operation on the volume rendering contribution weight and the seasonally adjusted albedo to obtain the seasonally adjusted rendering color, wherein the seasonally adjusted rendering color represents the seasonal difference of the pixel color in the target sparse remote sensing image; performing operation on the volume rendering contribution weight and the solar visibility to obtain the shadow mask, wherein the shadow mask represents the probability that the pixel in the target sparse remote sensing image is in the shadow; performing the shadow mask rendering and the season-specific rendering of the target sparse remote sensing image by using the seasonally adjusted rendering color, the shadow mask and the sky color to obtain the rendered target sparse remote sensing image.
[0049] The above embodiment realizes, for the first time in a remote sensing image, dynamic seasonal correction of surface reflectivity and time consistency maintenance (elimination of color jump in the same area in different seasons) of pixel-level spectral characteristics through joint operation of body rendering contribution weight and seasonal albedo; and obtains a shadow mask representing whether a pixel in the remote sensing image is in a shadow through interactive calculation of solar visibility and body rendering weight.
[0050] According to the embodiment of the present application, the above-mentioned trained sparse remote sensing image rendering model is obtained by: performing preprocessing on the remote sensing image sample to obtain a low-resolution depth map sample, wherein the resolution of the low-resolution depth map sample is lower than the resolution of the remote sensing image sample; performing scale adjustment and offset correction on the low-resolution depth map sample by using a dense depth supervision module to obtain a multi-view Figure 1 consistent accumulated depth map, and performing depth supervision on the training process of the sparse remote sensing image rendering model by using the multi-view Figure 1 consistent accumulated depth map; performing processing on the remote sensing image sample by using the sparse remote sensing image rendering model to obtain rendering prediction parameters, and performing processing on the rendering prediction parameters by using a seasonal change module to obtain shadow mask prediction information and seasonal adjustment prediction information; performing RGB supervision on the training process of the sparse remote sensing image rendering model by using the shadow mask prediction information and the seasonal adjustment prediction information; and performing parameter updating on the sparse remote sensing image rendering model by using the loss value in the depth supervision process and the loss value in the RGB supervision process to obtain the trained sparse remote sensing image rendering model.
[0051] The training method of the sparse remote sensing image rendering model provided by the present application uses a dense depth supervision module to guide sparse remote sensing image rendering (or three-dimensional reconstruction): the dense depth supervision module improves the three-dimensional reconstruction process of sparse remote sensing images. Traditional remote sensing image rendering methods usually rely on dense data sets for three-dimensional modeling, but in actual applications, only sparse remote sensing data can be obtained. Therefore, the training method provided by the present application uses the dense supervision mechanism in the depth learning model, so that even in the case of insufficient data, the model can effectively learn more rich spatial features from limited input, thereby improving the accuracy and reliability of three-dimensional reconstruction.
[0052] At the same time, the training method of the sparse remote sensing image rendering model provided by the present application is based on a new method of scale and offset correction, which corrects image distortion by performing geometric transformation on the depth image, i.e. adjusting the scale factor and position offset of the image. This training method can effectively improve the image quality, making subsequent feature extraction and analysis more accurate.
[0053] Furthermore, in order to make remote sensing image analysis more closely reflect reality, the training method of the sparse remote sensing image rendering model provided by this invention introduces time dimension information into the sparse remote sensing image rendering model. By modeling factors such as lighting conditions and seasonal changes in different time periods, it can more accurately reflect changes in the real world.
[0054] The following specific embodiments, in conjunction with the appendix, demonstrate this process. Figure 3 The training method for the sparse remote sensing image rendering model provided by this invention will be described in further detail.
[0055] Figure 3 This is a schematic diagram of the training architecture of a sparse remote sensing image rendering model according to an embodiment of the present invention.
[0056] like Figure 3 As shown, where, Indicates volume density, Albedo, which indicates seasonally adjusted albedo Indicates solar visibility. This represents the sky color; during model training, the goal of this invention is to calculate the color and density of each point in the scene. To this end, light samples are first performed on the input remote sensing image samples to obtain sampled light rays. Subsequently, the sampled light rays are used to sample points... Time of generating remote sensing image samples and the solar orientation angle associated with the remote sensing image samples. This data is input into a sparse remote sensing image rendering model. Specifically, it utilizes the time of year... and sampling points The color of each point is determined by its three-dimensional spatial coordinates. Simultaneously, the solar azimuth angle... and sampling points in the model Solar visibility is calculated using three-dimensional spatial coordinates, using only the solar orientation angle. To calculate sky color By combining solar visibility and sky color at a given point, the shadow adjustments required for rendering pixels can be calculated. Ultimately, by integrating these factors, the model can calculate the color and density information for each point, thereby achieving image rendering from a new perspective. This method not only improves the accuracy of 3D reconstruction but also enhances the realism of the scene, making it more consistent with actual lighting and seasonal conditions.
[0057] At the same time, such as Figure 3 As shown, during the model training process, depth supervision and RGB (Red, Green, Blue) supervision are performed to improve the training effect of the model.
[0058] According to an embodiment of the present application, the above-mentioned performing preprocessing on the remote sensing image sample to obtain a low-resolution depth map sample comprises: performing bundle adjustment on the remote sensing image sample to optimize rational polynomial coefficients of the remote sensing image sample to obtain an optimized remote sensing image sample; and performing multiple independent semi-global matching on the optimized remote sensing image sample to obtain the low-resolution depth map sample.
[0059] The processing procedure of the remote sensing image sample is further described in detail below through specific embodiments.
[0060] The present application uses bundle adjustment technology to optimize the rational polynomial coefficients (RPC, Rational Polynomial Coefficients) of the input remote sensing image sample. This operation aims to improve the geometric accuracy of the image and ensure the accuracy of subsequent processing. Then, for each input remote sensing image sample, an independent semi-global matching algorithm (SGM, Semi-Global Matching) is run, i.e., multiple independent semi-global matching is performed on each remote sensing image sample to obtain a series of low-resolution depth map samples. The low-resolution depth map sample is selected instead of the high-resolution depth map sample to prevent the model from overfitting the SGM solution and to avoid the problem of depth information loss due to insufficient texture information in the high-resolution image. In the depth prediction stage, the similarity between these low-resolution depth map samples is used as a supervision signal to guide the learning process of the depth prediction model. In this way, even in the case of only a small number of input images, high-quality new perspective images and three-dimensional surface models can be generated.
[0061] According to an embodiment of the present application, the above-mentioned performing scale adjustment and offset correction on the low-resolution depth map sample by using the dense depth supervision module to obtain a multi-view Figure 1 consistency accumulated depth map comprises: taking a line connecting a camera view angle and the low-resolution depth map sample as a given light ray, and calculating accumulated depth prediction information of the low-resolution depth map sample under the given light ray; performing scale adjustment and offset correction on the low-resolution depth map sample by using the dense depth supervision module to obtain a scale factor and an offset factor; and performing operation on the scale factor, the offset factor, and the accumulated depth prediction information based on a multi-view Figure 1 consistency constraint to obtain the multi-view Figure 1 consistency accumulated depth map.
[0062] According to an embodiment of the present application, the above-mentioned performing depth supervision on the training process of the sparse remote sensing image rendering model by using the multi-view Figure 1 consistency accumulated depth map comprises: taking the multi-view Figure 1 consistency accumulated depth map and a multi-view Figure 1The transmittance of the low-resolution depth map samples corresponding to the consistent cumulative depth map and the upsampling points of a given ray is calculated to obtain the distribution standard deviation of the low-resolution depth map samples along the given ray; the similarity between the low-resolution depth map samples and the preset normalized scaling factor and preset offset factor is calculated to obtain the equivalent uncertainty measure; the distribution standard deviation, equivalent uncertainty measure, and multi-view... Figure 1 The depth prior loss information is obtained by processing the consistent cumulative depth map and low-resolution depth map samples. The depth prior loss information is then used to supervise the ray sampling process of the remote sensing image samples and to perform depth supervision on the training process of the sparse remote sensing image rendering model.
[0063] The above embodiments involve a dense supervision module. The dense supervision module will be further described in detail below through specific embodiments.
[0064] The Dense Depth Supervision module aims to provide depth prior information for the model training process. This module predicts the depth of a given ray by accumulating the radiation field throughout the optimization volume and guides ray sampling. Along the ray... The depth prediction can be calculated using formula (1):
[0065] (1).
[0066] in, This represents the depth prediction information for ray r. Indicates the first Transmittance at each sampling point Indicates the first The absorption probability of each sampling point Indicates the first The distance from each sampling point to the camera's optical center. Indicates light rays The total number of sampling points.
[0067] Use N independent SGMs to obtain a low-resolution depth map sample for each image. As a depth guide, but these depth maps may not be multi-view. Figure 1 Therefore, this invention performs scale and offset corrections on each depth map, wherein, This indicates the scaling factor (or scaling adjustment factor). This represents the offset factor (or offset correction factor). Utilizing multi-view... Figure 1 Consistency constraints, aimed at restoring Multiview Figure 1 Cumulative depth map (i.e., along the light) The depth prediction information is shown in formula (2):
[0068] (2).
[0069] By jointly optimizing and and the model to achieve un-distorted multi-view Figure 1 consistent accumulated depth map and model rendered depth map between the consistency to achieve correction.
[0070] If the current sample point is opaque, its depth will contribute to the multi-view Figure 4 consistent accumulated depth map , while the previous transparent sample points will be ignored. To describe the distribution of samples along the ray, a standard deviation equation (i.e., distribution standard deviation) is defined as shown in equation (3):
[0071] (3).
[0072] where denotes the square of the standard deviation; a smaller standard deviation value means that the samples are concentrated around the estimated depth, which will produce a clearer edge on the object surface. Then, an equivalent uncertainty measure driven by the input data is defined, which is based on the similarity measure generated by SGM as shown in equation (4):
[0073] (4).
[0074] where denotes the cross-correlation similarity of the ray samples along the ray (i.e., ray r, same below) of the input depth, and are the preset normalization scaling factor and the preset offset factor (or offset parameter), respectively, which are empirically set to and in the experiment. This equivalent uncertainty measure can be used as a weight applied to the final depth loss, as a threshold to determine whether to activate the loss, and to guide the ray sampling at the same time.
[0075] All the above components together constitute the depth loss, which encourages to be close to and is guided by the input equivalent uncertainty measure as shown in equation (5):
[0076] (5).
[0077] where denotes the depth prior loss, A ray sub-region is defined as satisfying either of the following conditions: (1) ; (2) These conditions force the ray to terminate within the depth prior Outside this region, the depth loss either does not activate or is clipped. Throughout the training process, the depth loss is always in play.
[0078] According to an embodiment of the present application, the processing of the rendering prediction parameters by the seasonal variation module to obtain the shadow mask prediction information and the seasonal adjustment prediction information includes: determining a volume rendering prediction contribution weight using a predicted volume density in the rendering prediction parameters, wherein the volume rendering prediction contribution weight represents a color contribution degree of a sampling point on a sampling ray corresponding to the remote sensing image sample to the remote sensing image sample; and performing activation processing on the volume rendering prediction contribution weight and a predicted solar visibility in the rendering prediction parameters by the seasonal variation module to obtain the shadow mask prediction information.
[0079] According to an embodiment of the present application, the processing of the rendering prediction parameters by the seasonal variation module to obtain the shadow mask prediction information and the seasonal adjustment prediction information further includes: obtaining a seasonal category prediction probability corresponding to the remote sensing image sample, and obtaining an expected seasonal adjustment prediction probability representing that the remote sensing image sample belongs to different seasonal categories; performing operation on the seasonal category prediction probability and the expected seasonal adjustment prediction probability to obtain a seasonal adjustment prediction term; performing nonlinear activation on the seasonal adjustment prediction term and an albedo color corresponding to the remote sensing image sample to obtain a seasonally adjusted predicted albedo; and performing operation on the seasonally adjusted predicted albedo and the volume rendering prediction contribution weight by the seasonal variation module to obtain the seasonal adjustment prediction information.
[0080] According to an embodiment of the present application, the RGB supervision performed on the training process of the sparse remote sensing image rendering model by using the shadow mask prediction information and the seasonal adjustment prediction information includes: performing operation on the shadow mask prediction information, the seasonal adjustment prediction information, and a sky prediction color in the rendering prediction parameters to obtain a predicted rendering color; and performing RGB supervision on the training process of the sparse remote sensing image rendering model by using the predicted rendering color.
[0081] The above embodiments relate to the seasonal variation module, which will be further described in detail below through specific embodiments.
[0082] The sparse remote sensing image rendering model uses spatial coordinates and a solar angle and time as inputs, a volume density , a seasonally adjusted albedo , a solar visibility , and a sky color To better account for seasonal variations, this invention incorporates a seasonally adjusted albedo term when calculating albedo color: seasonally adjusted albedo. This adjustment term is based on geographical location and specific time of year. The albedo color is determined by combining it with the seasonal adjustment term, and then a sigmoid nonlinear transformation is applied to the combined seasonally adjusted albedo. This allows the combined term to take any value before the nonlinear transformation, while after the nonlinear transformation, the color value falls within the effective color range. Therefore, the seasonally adjusted albedo is expressed as shown in formula (6):
[0083] (6).
[0084] in, The three-dimensional spatial coordinates in the scene are: The intrinsic albedo of the sampling points, Indicates time The three-dimensional spatial coordinates in the time and scene are Seasonal adjustments for sampling points Indicates time The three-dimensional spatial coordinates in the time and scene are The seasonally adjusted albedo of the sampling points. This represents a non-linear activation function.
[0085] To calculate the seasonal adjustment term, this invention introduces two intermediate variables. and ,in, Indicates time category, Indicates time adjustment. Time category. It depends on the time of year, and time adjustment This depends on the location information in the model. Time category It is Matrix, where This is a hyperparameter representing the number of different seasonal categories the model can distinguish. It is processed using the softmax function. This ensures that its values form a discrete probability distribution with respect to different seasonal categories. Time adjustment. It is Matrix, where, This refers to the number of output channels in the rendered image. Seasonal adjustment. The calculation is shown in formula (7):
[0086] (7).
[0087] in, The An element represents time Corresponds to the probability of season , while The column represents the seasonal adjustment of each season category. Therefore, In fact, it is the expected seasonal adjustment based on the probability of different season categories in If Too big (that is, the model can generate too many possible seasons), the shadow effect will be absorbed into the seasonal adjustment. If Too small, that is, the model can generate too many seasonal possibilities, the shadow effect may be wrongly attributed to the seasonal adjustment. Conversely, if Too small, the model will not be able to fully express the characteristics of each season. In the test process, set To achieve a good balance between the sun and time change.
[0088] In the sparse remote sensing image rendering model, only the time Change the albedo value of the seasonal adjustment, while the density, solar visibility and sky color are assumed to be unaffected by Although it is unrealistic to assume that the density of an area will not change over time, as changes in vegetation do indeed cause changes in density, there are two reasons for choosing to set the model in this way: one is to prevent the model from using To change the density inside the scene or the shadows in the rendered image to explain the pure visual changes of fixed density objects such as buildings; two is to take into account the fact that each region has a single true digital elevation map, making it difficult to accurately assess the quality changes in the elevation map caused by changes in density. The rendering color adjusted by season can be represented by equation (8):
[0089] (8).
[0090] Where Indicates the rendering color of the corresponding region of light At time After seasonal adjustment, Indicates the time (season) corresponding to the color, Indicates the light or scene surface region corresponding to the color; Indicates the accumulation of contributions of all sampling points On light , where Indicates that the sampling point Belongs to the propagation path range of light ; Indicates the Sampling point On light At time Seasonally adjusted albedo; Light Upper sampling points The volume rendering contribution weight of the final rendered color of this ray. This indicates the ray corresponding to that weight. This indicates the sampling point corresponding to the weight.
[0091] To calculate the seasonally adjusted rendering colors, the time information for the year needs to be input into the network. However, this invention does not directly input the date and month, but instead encodes the time as shown in formula (9):
[0092] (9).
[0093] in, This is the proportion of a year that has been completed, that is, the time from the start of the year to the time when the sparse remote sensing image of the target is generated, which is the proportion of the total time of the year. For two time points... and , Indicates two time points and Corresponding The difference between values, their differences Distance is ,exist This means that the maximum value is reached at the other end of the year. This method also ensures that the beginning and end dates of the year have similar codes.
[0094] Solar visibility and the color of the sky It does not directly affect the seasonally adjusted albedo, but is used to calculate a shadow mask, which measures whether a rendered pixel is in shadow. The shadow mask is calculated as shown in formula (10):
[0095] (10).
[0096] in, It is the solar azimuth angle. It's the solar visibility of the network. It's the sigmoid activation function. It's important to note that... It means The probability that the surface is visible from the sun. Therefore, yes The probability that a surface is in shadow. (Hyperparameter) The hyperparameters control the smoothness of the transition between non-shadow and shadow areas. This determines the location where this transition (the transition between non-shaded and shaded areas) occurs.
[0097] With shadow masking and seasonally adjusted render colors, shadows and seasonally adjusted render colors can be calculated. As shown in formula (11):
[0098] (11).
[0099] Formula (11) assumes It is the color of the pixel under direct sunlight, and It is the color of the pixel when it is in shadow. The value represents the probability that a pixel is in shadow. Besides direct sunlight, other light sources can also illuminate shadow areas. Therefore, A value greater than zero ensures that the shadowed area is partially illuminated.
[0100] According to an embodiment of the present invention, the loss value in the above-mentioned RGB supervision process includes the loss value of the sampled light corresponding to the remote sensing image sample and the loss value of sunlight.
[0101] The loss function used in the model training process of this invention will be further explained in detail below through specific embodiments.
[0102] To ensure that at least one color channel exceeds a user-defined threshold in the seasonally adjusted rendered color, this invention introduces a loss term. By encouraging the model in At least one channel value is present, avoiding the use of seasonally adjusted albedo to represent shadows in dark areas. Loss term As shown in formula (12)
[0103] (12).
[0104] in, Indicates the rendered color based on seasonal adjustments. The designed albedo activation loss term, A 3D vector representing the seasonally adjusted rendered color, 𝔸 is a vector between Hyperparameters of the interval. Even if there are Even with a loss term, the model in this invention still tends to ignore the effect of shadows by setting the sky color to 1. This indicates the first color of the rendered image after seasonal adjustment. Each color channel value With preset threshold The smaller of the two values is taken. To avoid this situation, the present invention adds a loss function. As shown in equations (13) and (14):
[0105] (13).
[0106] (14).
[0107] where S is also a hyperparameter between intervals, represents the value of the th color channel of the sky color. The use of ensures that the color of the shadow area is at least 50% lower than normal, while still allowing some indirect light to illuminate the area. For a given image ray , the total loss associated with this ray is as shown in equation (15):
[0108] (15).
[0109] where represents the three-dimensional color vector of the pixel corresponding to the ray in the target sparse remote sensing image in the real scene, GT refers to the real color data of the pixel in the remote sensing image sample, and is the benchmark used to calculate the rendering error (loss value) during model training.
[0110] The loss function related to the sun ray contains two items: the first item minimizes the MSE (Mean-Square Error) between the estimated solar visibility and the solar visibility; the second item is the loss, applied to the sun ray to ensure that the sky color is correct even if the input sun angle is significantly different from the sun angle in the training set. The total loss of the sun ray is as shown in equation (16):
[0111] (16).
[0112] Through the combination of these loss functions, the present application ensures that the model can optimize its prediction results during the training process, making it as close as possible to the actual lighting conditions, and effectively simulating natural lighting and shadow effects. The total loss of the model is as shown in equation (17):
[0113] (17).
[0114] where is the camera parameter regularization loss term.
[0115] In the training process of the sparse remote sensing image rendering model provided by the present application, the use of dense depth supervision is proposed to guide the three-dimensional reconstruction of sparse remote sensing images: the three-dimensional reconstruction process of remote sensing images is improved through the dense depth supervision technology. Traditional rendering models usually rely on dense data sets for three-dimensional modeling, but in actual applications, only sparse remote sensing data can be obtained. Therefore, by using the dense supervision mechanism in the deep learning model, even in the case of insufficient data, the model can effectively learn more rich spatial features from limited input, thereby improving the accuracy and reliability of three-dimensional reconstruction.
[0116] At the same time, the present application proposes to use scale and offset correction for depth de-distortion: based on scale and offset correction, the present application corrects image distortion by performing geometric transformation on depth images, i.e. adjusting the scale factor and position offset of the image. This way can effectively improve the image quality, so that the subsequent feature extraction and analysis are more accurate.
[0117] In addition, the present application proposes to introduce time information to simulate real-world lighting and seasonal conditions: in order to make the remote sensing data analysis more close to the actual situation, the time dimension information is introduced into the model, and by modeling the lighting conditions, seasonal changes and other factors in different time periods, the changes in the real world can be more accurately reflected.
[0118] Figure 4 is a block diagram of an electronic device suitable for implementing a sparse remote sensing image rendering method based on season-specificity according to an embodiment of the present application.
[0119] As shown in , the electronic device 400 according to an embodiment of the present application includes a processor 401, which can perform various appropriate actions and processes according to programs stored in a ROM 402 (Read Only Memory) or loaded into a RAM 403 (Random Access Memory) from a storage section 408. The processor 401 may, for example, include a general-purpose microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a special-purpose microprocessor (such as an application-specific integrated circuit (ASIC)), etc. The processor 401 can also include on-board memory for cache use. The processor 401 can include a single processing unit or multiple processing units for performing different actions of the method processes according to embodiments of the present application.
[0120] In the RAM 403, various programs and data required for the operation of the electronic device 400 are stored. The processor 401, the ROM 402, and the RAM 403 are connected to each other via the bus 404. The processor 401 performs various operations of the method flow according to the embodiments of the present application by executing the programs in the ROM 402 and / or the RAM 403. It should be noted that the programs can also be stored in one or more memories other than the ROM 402 and the RAM 403. The processor 401 can also perform various operations of the method flow according to the embodiments of the present application by executing the programs stored in the one or more memories.
[0121] According to the embodiments of the present application, the electronic device 400 can further include an input / output (I / O) interface 405, which is also connected to the bus 404. The electronic device 400 can further include one or more of the following components connected to the input / output (I / O) interface 405: an input part 406 including a keyboard, a mouse, etc.; an output part 407 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage part 408 including a hard disk, etc.; and a communication part 409 including a network interface card such as a LAN card, a modem, etc. The communication part 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the input / output (I / O) interface 405 as necessary. A removable medium 411 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is attached to the drive 410 as necessary, so that a computer program read therefrom is installed in the storage part 408 as necessary.
[0122] The present application also provides a computer readable storage medium, which can be included in the device / apparatus / system described in the above embodiments; or can exist separately without being assembled into the device / apparatus / system. The above computer readable storage medium carries one or more programs, when the one or more programs are executed, the method according to the embodiments of the present application is implemented.
[0123] According to embodiments of the present application, the computer readable storage medium can be a non-transitory computer readable storage medium, such as, for example, without limitation, a portable computer diskette, a hard disk, random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the present application, a computer readable storage medium can be any tangible medium that can contain, or store a program for use by or in connection with an instruction execution system, apparatus, or device. For example, according to embodiments of the present application, the computer readable storage medium can include the ROM 402 and / or the RAM 403 described above, and / or one or more other memories not expressly described above.
[0124] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functional processes, and operational processes, according to various embodiments of the present application. In this regard, each block in the flow diagrams or block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may
[0125] Those skilled in the art will appreciate that the features recited in the various embodiments of the present application can be combined and / or integrated in a variety of ways, even if such combinations or integrations are not expressly noted in the present application. In particular, the features recited in the various embodiments of the present application can be combined and / or integrated in a variety of ways without departing from the spirit and scope of the present application. All such combinations and / or integrations are within the scope of the present application.
[0126] The embodiments of the present application have been described above. However, these embodiments are merely meant to be illustrative and not limiting of the scope of the present application. Although each of the embodiments has been described above, this does not mean that measures in each of the embodiments cannot be advantageously combined. Various alternatives and modifications can be made to the embodiments of the present application without departing from the scope of the present application, and all such alternatives and modifications are intended to fall within the scope of the present application.
Claims
1. A method for rendering sparse remote sensing images based on seasonal specificity, characterized in that, The method includes: Ray sampling is performed on the sparse remote sensing image of the target to obtain sampled rays; The three-dimensional coordinates of the sampling points on the sampling ray, the time of generating the target sparse remote sensing image, and the solar azimuth angle associated with the target sparse remote sensing image are input into the trained sparse remote sensing image rendering model for processing to obtain volume density, seasonally adjusted albedo, solar visibility, and sky color. The volume density is used to determine the contribution weight of the sampling points on the sampling ray to the volume rendering of the target sparse remote sensing image; The process involves performing shadow masking and seasonally specific rendering on the target sparse remote sensing image using the volume rendering contribution weight, the seasonally adjusted albedo, the solar visibility, and the sky color to obtain the rendered target sparse remote sensing image. This includes: calculating the volume rendering contribution weight and the seasonally adjusted albedo to obtain a seasonally adjusted rendering color, where the seasonally adjusted rendering color represents the seasonal difference in pixel color in the target sparse remote sensing image; calculating the volume rendering contribution weight and the solar visibility to obtain a shadow mask, where the shadow mask represents the probability that a pixel in the target sparse remote sensing image is in shadow; and performing shadow masking and seasonally specific rendering on the target sparse remote sensing image using the seasonally adjusted rendering color, the shadow mask, and the sky color to obtain the rendered target sparse remote sensing image. The trained sparse remote sensing image rendering model is obtained through the following operations: Preprocessing is performed on the remote sensing image samples to obtain low-resolution depth map samples, wherein the resolution of the low-resolution depth map samples is lower than that of the remote sensing image samples; The low-resolution depth map samples are scaled and offset corrected using a dense depth supervision module to obtain a multi-view consistent cumulative depth map, and the training process of the sparse remote sensing image rendering model is supervised using the multi-view consistent cumulative depth map. The sparse remote sensing image rendering model is used to process the remote sensing image samples to obtain rendering prediction parameters, and the seasonal change module is used to process the rendering prediction parameters to obtain shadow masking prediction information and seasonal adjustment prediction information. The training process of the sparse remote sensing image rendering model is supervised by RGB using the shadow masking prediction information and the seasonal adjustment prediction information. The parameters of the sparse remote sensing image rendering model are updated using the loss values from the deep supervision process and the RGB supervision process to obtain the trained sparse remote sensing image rendering model.
2. The method according to claim 1, characterized in that, Preprocessing the remote sensing image samples yields low-resolution depth map samples, including: The rational polynomial coefficients of the remote sensing image samples are optimized by performing bundle adjustment on the remote sensing image samples to obtain optimized remote sensing image samples. Multiple semi-global matching operations are performed independently on the optimized remote sensing image samples to obtain low-resolution depth map samples.
3. The method according to claim 1, characterized in that, The low-resolution depth map samples are scaled and offset corrected using a dense depth supervision module to obtain a multi-view consistent cumulative depth map, including: The line connecting the camera viewpoint and the low-resolution depth map sample is used as a given ray, and the cumulative depth prediction information of the low-resolution depth map sample is calculated under the given ray. The dense depth supervision module is used to perform scaling and offset correction on the low-resolution depth map samples to obtain scaling factor and offset factor. Based on the multi-view consistency constraint, the scaling factor, the offset factor, and the cumulative depth prediction information are calculated to obtain the multi-view consistent cumulative depth map.
4. The method according to claim 3, characterized in that, The training process of the sparse remote sensing image rendering model is performed using the multi-view consistent cumulative depth map as a depth supervision method, including: The standard deviation of the distribution of the low-resolution depth map sample along the given ray is obtained by calculating the transmittance of the sampled points on the given ray, based on the multi-view consistent cumulative depth map, the low-resolution depth map sample corresponding to the multi-view consistent cumulative depth map, and the transmittance of the sampled points on the given ray. The similarity between the preset normalized scaling factor, the preset offset factor, and the low-resolution depth map sample is calculated to obtain an equivalent uncertainty measure. The distribution standard deviation, the equivalent uncertainty measure, the multi-view consistent cumulative depth map, and the low-resolution depth map sample are calculated to obtain the depth prior loss information. The depth prior loss information is used to supervise the ray sampling process of the remote sensing image samples and to perform depth supervision on the training process of the sparse remote sensing image rendering model.
5. The method according to claim 1, characterized in that, The rendering prediction parameters are processed using the seasonal variation module to obtain shadow mask prediction information and seasonal adjustment prediction information, including: The volume rendering prediction contribution weight is determined using the predicted volume density in the rendering prediction parameters, wherein the volume rendering prediction contribution weight represents the color contribution of the upsampling point of the sampling ray corresponding to the remote sensing image sample to the remote sensing image sample. The seasonal variation module is used to activate the volume rendering prediction contribution weight and the predicted solar visibility in the rendering prediction parameters to obtain the shadow mask prediction information.
6. The method according to claim 5, characterized in that, Also includes: Obtain the seasonal category prediction probability corresponding to the remote sensing image sample, and obtain the expected seasonal adjustment prediction probability characterizing the remote sensing image sample belonging to different seasonal categories; The seasonal category prediction probability and the expected seasonal adjustment prediction probability are calculated to obtain the seasonal adjustment prediction term; The seasonal adjustment prediction term and the albedo color corresponding to the remote sensing image sample are nonlinearly activated to obtain the seasonally adjusted predicted albedo. The seasonal variation module is used to calculate the seasonally adjusted predicted albedo and the volume rendering prediction contribution weight to obtain the seasonally adjusted prediction information.
7. The method according to claim 1, characterized in that, The RGB supervision performed on the training process of the sparse remote sensing image rendering model using the shadow masking prediction information and the seasonal adjustment prediction information includes: The predicted rendering color is obtained by performing calculations on the shadow mask prediction information, the seasonal adjustment prediction information, and the sky prediction color in the rendering prediction parameters. The training process of the sparse remote sensing image rendering model is supervised by the predicted rendering color.
8. The method according to claim 1, characterized in that, The loss values in the RGB supervision process include the loss values of the sampled light corresponding to the remote sensing image sample and the loss values of sunlight.