Time-varying information decoupling method based on time feature guidance
Through the time-varying information decoupling method based on time characteristics, time-varying light and shadow and geometric information are decoupled, and the learning rules of space-time field function are optimized, improving the quality and accuracy of remote sensing image reconstruction.
Patent Information
- Application Number
- CN202510218081.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art fails to fully consider the correlation law between lighting elements and time, resulting in the negative impact of time-varying light and shadow information on geometric constraints, which in turn affects the quality of remote sensing image reconstruction.
The time-varying information decoupling method based on time characteristics is adopted, and the time-varying light and shadow information and geometric information are decoupled by constructing time-varying feature coding, integral space, multi-network learning and lighting intensity calculation, and the learning rules of space-time field function are optimized.
Effectively curb the negative impact of time-varying light and shadow information on geometric constraints, improving the quality and accuracy of remote sensing image reconstruction.
Smart Images

Figure CN120147527A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of remote sensing image reconstruction, and specifically to a time-varying information decoupling method guided by time features. Background Art
[0002] The acquisition times of the images that make up the super-generalized stereo image pair are inconsistent. Situations such as scene light and shadow changes, ground object movement, or geometric changes in different images will all cause different degrees of differences in the pixel values of corresponding image points between the images. This difference will cause different degrees of interference to the geometric constraints of the light samples, increasing the difficulty of time-varying three-dimensional reconstruction.
[0003] Due to differences in environmental factors such as scene illumination and shadow distribution, the pixel values in the shadow areas of the images are generally low. The traditional methods for the research and processing of shadow areas mainly focus on shadow detection and shadow removal, but this approach ignores the value of the shadow information itself. On the contrary, many NeRF methods attempt to utilize shadow information. In 2021, S-NeRF simulated the direct sunlight through a local light source visibility field, modeled the indirect illumination from the diffuse light source as a learnable global color field, as a function of the sun's position, to reduce the height and color errors in the shadow areas. Subsequently, Roger Marí et al. proposed Sat-NeRF based on S-NeRF, and for the first time introduced the rational polynomial imaging model corresponding to satellite images into NeRF, and processed the transient targets introduced due to multi-temporal phases according to the shadow-aware irradiance model and uncertainty weighting. Subsequently, based on Sat-NeRF, EO-NeRF introduced the SDF function to characterize the scene geometry, and inferred the shadows by learning the geometric shape to achieve the understanding of the shadows in the input images in three-dimensional space.
[0004] Although existing research considers the influence of light and shadow on geometric constraints, introduces light information, and decouples the scene into light elements such as light intensity, albedo, and sky diffuse light based on ray sample rendering, it has a certain positive guiding effect on the geometric constraints in the shadow-covered areas. However, the existing technologies do not fully consider the correlation law between the above light elements and time, and cannot suppress the negative impact of time-varying light and shadow information on geometric constraints, thus leading to the problem of poor quality of the reconstructed images. Summary of the Invention
[0005] The purpose of the present invention is: aiming at the problem that the existing methods cannot suppress the negative impact of time-varying light and shadow information on geometric constraints, which in turn leads to the problem of poor quality of the reconstructed images, to provide a time-varying information decoupling method guided by time features.
[0006] The technical solution adopted by the present invention to solve the above technical problems is:
[0007] Time-varying information decoupling method guided by time characteristics, comprising the following steps:
[0008] Step 1: Obtain a super-generalized stereo image pair, and randomly select an image I from the super-generalized stereo image pair i , and then obtain the observation time t i of the image I i , viewing direction v i and camera origin information, and obtain the time feature encoding Φ i of the observation time t i ;
[0009] Step 2: Randomly select a pixel point on the image I i , and then construct a ray passing through this pixel point along the viewing direction v i , perform importance sampling on this ray to obtain spatial sampling points, and finally, construct the integration space V i of the spatial sampling points;
[0010] Step 3: Obtain the coordinate p of the penetrated pixel point, perform position encoding on this coordinate, and then integrate the position-encoded coordinate p in the integration space V i to obtain the encoded integration result MSI(p);
[0011] Step 4: Input the encoded integration result MSI(p) into the MLP static network to obtain the output static geometric value g static ;
[0012] Step 5: Input the static geometric value g static and the time feature Φ i into the geometric field bias function MLP offset to obtain the output geometric value bias g offset ;
[0013] Step 6: Add the static geometric value g static and the geometric value bias g offset to obtain the time-varying signed distance function value g i at time t i ;
[0014] Step 7: Input the coordinate p, the time feature Φ i , the viewing direction v i and the time-varying signed distance function value g i at time t i into the MLP alb network to obtain the output radiance value r i ;
[0015] Step Eight: Repeat Steps Two to Seven to obtain the radiation value of each pixel, and perform volume rendering on the radiation value of each pixel along the light ray corresponding to the pixel. After the volume rendering of the radiation values of all pixels, obtain the albedo image A at time t i ; i ;
[0016] Step Nine: Perform geometric estimation on image I i to obtain the geometric entities in image I i ;
[0017] Step Ten: Randomly select a pixel on image I i , construct a light ray ray passing through the pixel based on the camera origin information, and obtain the intersection point of the light ray ray and the geometric entities in image I i , which is the object surface point C surf ;
[0018] Step Eleven: Read the parameter file of image I i to obtain the solar elevation angle, and obtain the solar direction v sun ,
[0019] Step Twelve: Generate a pseudo-light ray ray from the object surface point C surf , and project the pseudo-light ray ray fake towards the solar direction v fake to obtain the ray ray sun ; fake ;
[0020] Step Thirteen: Perform equally spaced sampling on the ray ray fake , and calculate the opacity of each sampling point on the ray ray fake . The opacity of all sampling points on the ray ray fake is the cumulative opacity, and the cumulative opacity is used as the light intensity S i at the observation time t i passing through this pixel in the viewing direction v i (pix);
[0021] Step Fourteen: Repeat Steps Ten to Thirteen to obtain the light intensity S i (pix) of each pixel on image I i , and then obtain the light intensity distribution map S i of image I i ;
[0022] Step Fifteen: Obtain the diffuse reflection information of image I i , that is, the illumination information other than direct sunlight, and encode the diffuse reflection information to obtain the illumination feature vector F sun;
[0023] Step Sixteen: Input the feature vector F sun , the sunlight direction v sun and the time feature Φ i into the ambient light estimation network MLP env to obtain the global ambient light variable E i at the output time t i ;
[0024] Step Seventeen: Based on the ambient light variable E i , the light intensity distribution map S i and the albedo image A i , construct the predicted image of the image I i
[0025] Furthermore, the time-varying signed distance function value g i at time t i is expressed as:
[0026] g i = g static + g offset
[0027] g static = MLP static (MSI(p))
[0028] g offset = MLP offset (g static , Φ i ).
[0029] Furthermore, the radiation value r i is expressed as:
[0030] r i = TVSRF(p, v i , Φ i , TVSDF(p, Φ i ))
[0031] Furthermore, the global ambient light variable E i at time t i is expressed as:
[0032] E i = MLP env (F sun , v sun , Φ i ).
[0033] Furthermore, the predicted image is expressed as:
[0034]
[0035] Furthermore, the time feature Φ i is encoded as:
[0036]
[0037] where represents the vector concatenation operator, respectively represent the normalization of the year, the date within the year, and the hour within the day, Y number represents the current year, Y min represents the minimum year of the input time series, Y max represents the maximum year of the input time series, M number represents the current month, M sum represents the total number of observable days within the month, D number represents the date within the corresponding month, H number represents the current hour, H sum represents the total number of observable hours within the day.
[0038] Furthermore, the integration space V i is expressed as:
[0039]
[0040] where V cube and V cone respectively represent the corresponding cubic and conical integration spaces, Ψ sat and Ψ air respectively represent satellite and airborne image sets.
[0041] Furthermore, the encoded integration result MSI(p) is expressed as:
[0042]
[0043] where V represents the integration space and dv represents the differential of the infinitesimal volume element.
[0044] Furthermore, the result of the position encoding pos_enc(p) is expressed as:
[0045] pos_enc(p) = [a 1 cos(2πω 1 p), a 1 sin(2πω 1 p), …, a n cos(2πω n p), a n sin(2πω n p)]
[0046] Among them, ω 1 ...ω n is the frequency coefficient, a 1 ...a n are the coding weights of different frequencies, and pos_enc(p) represents the result of the position encoding.
[0047] Furthermore, the light intensity S i (pix) is expressed as:
[0048]
[0049] Among them, δ i = ||p i+1 - p i || 2 δ i represents the spacing between adjacent sampling points, σ i represents the volume density value at the spatial point p i at the moment t i , a i represents an intermediate variable, and σ i is obtained by converting the time-varying signed distance function value g i through the following formula:
[0050] σ i = κ · σ(-g i )
[0051]
[0052] Among them, κ and λ represent learnable parameters.
[0053] The beneficial effects of the present invention are:
[0054] This application realizes the decoupling of time-varying light and shadow information and time-varying geometric information in the scene through the time-varying information decoupling technology guided by time characteristics, optimizes the learning rules of the spatio-temporal field function, and reduces the fitting difficulty of the spatio-temporal field function. Based on the time-varying geometric information decoupling strategy guided by temporal characteristics, this application drives the spatio-temporal field function to understand the imaging law before and after geometric changes, so that the spatio-temporal field function can dynamically fit different geometric forms at different times; a time-varying light and shadow information decoupling strategy based on the mining of time-varying light laws is proposed to strengthen the spatio-temporal field function's understanding of the imaging law under light and shadow changes, effectively suppressing the negative impact of time-varying light and shadow information on geometric constraints, and thus improving the quality of the reconstructed image. Description of the Drawings
[0055] Figure 1 is the flowchart of time-varying information decoupling;
[0056] Figure 2Flow chart of time-varying information decoupling guided by time features;
[0057] Figure 3 To construct a schematic diagram of the ray of light ray. Detailed implementation manners
[0058] It should be particularly noted that, without conflict, the various implementation manners disclosed in this application can be combined with each other.
[0059] Detailed implementation manner 1: The time-varying information decoupling method based on time feature guidance described in this implementation manner includes the following steps:
[0060] Step 1: Obtain a super-generalized stereo pair, and randomly select an image I from the super-generalized stereo pair i , and then obtain the observation time t i of the image I i , the viewing direction v i , and the camera origin information, and obtain the time feature encoding Φ i of the observation time t i ;
[0061] Step 2: Randomly select a pixel point on the image I i , and then construct a ray passing through this pixel point along the viewing direction v i , perform importance sampling on this ray to obtain spatial sampling points, and finally, construct the integration space V i of the spatial sampling points;
[0062] Step 3: Obtain the coordinates p of the penetrated pixel point, perform position encoding on this coordinate, and then integrate the position-encoded coordinates p in the integration space V i to obtain the encoded integration result MSI(p);
[0063] Step 4: Input the encoded integration result MSI(p) into the MLP static network to obtain the output static geometric value g static ;
[0064] Step 5: Input the static geometric value g static and the time feature Φ i into the geometric field bias function MLP offset to obtain the output geometric value bias g offset ;
[0065] Step 6: Add the static geometric value g static and the geometric value bias g offset to obtain the time-varying signed distance function value g i at time t i ;
[0066] Step Seven: Input the coordinate p, the time feature Φ i , the viewing direction v i and the value g of the time-varying signed distance function at time t i into the MLP i network to obtain the output radiance value r alb ; i
[0067] Step Eight: Repeat Steps Two to Seven to obtain the radiance value of each pixel, and perform volume rendering of the radiance value of each pixel along the ray corresponding to the pixel. After volume rendering the radiance values of all pixels, obtain the albedo image A i at time t i ;
[0068] Step Nine: Perform geometric estimation on the image I i to obtain the geometric entities in the image I i ;
[0069] Step Ten: Randomly select a pixel on the image I i , construct a ray ray passing through the pixel based on the camera origin information, and obtain the intersection point of the ray ray and the geometric entities in the image I i , that is, the object surface point C surf ;
[0070] Step Eleven: Read the parameter file of the image I i to obtain the solar elevation angle, and obtain the solar direction v sun according to the solar elevation angle
[0071] Step Twelve: Generate a pseudo-ray ray surf from the object surface point C fake , and project the pseudo-ray ray fake towards the solar direction v sun to obtain the ray ray fake ;
[0072] Step Thirteen: Perform equally spaced sampling on the ray ray fake , and calculate the opacity of each sampling point on the ray ray fake . The opacity of all sampling points on the ray ray fake , that is, the cumulative opacity, and use the cumulative opacity as the light intensity S i at the observation time t i passing through this pixel in the viewing direction v i (pix);
[0073] Step Fourteen: Repeat Steps Ten to Thirteen to obtain the image I i The illumination intensity S of each pixel point above i (pix), and then obtain the image I i The illumination intensity distribution map S of i ;
[0074] Step 15: Obtain the image I i The diffuse reflection information, that is, the illumination information that is not directly irradiated by the sun, and encode the diffuse reflection information to obtain the illumination feature vector F sun ;
[0075] Step 16: Input the feature vector F sun , the sunlight direction v sun and the time feature Φ i into the ambient light estimation network MLP env to obtain the global ambient light variable E at the output time t i ; i ;
[0076] Step 17: Based on the ambient light variable E i , the illumination intensity distribution map S i and the albedo image A i , construct the predicted image of the image I i
[0077] Super-generalized stereo pair: The combination of two or more optical remote sensing images with different observation perspectives obtained in any way covering the same scene is defined as a "super-generalized stereo pair". The composition form of the super-generalized stereo pair is free, and it can be a "standard stereo pair" or a "generalized stereo pair", or it can be general frontal or oblique remote sensing images. The number of images is not fixed, and it can be incrementally accumulated with the acquisition of new images to form an input increment; it can be obtained by different payload platforms at different times. There are unfixed differences in observation perspectives and resolutions among multiple images, and the intersection angle, base-height ratio, and overlap degree between two images do not necessarily meet the composition conditions of the "standard stereo pair" or the "generalized stereo pair". Once there are differences in observation perspectives among multiple optical remote sensing images, geometric constraints can be formed on the three-dimensional shape of the scene. Therefore, in theory, three-dimensional reconstruction of the scene can be achieved based on the super-generalized stereo pair. Compared with the standard or generalized stereo pair, the main advantage of the "super-generalized stereo pair" is reflected in the high-frequency and multi-perspective acquisition of image data, and it has the necessary data conditions to realize the three-dimensional reconstruction of the time-varying scene from "static" to "temporal", which better meets the requirements of real-scene three-dimensional.
[0078] This application makes full use of prior conditions such as time and ambient light, combines physical imaging principles, designs a time-varying information decoupling rule and its network architecture for scene light and shadow changes and geometric changes, strengthens the perception ability of the spatio-temporal field function to the physical imaging laws under light and shadow changes and geometric changes, and drives it to understand the light and shadow imaging laws and geometric imaging laws during the learning process. Thus, it effectively suppresses the negative impact of time-varying light and shadow information on geometric constraints, accurately perceives various geometric forms of the time-varying scene, and helps improve the reliability of spatio-temporal field function fitting in the case of light and shadow changes and geometric changes in the ultra-generalized stereo pair.
[0079] This application studies the time-varying information decoupling method of the ultra-generalized stereo pair, optimizes the learning rule of the spatio-temporal field function, and improves the interpretability while simplifying the learning difficulty. As Figure 1 shown, this application makes full use of prior conditions such as time and ambient light, combines physical imaging principles, designs a time-varying information decoupling rule and its network architecture for scene light and shadow changes and geometric changes, strengthens the perception ability of the spatio-temporal field function to the physical imaging laws under light and shadow changes and geometric changes, and drives it to understand the light and shadow imaging laws and geometric imaging laws during the learning process. Thus, it effectively suppresses the negative impact of time-varying light and shadow information on geometric constraints, accurately perceives various geometric forms of the time-varying scene, and helps improve the reliability of spatio-temporal field function fitting in the case of light and shadow changes and geometric changes in the ultra-generalized stereo pair.
[0080] This application proposes a time-varying information decoupling technology guided by time features to achieve the decoupling of time-varying light and shadow information and time-varying geometric information in the scene, optimizes the learning rule of the spatio-temporal field function, and reduces the fitting difficulty of the spatio-temporal field function. This application proposes a time-varying geometric information decoupling strategy guided by temporal feature to drive the spatio-temporal field function to understand the imaging laws before and after geometric changes, so that the spatio-temporal field function can dynamically fit different geometric forms at different times; this application proposes a time-varying light and shadow information decoupling strategy based on the excavation of time-varying light laws to strengthen the spatio-temporal field function's understanding of the imaging laws under light and shadow changes, and effectively suppress the negative impact of time-varying light and shadow information on geometric constraints.
[0081] The key technology of this application is the time-varying information decoupling technology guided by time features. This technology aims to design a decoupling mechanism for time-varying light and shadow information and time-varying geometric information in the scene, and optimize the learning rule of the spatio-temporal field function. The overall technical framework is as Figure 2 shown, mainly including a time-varying geometric information decoupling strategy guided by temporal features and a light and shadow time-varying information decoupling strategy based on the excavation of time-varying light laws.
[0082] (1) Time-varying geometric information decoupling strategy guided by temporal features
[0083] This strategy reduces the fitting difficulty of a single network by decoupling a complex single task into multiple subtasks and letting multiple networks learn the subtasks separately.
[0084] For the time-varying geometric function of the spatio-temporal field, this strategy decouples the learning task into static geometric value learning and geometric bias value learning, and assigns the learning of static geometry and geometric bias to two networks, MLP static and MLP offset respectively. The running logic of the geometric part of the spatio-temporal field function is as follows: Suppose an input ray along the viewing direction v i is given, and the position coordinates p of each sampling point on this ray. After obtaining the result MSI(p) of its multi-mode spatial integral encoding. First, input MSI(p) into MLP static to train and obtain the static geometric value g static , as shown in Equation (1). Then, input the time feature Φ i corresponding to a certain moment obtained by time encoding, together with the static geometric value g static into the geometric field bias function MLP offset to obtain the geometric value bias g offset , as shown in Equation (2). Finally, add the static geometric value and the geometric bias value to obtain the time-varying signed distance function value g i at time t i , as shown in Equation (3).
[0085] g static = MLP static (MSI(p)) (1)
[0086] g offset = MLP offset (g static , Φ i ) (2)
[0087] g i = g static + g offset (3)
[0088] The learning of the time-varying radiation function of the spatio-temporal field is separately responsible for by another network MLP alb . This network directly accepts the position information p of the sampling point, the time feature Φ i , the viewing direction information v i and the time-varying signed distance function value g i at time t i , and outputs the radiation value r i observed at position coordinate p in space at time t i along the viewing direction v i :
[0089] r i = TVSRF(p, v i , Φ i , TVSDF(p, Φ i )) (4)
[0090] Traverse all pixels of the input image, construct a ray sample from each pixel, and perform volume rendering along the ray according to the output of the time-varying radiation function, then the albedo image A corresponding to the input image at time t can be obtained. i at time t i .
[0091] Send the set of sampling points P into the MLP network to obtain TVSDF.
[0092] The purpose of spatial integration is to more accurately encode the three-dimensional information of spatial points. By constructing an integration space that conforms to the imaging characteristics of the image, the traditional single-scale position encoding is extended to scale-adaptive spatial integration encoding, thereby expanding the neural network's perception ability of spatial scales and reducing the geometric constraint deviation caused by spatial resolution differences. First, generate a ray passing through a certain pixel according to the spatio-temporal field imaging geometric model; second, perform ray sampling to obtain spatial sampling points and importance sampling. Then, according to the imaging payload platform and spatial resolution of the input image I i , construct the integration space V of the spatial sampling points in a multi-mode adaptive manner i .
[0093] When constructing the integration space, considering the differences in the imaging models and imaging distances of spaceborne images and airborne images, different spatial integration construction modes are adopted for the two respectively, so that the ray samples generated from images with different spatial resolutions can adapt to the spatial scale. In terms of the shape of the integration space, for spaceborne images, the satellite observation point is far from the observed ground surface and the field of view angle is narrow, so a cuboid-shaped integration space is adopted; while for airborne images compared with spaceborne images, the observation point is much closer to the observed ground object and the field of view angle is also larger, so a cone-shaped integration space is adopted.
[0094]
[0095] Among them, V cube , V cone respectively represent the corresponding cubic and conical integration spaces, and Ψ sat , Ψ air respectively represent the sets of spaceborne and airborne images.
[0096] In terms of the size of the integration space, the size of the integration space is determined according to the input image type, the corresponding spatial resolution of the image, and the interval between adjacent sampling points. For high-spatial-resolution images, the corresponding integration space V should be small, while for low-spatial-resolution images, the corresponding integration space V should be large to achieve the unification of the actual imaging three-dimensional space range.
[0097] Then, position encoding is performed on the input coordinate p = (x, y, z). The purpose of position encoding is to improve the ability of the neural network to distinguish high- and low-frequency signals. The encoding process is shown in Equation (6). Finally, integration is performed on the encoding result in space V, as shown in the following equation:
[0098] pos_enc(p) = [a 1 cos(2πω 1 p), a 1 sin(2πω 1 p), …, a n cos(2πω n p), a n sin(2πω n p)] (6)
[0099]
[0100] Among them, ω 1 ... ω n is the frequency coefficient, a 1 ... a n are the encoding weights of different frequencies, and pos_enc(p) represents the result of position encoding. Finally, according to the payload platform and spatial resolution of the input image, the result of position encoding is integrated in the space of the corresponding mode, that is, the Multimodal Spatial Integration (MSI) encoding result MSI(p) is obtained. The multimodal spatial integration encoding can adjust the encoding result to different degrees according to the actual covered spatial range of the spatial sampling points, thereby endowing the spatio-temporal field function with the perception ability of the spatial scale.
[0101] (2) Decoupling strategy for time-varying light and shadow information based on mining of light-time variation laws
[0102] The decoupling strategy for time-varying light and shadow information greatly reduces the learning difficulty of the network by splitting the complex and unruly irradiance learning task into a light intensity calculation task based on physical reasoning and an ambient light estimation task based on feature guidance.
[0103] The light intensity calculation module based on physical reasoning infers the light intensity at time t i from the time-varying geometric field. Specifically, assuming that the input image I is giveni a certain pixel pix in it, and a ray of light ray passing through this pixel. First, this module estimates the scene geometry according to the time-varying geometric function TVSDF at time t i , and searches for the intersection point of the ray and the geometric entity along the ray ray, that is, the object surface point C surf . Subsequently, a pseudo-ray ray surf is generated from C fake and projected towards the sun light source, and the probability that the pseudo-ray ray fake hits an object during the projection towards the sun light source is calculated, that is, the cumulative opacity S i (pix). And S i (pix) is the light intensity observed through the pixel pix in the viewing direction v i at time t i . Therefore, by repeating the above steps for all pixels of the input image I i , a prediction of the light intensity distribution map S i corresponding to the input image I i can be obtained.
[0104] After searching for the intersection point of the ray and the geometric entity along the ray ray, that is, the object surface point C surf (that is, Figure 3 x_s in fake ), a ray ray Figure 3 (that is, r_sun in 1 ) emitted from this point and pointing towards the sun light source is generated, and then equally spaced sampling is performed along this ray, and a set of sampling points P = {p n} is obtained. Then, the cumulative opacity is estimated from the sampling points and used as the light intensity S i observed through the pixel pix in the viewing direction v i at time t i (pix):
[0105]
[0106] where δ i = ||p i+1 - p i || 2 represents the spacing between adjacent sampling points, and σ i is the volume density value at the spatial point p i at time t i , which is obtained by converting the time-varying signed distance function value g i through the following formula:
[0107] σ i = κ·σ(-g i ) (10)
[0108]
[0109] Among them, κ and λ are learnable parameters, and σ(-g i ) is actually the cumulative distribution function based on the Laplace distribution, with a mean of 0 and a scale parameter of λ.
[0110] The feature-guided ambient light estimation module estimates the ambient light at the imaging time by fully exploiting the associations among the image content, imaging time, and solar illumination direction.
[0111] Specifically, the module first uses the feature network to obtain the illumination feature vector F sun from the image, and reads the image parameter file to obtain the solar elevation angle. Then, it calculates the solar light direction v sun , and inputs the feature vector F sun , the solar light direction v sun , and the time feature Φ i into the ambient light estimation network MLP env to obtain the global ambient light variable E i at time t i , as shown in Equation (12).
[0112] E i = MLP env (F sun , v sun , Φ i ) (12)
[0113] Based on the ambient light E i , the light intensity distribution map S i , and the albedo image A i of the object, the prediction i of the input image I can be constructed as shown in Equation (13) for fitting the network through pixel consistency constraints. Among them, (S i + E i (1 - S i )) is the output of the irradiance calculation module.
[0114]
[0115] Time encoding module: A time feature extraction method based on phase shift time encoding (TSE) is proposed, as shown in Equation (14), which introduces the time dimension and gives the observation time series corresponding to the input image in the order of year, month, day, and hour.
[0116]
[0117] In the formula Indicates a vector concatenation operator, which can concatenate multiple individual vectors into a one-dimensional vector. Respectively represent the normalization of the year, the date within the year, and the hour within the day. Among them, Y number Represents the current year, Y min Represents the minimum year of the input time series, Y max Represents the maximum year of the input time series; M number Is the current month, M sum Is the total number of observable days within a month, which is fixed and consistent for each month after being set, D number Corresponds to the date within the month; H number Is the current hour, H sum Is the total number of observable hours within a day, which is fixed and consistent for each day after being set, H sum Is fixed and consistent.
[0118] In view of the periodic law of time units, the sin(·) function and its phase-shifted form cos(·) are used to extract the time features in the time feature function. The time encoding function fully considers the irradiance difference of solar illumination within a day for encoding the intra-day time information, while fully considering the periodicity of seasonal changes for encoding the intra-year time information. Time features are helpful for the radiation field to output more accurate TVSRF values according to the illumination characteristics at different times during the fitting process of the spatio-temporal field function, thereby optimizing the learning rules and reducing the fitting difficulty of the spatio-temporal field function.
[0119] It should be noted that the specific implementation manners are only explanations and illustrations of the technical solutions of the present invention, and the scope of the right protection cannot be limited thereby. All those that are only partial changes made according to the claims and the specification of the present invention should still fall within the protection scope of the present invention.
Claims
1. A time-varying information decoupling method based on time feature guidance, characterized by The following steps are involved: Step 1: Obtain a super-generalized stereo pair and randomly select an image I from the super-generalized stereo pair i , then obtain image I i The observation time t i , viewing direction v i And the camera origin information, and obtain the observation time t i The temporal feature encoding Φ i ; Step 2: In image I i A pixel point is randomly selected on the image, and then a pixel along the viewing direction v is constructed. i The light that penetrates the pixel point is sampled and the importance of the light is sampled to obtain the spatial sampling point. Finally, the integral space V of the spatial sampling point is constructed. i ; Step 3: Get the coordinates p of the penetrated pixel point and encode the position of the coordinates, then in the integral space V i Integrate the position-encoded coordinate p to obtain the encoding integral result MSI(p); Step 4: Input the encoded integral result MSI(p) into MLP static Network, get the output static geometry value g static ; Step 5: Set the static geometry value g static and time characteristics Φ i Input geometry field bias function MLP offset In the output, the geometric value bias g is obtained offset ; Step 6: Set the static geometry value g static and the geometric value bias g offset Add together to get t i The time-varying signed distance function value g at time i ; Step 7: Coordinate p, time feature Φ i , viewing direction v i and t i The time-varying signed distance function value g at time i Enter MLP alb Network, get the output radiation value r i ; Step 8: Repeat steps 2 to 7 to obtain the radiation value of each pixel, and perform volume rendering on the radiation value of each pixel along the light corresponding to the pixel. After volume rendering of the radiation values of all pixels, t i Albedo image A at time i ; Step 9: Image I i Perform geometric estimation and obtain image I i The geometric entities in Step 10: In Image I i A pixel point is randomly selected, and a ray that penetrates the pixel point is constructed based on the camera origin information, and the ray and image I are obtained. i The intersection point of the geometric entity, that is, point C on the surface of the object surf ; Step 11: Read image I i Parameter file, get the sunlight pitch angle, and get the sunlight direction v according to the sunlight pitch angle sun , Step 12: From point C on the surface of the object surf Creating fake rays fake , and the pseudo ray ray fake Towards the sunlight v sun Projection, get ray fake ; Step 13: In the ray fake Perform equal-interval sampling on the fake The opacity of each sampling point, ray fake The opacity of all sampling points on the surface is the cumulative opacity, and the cumulative opacity is taken as the opacity at the observation time t i , from the viewing direction v i The light intensity S passing through the pixel i (pix); Step 14: Repeat steps 10 to 13 to obtain image I i The light intensity S of each pixel i (pix), and then get image I i Light intensity distribution diagram S i ; Step 15: Get image I i The diffuse reflection information, that is, the illumination information that is not directly illuminated by the sun, is encoded to obtain the illumination feature vector F sun ; Step 16: Transform the eigenvector F sun 、Direction of sunlight v sun And the time characteristic Φ i Input to the ambient light estimation network MLP env In the output, we get t i The global ambient light variable E at the moment i ; Step 17: Based on the ambient light variable E i , light intensity distribution diagram S i and albedo image A i , construct image I i The predicted image 2. The time-varying information decoupling method based on time feature guidance according to claim 1 is characterized in that The i The time-varying signed distance function value g at time i It is expressed as: g i =g static +g offset g static =MLP static (MSI(p)) G offset =MLP offset (G static ,Φ i )。 3. The time-varying information decoupling method based on time feature guidance according to claim 2 is characterized in that The radiation value r i It is expressed as: r i =TVSRF(p,v i ,Φ i ,TVSDF(p,Φ i ))。 4. The time-varying information decoupling method based on time feature guidance according to claim 3 is characterized in that The i The global ambient light variable E at the moment i It is expressed as: E i =MLP env (F sun ,v sun ,Φ i )。 5. The time-varying information decoupling method based on time feature guidance according to claim 4 is characterized in that The predicted image It is expressed as:
6. The time-varying information decoupling method based on time feature guidance according to claim 5 is characterized in that The time characteristic Φ i The encoding is represented as: in, Represents a vector connector. Represents the normalization of year, day of the year, and hour of the day, respectively. number Indicates the current year, Y min Indicates the minimum year of the input time series, Y max Indicates the maximum year of the input time series, M number Indicates the current month, M sum represents the total number of observable days in a month, D number Indicates the corresponding date within the month, H number Indicates the current hour, H sum Represents the total number of observable hours in a day.
7. The time-varying information decoupling method based on time feature guidance according to claim 6 is characterized in that The integration space V i It is expressed as: Among them, V cube 、V cone denote the corresponding cubic and pyramidal integral spaces, Ψ sat , air Represent satellite and airborne image collections respectively.
8. The time-varying information decoupling method based on time feature guidance according to claim 7 is characterized in that The coding integration result MSI(p) is expressed as: Among them, V represents the integration space, and dv represents the differential of the small volume element.
9. The time-varying information decoupling method based on time feature guidance according to claim 8 is characterized in that The position encoding result pos_enc(p) is expressed as: pos_enc(p)=[a1 cos(2πω1p),a1 sin(2πω1p),…,a n cos(2πω n p),a n sin(2πω n p)] Among them, ω1...ω n is the frequency coefficient, a1...a n are the encoding weights of different frequencies, and pos_enc(p) represents the result of position encoding.
10. The time-varying information decoupling method based on time feature guidance according to claim 9 is characterized in that The light intensity S i (pix) is represented as: Among them, δ i =||p i+1 -p i ||2,δ i represents the distance between adjacent sampling points, σ i Indicates t i At the point p in space i The volume density value at a i represents the intermediate variable, σ i From the time-varying signed distance function value g i By converting the following formula: s i =κ·σ(-g i ) Among them, κ and λ represent learnable parameters.