Self-supervised space-time field reconstruction method based on multi-mode spatial integral coding and spectral band optimization, storage medium and equipment

Through a self-supervised space-time field reconstruction method based on multimodal spatial integral coding and spectrum segment optimization, the space-time field is characterized and optimized by using the time-varying radiation function and the time-varying symbol distance function, the problem of the difference in the resolution of the hollow spectrum radiation of supergeneous stereoscopic images for time-varying scenes is solved, and the efficient three-dimensional reconstruction effect is achieved.

CN120147526AActive Publication Date: 2025-06-13THE INST OF AUTOMATION HEILONGJIANG ACADEMY OF SCI
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510218074.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-13
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

The three-dimensional reconstruction of time-varying scenes based on ultra-generalized stereo image pairs faces severe challenges such as the complexity of scene space-time representation, the imaging difference of multi-source images, the imaging difference of time-varying scenes, and the complexity of image incremental methods. In particular, the difference in the resolution of space spectral radiation affects the three-dimensional reconstruction effect.

Method used

A self-supervised space-time field reconstruction method based on multimodal spatial integral coding and spectrum segment preference is adopted to characterize the space-time field through time-varying radiation function and time-varying symbol distance function, and to realize the space-time field reconstruction by optimizing these functions. The method includes constructing light emitted by the camera and passing through the image pixels, constructing the integral space based on the spatial sampling points on the light, performing position coding and integrating, generating multi-mode spatial integral coding results, and optimizing the time-varying radiation function and the time-varying symbol distance function through volume rendering and self-supervised spectral loss.

Benefits of technology

The impact of the difference in spatial spectrum resolution on the three-dimensional reconstruction of time-varying scenes is effectively solved, the effect and accuracy of three-dimensional reconstruction is improved, and the accurate reconstruction of the space-time field is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147526A_ABST
    Figure CN120147526A_ABST
Patent Text Reader

Abstract

The invention discloses a self-supervised space-time field reconstruction method based on multi-mode space integral coding and spectral band optimization, a storage medium and equipment, and belongs to the technical field of remote sensing three-dimensional reconstruction. The objective of the invention is to solve the problem that the time-varying scene three-dimensional reconstruction effect based on a super-generalized stereopair is affected by the spatial-spectral-radiance resolution difference. According to the method, light is generated through each image of a super-generalized stereopair and a space-time field imaging geometric model of the super-generalized stereopair, basic self-supervision information is given to the light on the basis of image pixels, light samples are formed, the pixels with different spatial resolutions are bound with the actual constraint space range of the pixels through the spatial integration strategy of the light, and a super-generalized stereopair image is obtained. Meanwhile, a multi-channel joint self-supervision strategy is optimally designed, and a supporting condition is provided for establishing geometric constraints based on light sample joint with different wave band numbers and numerical precision, so that a function for supervising a space-time field is developed, and space-time field reconstruction is finally realized based on the space-time field function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of remote sensing three-dimensional reconstruction, and particularly relates to a self-supervised spatio-temporal field reconstruction method, a storage medium, and a device. Background Art

[0002] Remote sensing three-dimensional reconstruction technology is the main way to obtain three-dimensional information of scenes on a large scale. Compared with traditional remote sensing three-dimensional reconstruction, time-varying scene three-dimensional reconstruction is not for three-dimensional reconstruction at a certain point in time, but to reconstruct the different geometric forms of the scene at different time periods and the geometric change relationships between different time periods, which is an important means to obtain three-dimensional information of time-varying scenes.

[0003] Ultra-generalized stereo pair: The combination of two or more optical remote sensing images with different observation perspectives obtained in any way covering the same scene is defined as an "ultra-generalized stereo pair". The composition form of the ultra-generalized stereo pair is free, and it can be a "standard stereo pair" or a "generalized stereo pair", or it can be general front-view or oblique remote sensing images. The number of images is not fixed, and it can be incrementally accumulated with the shooting of new images to form an input increment; it can be obtained by different payload platforms at different time phases. There are unfixed differences in observation perspectives and resolutions between multiple images, and the intersection angles, base-height ratios, overlap degrees, etc. between two images do not necessarily meet the composition conditions of the "standard stereo pair" or the "generalized stereo pair". Once there are differences in observation perspectives between multiple optical remote sensing images, geometric constraints can be formed on the three-dimensional shape of the scene. Therefore, in theory, three-dimensional reconstruction of the scene can be achieved based on the ultra-generalized stereo pair. Compared with the standard or generalized stereo pair, the main advantage of the "ultra-generalized stereo pair" is reflected in the high-frequency and multi-perspective acquisition of image data, which provides the necessary data conditions for realizing three-dimensional reconstruction of time-varying scenes from "static" to "temporal", and better meets the requirements of real-scene three-dimensional.

[0004] However, the three-dimensional reconstruction of time-varying scenes based on the "ultra-generalized stereo pair" needs to face severe challenges in many aspects such as the complexity of scene spatio-temporal representation, the imaging differences of multi-source images, the imaging differences of time-varying scenes, and the complexity of image increment methods. And the three-dimensional reconstruction of time-varying scenes based on the ultra-generalized stereo pair needs to face more complex problems formed by the coupling of multiple challenges, which is a technical bottleneck that is difficult to break through currently.

[0005] Among them, the imaging differences of multi-source images are mainly due to certain differences in the imaging geometric models, spatial resolutions, spectral resolutions, and radiation resolutions of spaceborne and airborne images from different sources, which greatly affect the reliability of three-dimensional reconstruction based on the ultra-generalized stereo pair.

[0006] First, there are differences in imaging geometric models in the images:

[0007] For the imaging geometric model of spaceborne images, the Rational Function Model (RFM) has been widely used as a substitute for the strict spaceborne imaging geometric model in most application scenarios. It describes the mapping relationship between the row and column coordinates of image pixels and the three-dimensional coordinates of longitude, latitude, and altitude with a set of Rational Polynomial Coefficients (RPC). It has many advantages such as high computational efficiency, confidentiality of geometric parameters of satellites and sensors, and no need for iteration in coordinate back-calculation. Currently, the sensor calibration products released by high-resolution satellites at home and abroad all come with RPC-related files. However, the generation methods and accuracies of RPCs for different satellites vary, resulting in uncertain coordinate mapping deviations in the imaging geometric model. It is difficult to directly establish accurate joint geometric constraints based on multiple images from different satellites, which further leads to low accuracy in scene three-dimensional reconstruction or even large overall morphological deviations. For the imaging geometric model of airborne images, the imaging geometric model of airborne images is generally considered to satisfy the perspective projection model, also known as the pinhole model. On the one hand, there are differences in the processes of different lenses, and there are inherent differences in the accuracy of the imaging geometric model. On the other hand, the parameters of the imaging geometric model of airborne images are usually unknown, and the approximate imaging geometric models estimated by different types of methods have different degrees of errors.

[0008] It can be seen that for ultra-generalized stereo image pairs, whether for spaceborne or airborne images, there are certain differences between the imaging geometric models of sensors, and it is difficult to ensure the consistency of geometric constraints in the three-dimensional space. On the one hand, the parameter representation forms of the imaging geometric models are not unified. For example, the parameter representation forms and quantities of the RFM model and the pinhole model are inconsistent. On the other hand, the accuracies of the imaging geometric models of different images are inconsistent, resulting in spatial mapping deviations. Therefore, to effectively establish geometric constraints for ultra-generalized stereo image pairs in the spatio-temporal field and fit the spatio-temporal field function of the scene, it is necessary to reconstruct the imaging geometric model of the image in the spatio-temporal field coordinate system, standardize its parameter representation form, form the spatio-temporal field imaging geometric model, and at the same time, the model parameters should be further optimized based on the spatial geometric constraint relationship to improve the spatial mapping accuracy.

[0009] Second, there are differences in spatial resolution of images:

[0010] In the early days of defining generalized stereo image pairs in the field of remote sensing, the differences in spatial resolution were the key issues of concern. The main research focused on spaceborne images, and the technical framework adopted was still based on traditional stereo matching methods. The common idea was to make multiple images reach a unified resolution through sampling or super-resolution strategies, or to set some weight strategies for spatial resolution and then perform stereo matching. Typical works include: the Spanish scholar Manuel Aguilar collected multiple sets of GeoEye-1 and WorldView-2 satellite data and carried out research on the three-dimensional coordinate calculation problem of generalized stereo image pairs. Considering image characteristics such as radiation and illumination, the relationship between three-dimensional reconstruction accuracy and various factors such as ground control points, intersection angles between image pairs, base-height ratios, and resolutions was comprehensively analyzed. Professor Gu Yanjianfeng from Harbin Institute of Technology proposed adding a scale difference factor to the traditional calculation model to reduce the influence of resolution differences on the accuracy of three-dimensional coordinate calculation. However, most related studies avoided the problem of dense stereo matching and only focused on the calculation accuracy of a few manually selected matching points, and could not achieve three-dimensional reconstruction of the complete scene.

[0011] The actual spatial ranges covered by pixels of images with different spatial resolutions are different when imaging. When there are differences in image spatial resolution, the light from the low-resolution image actually covers a larger spatial range, while the light from the high-resolution image actually covers a smaller spatial range. The NeRF series of methods model the relationship between light samples and image pixels as a one-to-one correspondence. Unfortunately, most of the NeRF series of methods pay more attention to natural scene images and do not attach importance to distinguishing the actual spatial scales corresponding to pixels. A single scale is used for all pixels, resulting in inaccurate actual spatial ranges corresponding to the light collected according to images with different spatial resolutions, and causing errors in the spatial information of the generated light samples. MipNeRF improved the aliasing problem caused by image spatial resolution, making pixels no longer correspond to a single spatial position but to a spatial frustum. However, its processing method is only limited to the perspective pinhole model for close views and is not suitable for remote sensing images. For ultra-generalized stereo image pairs, due to the different actual spatial scales corresponding to the image pixels with different spatial resolutions, the existing light sample generation methods based on a single scale processing method will cause geometric constraint errors of varying degrees; while the processing method based on the spatial frustum is difficult to uniformly process images with different observation distances and imaging methods in the remote sensing field. Therefore, there is an urgent need for a light sample generation strategy with adaptive spatial resolution.

[0012] Third, there are differences in spectral and radiation resolutions in the images:

[0013] Research in the remote sensing field on spectral resolution differences mainly focuses on directions such as image super-resolution and fusion. According to research, there is currently no research on joint three-dimensional reconstruction based on visible light, multi-spectral, and hyperspectral images. There are currently two main methods for three-dimensional reconstruction based on hyperspectral images. The first is the method of combining lidar. Harbin Institute of Technology

[23] The research group led by Professor Gu Yanjun integrated hyperspectral images and LiDAR point clouds to generate three-dimensional hyperspectral point clouds for three-dimensional scene reconstruction. The second method is to directly utilize multiple hyperspectral images. Ali Zia et al. used VisualSFM software to establish different three-dimensional models of an object at each wavelength and then merge them to complement each other in terms of structure. Ali Can Karaca et al. from Kocaeli University in Turkey used two parallel hyperspectral cameras and proposed the SegmentSpectra algorithm to achieve a reconstruction accuracy comparable to that of LiDAR point clouds. The above research on three-dimensional reconstruction methods based on multi-spectral images does not handle the case where there are differences in spectral resolution. In fact, the images that make up a super-generalized stereo pair may be of various types such as panchromatic, RGB, multi-spectral, and hyperspectral images, and there may be differences in the range and number of their imaging bands. Such differences will cause feature differences between images, resulting in unsatisfactory traditional stereo matching effects. At the same time, since NeRF series methods obtain self-supervised information by the pixel values of the corresponding images for each ray, the difference in the number of spectral bands leads to a dimensional difference in the self-supervised information of the rays generated by images with different numbers of bands, making it difficult to construct joint geometric constraints. Although common remote sensing images can have single-band, 3-band, 4-band, 8-band, or even hundreds of bands, the number of bands in an image is not the key information determining the geometric constraint ability. The viewing angle of the image is more important for geometric constraints. Involving too many bands does not necessarily improve the geometric constraint ability and may even waste computing resources. Therefore, adaptive self-supervised information dimensional processing is the prerequisite for establishing effective geometric constraints based on the ray self-supervised learning method. As shown in Table 1, different bands have different application scenarios. Therefore, for three-dimensional reconstruction of a specific scene, appropriate band selection can be carried out to limit the impact of dimensional differences.

[0014] In addition, the radiometric resolution of remote sensing images determines the precision bit width of each pixel value, and remote sensing images with 10-bit to 12-bit are quite common. The NeRF-based neural implicit representation series methods construct a loss function based on the accumulation of rendering errors for each ray. The difference in radiometric resolution will cause differences in the precision of the self-supervised information of the rays, resulting in a loss preference for pixel values of a certain bit width during the iterative process, which is not conducive to joint geometric constraints between rays and is even less conducive to the accurate fitting of the spatio-temporal field function. Therefore, it is necessary to normalize the bit width of the input image to avoid the occurrence of loss preference.

[0015] Table 1 Characteristics of common bands in the field of remote sensing technology

[0016]

[0017] In summary, the differences in image spectrum and radiation resolution result in inconsistent dimensions and accuracies of pixel values at the same spatial position in different types of images, making it difficult to backpropagate errors through a common neural network structure and prone to causing biases in joint self-supervised learning. Therefore, it is necessary to study self-supervised learning strategies that adapt to spectral and radiation resolution differences or introduce other auxiliary geometric constraint rules. Summary of the Invention

[0018] The present invention aims to solve the problem that the three-dimensional reconstruction effect of a time-varying scene based on a super-generalized stereo pair is affected by the differences in spatial, spectral, and radiation resolutions.

[0019] A self-supervised spatio-temporal field reconstruction method based on multi-mode spatial integral coding and spectral band optimization, which characterizes the spatio-temporal field based on the time-varying radiation function TVSRF and the time-varying signed distance function TVSRF, and realizes spatio-temporal field reconstruction through the optimized time-varying radiation function TVSRF and the time-varying signed distance function TVSDF. The optimization processes of the time-varying radiation function TVSRF and the time-varying signed distance function TVSDF include:

[0020] Construct the light rays emitted by the camera and passing through the image pixels according to the spatio-temporal field imaging geometric model; construct the integral space V of the spatial sampling points based on the spatial sampling points i on the light rays i ;

[0021] Perform position encoding on the coordinates p = (x, y, z) in the space to obtain the encoding result pos_enc(p); perform integration on the encoding result in the integral space V to obtain the multi-mode spatial integral coding result MSI(p);

[0022] Convert the pixel value I(i, j) in the image to m-bit numerical precision to obtain I * (i, j), then round the conversion result I * (i, j), convert it to m-bit binary representation; subsequently divide the obtained I * (i, j) by 2 m to obtain I'(i, j);

[0023] For a band set Band out = {B, G, R, IR, Pan}, perform volume rendering along the light ray according to the output of the time-varying radiation function TVSRF, including the following steps:

[0024] Given a certain pixel pix and the light ray Ray = {r(d) = o + dv, d ≥ 0} passing through this pixel, and the depth range [d n , t f of the reconstructed target scene, first sample along the ray within [d n , d f and obtain a set of sampling points P = {pi |p i = r(d i ), i = 1, 2…, N}; where v is the unit direction vector of the light ray, d is the depth along the light ray starting from the camera origin o, and N represents the number of sampling points;

[0025] Then, the volume rendering result corresponding to the pixel through which the light ray passes is obtained according to the following formula:

[0026]

[0027] w(p i ) = T(p i )a(p i ),

[0028] δ(p i ) = d i+1 - d i

[0029] σ(p i ) = κ·σ’(-TVSDF(MSI(p i ), t i ))

[0030] where δ(p i ) represents the depth interval between adjacent sampling points, a(p i ) is the opacity along the ray segment r(d i+1 ) - r(d i ), T(p i ) is the cumulative transparency when the light ray starts from the origin o and reaches the spatial point p i , w(p i ) is the corresponding volume rendering integration weight at the spatial point p i ; σ(p i ) is the volume density at the spatial point p i at time t i ; k, λ are learnable parameters, and σ’(x) is the cumulative distribution function based on the Laplace distribution, with a mean of 0 and a scale parameter of λ; TVSRF uses a neural network model, which obtains i and t i based on the input MSI(p are the radiation values corresponding to the respective bands, and TVSRF(MSI(p i ), t i ) is at the spatial point p i at time t iOutput result of the time-varying radiation function; the time-varying signed distance function TVSDF uses a neural network model, which is based on the input MSI(p i ) and t i to obtain the signed distance value, which refers to the minimum value of the coordinate distance between the spatial point coordinate p and a point on the surface of the ground object. TVSDF(MSI(p i ),t i ) is the output result of the time-varying signed distance function at the spatial point p i at time t i ;

[0031] Assume that the set of bands included in the input image I i corresponding to the pixel is Band in . The channel preference mechanism will also construct a 5D vector R in =(r B ,r G ,r R ,r IR ,r Pan ) corresponding to the pixel. The value r k of one dimension is the I'(i,j) value corresponding to the corresponding band, k∈{B,G,R,IR,Pan}; r k satisfies the rule: r k =0, if

[0032] Construct a self-supervised spectral loss in N = ||R out || in at the pixel level by R 0 , where ||·|| 1 and ||·|| 0 correspond to the 1-norm and 0-norm respectively; cube Optimize the time-varying radiation function TVSRF and the time-varying signed distance function TVSDF according to the spectral loss.

[0033] Furthermore, the integral spaces

[0034] V cone and V i represent the corresponding cuboid-shaped and cone-shaped integral spaces respectively; I sat represents the image corresponding to the spatial sampling point i, and Ψ air and Ψ 1 represent the satellite and airborne image sets respectively.

[0035] Furthermore, the method of position encoding for the coordinate p=(x,y,z) in the integral space is as follows;

[0036] ​pos_enc(p) = [a 1 cos(2πω 1 p), a 1 sin(2πω 1 p), …, a n cos(2πω n p), a n sin(2πω n p)](2)

[0037] where ω 1 ...ω n is the frequency coefficient, a 1 ...a n is the encoding weight of different frequencies, and pos_enc(p) represents the result of position encoding.

[0038] Furthermore, the pixel value I(i, j) in the image is converted to m-bit numerical precision to obtain I * (i, j) with the following formula:

[0039]

[0040] where I(i, j) represents the pixel at the i-th row and j-th column in the image, and n represents the original quantization bits of the image.

[0041] Furthermore, the position encoding is performed on the coordinates p = (x, y, z) in the space to obtain the encoding result pos_enc(p); the encoding result is integrated in the integration space V to obtain the multi-mode space integration encoding result MSI(p):

[0042]

[0043] where V represents the integration space corresponding to p.

[0044] Furthermore, x in σ’(x) represents the input of the cumulative distribution function.

[0045] A computer storage medium stores at least one instruction, and the at least one instruction is loaded and executed by a processor to implement the described self-supervised spatio-temporal field reconstruction method based on multi-mode space integration encoding and spectral band optimization.

[0046] A self-supervised spatio-temporal field reconstruction device based on multi-mode space integration encoding and spectral band optimization, the device includes a processor and a memory, and the memory stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement the described self-supervised spatio-temporal field reconstruction method based on multi-mode space integration encoding and spectral band optimization.

[0047] Beneficial effects:

[0048] The present invention generates light through each image of a super-generalized stereo pair and its space-time field imaging geometric model, gives basic self-supervision information to the light based on the image pixels, forms light samples, and bundles pixels of different spatial resolutions with their actual constraint space ranges through the spatial integration strategy of the light. At the same time, the self-supervision strategy of multi-channel joint is optimized and designed to provide support conditions for the joint establishment of geometric constraints based on light samples with different band numbers and numerical precision, thereby developing a function for supervising the space-time field and finally realizing space-time field reconstruction. The present invention can effectively solve the influence of the difference in spatial spectral resolution on the three-dimensional reconstruction effect of time-varying scenes based on super-generalized stereo pairs, and improve the effect of three-dimensional reconstruction. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 Schematic diagram of the resolution difference of multi-source images.

[0050] Figure 2 Flowchart for self-supervised ray sampling generation for space-time fields.

[0051] Figure 3 Flowchart for self-supervised light sample generation based on multi-modal spatial quadrature coding and channel optimization. DETAILED DESCRIPTION

[0052] In view of the problem of spatial spectral resolution differences of multi-source images, the present invention is mainly to solve the following specific problems:

[0053] ① The actual imaging three-dimensional space ranges corresponding to the light samples of images with different spatial resolutions are inconsistent, such as Figure 1 As shown in (a) in the figure, this causes the problem of spatial position deviation in the joint geometric constraints.

[0054] ② The dimensions and precision of the image values ​​of images with different spectral and radiation resolutions are inconsistent, such as Figure 1 As shown in (b) in the figure, the form of the self-supervisory information of the light samples is not uniform and there are large differences in accuracy, which makes it difficult to effectively conduct joint self-supervised learning.

[0055] Aiming at ultra-generalized stereo image pairs with spatial, spectral, and radiometric resolution differences, the present invention proposes a self-supervised light sample generation method for supervising the fitting of spatiotemporal field functions. Figure 2As shown in the figure, the present invention generates rays through each image of the ultra-generalized stereo image pair and its spatio-temporal field imaging geometric model, endows the rays with basic self-supervised information based on the image pixels, and forms ray samples. Deeply study the spatial integration strategy of the rays, and bundle pixels with different spatial resolutions with their actual constrained spatial ranges. Optimize the design of a multi-channel joint self-supervised strategy to provide support conditions for jointly establishing geometric constraints based on ray samples with different numbers of bands and numerical precisions, and provide a necessary ray sample set for spatio-temporal field function fitting. Thus, a method for supervising the spatio-temporal field function is developed, and finally spatio-temporal field reconstruction is realized. Next, the present invention will be further described in combination with specific embodiments. Specific Embodiment 1:

[0057] This embodiment is a self-supervised spatio-temporal field reconstruction method based on multi-mode spatial integration coding and spectral band optimization. Ray samples are generated based on the pixel coordinates of the image and the spatio-temporal field imaging geometric model, and self-supervised information is given to the ray samples based on the pixel values; a multi-mode spatial integration coding strategy is proposed, and the spatial information processing method is optimized in combination with the imaging characteristics of spaceborne and airborne images to reduce the geometric constraint deviation caused by the difference in spatial resolution; a self-supervised learning strategy based on channel optimization is proposed, and ray samples generated from images with different spectral and radiation resolutions are fully utilized to jointly establish geometric constraint rules.

[0058] As Figure 3 shown, a self-supervised spatio-temporal field reconstruction method based on multi-mode spatial integration coding and spectral band optimization described in this embodiment mainly includes a multi-mode spatial integration coding strategy and a self-supervised learning strategy based on spectral band optimization. Among them,

[0059] The multi-mode spatial integration coding strategy is as follows:

[0060] First, according to the ultra-generalized stereo image pair and its spatio-temporal field imaging geometric model, a corresponding ray can be generated for any pixel of each image, and then multi-mode spatial integration coding is performed. The purpose of spatial integration is to more accurately encode the three-dimensional information of spatial points. By constructing an integration space that conforms to the imaging characteristics of the image, the traditional single-scale position coding is extended to scale-adaptive spatial integration coding, thereby expanding the neural network's perception ability of spatial scales and reducing the geometric constraint deviation caused by the difference in spatial resolution.

[0061] As Figure 3 shown in the upper part, it includes the following steps:

[0062] S1. Construct rays emitted by the camera and passing through the image pixels according to the spatio-temporal field imaging geometric model.

[0063] The purpose is to imitate the process of the camera emitting radiation during imaging. Satellite imaging models often rely on the RFM imaging model, and the pinhole camera model approximates the RFM imaging model. Therefore, in some embodiments, the pinhole model can be used as a normalized representation form of the imaging geometry model parameters, and the on-board image imaging geometry model is approximated as the pinhole model to facilitate the unification of the form with the airborne imaging geometry model. Specifically, the imaging geometry model P of the spatio-temporal field is the projection matrix of perspective imaging, and its form is a 4x4 matrix. Given a coordinate point p=(x, y, z) in three-dimensional space; P can map p to a pixel coordinate pix=(u′, v′) in the image I:

[0064] P[x,y,z,1] T =[ku′,kv′,k,1] T

[0065] And given a certain pixel coordinate (u, v), through the matrix P -1 it can be mapped to three-dimensional space:

[0066] P -1 [u′,v′,1,1]=[x′,y′,z′,1]

[0067] Therefore, in some embodiments, the above formula is used to construct the imaging geometry model of the spatio-temporal field. The image I={I(u, v)|u∈1…m, v∈1…n}, indicating that the image I is an m-row and n-column matrix.

[0068] S2. According to the imaging payload platform and spatial resolution of the input image I i the integration space V of the spatial sampling points is constructed adaptively in multiple modes i .

[0069] In step S1, the light rays have been obtained, but relying on a single form of light rays cannot reflect the differences in spatial resolution between on-board images and airborne images in the subsequent rendering process. Therefore, different spatial integration construction modes are adopted for the two respectively. By constructing the integration space, the light rays are expanded into "cylindrical bodies" with different shapes, so as to promote the light ray samples generated from images with different spatial resolutions to adapt to the spatial scale.

[0070] In terms of the form of the integration space, for on-board images, the satellite observation point is far from the observed ground surface and the field of view angle is narrow, so a cuboid-shaped integration space is adopted; while for airborne images compared with on-board images, the observation point is much closer to the observed ground objects and the field of view angle is also larger, so a cone-shaped integration space is adopted.

[0071] Integration space:

[0072]

[0073] Among them, \(V_i\) represents the integration space corresponding to the spatial sampling point \(i\), and \(V\) cube , \(V\) cone respectively represent the corresponding cuboid-shaped and pyramid-shaped integration spaces; \(I\) i represents the image corresponding to the spatial sampling point \(i\), and \(\varPsi\) sat , \(\varPsi\) air respectively represent satellite and airborne image sets.

[0074] In terms of the size of the integration space, the size of the integration space is determined according to the input image type, the spatial resolution corresponding to the image, and the interval between adjacent sampling points. For high-spatial-resolution images, the corresponding integration space \(V\) should be smaller, while for low-spatial-resolution images, the corresponding integration space \(V\) should be larger to achieve the unification of the actual imaging three-dimensional space range.

[0075] Then, position encoding is performed on the input coordinate \(p=(x,y,z)\). The purpose of position encoding is to improve the ability of the neural network to distinguish high- and low-frequency signals. The encoding process is shown in Equation (2).

[0076] \(\text{pos\_enc}(p)=[a\) 1 \cos(2\pi\omega\) 1 \(p),a\) 1 \sin(2\pi\omega\) 1 \(p),\cdots,a\) n \cos(2\pi\omega\) n \(p),a\) n \sin(2\pi\omega\) n \(p)](2)\)

[0077] Among them, \(\omega\) 1 \(\cdots\omega\) n are frequency coefficients, and \(a\) 1 \(\cdots a\) n are encoding weights of different frequencies. \(\text{pos\_enc}(p)\) represents the result of position encoding.

[0078] Finally, integration is performed on the encoding result in the integration space \(V\), that is, according to the payload platform and spatial resolution of the input image, the result of position encoding is integrated in the space of the corresponding mode, and the multi-modal spatial integration (MSI) encoding result \(\text{MSI}(p)\) is obtained:

[0079]

[0080] Among them, \(V\) represents the integration space corresponding to \(p\).

[0081] The spatial resolution of MSI processing during the rendering process determines the actual spatial size corresponding to one rendered pixel. The multi-mode spatial integral encoding can adjust the encoding results to varying degrees according to the actual covered spatial range of the spatial sampling points, thus endowing the spatio-temporal field function with the ability to perceive the spatial scale.

[0082] The self-supervised learning strategy based on channel preference:

[0083] The self-supervised learning strategy based on channel preference determines the number of channels and the numerical precision of the rendered output image. To achieve self-supervised learning based on images with different spectral and radiation resolutions, it is necessary to unify the dimension and precision of the pixel values.

[0084] First, for the problem of radiation resolution difference, the high-quantization-precision images can be uniformly converted to the m-bit numerical precision according to the data scaling ratio method, as shown in Equation (4):

[0085]

[0086] Among them, I(i,j) represents the pixel value of the i-th row and the j-th column in the image, and n represents the original quantization bit number of the image.

[0087] For the spectral values of each band, I * (i,j) is obtained by conversion according to the above formula, and then the conversion result I * (i,j) is rounded and converted to the m-bit binary representation (the specific value of m can be freely set according to the specific input image) to achieve the normalization processing of the radiation resolution. Subsequently, I * (i,j) obtained from Equation (4) is divided by 2 m to obtain I'(i,j), achieving the normalization of the band numerical range to avoid the phenomenon of gradient explosion during the training process.

[0088] Secondly, for the problem of spectral resolution difference, as mentioned above, the observation perspective of the image is the key to constructing the spatial geometric constraint, and the number of image bands is not the decisive information for the geometric constraint ability. Too many bands participating in the supervision may not necessarily improve the geometric constraint ability and even waste computing resources. The channel preference mechanism proposed by the present invention can perform appropriate band selection for the three-dimensional reconstruction of a specific scene to limit the influence of the dimension difference. According to the spectral range characteristics of typical high-resolution satellite sensors in China listed in Table 1, single-band panchromatic (Pan) and multi-spectral images covering red (R), green (G), blue (B), and near-infrared (IR) are the most common and have the widest application scenarios, and airborne images also cover the above bands. Therefore, it is preset that for all images, the following process is adopted for corresponding self-supervised learning in these five channels. Such as Figure 3As shown, after the radiation resolution of the pixels is normalized, the channel optimization of the image is first performed, and the inherent bands of the image are screened according to the above five preset channels. Then, each image performs self-supervised learning based on the pixel consistency loss only on its own inherent channels corresponding to the above preset channels.

[0089] Specifically, first define a band set Band out ={B, G, R, IR, Pan}, and set the dimension of the radiation value output by TVSRF to 5 dimensions as well, and stipulate that the meaning of each dimension corresponds one-to-one with the meaning of each element in the band set, that is When a certain pixel pix and the light ray ray passing through this pixel are given, volume rendering is performed along this light ray according to the output of TVSRF, and a 5D vector output corresponding to the input pixel pix can be obtained. The process of volume rendering along this light ray according to the output of TVSRF includes:

[0090] Given a certain pixel pix and the light ray Ray = {r(d)=o + dv, d≥0} passing through this pixel, and the depth range [d n , t f of the reconstructed target scene, first sample along the ray within [d n , d f and obtain a set of sampling points P = {p i |p i =r(d i ), i = 1, 2…, N}.

[0091] Among them, v is the unit direction vector of the light ray, d is the depth along the light ray starting from the camera origin o, and N represents the number of sampling points.

[0092] Then, the volume rendering result corresponding to the pixel passed through by the light ray is obtained according to the following formula:

[0093]

[0094] w(p i ) = T(p i )a(p i )

[0095]

[0096] δ(p i ) = d i+1 -d i

[0097] Among them, δ(p i ) represents the depth interval between adjacent sampling points, and a(p i ) is along the length in the ray segment r(di+1 ) - r(d i ) Opacity at T(p i ) is the cumulative opacity when the light travels from the origin o to the spatial point p i at point w(p i ) is the corresponding volume rendering integral weight at the spatial point p i ; TVSRF represents the time-varying radiation function, which obtains the radiation intensity based on the input MSI(p i ) and t i and can be implemented through a neural network model. TVSRF(MSI(p i ), t i ) is the output result of the time-varying radiation function at time t i at the spatial point p i , that is, the radiation intensity; σ’(p i ) is the volume density at the spatial point p i at time t i .

[0098] σ(p i ) is the volume density at the spatial point p i at time t i , which is obtained by converting the output value TVSDF(MSI(p i ) at time t i at the spatial point p i ) through the following formula: i σ(p

[0099] ) = κ · σ’(-TVSDF(MSI(p i ), t i )) i

[0100]

[0101] where κ and λ are learnable parameters, and σ’(x) is actually the cumulative distribution function based on the Laplace distribution, with a mean of 0 and a scale parameter of λ; x in σ’(x) represents the input of the cumulative distribution function; TVSDF is the time-varying signed distance function, which obtains the signed distance value based on the input MSI(p i ) and t i . The signed distance value refers to the minimum value of the coordinate distance between the spatial point coordinate p and a point on the ground object surface, and can be implemented through a neural network model. TVSDF(MSI(p i ), t i ) is the output result of the time-varying signed distance function at time t i at the spatial point p i , that is, the signed distance value. ​

[0102] Assume the input image I corresponding to the pixel i The set of bands it contains is Band in , and the channel preference mechanism will also construct a 5D vector R corresponding to this pixel in =(r B , r G , r R , r IR , r Pan ), where the value r k of one dimension is the I'(i, j) value corresponding to the corresponding band, k ∈ {B, G, R, IR, Pan}; r k satisfies the following rules:

[0103]

[0104] For example, if an image Band in ={B, G, R,} ∪ Band other and then R in =(r B , r G , r R , 0, 0); Band other represents other bands except B, G, and R

[0105] Finally, from R in , R out a self-supervised spectral loss can be constructed at the pixel level:

[0106]

[0107] where, ||·|| 1 , ||·|| 0 correspond to the 1-norm and 0-norm respectively

[0108] The role of the loss is to indirectly optimize the spatio-temporal field function by fitting the spectral values. R in and R out correspond to the input real image and the output rendered image of the spatio-temporal field respectively. In the present invention, the spatio-temporal field is mainly characterized by two functions TVSDF and TVSRF encoded by a neural network. TVSDF characterizes the geometry of the spatio-temporal field at a certain moment through the signed distance field value, while TVSRF characterizes the appearance of the spatio-temporal field at a certain moment through the radiation intensity value. By performing differentiable rendering on the three-dimensional geometry and appearance estimated by the two functions in a specific space, the present invention obtains the imaging R out at a specific moment and specific viewing angle, and through R in and R outThe resulting loss can then be used to optimize the estimation of the spatio-temporal field functions TVSDF and TVSRF in the reverse direction. Finally, spatio-temporal field reconstruction is achieved using the optimized TVSDF and TVSRF. Specific Embodiment 2:

[0110] This embodiment is a computer storage medium, in which at least one instruction is stored, and the at least one instruction is loaded and executed by a processor to implement the self-supervised spatio-temporal field reconstruction method based on multi-mode spatial integral coding and spectral band optimization.

[0111] It should be understood that the instructions include computer program products, software, or computerized methods corresponding to any method described in the present invention; the instructions can be used to program a computer system or other electronic devices. A computer storage medium can include a readable medium on which instructions are stored, and can include, but are not limited to, magnetic storage media, optical storage media; magneto-optical storage media include read-only memory ROM, random access memory RAM, erasable programmable memory (e.g., EPROM and EEPROM), and flash memory layers, or other types of media suitable for storing electronic instructions. Specific Embodiment 3:

[0113] This embodiment is a self-supervised spatio-temporal field reconstruction device based on multi-mode spatial integral coding and spectral band optimization. The device includes a processor and a memory. It should be understood that it includes any device including a processor and a memory described in the present invention. The device may further include other units and modules for display, interaction, processing, control, etc. through signals or instructions, as well as other functions.

[0114] At least one instruction is stored in the memory, and the at least one instruction is loaded and executed by the processor to implement the self-supervised spatio-temporal field reconstruction method based on multi-mode spatial integral coding and spectral band optimization.

[0115] Those skilled in the art should understand that the stored at least one instruction is a computer program product corresponding to the method or system. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The solutions in the embodiments of the present application can be implemented using various computer languages, for example, object-oriented programming languages such as Java and interpreted scripting languages such as JavaScript.

[0116] This application is described with reference to the flowcharts and / or block diagrams of methods, systems, and computer program products according to embodiments of the present application, and can also be used for corresponding devices. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one process Figure 1 or more processes and / or blocks Figure 1 or a device for implementing the functions specified in one block or more blocks.

[0117] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in one process Figure 1 or more processes and / or blocks Figure 1 or a device for implementing the functions specified in one block or more blocks.

[0118] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 or more processes and / or blocks Figure 1 or a device for implementing the functions specified in one block or more blocks.

[0119] Although the preferred embodiments of the present application have been described, those skilled in the art can make additional changes and modifications once they know the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments as well as all changes and modifications falling within the scope of the present application.

[0120] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.

[0121] The above calculation examples of the present invention are only for explaining in detail the calculation model and calculation process of the present invention, rather than limiting the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or variations can be made based on the above description. It is impossible to list all the implementation manners here. Any obvious changes or variations derived from the technical solutions of the present invention still fall within the protection scope of the present invention.

Claims

1. A self-supervised space-time field reconstruction method based on multi-mode spatial integral coding and spectrum segment optimization, characterized in that: The space-time field is characterized based on the time-varying radiation function TVSRF and the time-varying signed distance function TVSDF, and the space-time field is reconstructed by optimizing the time-varying radiation function TVSRF and the time-varying signed distance function TVSDF. The optimization process of the time-varying radiation function TVSRF and the time-varying signed distance function TVSDF includes: According to the spatial-temporal imaging geometric model, the light emitted by the camera and passing through the image pixels is constructed; based on the spatial sampling point i on the light, the integral space V of the spatial sampling point is constructed i ; Position encoding is performed on the coordinate p=(x, y, z) in the space to obtain the encoding result pos_enc(p); the encoding result is integrated in the integral space V to obtain the multi-mode spatial integral encoding result MSI(p); Convert the pixel value I(i,j) in the image to m-bit numerical precision to obtain I * (i,j), and then convert the result I * (i,j) is rounded and converted to m-bit binary representation; then the obtained I * (i,j) divided by 2 m Get I'(i,j); For a band set Band out = {B, G, R, IR, Pan}, and volume rendering is performed along the ray according to the output of the time-varying radiation function TVSRF, including the following steps: Given a pixel pix and the ray Ray passing through the pixel, Ray = {r(d) = o + dv, d ≥ 0}, and the depth range [d n , t f ], first in [d n , d f ] and obtain a sampling point set P = {p i |p i = r(d i ), i = 1, 2…, N}; where v is the unit direction vector of the light, d is the depth along the light starting from the camera origin o, and N represents the number of sampling points; Then, the volume rendering result corresponding to the pixel passed by the light is obtained according to the following formula: σ(p i )=k·s'(-TVSDF(MSI(p i ),t i )), Among them, δ(p i ) represents the depth interval between adjacent sampling points, a(p i ) is the length of the ray segment r(d i+1 )-r(d i ) on the opacity, T(p i ) is the light from the origin o to the space point p i The cumulative transparency at the time of treatment, w(o i ) is the space point p i The corresponding volume rendering integral weight at σ(p i ) is t i At the point p in space i The volume density at the location; κ, λ are learnable parameters, and σ'(x) is the cumulative distribution function based on the Laplace distribution, with a mean of 0 and a scale parameter of λ; TVSRF uses a neural network model, which is based on the input MSI (p i ) and t i get are the radiation values ​​of the corresponding bands, TVSRF(MSI(p i ),t i ) is t i At the point p in space i The output result of the time-varying radiation function at the time-varying signed distance function TVSDF adopts a neural network model, which is based on the input MSI (p i ) and t i The signed distance value is the minimum distance between the coordinates of a spatial point p and a point on the surface of the object. TVSDF(MSI(p i ),t i ) is t i At the point p in space i The output result of the time-varying signed distance function at ; Assume that the pixel corresponds to the input image I i The included band set is Band in , the channel optimization mechanism will also construct a 5-dimensional vector R corresponding to the pixel in =(r B ,r G ,r R ,r IR ,r Pan ), a dimension value r k is the I'(i,j) value corresponding to the corresponding band, k∈{B,G,R,IR,Pan}; r k Satisfy the rules: r k =0, if By R in , R out Constructing a self-supervised spectral loss at the pixel level Among them, ||·||1 and ||·||0 correspond to the 1-norm and 0-norm respectively; The time-varying radiance function TVSRF and the time-varying signed distance function TVSDF are optimized according to the spectral loss.

2. According to claim 1, a self-supervised space-time field reconstruction method based on multi-mode spatial integral coding and spectrum segment optimization is characterized in that: The integration space V cube 、V cone Respectively represent the corresponding rectangular and cone-shaped integral spaces; I i represents the image corresponding to the spatial sampling point i, Ψ sat , air Represent satellite and airborne image collections respectively.

3. The self-supervised space-time field reconstruction method based on multi-mode spatial integral coding and spectrum segment optimization according to claim 2 is characterized in that: The method of position encoding for the coordinate p=(x, y, z) in the integral space is as follows: pos_enc(p)=[a1cos(2πω1p),a1sin(2πω1p),…,a n cos(2πω n p),a n sin(2πω n p)] (2) Among them, ω1...ω n is the frequency coefficient, a1...a n are the encoding weights of different frequencies, and pos_enc(p) represents the result of position encoding.

4. A self-supervised space-time field reconstruction method based on multi-mode spatial integral coding and spectrum segment optimization according to any one of claims 1 to 3, characterized in that: Convert the pixel value I(i,j) in the image to m-bit numerical precision to obtain I * The formula for (i,j) is as follows: Among them, I(i,j) represents the i-th row and j-th column in the image, and n represents the number of original quantization bits of the image.

5. The self-supervised space-time field reconstruction method based on multi-mode spatial integral coding and spectrum segment optimization according to claim 4 is characterized in that: The coordinate p = (x, y, z) in the space is position-encoded to obtain the encoding result pos_enc(p); the encoding result is integrated in the integral space V to obtain the multi-mode spatial integral encoding result MSI(p): Among them, V represents the integral space corresponding to p.

6. The self-supervised space-time field reconstruction method based on multi-mode spatial integral coding and spectrum segment optimization according to claim 5 is characterized in that: Said The x in σ'(x) represents the input of the cumulative distribution function.

7. A computer storage medium, characterized in that: The storage medium stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement a self-supervised space-time field reconstruction method based on multi-mode spatial integral coding and spectral segment optimization as described in any one of claims 1 to 6.

8. A self-supervised space-time field reconstruction device based on multi-mode spatial integral coding and spectrum segment optimization, characterized in that: The device includes a processor and a memory, wherein the memory stores at least one instruction, and the at least one instruction is loaded and executed by the processor to implement a self-supervised space-time field reconstruction method based on multi-mode spatial integral coding and spectral segment optimization as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Pairs coding compression hyperspectral imaging device

    CN103808410A

  • Method and device for three-dimensional reconstruction of remote sensing image

    CN117765172A

  • Systems and methods for magnetic resonance imaging

    US20200202586A1

  • System and method for topological characterization of tissue

    WO2023141216A2