Incremental three-dimensional reconstruction method based on time-varying scene geometric reasoning

By introducing time-varying scene geometric inference into the incremental three-dimensional reconstruction method, dynamically fitting the space-time field function, and optimizing the incremental investment strategy of image, the accuracy problem of ultra-generalized stereoscopic images for time-varying scenes is solved, and efficient three-dimensional reconstruction effect is achieved.

CN120147548APending Publication Date: 2025-06-13THE INST OF AUTOMATION HEILONGJIANG ACADEMY OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510304646.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

When the existing incremental three-dimensional reconstruction method deals with ultra-generalized stereo image pairs, it is difficult to ensure the accuracy of time-varying scene three-dimensional reconstruction. Especially when the image increment method is complex, it is easy to have small sample problems, sample imbalance and large viewing angle differences.

Method used

An incremental three-dimensional reconstruction method based on time-varying scene geometric reasoning is proposed. By initially determining whether the newly added image is a historical image, calculating image differences, generating image sample sets, fitting space-time field functions, and dynamically adjusting the image incremental investment strategy, optimizing the learning rules of space-time field functions to realize three-dimensional reconstruction.

Benefits of technology

The three-dimensional accurate reconstruction of time-varying scenes is realized, the fitting inaccurate problem caused by the complexity of image incremental methods is solved, and the reliability and efficiency of space-time field function fitting are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147548A_ABST
    Figure CN120147548A_ABST
Patent Text Reader

Abstract

The invention discloses an incremental three-dimensional reconstruction method based on time-varying scene geometric reasoning, and belongs to the technical field of incremental three-dimensional reconstruction. According to the method, the problem of poor accuracy of three-dimensional reconstruction of the time-varying scene of the super-generalized stereopair by adopting an existing incremental three-dimensional reconstruction method is solved. According to the method, aiming at different geometric forms of a time-varying scene in different time periods, incremental processing of a newly-added image is taken as a main line, a newly-added image incremental input strategy combining time sequence and visual angle reasoning is researched, and the necessity of dividing a newly-added time period is analyzed and decided through image change; and optimizing the image increment input modes of the newly added time period and different historical time periods. And dynamically fitting a space-time field function based on a self-supervised light sample generated by inputting an image in each time period in combination with an optimized space-time field function learning rule. And three-dimensional accurate reconstruction of the time-varying scene is realized based on the fully-fitted space-time field function. The method can be applied to incremental three-dimensional reconstruction of a time-varying scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of incremental three-dimensional reconstruction, and particularly relates to an incremental three-dimensional reconstruction method based on geometric reasoning of time-varying scenes. Background Art

[0002] The combination of two or more optical remote sensing images obtained from different observation perspectives in any way covering the same scene is defined as a "super-generalized stereo pair". The composition form of the super-generalized stereo pair is free, the number of images is not fixed, and an input increment can be incrementally accumulated with the shooting of new images; it can be obtained by different payload platforms at different times, and there are unfixed differences in observation perspectives and resolutions among multiple images. The intersection angle, base-height ratio, overlap degree, etc. between any two of them do not necessarily meet the composition conditions of a "standard stereo pair" or a "generalized stereo pair". Once there are differences in observation perspectives among multiple optical remote sensing images, geometric constraints can be formed on the three-dimensional shape of the scene. Therefore, in theory, three-dimensional reconstruction of the scene can be achieved based on the super-generalized stereo pair. Compared with the standard or generalized stereo pair, the main advantage of the "super-generalized stereo pair" lies in the high-frequency and multi-perspective acquisition of image data, which provides the necessary data conditions for realizing three-dimensional reconstruction of time-varying scenes from "static" to "temporal", and is more in line with the requirements of real-scene three-dimensional.

[0003] The super-generalized stereo pair is collected successively over time. The newly added images constitute an image increment, and it also determines that an incremental processing is required for the time-varying three-dimensional reconstruction of the scene based on the super-generalized stereo pair. Incremental three-dimensional reconstruction generally refers to starting from the initial two or three images, incrementally adding new images until all images are processed, and gradually realizing three-dimensional reconstruction. Traditional incremental three-dimensional reconstruction mostly refers to the incremental (SfM, Structure from Motion) series of methods. Since it is a method based on stereo matching, the incremental method is sensitive to the selection of the initial images and the order of image addition, and there are problems such as scene drift, cumulative error in large-scale scene reconstruction, and low efficiency, which are not suitable for the task of time-varying three-dimensional reconstruction of the scene based on the super-generalized stereo pair.

[0004] Different from the incremental 3D reconstruction in the traditional sense, the complexity of the image increment method of the ultra-generalized stereo pair is mainly reflected in that the newly added images may be newly captured images or historical images that were not collected before but are newly collected. At the same time, the scene may undergo uncertain geometric changes over time, and the observation time, viewing angle, and quantity of the images that can be obtained under each different geometric form are all uncertain. This complexity of the image increment method will lead to the following three typical specific situations: First, in the early stage of reconstruction, when the newly added changing scene or the viewing angles of the existing images are relatively close, there are too few images that can effectively form geometric constraints, resulting in a small sample problem in the fitting process and a significant decrease in the reliability of the spatio-temporal field function fitting. Second, when there are new geometric changes in a scene in a newly added image, due to the small number of light samples in the changing area, when fitting the spatio-temporal field function together with a large number of images with unchanged scenes, the changing area is easily mistaken for an outlier sample and ignored. In addition, after the number of images increases, there is serious redundant information of light samples due to excessive input of images with small viewing angle differences, which affects the efficiency of the spatio-temporal field function fitting; images with large viewing angle differences cannot be strictly registered, making it difficult to accurately analyze the geometric changes of the scene.

[0005] There are mainly two ways of common scene change analysis in the field of remote sensing: change detection based on orthographic views and change detection based on 3D data. The former usually requires two orthographic views that are strictly registered at different times for comparative analysis. If the images cannot be strictly registered, only rough qualitative analysis can be carried out, and accurate quantitative analysis cannot be performed; the latter uses remote sensing 3D data at different times to analyze the geometric changes of the scene. Typical 3D data includes DEM (Digital Elevation Model) data, DSM data, or 3D point cloud data of the scene obtained by methods such as oblique photogrammetry and LiDAR. Usually, 3D registration is also required, and then qualitative and quantitative analysis of the changing area is carried out. On the one hand, due to the differences in the imaging characteristics of the images, they usually cannot be strictly registered, making it difficult to analyze the quantitative geometric changes of the scene based on the images; on the other hand, the 3D geometric changes of the scene and 3D reconstruction are both realized based on the dynamic fitting of the spatio-temporal field function, and its reliability depends more on the incremental 3D reconstruction strategy.

[0006] To sum up, in order to realize the 3D reconstruction of the time-varying scene based on the ultra-generalized stereo pair, under each different geometric form, the incremental processing method of the ultra-generalized stereo pair needs to be optimized by considering the observation time, viewing angle of the images and the geometric change situation of the scene. Otherwise, it is difficult to ensure the sufficiency of geometric constraints and the accuracy of time-varying analysis, and thus the accuracy of 3D reconstruction cannot be guaranteed. Summary of the Invention

[0007] The object of the present invention is to solve the problem of poor accuracy in three-dimensional reconstruction of a time-varying scene of a super-generalized stereo pair by using the existing incremental three-dimensional reconstruction method, and an incremental three-dimensional reconstruction method based on geometric reasoning of a time-varying scene is proposed.

[0008] The technical solution adopted by the present invention to solve the above technical problems is: an incremental three-dimensional reconstruction method based on geometric reasoning of a time-varying scene, and the method specifically includes the following steps:

[0009] Step 1: For the newly added image I at time t new , initially determine whether the newly added image I new is a historical image according to the shooting time of the newly added image I new ;

[0010] If it is initially determined that the newly added image I new is not a historical image, then continue to execute Step 2 on the newly added image I new ;

[0011] If it is initially determined that the newly added image I new is a historical image, then directly execute Step 3;

[0012] Step 2: Calculate the difference between the newly added image I new and the rendered simulation image, and finally determine whether the newly added image I new is a historical image according to the calculated difference;

[0013] Execute Step 3 according to the final determination result;

[0014] Step 3: Generate an image sample set Train_samples according to the newly added image I new , and fit a spatio-temporal field function according to the image sample set Train_samples;

[0015] Step 4: Perform three-dimensional reconstruction of the scenes in each historical period according to the spatio-temporal field function obtained in Step 3, and return to execute Step 1 at time t+1.

[0016] Furthermore, the specific process of Step 2 is:

[0017] Step 2-1: Render a simulation image I new with the same viewing angle as the newly added image I new according to the spatio-temporal field function fitted at time t-1 and the imaging geometric model of the newly added image I sim :

[0018] I sim = {pix|pix = render(v i )} (1)

[0019] Among them, v i is any view direction of the newly added image;

[0020] pix is the pixel value of the light ray corresponding to the view direction v i ;

[0021] render(·) represents the image rendering process;

[0022] {pix} is the set composed of the pixel values corresponding to each light ray;

[0023] Step Two: Calculate the difference map between the newly added image I new and the simulation image I sim , and calculate the difference size Dis new between the newly added image I sim and the simulation image I 2D ;

[0024] If the difference Dis new between the newly added image I sim and the simulation image I 2D is less than the preset threshold, then process the newly added image as a historical image;

[0025] If the difference Dis new between the newly added image I sim and the simulation image I 2D is greater than or equal to the preset threshold, then take the newly added image as the first image of the new period.

[0026] Furthermore, the specific process of the said Step Two is:

[0027]

[0028] Among them, the sizes of the newly added image I new and the simulation image I sim are the same, W is the width of the image, and H is the height of the image;

[0029] diff(I new , I sim ) represents the difference map between the newly added image I new and the simulation image I sim ;

[0030] sum(·) represents the sum of the values of all pixel points in the difference map;

[0031] Dis 2D is the difference between the newly added image I new and the simulation image I sim .

[0032] Furthermore, in the said Step Three, according to the newly added image Inew Generate an image sample set Train_samples. The specific process is as follows:

[0033] (1) If the newly added image I new is not a historical image, extract one image from each historical period respectively, and use the newly added image I new and the extracted images to form the training image sample set Train_samples:

[0034] Train_samples = I s1 , I s2 ,..., I sn , I new (3)

[0035] Among them, I s1 , I s2 ,..., I sn respectively represent the images extracted from each historical period;

[0036] n represents the number of historical periods;

[0037] (2) If the newly added image I new is a historical image, determine whether the number of historical images within the historical period is sufficient:

[0038] If the number of historical images within the historical period is not sufficient, use the newly added image I new and all the historical images within the historical period to form the training image sample set Train_samples:

[0039] Train_samples = I 1 , I 2 ... I n , I new (4)

[0040] Among them, I 1 , I 2 ... I n respectively represent all the images in each historical period;

[0041] If the number of historical images within the historical period is sufficient, obtain the historical image with the closest perspective to the newly added image I new from all the historical images, remove the obtained historical image from the dataset composed of all the historical images, and use the remaining historical images and the newly added image I new to form the training image sample set Train_samples.

[0042] Furthermore, the specific method for fitting the spatio-temporal field function according to the image sample set Train_samples is as follows:

[0043] (1) When the newly added image I new is a historical image:

[0044] Judge whether the number of viewpoints corresponding to the images in the training image sample set Train_samples is less than the threshold;

[0045] If the number of viewpoints is less than the threshold, each image in the training image sample set Train_samples is processed by means of position encoding and frequency encoding, and the spatio-temporal field function fitting result is obtained according to the processing results;

[0046] If the number of viewpoints is greater than or equal to the threshold, directly use each image in the training image sample set Train_samples to obtain the spatio-temporal field function fitting result;

[0047] (2) When the newly added image I new is not a historical image:

[0048] Each image in the training image sample set Train_samples is processed by means of position encoding and frequency encoding, and the spatio-temporal field function fitting result is obtained according to the processing results.

[0049] Furthermore, the specific process of the position encoding is as follows:

[0050] For any image in the training image sample set Train_samples, record the position coordinates of a sampling point on the input light ray along the viewpoint direction v i as p=(x, y, z), and perform position encoding on the input coordinates p=(x, y, z):

[0051] pos_enc(p)=[a 1 cos(2πω 1 p), a 1 sin(2πω 1 p), …, a m cos(2πω m p), a m sin(2πω m p)] (5)

[0052] where ω 1 ...ω m is the frequency;

[0053] a 1 ...a m are the encoding weights of different frequencies;

[0054] pos_enc(p) represents the result of position encoding for the sampling point p = (x, y, z).

[0055] Furthermore, the method of frequency encoding is as follows:

[0056] Step 1: Abbreviate pos_enc(p) as p′ and initialize k = 0;

[0057] Step 2: According to the frequency f r = kf d perform frequency encoding on p′, where f d is the frequency sampling interval;

[0058]

[0059] wherein, represents the encoding weight of the frequency kf d and the superscript T represents transpose;

[0060]

[0061] where V represents the integration space and dv represents the integration variable;

[0062] Step 3: Input MSI(p) into the first multi-layer perceptron MLP static , and output the static geometric value g static through the first multi-layer perceptron MLP static :

[0063] g static = MLP static (MSI(p)) (8)

[0064] Denote the time feature of the shooting time of the current image as Φ i , and input Φ i and the static geometric value g static together into the second multi-layer perceptron MLP offset to obtain the geometric value bias g offset :

[0065] g offset = MLP offset (g static , Φ i ) (9)

[0066] Add the static geometric value and the geometric bias value, and use the addition result as the time-varying signed distance function value g i :

[0067] g i = g static + g offset (10)

[0068] Take the position information p, time feature Φ of the sampling point i , viewing angle information v i and the time-varying signed distance function value g i as the input of the third multi-layer perceptron MLP alb , and output the radiation value r through the MLP alb : i :

[0069] r i = MLP alb (p, v i , Φ i , g i ) (11)

[0070] Step 4. Determine whether k = K is satisfied, where K is a set threshold:

[0071] If not satisfied, let k = k + 1 and return to execute Step 2;

[0072] If satisfied, stop the iteration and obtain the time-varying signed distance function values and radiation values of the current image in different viewing angle directions and different frequencies.

[0073] Furthermore, the specific process of Step 4 is as follows:

[0074] Step 4-1. Statistically organize the time information of all historical input images to form a time line, sample the time information on the time line to obtain each sampling time point, and use each sampling time point to form a sampling time series T′ = (t 1 , t 2 , …, t L ), t 1 , t 2 , …, t L represents the 1st, 2nd, …, Lth sampling time point;

[0075] Step 4-2. Use the spatio-temporal field function fitted in Step 3 to reconstruct the scene model of the sampling time point t j to obtain the reconstruction result Mesh of the scene model of the sampling time point t j :

[0076] Mesh = MC(g)

[0077] where MC(·) is the Marching Cube function, and g represents the set of time-varying signed distance function values corresponding to the images at all viewing angle directions at the sampling time point t j ;

[0078] Step 4-3. Calculate the sampling time point t j and the sampling time point tj+1 Degree of difference of the reconstructed scene model:

[0079]

[0080] Wherein, represents the reconstruction result M of the scene model at the sampling time point t j and the degree of difference from the reconstruction result M of the scene model at the sampling time point t j at the sampling time point t j+1 and the reconstruction result M of the scene model at the sampling time point t j+1 ;

[0081] If is less than the threshold τ, then the sampling time point t j and the sampling time point t j+1 belong to the same historical period;

[0082] If is greater than or equal to the threshold τ, then the sampling time point t j and the sampling time point t j+1 do not belong to the same historical period;

[0083] Step Four: After performing Step Three on the reconstruction results of the scene models at every two adjacent sampling time points respectively, divide the sampling time series T′ into n independent historical periods;

[0084] For any one of the divided historical periods, perform fusion processing on all the reconstruction results of the scene models included in this historical period, and use the fusion result as the three-dimensional scene reconstruction result of this historical period. Similarly, obtain the three-dimensional scene reconstruction results of each historical period respectively;

[0085] Combine the three-dimensional scene reconstruction results of the n historical periods to obtain the three-dimensional reconstruction result of the time-varying scene.

[0086] Furthermore, in the above Step Four, the fusion processing adopts an equalization method.

[0087] The beneficial effects of the present invention are as follows:

[0088] Aiming at the different geometric forms of the time-varying scene in different periods, the present invention takes the incremental processing of new images as the main line, studies the incremental input strategy of new images by jointly considering time series and perspective reasoning, analyzes the necessity of dividing new periods through image change analysis and decision-making, and optimizes the image incremental input methods for new periods and different historical periods. Based on the self-supervised light samples generated from the images input in each period, combined with the optimized learning rules of the spatio-temporal field function, the spatio-temporal field function is dynamically fitted. Based on the fully fitted spatio-temporal field function, the geometric forms of the scene in each period and the geometric change relationships between different periods are inferred and generated, so as to realize the accurate three-dimensional reconstruction of the time-varying scene and solve the complexity problem of the image increment method. Description of the Drawings

[0089] Figure 1 is the framework diagram of the incremental time-varying scene three-dimensional reconstruction method of the present invention;

[0090] Figure 2 is a schematic diagram of large perspective differences that cannot be strictly registered;

[0091] Figure 3 is a schematic diagram of the newly added geometric change information being submerged due to sample imbalance;

[0092] Figure 4 is a schematic diagram of inaccurate fitting of the spatio-temporal field function due to the small sample problem;

[0093] Figure 5 is the flowchart of the incremental time-varying scene three-dimensional reconstruction method of the present invention. Detailed Implementation Manner

[0094] Detailed Implementation Manner 1: Combining Figure 1 and Figure 5 to illustrate this implementation manner. An incremental three-dimensional reconstruction method based on time-varying scene geometric reasoning described in this implementation manner specifically includes the following steps:

[0095] Step 1. For the newly added image I new at time t, initially determine whether the newly added image I new is a historical image according to the shooting time of the newly added image I new ;

[0096] If it is initially determined that the newly added image I new is not a historical image, then continue to execute Step 2 for the newly added image I new ;

[0097] If it is initially determined that the newly added image I new is a historical image, then directly execute Step 3;

[0098] Step 2. Calculate the difference between the newly added image I new and the rendered simulation image, and finally determine whether the newly added image I new is a historical image according to the calculated difference;

[0099] Execute Step 3 according to the final determination result;

[0100] Step 3. Generate an image sample set Train_samples according to the newly added image I new and fit a spatio-temporal field function according to the image sample set Train_samples;

[0101] Step 4: Perform 3D reconstruction of historical scenes based on the spatio-temporal field function obtained in Step 3, and return to execute Step 1 at time t+1.

[0102] The image increment method of the ultra-generalized stereo pair has the following complexity problems:

[0103] (1) As Figure 2 shown, the viewing angle of the newly added image may be significantly different from that of the historical image, and the images cannot be strictly registered, resulting in difficulties in accurately discriminating geometric changes in the two-dimensional image level.

[0104] (2) As Figure 3 shown, there are indeed geometric changes in the scene of the newly added image, but due to insufficient light samples in the changed area, a sample imbalance problem is caused, resulting in the changed area being ignored as an outlier sample during the fitting process of the spatio-temporal field function.

[0105] (3) As Figure 4 shown, when the number of images is small or the viewing angle difference between images is small, only a small number of view geometric constraints can be provided, resulting in a small sample problem at the observation viewing angle level, and the fitting of the spatio-temporal field function is inaccurate.

[0106] Due to the above problems, the present invention proposes an incremental 3D reconstruction technology based on multi-dimensional time-varying scene geometric reasoning, dynamically fitting the spatio-temporal field function with the input of newly added images, and realizing 3D reconstruction of time-varying scenes. An image increment input strategy based on joint reasoning of time series and viewing angle is proposed to solve the complexity problem of the image increment method. Specifically, a geometric change discrimination strategy based on new view rendering is proposed to solve the problem that it is difficult to perform time-varying analysis due to the large viewing angle difference between the newly added image and the historical image; at the same time, an image sample equalization strategy is introduced to solve the problem that the changed area samples are few and are ignored; in addition, a spatio-temporal information adaptive frequency encoding method is proposed to jointly improve the fitting reliability of the spatio-temporal field function in the case of few views with multiple geometric constraints. Finally, based on the fully fitted spatio-temporal field function, the different geometric forms of the scene at different time periods and the geometric changes between time periods are inferred to realize 3D reconstruction of time-varying scenes.

[0107] Moreover, the method of the present invention is applicable to various types of images such as visible light images, hyperspectral images, multispectral images, and infrared images.

[0108] Specific Embodiment 2: The difference between this embodiment and Specific Embodiment 1 is that the specific process of Step 2 is as follows:

[0109] Step 2-1: According to the spatio-temporal field function fitted at time t-1 (i.e., the spatio-temporal field function obtained by fitting at time t-1) and the imaging geometric model of the newly added image I new , render and generate a simulation image I with the same viewing angle as the newly added image I new ​sim :

[0110] I sim = {pix|pix = render(v i )} (1)

[0111] where v i is any view direction of the newly added image;

[0112] pix is the pixel value of the ray corresponding to the view direction v i ;

[0113] render(·) represents the image rendering process;

[0114] {pix} is the set composed of the pixel values corresponding to each ray (the set of all pixel values is the simulation image I sim );

[0115] Step Two: Calculate the difference map between the newly added image I new and the simulation image I sim , and calculate the difference size Dis new between the newly added image I sim and the simulation image I 2D ;

[0116] If the difference Dis new between the newly added image I sim and the simulation image I 2D is less than the preset threshold, then process the newly added image as a historical image (that is, it is considered that the scene of the newly added image has not changed);

[0117] If the difference Dis new between the newly added image I sim and the simulation image I 2D is greater than or equal to the preset threshold, then regard the newly added image as the first image of a new time period (that is, it is considered that the scene of the newly added image has changed at this moment).

[0118] Other steps and parameters are the same as those in the specific implementation manner one.

[0119] This implementation manner fully explores the relationships such as temporal sequence, view angle, and geometric changes between historical images and newly added images around different newly added image situations, and proposes a discriminant strategy for scene geometric changes based on new view rendering according to the rendered simulation images, solving the problem that the view angle of the newly added image may be greatly different from that of the historical image, and it is difficult to accurately judge the state of scene geometric changes due to the inability to strictly register.

[0120] Specific implementation manner three: The difference between this implementation manner and the first or second specific implementation manner is that the specific process of the said Step Two is as follows:

[0121]

[0122] Among them, the newly added image I new has the same size as the simulated image I sim , where W is the width of the image and H is the height of the image;

[0123] diff(I new , I sim ) represents the difference map between the newly added image I new and the simulated image I sim (after taking the difference between the pixel values at the corresponding positions in the newly added image I new and the simulated image I sim and then taking the absolute value of the difference result, the absolute values corresponding to each pixel position form the difference map);

[0124] sum(·) represents the sum of the values of all pixel points in the difference map;

[0125] Dis 2D is the difference between the newly added image I new and the simulated image I sim .

[0126] Other steps and parameters are the same as those in the first or second specific implementation manner.

[0127] Specific implementation manner four: The difference between this implementation manner and one of the first to third specific implementation manners is that in step three, according to the newly added image I new , an image sample set Train_samples is generated, and the specific process is as follows:

[0128] (1) If the newly added image I new is not a historical image, then one image is extracted from each historical period, and the newly added image I new and the extracted images are used to form the training image sample set Train_samples:

[0129] Train_samples = I s1 , I s2 ,..., I sn , I new (3)

[0130] where I s1 , I s2 ,..., I sn respectively represent the images extracted from each historical period;

[0131] n represents the number of historical periods;

[0132] (3) If the newly added image I newIf it is a historical image, determine whether the number of historical images within the historical period is sufficient:

[0133] If the number of historical images within the historical period is insufficient, use the newly added image I new and all the historical images within the historical period to form a training image sample set Train_samples:

[0134] Train_samples = I 1 , I 2 ... I n , I new (4)

[0135] where I 1 , I 2 ... I n respectively represent all the images in each historical period;

[0136] If the number of historical images within the historical period is sufficient, obtain the historical image with the closest perspective to the newly added image I new from all the historical images, remove the obtained historical image from the dataset composed of all the historical images, and use the remaining historical images and the newly added image I new to form a training image sample set Train_samples.

[0137] Other steps and parameters are the same as those in any one of the specific embodiments one to three.

[0138] When the number of historical images is sufficient, to ensure the training efficiency, the present invention does not directly add the newly added image to the training image sample set, but replaces the image with the closest perspective to the newly added image among the historical images.

[0139] Using the image sample equalization strategy proposed in this embodiment, it is possible to solve the problem that in the process of fitting the spatio-temporal field function, due to sample imbalance, the changing area is misidentified as an outlier sample and ignored.

[0140] Specific embodiment five: The difference between this embodiment and any one of the specific embodiments one to four is that the spatio-temporal field function is fitted according to the image sample set Train_samples, specifically as follows:

[0141] (1) When the newly added image I new is a historical image:

[0142] Judge whether the number of perspectives corresponding to the images in the training image sample set Train_samples is less than the threshold;

[0143] If the number of viewpoints is less than the threshold, each image in the training image sample set Train_samples is processed by means of position encoding and frequency encoding, and the fitting result of the spatio-temporal field function is obtained according to the processing results;

[0144] If the number of viewpoints is greater than or equal to the threshold, the fitting result of the spatio-temporal field function is directly obtained by using each image in the training image sample set Train_samples;

[0145] (2) When the new image I new is not a historical image:

[0146] Each image in the training image sample set Train_samples is processed by means of position encoding and frequency encoding, and the fitting result of the spatio-temporal field function is obtained according to the processing results.

[0147] That is, the fitting result of the spatio-temporal field function is obtained according to the time-varying signed distance function values and radiation values of each image in the training set in different viewpoint directions and different encoding frequencies.

[0148] Other steps and parameters are the same as those in any one of the specific embodiments 1 to 4.

[0149] When the new image is a historical image and the number of viewpoints is greater than or equal to the threshold, the spatio-temporal field function fitting method is:

[0150] For any image in the training image sample set Train_samples, the position coordinates of a sampling point on the input light ray of the image along the viewpoint direction v i are denoted as p = (x, y, z):

[0151]

[0152] where V represents the integration space;

[0153] dv represents the integration variable;

[0154] Input MSI(p) into the first multi-layer perceptron MLP static , and output the static geometric value g static through the first multi-layer perceptron MLP static :

[0155] g static = MLP static (MSI(p))

[0156] Denote the time feature of the shooting time of the current image as Φ i , and input Φ i and the static geometric value g static into the second multi-layer perceptron MLP offsetObtain the geometric value offset g offset :

[0157] g offset = MLP offset (g static , Φ i )

[0158] Add the static geometric value and the geometric offset value, and use the added result as the time-varying signed distance function value g i :

[0159] g i = g static + g offset

[0160] Use the position information p of the sampling point, the time feature Φ i , the viewing angle information v i and the time-varying signed distance function value g i as the input of the third multi-layer perceptron MLP alb , and output the radiation value r alb through MLP i :

[0161] r i = MLP alb (p, v i , Φ i , g i )

[0162] The time-varying signed distance function values and radiation values of all images in the training set at different viewing angle directions are the fitting results of the spatio-temporal field function.

[0163] Specific Embodiment Six: Different from one of Specific Embodiments One to Five, the specific process of the position encoding is as follows:

[0164] For any image in the training image sample set Train_samples, record the position coordinate of a sampling point on the input light ray along the viewing angle direction v i as p = (x, y, z), and perform position encoding on the input coordinate p = (x, y, z):

[0165] pos_enc(p) = [a 1 cos(2πω 1 p), a 1 sin(2πω 1 p), …, a m cos(2πω m p), a m sin(2πω m p)] (5)

[0166] Among them, ω 1 ...ω m is the frequency;

[0167] a 1 ...a m are the encoding weights for different frequencies;

[0168] pos_enc(p) represents the result of position encoding for the sampling point p = (x, y, z).

[0169] Other steps and parameters are the same as those in any one of the first to fifth specific embodiments.

[0170] Specific embodiment seven: The difference between this embodiment and any one of the first to sixth specific embodiments is that the method of frequency encoding is as follows:

[0171] Step 1: Abbreviate pos_enc(p) as p′ and initialize k = 0;

[0172] Step 2: Perform frequency encoding on p′ according to the frequency f r = kf d , where f d is the frequency sampling interval;

[0173]

[0174] Among them, represents the encoding weight of the frequency kf d , and the superscript T represents transpose;

[0175]

[0176] Among them, V represents the integration space and dv represents the integration variable;

[0177] Step 3: Input MSI(p) into the first multi-layer perceptron MLP static , and output the static geometric value g static through the first multi-layer perceptron MLP static :

[0178] g static = MLP static (MSI(p)) (8)

[0179] Record the time feature of the shooting time of the current image as Φ i , and input Φ i and the static geometric value g static into the second multi-layer perceptron MLP offset together to obtain the geometric value bias g offset :

[0180] goffset = MLP offset (g static , Φ i ) (9)

[0181] Add the static geometric value to the geometric offset value, and use the added result as the time-varying signed distance function value g i :

[0182] g i = g static + g offset (10)

[0183] Use the position information p of the sampling point, the time feature Φ i , the viewing angle information v i , and the time-varying signed distance function value g i as the input of the third multi-layer perceptron MLP alb , and output the radiation value r alb through the MLP i :

[0184] r i = MLP alb (p, v i , Φ i , g i ) (11)

[0185] Step 4: Determine whether k = K is satisfied, where K is a set threshold:

[0186] If not satisfied, let k = k + 1, and return to execute Step 2;

[0187] If satisfied, stop the iteration and obtain the time-varying signed distance function values and radiation values of the current image at different viewing angle directions and different frequencies.

[0188] Other steps and parameters are the same as those in any one of the first to sixth specific embodiments.

[0189] Specific Embodiment Eight: The difference between this embodiment and any one of the first to seventh specific embodiments is that the specific process of Step Four is as follows:

[0190] Step Four One: Statistically analyze the time information of all historical input images to form a timeline, sample the time information on the timeline to obtain each sampling time point, and use each sampling time point to form a sampling time series T' = (t 1 , t 2 , …, t L ), where t 1 , t 2 , …, t L represents the 1st, 2nd, …, Lth sampling time point;

[0191] Step 42: Use the spatio-temporal field function obtained by fitting in Step 3 to reconstruct the scene model at the sampling time point t j to obtain the reconstruction result Mesh of the scene model at the sampling time point t j :

[0192] Mesh = MC(g)

[0193] where MC(·) is the Marching Cube function, and g represents the set of time-varying signed distance function values corresponding to the image at the sampling time point t j in all viewing directions;

[0194] Step 43: Calculate the difference degree between the sampling time point t j and the reconstructed scene model at the sampling time point t j+1 :

[0195]

[0196] where represents the difference degree between the reconstruction result M j of the scene model at the sampling time point t j and the reconstruction result M j+1 of the scene model at the sampling time point t j+1 ;

[0197] If is less than the threshold τ, then the sampling time point t j and the sampling time point t j+1 belong to the same historical period (no scene geometry change occurred between the moments corresponding to the two scene models);

[0198] If is greater than or equal to the threshold τ, then the sampling time point t j and the sampling time point t j+1 do not belong to the same historical period (scene geometry change occurred between the moments corresponding to the two scene models. At this time, the scene model at the previous moment belongs to the previous historical period, and the scene model at the later moment belongs to the new period);

[0199] Step 44: After performing Step 43 on the reconstruction results of the scene models at every two adjacent sampling time points respectively, divide the sampling time series T′ into n independent historical periods;

[0200] For any one of the divided historical periods, perform fusion processing on all the reconstruction results of the scene models included in this historical period, and use the fusion result as the three-dimensional scene reconstruction result of this historical period. Similarly, obtain the three-dimensional scene reconstruction results of each historical period;

[0201] Combine the three-dimensional scene reconstruction results of n historical periods to obtain the three-dimensional reconstruction result of the time-varying scene.

[0202] The other steps and parameters are the same as those in any one of the first to seventh specific embodiments.

[0203] Specific Embodiment Nine: The difference between this embodiment and any one of the first to eighth specific embodiments is that in Step Four, the fusion process adopts an equalization method.

[0204] The other steps and parameters are the same as those in any one of the first to eighth specific embodiments.

[0205] The above examples of the present invention are only for explaining in detail the calculation model and calculation process of the present invention, rather than limiting the embodiments of the present invention. For those of ordinary skill in the art, other different forms of changes or variations can be made on the basis of the above description. It is impossible to list all the embodiments here. Any obvious changes or variations derived from the technical solutions of the present invention are still within the protection scope of the present invention.

Claims

1. An incremental 3D reconstruction method based on time-varying scene geometry reasoning, characterized in that: The method specifically comprises the following steps: Step 1: For the newly added image I at time t new , according to the newly added image I new Preliminary determination of the shooting time of the newly added image I new Whether it is a historical image; If the initial judgment is that the newly added image I new If it is not a historical image, continue to add the image I new Execute step 2; If the initial judgment is that the newly added image I new If it is a historical image, go directly to step 3; Step 2: Calculate the newly added image I new The difference between the simulated image and the rendered image is finally determined based on the calculated difference. new Whether it is a historical image; Execute step 3 according to the final result; Step 3: Based on the newly added image I new Generate an image sample set Train_samples, and fit the space-time field function according to the image sample set Train_samples; Step 4: Perform three-dimensional reconstruction of scenes in each historical period based on the space-time field function obtained in step 3, and return to execute step 1 at time t+1.

2. The incremental 3D reconstruction method based on time-varying scene geometry reasoning according to claim 1, characterized in that: The specific process of step 2 is as follows: Step 2.1: Based on the space-time field function fitted at time t-1 and the newly added image I new Imaging geometry model, rendering generation and new image I new Simulation image with consistent viewing angle I sim : I sim ={pix|pix=render(v i )} (1) Among them, v i Any viewing direction of the newly added image; pix is ​​the viewing direction v i The pixel value of the corresponding light; render(·) indicates the image rendering process; {pix} is the set of pixel values ​​corresponding to each ray; Step 22: Calculate the newly added image I new With the simulation image I sim The difference map is used to calculate the new image I new With the simulation image I sim The difference in size Dis 2D ; If you add image I new With the simulation image I sim The difference 2D If it is less than the preset threshold, the newly added image will be processed as a historical image; If you add image I new With the simulation image I sim The difference 2D If it is greater than or equal to the preset threshold, the newly added image will be used as the first image of the new period.

3. The incremental 3D reconstruction method based on time-varying scene geometry reasoning according to claim 2, characterized in that: The specific process of step 22 is as follows: Among them, the newly added image I new With the simulation image I sim The size of is the same, W is the width of the image, and H is the height of the image; diff(I new ,I sim ) indicates the newly added image I new With the simulation image I sim Difference map of sum(·) means summing the values ​​of all pixels in the difference map; Dis 2D To add image I new With the simulation image I sim difference.

4. The incremental 3D reconstruction method based on time-varying scene geometric reasoning according to claim 3, characterized in that: In the step 3, according to the newly added image I new Generate the image sample set Train_samples. The specific process is as follows: (1) If a new image I is added new If it is not a historical image, then an image is extracted from each historical period and the newly added image I new And the extracted images constitute the training image sample set Train_samples: Train_samples=I s1 ,I s2 ,...,I sn ,I new (3) Among them, I s1 ,I s2 ,...,I sn They represent images extracted from various historical periods respectively; n represents the number of historical periods; (2) If a new image I is added new If it is a historical image, determine whether the number of historical images in the historical period is sufficient: If the number of historical images in the historical period is insufficient, the newly added images I new And all historical images in the historical period constitute the training image sample set Train_samples: Train_samples=I1,I2...I n ,I new (4) Among them, I1, I2…I n Respectively represent all images in each historical period; If the number of historical images in the historical period is sufficient, then the newly added image I is obtained from all historical images. new The historical image with the closest perspective is removed from the dataset composed of all historical images, and the remaining historical images and the newly added images I are used new Form the training image sample set Train_samples.

5. The incremental 3D reconstruction method based on time-varying scene geometry reasoning according to claim 4, characterized in that: The space-time field function is fitted according to the image sample set Train_samples, specifically: (1) When a new image I is added new For historical images: Determine whether the number of viewing angles corresponding to the images in the training image sample set Train_samples is less than a threshold; If the number of viewing angles is less than the threshold, each image in the training image sample set Train_samples is processed separately by position coding and frequency coding, and the spatiotemporal field function fitting result is obtained according to the processing result; If the number of viewing angles is greater than or equal to the threshold, directly use each image in the training image sample set Train_samples to obtain the spatiotemporal field function fitting result; (2) When a new image I is added new When it is not a historical image: Each image in the training image sample set Train_samples is processed separately by position coding and frequency coding, and the fitting result of the space-time field function is obtained according to the processing result.

6. The incremental 3D reconstruction method based on time-varying scene geometric reasoning according to claim 5, characterized in that: The specific process of the position encoding is: For any image in the training image sample set Train_samples, the image is moved along the viewing direction v i The position coordinate of a sampling point on the input light is marked as p = (x, y, z), and the input coordinate p = (x, y, z) is position encoded: pos_enc(p)=[a1 cos(2πω1p),a1 sin(2πω1p),…,a m cos(2πω m p),a m sin(2πω m p)] (5) Among them, ω1…ω m is the frequency; a1…a m is the coding weight of different frequencies; pos_enc(p) represents the result of position encoding of the sampling point p=(x, y, z).

7. The incremental 3D reconstruction method based on time-varying scene geometric reasoning according to claim 6, characterized in that: The frequency encoding method is: Step 1: abbreviate pos_enc(p) to p′ and initialize k=0; Step 2: According to the frequency f r =kf d Frequency encode p′, f d is the frequency sampling interval; in, Indicates frequency kf d The encoding weight of , the superscript T indicates transposition; Where V represents the integration space, dv represents the integration variable; Step 3: Input MSI(p) into the first multi-layer perceptron MLP static , through the first multi-layer perceptron MLP static Output static geometry value g static : g static =MLP static (MSI(p)) (8) The time characteristic of the shooting time of the current image is recorded as Φ i , Φ i and the static geometry value g static Input to the second multi-layer perceptron MLP offset Get the geometric value bias g offset : G offset =MLP offset (G static ,Φ i ) (9) Add the static geometry value to the geometry bias value and use the result as the time-varying signed distance function value g i : g i =g static +g offset (10) The location information p of the sampling point and the time feature Φ i , perspective information v i And the time-varying signed distance function value g i As the third multi-layer perceptron MLP alb The input is passed through the MLP alb Output radiation value r i : r i =MLP alb (p,v i ,Φ i ,h i ) (11) Step 4: Determine whether k=K is satisfied, where K is the set threshold: If not, set k=k+1 and return to step 2; If satisfied, the iteration is stopped to obtain the time-varying signed distance function value and the radiation value of the current image at different viewing angles and frequencies.

8. The incremental 3D reconstruction method based on time-varying scene geometric reasoning according to claim 7, characterized in that: The specific process of step 4 is as follows: Step 41: Count the time information of all historical input images to form a timeline, sample the time information on the timeline to obtain each sampling time point, and use each sampling time point to form a sampling time series T′=(t1, t2, …, t L ), t1, t2, …, t L represents the 1st, 2nd, …, Lth sampling time points; Step 4.2: Use the space-time field function obtained by fitting in step 3 to calculate the sampling time point t j The scene model is reconstructed to obtain the sampling time point t j The scene model reconstruction result Mesh: Mesh=MC(g) Where MC(·) is the Marching Cube function, g represents the sampling time point t j The set of time-varying signed distance function values ​​corresponding to the image in all viewing directions; Step 43: Calculate the sampling time point t j With sampling time t j+1 Difference of the reconstructed scene model: in, represents the sampling time point t j The scene model reconstruction result M j With sampling time t j+1 The scene model reconstruction result M j+1 The difference between like is less than the threshold τ, then the sampling time point t j With sampling time t j+1 belong to the same historical period; like is greater than or equal to the threshold τ, then the sampling time point t j With sampling time t j+1 They do not belong to the same historical period; Step 44: After executing step 43 for the scene model reconstruction results of every two adjacent sampling time points, the sampling time series T′ is divided into n independent historical periods; For any historical period obtained by division, the reconstruction results of all scene models included in the historical period are fused, and the fusion result is used as the 3D scene reconstruction result of the historical period. Similarly, the 3D scene reconstruction results of each historical period are obtained respectively; The three-dimensional scene reconstruction results of n historical periods are combined to obtain the three-dimensional reconstruction result of the time-varying scene.

9. The incremental 3D reconstruction method based on time-varying scene geometry reasoning according to claim 8, characterized in that: In the step 44, the fusion processing adopts an equalization method.