Three-dimensional scene reconstruction method and system based on mixed background and illumination correction

By combining basic prediction and illumination correction models to generate high-frequency illumination details, and utilizing a mixed background model and spatiotemporal consistency processing algorithm, the image quality problem under complex illumination in 3D scene reconstruction is solved, achieving high-quality 3D scene reconstruction.

CN120807803APending Publication Date: 2025-10-17NO 30 INST OF CHINA ELECTRONIC TECH GRP CORP +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511049362.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing 3D scene reconstruction technology has poor reconstructed image quality when dealing with complex lighting effects, especially in terms of high-frequency lighting details and background realism, resulting in blurred details and low integrity of static scenes.

Method used

A basic prediction model combined with an illumination correction model is used to perform foreground modeling to generate high-frequency illumination details. A background image containing high-frequency details is generated through a mixed background model. At the same time, a spatiotemporal consistency transient processing algorithm is introduced to distinguish transient objects from static scenes.

Benefits of technology

The realism and detail richness of the reconstructed images are significantly improved, the structural integrity of the static scenes is ensured, and the quality and immersiveness of the reconstructed images are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120807803A_ABST
    Figure CN120807803A_ABST
Patent Text Reader

Abstract

The invention discloses a mixed background and illumination correction-based three-dimensional scene reconstruction method and system, and the method comprises the steps: generating a rendering foreground image through a double-path appearance model of basic prediction and illumination correction, and effectively capturing and reconstructing high-frequency illumination details, such as highlight and hard shadow, in a scene. Therefore, the problem that details are fuzzy and distorted under complex illumination in the prior art is solved; meanwhile, a mixed background model is adopted, and on the basis that a traditional background low-frequency SH coefficient is generated, high-frequency detail parameters used for controlling a programmed noise function are additionally obtained; therefore, a background image containing abundant details such as cloud layers and textures can be generated; and finally, a real transient object and a difficult static scene part can be distinguished more accurately by introducing a transient processing mechanism for space-time consistency inspection, so that the static scene details are prevented from being shielded by mistake, and the finally reconstructed static scene structure is more accurate. Therefore, the quality of the reconstructed image is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of three-dimensional scene reconstruction, and particularly relates to a three-dimensional scene reconstruction method and system based on mixed background and illumination correction. BACKGROUND

[0002] Three-dimensional scene reconstruction aims to synthesize new images under a new three-dimensional view from a set of in-the-wild two-dimensional images, which usually contains dramatic appearance changes caused by different times, weather or transient objects (such as pedestrians and vehicles); where the current mainstream technical route can be divided into two categories, namely the method based on neural radiation field (NeRF) and the method based on three-dimensional Gaussian splash (3DGS).

[0003] In specific applications, the method based on three-dimensional Gaussian splash (3DGS) uses an explicit three-dimensional Gaussian point cloud to represent the scene, which realizes real-time rendering, but it has inherent deficiencies in dealing with in-the-wild image sets; where in the subsequent improvement work, such as Splatfacto-W, the technology introduces appearance features and background models, which to some extent solves this problem.

[0004] In specific implementation, the specific scheme of Splatfacto-W is as follows: appearance modeling: give each Gaussian point an appearance feature, and combine the appearance embedding vector of each image to predict its color (represented by spherical harmonic function SH coefficient) through a multi-layer perception (MLP) network; transient object processing: by calculating the loss between the rendered image and the real image, it is heuristically considered that the high loss area is a transient object and is masked, so as to ignore these areas in optimization; background modeling: use another MLP to predict the spherical harmonic function coefficients of the background (such as the sky) according to the appearance embedding vector of the image, and consider the background as infinitely far away.

[0005] Although the existing technology represented by Splatfacto-W has made significant progress, it still has the following objective shortcomings: (1) The existing method cannot handle complex lighting effects well. Although the appearance model of the existing method can handle the overall lighting atmosphere change, it has slow convergence speed and poor reconstruction effect for high-frequency and localized lighting effects (such as sharp shadow edges and highlight reflections) caused by direct sunlight at a specific time. The model is prone to misclassify the large loss of these areas as transient objects, resulting in loss of details; (2) The background model is too simplified and lacks realism. The existing technology uses low-order spherical harmonics to model the background, which can only represent low-frequency information (such as uniform sky on a sunny or cloudy day) with smooth color changes. However, for backgrounds containing complex and high-frequency details (such as cloud layers of different shapes and the outlines of distant mountains), the model cannot effectively represent them, resulting in a lack of realism and details in the synthesized background; (3) The robustness of the transient object mask is insufficient. When facing the above complex lighting areas, the heuristic mask strategy based on rendering loss will misjudge the hard shadows or highlights in the static scene as transient objects, thereby incorrectly "shielding" them. This affects the integrity and accuracy of the static scene.

[0006] Therefore, based on the foregoing problems, how to provide a three-dimensional scene reconstruction method based on mixed background and lighting correction to improve the quality of the reconstructed image has become a problem to be solved. SUMMARY

[0007] The technical problem to be solved by the present application is the poor quality of the reconstructed image in a three-dimensional scene. The present application provides a three-dimensional scene reconstruction method and system based on mixed background and lighting correction, which solves the problem of poor quality of the reconstructed image caused by poor handling of complex lighting effects, lack of realism and details in the synthesized background, and low integrity and accuracy of the static scene reconstruction in the traditional technology.

[0008] The present application is implemented by the following technical solutions: In a first aspect, a three-dimensional scene reconstruction method based on mixed background and lighting correction is provided, comprising: initializing reconstruction parameters, wherein the reconstruction parameters include attribute information of each Gaussian point in the scene and image information of each training image in the training set; selecting a training image from the training set to obtain the foreground SH coefficient of each Gaussian point and the color correction amount based on the basic color model and the lighting correction model using the attribute information of each Gaussian point and the image information of the selected training image, wherein the color correction amount is used to represent the color deviation caused by high-frequency effects; generating a rendered foreground image according to the color correction amount and the foreground SH coefficient of each Gaussian point; generating a background low-frequency SH coefficient and a high-frequency detail parameter for controlling at least one noise function based on the image information of the selected training image and using a mixed background model; generate a rendered background image containing high-frequency details by using the background low-frequency SH coefficients and the high-frequency detail parameters; generate an initial rendering image based on the rendered foreground image and the rendered background image; obtain a neighbor image of the selected training image, and use the neighbor image and a spatiotemporal consistency transient processing algorithm to distinguish interference pixel points caused by transient objects in the initial rendering image to generate a transient object mask; generate a final rendering image by using the transient object mask and the initial rendering image; update the reconstruction parameters and the model parameters of the reconstruction model by using the image loss between the final rendering image and the selected training image, and after the update, reselect a training image from the training set until the iteration stopping condition is met to obtain optimal reconstruction parameters and an optimal reconstruction model, so as to generate a reconstruction rendering image under different perspectives by using the optimal reconstruction parameters and the optimal reconstruction model, wherein the reconstruction model includes a basic color model, an illumination correction model and a mixed background model.

[0009] Based on the above disclosure, when foreground modeling is performed, the present application uses a basic prediction model to derive the foreground SH coefficients of each Gaussian point, and then introduces an illumination correction model to determine the color correction amount of each Gaussian point to compensate for color deviation caused by high-frequency effects such as hard shadows and highlights. Then, based on the foreground SH coefficients of each Gaussian point and the corresponding color correction amount, a rendered foreground image is generated. Meanwhile, the present application also uses a brand-new mixed background model to determine the high-frequency detail parameters for controlling one or more procedural noise functions (such as Perlin noise or Simplex noise) while generating background low-frequency SH coefficients to generate high-frequency details such as cloud textures. In this way, based on the background low-frequency SH coefficients and the high-frequency detail parameters, a rendered background image with rich details and reality can be generated. Then, after the present application generates an initial rendering image based on the foregoing foreground and background images, a spatiotemporal consistency checking transient processing mechanism is introduced, that is, by using the image information of the neighbor image of the current image and combining a spatiotemporal consistency transient processing algorithm, interference pixel points caused by transient objects in the initial rendering image are distinguished to generate a transient object mask. Then, based on the transient object mask, a final rendering image can be obtained. Then, based on the image loss between the final rendering image and the real image, the reconstruction parameters and the model parameters of each model are updated in reverse until the iteration stopping condition is met, and then the optimal reconstruction parameters and the optimal reconstruction model can be obtained. Finally, by using the optimal reconstruction parameters and the optimal reconstruction model, a reconstruction rendering image under different perspectives can be generated.

[0010] Through the above design, the application generates a rendered foreground image through a two-way appearance model of basic prediction + illumination correction, can effectively capture and reconstruct high-frequency lighting details such as highlights and hard shadows in the scene, so that the synthesized foreground image is more realistic, thereby solving the problem of blurred and distorted details under complex lighting in the prior art; at the same time, a hybrid background model is adopted, and on the basis of generating traditional background low-frequency SH coefficients, high-frequency detail parameters for controlling a programmed noise function are additionally obtained; in this way, a background image containing rich details such as cloud layers and textures can be generated; finally, through the introduction of a transient processing mechanism of spatiotemporal consistency checking, a truly transient object and a difficult static scene part can be more accurately distinguished, thereby avoiding the false shielding of static scene details, and further making the finally reconstructed static scene structure more complete and cleaner; thus, the application significantly improves the quality of the reconstructed image, and is very suitable for large-scale application and promotion.

[0011] In a possible design, the attribute information of any Gaussian point includes appearance feature vectors and positions of the any Gaussian point, and the image information of any training image includes a camera pose and appearance embedding vectors corresponding to the any training image; Wherein, the foreground SH coefficients of each Gaussian point and the color correction amount of each Gaussian point are obtained based on the basic color model and the illumination correction model by using the attribute information of each Gaussian point and the image information of the selected training image, including: The appearance embedding vectors of the selected training image and the appearance feature vectors of each Gaussian point are input into the basic color model to obtain the foreground SH coefficients corresponding to the foreground color of each Gaussian point. The foreground SH coefficients of each Gaussian point, the appearance feature vectors of each Gaussian point and the observation direction vectors of each Gaussian point are input into the illumination correction model to obtain the color correction amount of each Gaussian point, wherein the observation direction vector of any Gaussian point is obtained according to the position of the any Gaussian point and the camera pose.

[0012] In a possible design, the illumination correction model adopts a lightweight MLP model, wherein the lightweight MLP model includes 2 hidden layers, and each hidden layer contains 64 neurons.

[0013] In a possible design, the image information of any training image includes a camera pose and appearance embedding vectors corresponding to the any training image, and the hybrid background model includes: a backbone network layer, a first output head and a second output head. Wherein, the backbone network layer is used to encode the appearance embedding vectors of the selected training image into 128-dimensional feature vectors, and input the feature vectors into the first output head and the second output head respectively. The first output head is used to predict the background low-frequency SH coefficients based on the input feature vectors. a second output head configured to predict, based on the input feature vector, a base frequency for controlling maximum cloud cluster stretch and size, an octave for controlling cloud layer hierarchy, a frequency gain for controlling frequency increase ratio between adjacent octaves, an amplitude gain for controlling amplitude decay rate between adjacent octaves, a noise offset for controlling cloud layer layout, a first color gradient for controlling starting color of cloud layer color mapping, and a second color gradient for controlling ending color of cloud layer color mapping, so as to generate the high frequency detail parameters by using the base frequency, the octave, the frequency gain, the amplitude gain, the noise offset, the first color gradient, and the second color gradient.

[0014] In one possible design, the backbone network layer employs an MLP network, where the MLP network includes 3 hidden layers, each of which includes 128 neurons, and each of which corresponds to a ReLU activation function, and the first output head employs a first linear fully connected layer, and the second output head employs a second linear fully connected layer including 15 output neurons, so as to correspond to 7 types of high frequency detail parameters respectively.

[0015] In one possible design, a rendered background image including high frequency details is generated by using background low frequency SH coefficients and high frequency detail parameters, including: obtaining a direction vector of a scene light; generating a background low frequency base color based on the direction vector of the scene light and the background low frequency SH coefficients; generating a high frequency background color according to the high frequency detail parameters; generating the rendered background image by using the background low frequency base color and the high frequency background color.

[0016] In one possible design, by using a nearest neighbor image and a spatiotemporal consistency transient processing algorithm, interference pixel points caused by a transient object in a selected training image are distinguished and selected to generate a transient object mask, including: calculating an L1 loss map between the initial rendered image and the selected training image pixel by pixel, where each pixel point in the L1 loss map corresponds to a pixel loss value; calculating a neighboring L1 loss map between the initial rendered image and each nearest neighbor image pixel by pixel; for any pixel point in the L1 loss map, determining pixel points corresponding to the position of the any pixel point in each neighboring L1 loss map as the nearest neighbor pixel points of the any pixel point; calculating the loss mean value of the nearest neighbor pixel points of the any pixel point, and obtaining a neighboring average loss map after polling all pixel points in the L1 loss map; performing a weighted sum on the L1 loss map and the neighboring average loss map to obtain a consistency loss map; based on the consistency loss map, calculating a mask threshold at a current iteration; for any one pixel point in the initial rendering map, if the consistency loss value of the any one pixel point is greater than the mask threshold, taking the any one pixel point in the initial rendering map as an interference pixel point, and obtaining a plurality of interference pixel points after polling all the pixel points in the initial rendering map; changing the pixel value of each interference pixel point in the initial rendering map to 0, and changing the pixel value of each remaining pixel point to 1 to obtain the transient object mask.

[0017] In one possible design, calculating the mask threshold at the current iteration includes: calculating an average value of all consistency loss values in the consistency loss map to obtain a current overall consistency loss; obtaining a maximum overall consistency loss and a minimum overall consistency loss in all iteration processes before the current iteration; calculating a mask percentage according to the current overall consistency loss, the maximum overall consistency loss and the minimum overall consistency loss; based on the mask percentage, determining the mask threshold from the consistency loss map.

[0018] In one possible design, calculating the mask percentage according to the current overall consistency loss, the maximum overall consistency loss and the minimum overall consistency loss includes: calculating the mask percentage according to the following formula; k = [(Lconsistent_current-Lconsistent_min) / (Lconsistent_max-Lconsistent_min)] × (Permax-Permin) + Permin In the formula, k represents the mask percentage, Lconsistent_current represents the overall consistency loss, Lconsistent_min represents the minimum overall consistency loss, Lconsistent_max represents the maximum overall consistency loss, and Permax and Permin represent the maximum mask percentage and the minimum mask percentage, respectively.

[0019] In a second aspect, a three-dimensional scene reconstruction system based on mixed background and illumination correction is provided, including: The initialization module is configured to initialize reconstruction parameters, wherein the reconstruction parameters comprise attribute information of each Gaussian point in the scene and image information of each training image in the training set. The enhanced appearance model module is configured to select a training image from the training set, to obtain foreground SH coefficients of each Gaussian point and color correction amounts based on the base color model and the illumination correction model by using the attribute information of each Gaussian point and the image information of the selected training image, wherein the color correction amounts are used to represent color deviations caused by high-frequency effects. The enhanced appearance model module is further configured to generate a rendered foreground image according to the color correction amounts and the foreground SH coefficients of each Gaussian point. The hybrid background model module is configured to generate background low-frequency SH coefficients and high-frequency detail parameters for controlling at least one noise function by using a hybrid background model based on the image information of the selected training image. The hybrid background model module is further configured to generate a rendered background image containing high-frequency details by using the background low-frequency SH coefficients and the high-frequency detail parameters. The rasterization module is configured to generate an initial rendered image based on the rendered foreground image and the rendered background image. The spatio-temporal consistency transient processing module is configured to obtain a nearest neighbor image of the selected training image, and to distinguish interference pixel points caused by a transient object in the initial rendered image by using the nearest neighbor image and a spatio-temporal consistency transient processing algorithm, to generate a transient object mask. The spatio-temporal consistency transient processing module is configured to generate a final rendered image by using the transient object mask and the initial rendered image. The update module is configured to update the reconstruction parameters and model parameters of a reconstruction model by using image loss between the final rendered image and the selected training image, and to reselect a training image from the training set after the update until an iteration stop condition is met, to obtain optimal reconstruction parameters and an optimal reconstruction model, and to generate reconstruction rendered images under different viewing angles by using the optimal reconstruction parameters and the optimal reconstruction model, wherein the reconstruction model comprises the base color model, the illumination correction model, and the hybrid background model.

[0020] In a third aspect, a device for three-dimensional scene reconstruction based on a hybrid background and illumination correction is provided. The device is taken as an example of an electronic device, and comprises a memory, a processor, and a transceiver connected in sequence and in communication. The memory is configured to store a computer program, the transceiver is configured to receive and send messages, and the processor is configured to read the computer program and execute the three-dimensional scene reconstruction method based on a hybrid background and illumination correction according to the first aspect or any possible design of the first aspect.

[0021] In a fourth aspect, a storage medium is provided, and the storage medium has stored thereon instructions which, when executed on a computer, perform the method for reconstructing a three-dimensional scene based on mixed background and illumination correction according to the first aspect or any possible design of the first aspect.

[0022] In a fifth aspect, a computer program product is provided, and the computer program product has instructions which, when executed on a computer, cause the computer to perform the method for reconstructing a three-dimensional scene based on mixed background and illumination correction according to the first aspect or any possible design of the first aspect.

[0023] Compared with the prior art, the present application has the following advantages and beneficial effects: (1) By means of the dual-path appearance model of "base + correction", the present application can effectively capture and reconstruct high-frequency illumination details such as highlights and hard shadows in the scene, so that the synthesized image is more realistic, thereby solving the problem of blurred and distorted details under complex illumination in the prior art.

[0024] (2) The mixed background model can generate a background containing rich details such as clouds and textures without significantly increasing the computational burden, and the appearance of the background can change with the embedding vectors (representing different weather or time) of different images, so that the sense of immersion and multi-view consistency are improved.

[0025] (3) By introducing the transient processing mechanism of spatiotemporal consistency checking, the present application can more accurately distinguish between real transient objects and difficult static scene parts, so that the error shielding of static scene details is avoided, and the finally reconstructed static scene structure is more complete and cleaner. BRIEF DESCRIPTION OF DRAWINGS

[0026] In order to more clearly illustrate the technical solutions of the example embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be considered as limiting the scope. For those skilled in the art, other related drawings can also be obtained without creative labor. In the drawings: Figure 1 The flowchart of the method for reconstructing a three-dimensional scene based on mixed background and illumination correction provided by the embodiments of the present application; Figure 2 The structural schematic diagram of the system for reconstructing a three-dimensional scene based on mixed background and illumination correction provided by the embodiments of the present application; Figure 3 The structural schematic diagram of the enhanced appearance model module provided by the embodiments of the present application; Figure 4 The structural schematic diagram of the mixed background model module provided by the embodiments of the present application; Figure 5 A structural schematic diagram of an electronic device provided by an embodiment of the present application is shown. DETAILED DESCRIPTION

[0027] In order to make the objectives, technical solutions and advantages of the present application clearer, further detailed description will be given below in combination with embodiments and drawings. The exemplary embodiments of the present application and the description thereof are only used to explain the present application and do not limit the present application. It should be understood that although the terms first, second, etc. can be used herein to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another unit. For example, the first unit can be called the second unit, and similarly, the second unit can be called the first unit, without departing from the scope of the exemplary embodiments of the present application.

[0028] Embodiment: Referring to Figure 1 As shown in the drawings, the three-dimensional scene reconstruction method based on mixed background and illumination correction provided by the present embodiment makes key improvements to the appearance model, the background model and the transient processing mechanism on the basis of the prior art (such as Splatfacto-W), so that the present application can effectively capture and reconstruct high-frequency lighting details such as highlights and hard shadows in the scene when generating the foreground image, thereby improving the fidelity of the generated foreground image. At the same time, based on the mixed background model, the background image containing rich details such as cloud layers and textures can be generated, and through the introduction of the transient processing mechanism of the spatiotemporal consistency test, the truly transient objects and the difficult static scene parts can be more accurately distinguished, avoiding the false shielding of the static scene details, so that the finally reconstructed static scene structure is more complete and cleaner. In this way, the present application significantly improves the quality of the reconstructed image, so that it is very suitable for large-scale application and promotion. For example, the present method can be but is not limited to running on the scene reconstruction end side. Optionally, the scene reconstruction end can be but is not limited to a computer or a server. It can be understood that the foregoing execution subject does not constitute a limitation on the embodiments of the present application. Correspondingly, the running steps of the present method can be but are not limited to the steps S1-S9 shown below.

[0029] S1. Initialize the reconstruction parameters, wherein the reconstruction parameters include attribute information of each Gaussian point in the scene and image information of each training image in the training set; in specific implementation, step S1 is an initialization step, i.e., the initialization of the attribute information of the Gaussian points and the image information of the training images is performed. Optionally, the attribute information of any Gaussian point includes the appearance feature vector and the position of the any Gaussian point, and the image information of any training image includes the camera pose and the appearance embedding vector corresponding to the any training image; wherein, when the initialization is performed, an appearance feature vector (for example, the vector is 72-dimensional) and a position are randomly generated for each Gaussian point in the scene. The vector and the position are learnable, and their values are continuously updated through the back propagation algorithm in the subsequent training process, so as to achieve the optimal values after the training iteration is completed.

[0030] Similarly, before the training starts, a unique appearance embedding vector (the vector is 48-dimensional) is randomly initialized for each training image in the training set. The vector is also learnable and is optimized in the training to capture the unique global information of the corresponding image, such as overall illumination (noon, dusk), weather (sunny, cloudy), etc. The camera pose of the corresponding view of each training image is obtained by a structure from motion (SFM) algorithm. In this way, after the initialization of the reconstruction parameters is completed, the training of the reconstruction model (which includes the basic color model, the illumination correction model and the mixed background model) and the update of the reconstruction parameters can be performed by using the aforementioned training images and the attribute information of the Gaussian points, so that the optimal reconstruction parameters and the optimal reconstruction model are obtained in the iteration process, and then the optimal reconstruction parameters and the optimal reconstruction model are used to generate the reconstruction rendering images of the same object under different views in the subsequent actual use.

[0031] Further, after the initialization of the reconstruction parameters is completed, an iterative optimization cycle can be performed. Specifically, foreground modeling is first performed, and the process is shown in the following step S2.

[0032] S2. Select a training image from the training set to obtain the foreground SH coefficients of each Gaussian point and the color correction amount based on the basic color model and the illumination correction model by using the attribute information of each Gaussian point and the image information of the selected training image, wherein the color correction amount is used to represent the color deviation caused by the high-frequency effect.

[0033] In specific application, step S2 aims to accurately model the high-frequency illumination effect. Unlike the existing technology of predicting color by a single MLP, the embodiment adopts a dual-channel modeling structure of "basic + correction", and the implementation process is shown in the following steps S21 and S22.

[0034] S21. The appearance embedding vector of the selected training image and the appearance feature vector of each Gaussian point are input into the base color model to obtain the foreground SH coefficient corresponding to the foreground color of each Gaussian point; in specific applications, the base color model is responsible for capturing the overall, low-frequency lighting atmosphere of the scene, and adopts a multi-layer perception (MLP) model, for example, which can be an MLP model containing 3 hidden layers, each with a width of 256 neurons, and using ReLU as the activation function; wherein the base color model is the same as the traditional MLP color prediction principle, and will not be described here.

[0035] After completing the prediction of the SH coefficient of the foreground base color, the embodiment further adds a light correction model to generate a color correction amount for each Gaussian point for precisely modeling the high-frequency lighting effect, wherein the specific process of generating the color correction amount based on the light correction model is shown in the following step S22.

[0036] S22. The foreground SH coefficient of each Gaussian point, and the appearance feature vector and observation direction vector of each Gaussian point are input into the light correction model to obtain the color correction amount of each Gaussian point, wherein the observation direction vector of any Gaussian point is obtained according to the position of the any Gaussian point and the camera pose.

[0037] In the embodiment, the observation direction refers to a normalized unit vector from the center position (μ) of a Gaussian point in three-dimensional space to the optical center of the camera (i.e. the optical center of the camera at the camera pose corresponding to the selected training image), which indicates the angle at which the camera views the three-dimensional Gaussian point; therefore, the observation direction vector of each Gaussian point can be pre-set according to its position and camera pose.

[0038] Furthermore, the light correction model can adopt, but is not limited to, a lightweight MLP model with 2 hidden layers, each containing 64 neurons, wherein the input of the lightweight MLP model is the output of the main MLP (i.e. the base color model), the appearance feature of each Gaussian point and the observation direction vector, and the output is a color correction amount ΔC; therefore, the lightweight MLP model learns and compensates for color deviations caused by high-frequency effects such as hard shadows and highlights, so that the model can converge more quickly and accurately in these difficult areas.

[0039] Based on this, the main MLP undertakes the main work of scene lighting modeling, needs to understand the high-dimensional input and output a complete set of SH coefficients describing the distribution of light field, while the light correction MLP undertakes the “expert correction” work, which fine-tunes the high-frequency, localized lighting effects such as highlights and hard shadows based on the former, and because its task target is more single and clear, it does not need a complex network structure.

[0040] Thus, by means of the primary-secondary model combination and the coarse-fine adjustment, the embodiment can achieve accurate modeling of complex lighting effects without significantly increasing the overall computational burden.

[0041] After obtaining the color correction amount and the foreground SH coefficient of each Gaussian point based on the foregoing steps S21 and S22, a foreground image capable of effectively capturing high-frequency lighting details such as highlights and hard shadows in the scene can be generated based on the foregoing, and the rendering process is shown in the following step S3.

[0042] S3. Generate a rendered foreground image according to the color correction amount and the foreground SH coefficient of each Gaussian point; in a specific implementation, the foreground SH coefficient output by the primary MLP is first combined with the observation direction vector to calculate a basic color in a standard manner; then, the color correction amount ΔC output by the lighting correction model is added to the basic color to obtain the final color of each Gaussian point, and a rendered foreground image is generated.

[0043] Specifically, the process of generating the basic color of each Gaussian point is as follows: based on the observation direction vector of each Gaussian point, each order of spherical harmonic function basis function (spherical harmonic function basis function obtained based on the foreground SH coefficient) is evaluated and weighted summed; of course, using the SH coefficient to generate the basic color is a common way in the Splatfacto-W scene reconstruction technology, and the principle will not be repeated.

[0044] Thus, after the rendered foreground image is generated, background modeling can be performed, and the process is shown in the following step S4.

[0045] S4. Based on the image information of the selected training image, and using a hybrid background model, generate a background low-frequency SH coefficient and a high-frequency detail parameter for controlling at least one noise function.

[0046] In a specific application, the embodiment discards the single low-frequency SH model of the prior art and adopts a hybrid strategy of “low-frequency base + programmed high-frequency details” to generate a rendered background image. Specifically, the embodiment provides a new background model to generate a background containing clouds, textures and other rich details.

[0047] In a specific implementation, the aforementioned hybrid background model can include, but is not limited to, a backbone network layer, a first output head and a second output head, wherein the input is still the appearance embedding vector of the selected training image, but the output is expanded into two parts, i.e., the background low-frequency SH coefficient (the same as the prior art, used to generate the overall smooth color tone of the background) and the high-frequency detail parameter (i.e., the procedural noise parameter), which is a series of parameters used to control one or more procedural noise functions (such as Perlin noise or Simplex noise), such as the frequency, amplitude, level (octaves) of the noise and the color gradient key points used to color the noise.

[0048] Based on this, the backbone network layer is responsible for encoding the input appearance embedding vector into a high-dimensional feature representation, and the structure adopts an MLP network with 3 hidden layers, each containing 128 neurons, and the corresponding activation function of each hidden layer is a ReLU activation function. The output of the last hidden layer is a 128-dimensional feature item, which will be the common input of the subsequent two output heads, i.e., the backbone network layer is used to encode the appearance embedding vector of the selected training image into a 128-dimensional feature vector, and the feature vector is input into the first output head and the second output head, respectively.

[0049] Then, the 128-dimensional feature vector is sent into two independent output heads in parallel to predict different types of parameters; wherein the first output head is used to predict the background low-frequency SH coefficient based on the input feature vector.

[0050] In this embodiment, the first output head is the background low-frequency SH coefficient prediction head, which predicts the spherical harmonic (SH) coefficient used to generate the low-frequency, smooth background color, and the structure adopts a linear fully connected layer (i.e., the first linear fully connected layer); specifically, the first linear fully connected layer maps the 128-dimensional feature vector to the required dimension of the output SH coefficient, using 3-order SH (lmax=3), so there are a total of 16 coefficients; and for the RGB three color channels, the total dimension is 16x3=48; therefore, the dimension of this layer is changed from 128 to 48.

[0051] At the same time, the second output head is the prediction head of the procedural noise parameter, which predicts the parameter used to control the procedural noise function to generate high-frequency details such as clouds; specifically, the second output head adopts another independent linear fully connected layer (i.e., the second linear fully connected layer) containing 15 output neurons, and the dimension is changed from the input 128-dimensional to 15-dimensional (each dimension corresponds to an output neuron), corresponding to the following 7 different types of high-frequency detail parameters.

[0052] Thus, the second output head is used to predict, based on the input feature vector, a base frequency for controlling the maximum cloud cluster stretch and size, an octave for controlling the cloud layer hierarchy, a frequency gain for controlling the frequency increase ratio between adjacent octaves, an amplitude gain for controlling the amplitude decay rate between adjacent octaves, a noise offset for controlling the cloud layer layout, a first color gradient for controlling the starting color of the cloud layer color mapping, and a second color gradient for controlling the ending color of the cloud layer color mapping, so as to generate the high-frequency detail parameters using the base frequency, the octave, the frequency gain, the amplitude gain, the noise offset, the first color gradient, and the second color gradient.

[0053] Further, the following Table 1 shows the specific description table of the aforementioned 7 high-frequency detail parameters.

[0054] Table 1 is a high-frequency detail parameter description table.

[0055] Table 1

[0056] In the present embodiment, the aforementioned description, the second output head is a linear fully connected layer from 128-dimensional input to 15-dimensional output, thus different output neurons of the output head need to be applied with corresponding activation functions to ensure that the predicted parameter values fall within the effective range required by their physical meaning; wherein the activation function corresponding to each high-frequency detail parameter is shown in Table 2 below.

[0057] Table 2 is a high-frequency detail parameter corresponding activation function table.

[0058] Table 2

[0059] Thus, based on the aforementioned hybrid background model using "low-frequency base + programmed high-frequency detail", after generating the corresponding background low-frequency SH coefficients and high-frequency detail parameters, the two can be combined to generate a background containing cloud layers, textures and other rich details, the process of which is shown in the following step S5.

[0060] S5. Utilize the background low-frequency SH coefficients and high-frequency detail parameters to generate a rendered background image containing high-frequency details; in specific applications, for example, but not limited to, the following steps S51-S54 can be used to generate the aforementioned rendered background image.

[0061] S51. Obtain a direction vector of a scene light; in a specific implementation, each training image actually corresponds to a virtual camera (i.e., camera pose), which generates a light for each pixel point on the image, i.e., a unique and straight ray is emitted from the camera center, which passes through the center point of the pixel on the image plane and is directed to the three-dimensional scene far away; therefore, each light is defined by two parts, the light origin (Origin): the position of the camera center; the light direction (Direction): the unit vector from the camera center to the center point of the specific pixel, denoted as dray; in this way, the direction vector of the scene light can be pre-set; and after obtaining the direction vector, the background low-frequency base color can be generated in combination with the direction vector and the background low-frequency coefficient, and the process is shown in the following step S52.

[0062] S52. Generate a background low-frequency base color based on the direction vector of the scene light and the background low-frequency SH coefficient; in a specific implementation, the low-frequency background SH coefficient output by the backbone network layer in the mixed background model is , and for the direction vector of any scene light from the camera, the background low-frequency SH coefficient can be weighted and summed based on the direction vector, and the calculation formula is:

[0063] In the formula, represents the background low-frequency base color, is an activation function, represents the order, represents the maximum order (usually 3), represents the index of the order, represents the value of the SH basis function (the SH basis function with the background low-frequency SH coefficient) at the direction vector of any scene light.

[0064] In this way, based on the foregoing formula, the background low-frequency base color corresponding to each scene light can be calculated; and then, the high-frequency background color can be generated using the high-frequency detail parameter, and the process is shown in the following step S53.

[0065] S53. Generate a high-frequency background color according to the high-frequency detail parameter; in a specific application, the foregoing high-frequency background color can be generated by using the following steps S53a-S53c.

[0066] S53a. According to the direction vector of the scene light and the noise offset, the input coordinates of the three-dimensional noise function are determined; in a specific implementation, the same operation is performed for each light in the scene, that is, the direction vector of each light is added to the predicted noise offset to obtain the input coordinates (i.e., sampling coordinates) of the three-dimensional noise function Simplex noise function; then, based on this, the scalar noise value is calculated, the process of which is shown in the following step S53b.

[0067] S53b. Based on the base frequency, octave, frequency gain, amplitude gain parameters and the input coordinates, the three-dimensional noise function is called repeatedly to perform fBm superposition to obtain the scalar noise value; in a specific implementation, a standard three-dimensional procedural noise function such as Perlin noise or Simplex noise is called, and the aforementioned hierarchical parameters and the input coordinates corresponding to each light are input, and the function outputs a smooth random value (scalar) in the range [-1, 1], and the smooth random value is the scalar noise value of each light; then, using the scalar noise value, the high-frequency background color corresponding to each light can be obtained, the process of which is shown in the following step S53c.

[0068] S53c. Based on the scalar noise value, the first color gradient and the second color gradient are linearly interpolated to obtain the high-frequency background color after linear interpolation; in this embodiment, linear interpolation is used, and each scalar noise value is used to sample between the predicted color gradients (Color0 to Color1) to obtain the high-frequency background color corresponding to each light.

[0069] In this way, after generating the high-frequency background color through the foregoing steps S53a-S53c, it is synthesized with the low-frequency background color to obtain the rendered background image.

[0070] Thus, through the foregoing step S5 and its sub-steps, the embodiment adopts the hybrid strategy of "low-frequency base + procedural high-frequency details", and can generate a rendered background image containing cloud layers, textures and other rich details; at the same time, the appearance of the background can change with the embedding vector of different images (representing different weather or time), based on which the immersion and multi-view consistency are also improved.

[0071] After completing the foreground modeling and background modeling, the initial rendering image can be generated, the process of which is shown in the following step S6.

[0072] S6. Based on the rendered foreground image and the rendered background image, an initial rendering image is generated; in specific implementation, a differential rasterizer may be used to alpha blend the rendered foreground image and the rendered background image to generate the initial rendering image; then, the embodiment introduces a spatiotemporal consistency transient processing mechanism to solve the problem that the heuristic mask strategy based on rendering loss in the prior art may misjudge the hard shadow or highlight in the static scene as a transient object, thereby incorrectly shielding it and affecting the accuracy and integrity of the static scene.

[0073] The spatiotemporal consistency transient processing process is shown in the following step S7.

[0074] S7. The neighbor images of the selected training image are obtained, and the neighbor images and the spatiotemporal consistency transient processing algorithm are used to distinguish the interference pixel points caused by the transient object in the initial rendering image to generate a transient object mask; in the embodiment, the spatial positions of the current view (i.e., the selected training image) and the remaining training images are calculated according to the camera pose of the current view, and then the remaining training images are sorted in ascending order of spatial distance, and the first N training images in the sorted order are used as the neighbor images; in this way, after obtaining the neighbor images, the embodiment distinguishes the interference pixel points caused by the transient object based on the loss of the current rendering image and the loss of the neighbor images, and then generates an accurate transient object mask; the generation process of the foregoing transient object mask may include but is not limited to the following steps S71-S78.

[0075] S71. The L1 loss map between the initial rendering image and the selected training image is calculated, wherein each pixel point in the L1 loss map corresponds to a pixel loss value; in specific implementation, the pixel values of the selected training image and the initial rendering image are subtracted pixel by pixel, and the absolute value is taken to obtain the L1 loss map; then, the L1 loss map between the initial rendering image and each neighbor image can be calculated, and the process is shown in the following step S72.

[0076] S72. The neighbor L1 loss map between the initial rendering image and each neighbor image is calculated; in specific implementation, the calculation process of the neighbor L1 loss map can refer to the foregoing step S71, and the principle is not described again.

[0077] In this way, after the neighbor L1 loss map between the initial rendering image and each neighbor image is calculated, the neighbor average loss map can be generated based thereon, and the process is shown in the following steps S73 and S74.

[0078] S73. For any pixel point in the L1 loss map, determine the pixel points in each adjacent L1 loss map corresponding to the position of the any pixel point as the near neighbor pixel points of the any pixel point; in this embodiment, in each adjacent L1 loss map, after finding the pixel points corresponding to the position of the any pixel point, the mean value of the loss values of the found pixel points can be calculated to generate an adjacent average loss map, the process of which is shown in the following step S74.

[0079] S74. Calculate the loss mean value of the near neighbor pixel points of the any pixel point, and after all the pixel points in the L1 loss map are polled, an adjacent average loss map is obtained.

[0080] After obtaining the adjacent average loss map, it can be weighted and summed with the L1 loss map to obtain a consistency loss map, the process of which is shown in the following step S75.

[0081] S75. Weighted sum the L1 loss map and the adjacent average loss map to obtain a consistency loss map; in specific implementation, for example, the weights of the L1 loss map and the adjacent average loss map are both 0.5; thus, after weighted sum to obtain the consistency loss map, the shielding threshold of this iteration can be determined to distinguish the interference pixel points caused by the transient object in the initial rendering map according to the shielding threshold; wherein, the calculation process of the shielding threshold is shown in the following step S76.

[0082] S76. Based on the consistency loss map, calculate the shielding threshold of the current iteration; in specific application, for example, but not limited to, the following steps S76a-S76d can be used to calculate the shielding threshold of the current iteration.

[0083] S76a. Calculate the average value of all consistency loss values in the consistency loss map to obtain the current overall consistency loss; in this embodiment, it is equivalent to summing and averaging the consistency loss values of all pixel points in the consistency loss map to obtain the current overall consistency loss; then, the historical extreme value can be obtained, the process of which is shown in the following step S76b.

[0084] S76b. Obtain the maximum overall consistency loss and the minimum overall consistency loss in all iteration processes before the current iteration; in specific application, it is equivalent to obtaining the maximum and minimum overall consistency losses in all previous iteration steps; then, combined with the current overall consistency loss, the mask percentage is calculated, the calculation process of which is shown in the following step S76c.

[0085] S76c. According to the current overall consistency loss, the maximum overall consistency loss and the minimum overall consistency loss, calculate the mask percentage; in this embodiment, for example, the calculation formula of the mask percentage is: k = [(Lconsistent_current - Lconsistent_min) / (Lconsistent_max - Lconsistent_min)] x (Permax - Permin) + Permin In the formula, k represents the mask percentage, Lconsistent_current represents the total consistency loss, Lconsistent_min represents the minimum total consistency loss, Lconsistent_max represents the maximum total consistency loss, and Permax and Permin represent the maximum mask percentage and the minimum mask percentage, respectively.

[0086] Thus, based on the foregoing formula, after the mask percentage is calculated, the shielding threshold can be determined, and the process is shown in the following step S76d.

[0087] S76d. Based on the mask percentage, the shielding threshold is determined from the consistency loss map; in this embodiment, the shielding threshold is the (1-k)th percentile of all pixel consistency loss values.

[0088] Specifically, assuming k = 0.1, i.e., it is desired to shield 10% of the pixels with the highest loss in the image, then the consistency loss values of each pixel point in the consistency loss map are sorted in order of consistency loss from low to high. Then, the consistency loss value at the position with a sorting length ratio of 90% is found as the shielding threshold. Of course, the foregoing example is only illustrative, and when the mask percentage is different, the determination process of the shielding threshold is also the same, which will not be described here.

[0089] Thus, based on the foregoing step S76 and its sub-steps, after the shielding threshold at the current iteration is determined, the transient object corresponding pixel points can be identified, and the process is shown in the following step S77.

[0090] S77. For any one pixel point in the initial rendering map, if the consistency loss value of the any one pixel point is greater than the shielding threshold, the any one pixel point in the initial rendering map is taken as a disturbance pixel point, and after all the pixel points in the initial rendering map are polled, a plurality of disturbance pixel points are obtained.

[0091] In the embodiment, since the consistency loss map is obtained by weighted sum of the L1 loss map and the neighboring average loss map, the size of the consistency loss map is the same as that of the initial rendering map, and each pixel point in the initial rendering map corresponds to a pixel point in the consistency loss map, that is, each pixel point in the initial rendering map corresponds to a consistency loss value; therefore, if the consistency loss value of any pixel point in the initial rendering map is greater than the shielding threshold, it is determined that the any pixel point is a transient object and is regarded as an interference pixel point, which needs to be shielded; in this way, all interference pixel points in the initial rendering map can be distinguished based on the foregoing principle; finally, an accurate transient object mask can be generated based on the distinguished interference pixel points, and the process is shown in the following step S78.

[0092] S78. The pixel value of each interference pixel point in the initial rendering map is changed to 0, and the pixel value of each other pixel point is changed to 1, so as to obtain the transient object mask.

[0093] Therefore, by the steps S71-S78, the transient processing mechanism of the spatio-temporal consistency test is introduced to generate the transient object mask in the embodiment, which can more accurately distinguish the real transient object and the difficult static scene part, so that the error shielding of the static scene details can be avoided, and thus the finally reconstructed static scene structure is more complete and cleaner; meanwhile, the shielding threshold is dynamically adjusted by the consistency loss map in each iteration, so that the generation of the mask is more intelligent and robust.

[0094] In this way, after obtaining the transient object mask, the final rendering map can be generated in combination with the initial rendering map, and the process is shown in the following step S8.

[0095] S8. The final rendering map is generated by using the transient object mask and the initial rendering map; in a specific application, the initial rendering map is multiplied point by point with the transient object mask, so as to obtain the final rendering map; then, the loss between the final rendering map and the original image is calculated to perform back propagation, so as to synchronously optimize all learnable parameters.

[0096] The back propagation process can be but is not limited to the following step S9.

[0097] S9. The reconstruction parameters and the model parameters of the reconstruction model are updated by using the image loss between the final rendering map and the selected training image, and after the update, a training image is selected again from the training set until the iteration stopping condition is met, so as to obtain the optimal reconstruction parameters and the optimal reconstruction model; the optimal reconstruction parameters and the optimal reconstruction model are used to generate the reconstructed rendering map under different viewing angles, wherein the reconstruction model includes a basic color model, an illumination correction model and a mixed background model.

[0098] In a specific implementation, the calculation process of the image loss is as follows: (1) the sum of the absolute values of the color difference of each pixel point between the final rendering image and the selected training image is calculated as the LI loss; (2) the SSIM loss (i.e., the structural similarity index loss) of the final rendering image and the selected training image is calculated, and the D-SSIM loss is calculated based on the D-SSIM loss; and (3) the LI loss and the D-SSIM loss are weighted and summed to obtain the image loss.

[0099] The calculation formula of the image loss is as follows:

[0100] In the formula, represents the image loss, represents the LI loss, represents the D-SSIM loss, represents a hyperparameter (with a value of 0.2); wherein, , and is the SSIM loss between the final rendering image and the selected training image.

[0101] Thus, after the image loss between the final rendering image and the selected training image is calculated based on the foregoing formula, the gradient of the loss with respect to all learnable parameters (including the reconstruction parameters and the model parameters of the foregoing three models, which can be set as the weights of the MLP) is calculated by the back propagation algorithm, and then the optimizer of the gradient (such as the Adma optimizer) is used to update the foregoing parameters according to the calculated gradient, so that the loss moves in the direction of reduction; thus, when the preset number of iterations or model convergence is reached, the optimal reconstruction parameters and the optimal reconstruction model can be obtained.

[0102] Based on this, when a new rendering image of a new view angle needs to be generated, a new camera pose is provided, and then the optimal reconstruction parameters and the optimal reconstruction model are used to render the reconstruction rendering image in real time under the view angle.

[0103] The three-dimensional scene reconstruction method based on mixed background and light correction described in detail by the foregoing steps S1-S9 significantly improves the light realism, greatly enhances the background details and dynamics, and improves the accuracy and robustness of scene decomposition, thereby improving the quality of the reconstructed image. In addition, since the model has a more accurate understanding of the scene (especially in terms of light and background), the overall training process converges faster, and the training efficiency and convergence quality are also improved.

[0104] As shown in FIG. 1, Figures 2 to 4 The second aspect of the present embodiment provides a hardware system for implementing the three-dimensional scene reconstruction method based on mixed background and light correction described in the first aspect of the embodiment. The initialization module is used to initialize the reconstruction parameters, wherein the reconstruction parameters include the attribute information of each Gaussian point in the scene and the image information of each training image in the training set.

[0105] The enhanced appearance model module is used to select a training image from the training set to utilize the attribute information of each Gaussian point and the image information of the selected training image, and based on the basic color model and the illumination correction model, respectively obtain the foreground SH coefficient and color correction amount of each Gaussian point, where the color correction amount is used to characterize the color deviation caused by high-frequency effects.

[0106] The enhanced appearance model module is also used to generate a rendered foreground image based on the color correction amount and foreground SH coefficient of each Gaussian point; in this implementation, see Figure 3 As shown in the figure, the enhanced appearance model module uses the basic color prediction MLP (i.e., basic color model) + illumination correction MLP (i.e., illumination correction model) to generate the foreground color.

[0107] The hybrid background model module is used to generate background low-frequency SH coefficients and high-frequency detail parameters for controlling at least one noise function based on image information of the selected training image and using the hybrid background model.

[0108] The hybrid background model module is also used to generate a rendered background image containing high-frequency details using the background low-frequency SH coefficient and high-frequency detail parameters; in this design example, see Figure 4 As shown in the figure, the hybrid background model module outputs low-frequency basic color + high-frequency details to generate background colors with rich details such as clouds and textures.

[0109] The rasterization module is used to generate an initial rendering image based on the rendered foreground image and the rendered background image.

[0110] The spatiotemporal consistency transient processing module is used to obtain neighboring images of the selected training image and use the neighboring images and the spatiotemporal consistency transient processing algorithm to distinguish the interfering pixels caused by transient objects in the initial rendering image to generate a transient object mask.

[0111] The spatiotemporal consistency transient processing module uses the transient object mask and the initial rendering to generate the final rendering.

[0112] The updating module is configured to update the reconstruction parameter and the model parameter of the reconstruction model by using the image loss between the final rendering image and the selected training image, and after the update, reselect a training image from the training set until the iteration stopping condition is met, to obtain the optimal reconstruction parameter and the optimal reconstruction model, so as to generate the reconstruction rendering image under different perspectives by using the optimal reconstruction parameter and the optimal reconstruction model, wherein the reconstruction model comprises a basic color model, an illumination correction model and a mixed background model.

[0113] The working process, working details and technical effects of the system provided by the embodiment can be referred to the first aspect of the embodiment, and will not be repeated here.

[0114] As shown in Figure 5 The third aspect of the embodiment provides a three-dimensional scene reconstruction device based on mixed background and illumination correction. Taking the device as an electronic device, the device comprises a memory, a processor and a transceiver connected in sequence, wherein the memory is configured to store a computer program, the transceiver is configured to receive and send messages, and the processor is configured to read the computer program and execute the three-dimensional scene reconstruction method based on mixed background and illumination correction as described in the first aspect of the embodiment.

[0115] For example, the memory can include, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a flash memory, a first-in first-out memory (FIFO) and / or a first-in last-out memory (FILO), etc.; specifically, the processor can include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor can be implemented in at least one of the following hardware forms: a DSP (Digital Signal Processing), a FPGA (Field-Programmable Gate Array), a PLA (Programmable Logic Array), and a CPU (Central Processing Unit). Meanwhile, the processor can include a main processor and a coprocessor. The main processor is a processor for processing data in an awake state, also known as a CPU (Central Processing Unit). The coprocessor is a low-power processor for processing data in a standby state.

[0116] In some embodiments, the processor can be integrated with a GPU (Graphics Processing Unit) that is responsible for rendering and drawing the content required to be displayed on the display screen, for example, the processor can not be limited to a microprocessor of STM32F105 series, a RISC (reduced instruction set computer) microprocessor, an X86 architecture processor, or a processor integrated with an embedded NPU (neural-network processing unit); the transceiver can be but not limited to a WIFI transceiver, a Bluetooth transceiver, a GPRS (General Packet Radio Service) transceiver, a ZigBee transceiver, a 3G transceiver, a 4G transceiver, and / or a 5G transceiver, etc. In addition, the device can further include but not limited to a power module, a display screen, and other necessary components.

[0117] The working process, working details and technical effects of the electronic device provided in the embodiment can be referred to the first aspect of the embodiment, and will not be repeated here.

[0118] The fourth aspect of the embodiment provides a storage medium storing instructions of the three-dimensional scene reconstruction method based on mixed background and light correction according to the first aspect of the embodiment, that is, the storage medium stores instructions, and when the instructions run on a computer, the three-dimensional scene reconstruction method based on mixed background and light correction according to the first aspect of the embodiment is executed.

[0119] The storage medium refers to a carrier for storing data, which can include but is not limited to floppy disks, optical disks, hard disks, flash memories, USB flash disks, and / or Memory Sticks, etc., and the computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.

[0120] The working process, working details and technical effects of the storage medium provided in the embodiment can be referred to the first aspect of the embodiment, and will not be repeated here.

[0121] The fifth aspect of the embodiment provides a computer program product containing instructions, which, when running on a computer, causes the computer to execute the three-dimensional scene reconstruction method based on mixed background and light correction according to the first aspect of the embodiment, wherein the computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices.

[0122] The above detailed description of the specific embodiments of the present application has been given to understand the purpose, technical solutions and beneficial effects of the present application. It should be understood that the above description is only a specific embodiment of the present application and is not used to limit the protection scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A three-dimensional scene reconstruction method based on mixed background and illumination correction, characterized in that: include: Initializing reconstruction parameters, where the reconstruction parameters include attribute information of each Gaussian point in the scene and image information of each training image in the training set; A training image is selected from the training set to utilize the attribute information of each Gaussian point and the image information of the selected training image, and based on the basic color model and the illumination correction model, the foreground SH coefficient and color correction amount of each Gaussian point are obtained respectively. The color correction amount is used to characterize the color deviation caused by high-frequency effects; Generate a rendered foreground image based on the color correction amount and foreground SH coefficient of each Gaussian point; Based on the image information of the selected training image and using the mixed background model, a background low-frequency SH coefficient and a high-frequency detail parameter for controlling at least one noise function are generated respectively; Using the background low-frequency SH coefficient and high-frequency detail parameters, a rendered background image containing high-frequency details is generated; Generate an initial rendering image based on the rendered foreground image and the rendered background image; Obtain neighboring images of the selected training image, and use the neighboring images and a spatiotemporal consistency transient processing algorithm to distinguish interfering pixels caused by transient objects in the initial rendering image to generate a transient object mask; Generate a final rendering using the transient object mask and the initial rendering; The image loss between the final rendering image and the selected training image is used to update the reconstruction parameters and the model parameters of the reconstruction model. After the update, a training image is reselected from the training set until the iteration stop condition is met, and the optimal reconstruction parameters and the optimal reconstruction model are obtained. The optimal reconstruction parameters and the optimal reconstruction model are used to generate reconstructed renderings under different perspectives, where the reconstruction model includes a basic color model, a lighting correction model, and a mixed background model.

2. The method according to claim 1, characterized in that The attribute information of any Gaussian point includes the appearance feature vector and position of the Gaussian point, and the image information of any training image includes the camera pose and appearance embedding vector corresponding to the training image; Among them, using the attribute information of each Gaussian point and the image information of the selected training image, and based on the basic color model and the illumination correction model, the foreground SH coefficient and color correction amount of each Gaussian point are obtained respectively, including: The appearance embedding vector of the selected training image and the appearance feature vector of each Gaussian point are input into the basic color model to obtain the foreground SH coefficient corresponding to the foreground color of each Gaussian point; The foreground SH coefficient of each Gaussian point, as well as the appearance feature vector and observation direction vector of each Gaussian point, are input into the illumination correction model to obtain the color correction amount of each Gaussian point, wherein the observation direction vector of any Gaussian point is obtained based on the position of the any Gaussian point and the camera pose.

3. The method according to claim 2, characterized in that The illumination correction model adopts a lightweight MLP model, wherein the lightweight MLP model includes 2 hidden layers, and each hidden layer contains 64 neurons.

4. The method according to claim 1, wherein The image information of any training image includes a camera pose and an appearance embedding vector corresponding to the training image, and the mixed background model includes: a backbone network layer, a first output head and a second output head; The backbone network layer is used to encode the appearance embedding vector of the selected training image into a 128-dimensional feature vector, and input the feature vector into the first output head and the second output head respectively; A first output head is used for predicting the background low-frequency SH coefficient based on the input feature vector; The second output head is used to predict, based on the input feature vector, a fundamental frequency for controlling the maximum cloud stretching and size, an octave for controlling the cloud layer level, a frequency gain for controlling the frequency increase ratio between adjacent octaves, an amplitude gain for controlling the amplitude decay rate between adjacent octaves, a noise offset for controlling the cloud layer layout, a first color gradient for controlling the starting color of the cloud color mapping, and a second color gradient for controlling the ending color of the cloud color mapping, so as to generate the high-frequency detail parameters using the fundamental frequency, the octave, the frequency gain, the amplitude gain, the noise offset, the first color gradient, and the second color gradient.

5. The method according to claim 4, characterized in that The backbone network layer adopts an MLP network, wherein the MLP network includes 3 hidden layers, each hidden layer includes 128 neurons, the activation function corresponding to each hidden layer is a ReLU activation function, and the first output head adopts a first linear fully connected layer, and the second output head adopts a second linear fully connected layer including 15 output neurons, to correspond to 7 types of high-frequency detail parameters respectively.

6. The method according to claim 1, characterized in that Using the background low-frequency SH coefficient and high-frequency detail parameters, a rendered background image containing high-frequency details is generated, including: Get the direction vector of the scene light; Generate a background low-frequency basic color based on the direction vector of the scene light and the background low-frequency SH coefficient; generating a high-frequency background color according to the high-frequency detail parameters; The rendered background image is generated using the background low-frequency basic color and the high-frequency background color.

7. The method according to claim 1, characterized in that Using neighbor images and spatiotemporal consistency transient processing algorithms, we distinguish the interfering pixels caused by transient objects in the selected training images to generate transient object masks, including: Calculating a pixel-by-pixel L1 loss map between the initial rendering image and the selected training image, wherein each pixel in the L1 loss map corresponds to a pixel loss value; Calculate a pixel-by-pixel neighboring L1 loss map between the initial rendering image and each neighboring image; For any pixel point in the L1 loss map, determine the pixel points corresponding to the position of the any pixel point in each adjacent L1 loss map as the neighboring pixel points of the any pixel point; Calculate the average loss of neighboring pixels of any pixel, and after polling all pixels in the L1 loss map, obtain a neighboring average loss map; Performing a weighted summation on the L1 loss map and the neighboring average loss map to obtain a consistency loss map; Based on the consistency loss graph, calculate the shielding threshold for the current iteration; For any pixel in the initial rendering image, if the consistency loss value of the any pixel is greater than the shielding threshold, the any pixel in the initial rendering image is used as an interference pixel, and after all pixels in the initial rendering image are polled, a plurality of interference pixels are obtained; The pixel value of each interfering pixel in the initial rendering image is changed to 0, and the pixel value of each other pixel is changed to 1, so as to obtain the transient object mask.

8. The method according to claim 7, characterized in that Calculate the shielding threshold for the current iteration, including: Calculate the average of all consistency loss values ​​in the consistency loss graph to obtain the current overall consistency loss; Obtain the maximum and minimum overall consistency losses in all iterations before the current iteration; Calculate the mask percentage based on the current overall consistency loss, the maximum overall consistency loss, and the minimum overall consistency loss; The masking threshold is determined from the consistency loss map based on the mask percentage.

9. The method according to claim 8, characterized in that The mask percentage is calculated based on the current overall consistency loss, the maximum overall consistency loss, and the minimum overall consistency loss, including: The mask percentage is calculated according to the following formula: k=[(Lconsistent_current-Lconsistent_min) / (Lconsistent_max-Lconsistent_min)]×(Permax-Permin)+Permin Wherein, k represents the mask percentage, Lconsistent_current represents the overall consistency loss, Lconsistent_min represents the minimum overall consistency loss, Lconsistent_max represents the maximum overall consistency loss, Permax and Permin represent the maximum mask percentage and the minimum mask percentage respectively.

10. A 3D scene reconstruction system based on mixed background and illumination correction, characterized in that: include: An initialization module is used to initialize reconstruction parameters, wherein the reconstruction parameters include attribute information of each Gaussian point in the scene and image information of each training image in the training set; An enhanced appearance model module is used to select a training image from the training set, utilize the attribute information of each Gaussian point and the image information of the selected training image, and obtain the foreground SH coefficient and color correction amount of each Gaussian point based on the basic color model and the illumination correction model, wherein the color correction amount is used to characterize the color deviation caused by high-frequency effects; The enhanced appearance model module is also used to generate a rendered foreground image based on the color correction amount and foreground SH coefficient of each Gaussian point; A hybrid background model module is used to generate background low-frequency SH coefficients and high-frequency detail parameters for controlling at least one noise function based on image information of the selected training image and using the hybrid background model; The hybrid background model module is also used to generate a rendered background image containing high-frequency details using the background low-frequency SH coefficient and high-frequency detail parameters; A rasterization module, configured to generate an initial rendering image based on a rendered foreground image and a rendered background image; A spatiotemporal consistency transient processing module is used to obtain neighboring images of the selected training image and use the neighboring images and the spatiotemporal consistency transient processing algorithm to distinguish the interfering pixels caused by transient objects in the initial rendering image to generate a transient object mask; The spatiotemporal consistency transient processing module generates the final rendering using the transient object mask and the initial rendering; The update module is used to update the reconstruction parameters and the model parameters of the reconstruction model by using the image loss between the final rendering image and the selected training image, and after the update, reselect a training image from the training set until the iteration stop condition is met, so as to obtain the optimal reconstruction parameters and the optimal reconstruction model, so as to use the optimal reconstruction parameters and the optimal reconstruction model to generate reconstructed renderings under different perspectives, wherein the reconstruction model includes a basic color model, an illumination correction model and a mixed background model.

Citation Information

Cited By

  • Global nerve drawing method and system for end-to-end mixed representation of full frequency domain illumination

    CN122156539A

  • A global neural rendering method and system for end-to-end hybrid representation of full frequency domain lighting

    CN122156539B