3D Reconstruction Method Based on 3D Gaussian Splashing with Event Cameras

By combining the event stream and the three-dimensional Gaussian splatter model of the event camera, the difference adjustment of the time surface information and rendered prediction values is used to solve the contradiction between density and speed in three-dimensional reconstruction, and a high-precision and efficient three-dimensional reconstruction effect is achieved.

CN119942004BActive Publication Date: 2025-07-29BEIJING BIG DATA ADVANCED TECH RES INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510429811.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-29
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

Existing three-dimensional reconstruction methods are difficult to achieve high-precision and high-density dense reconstruction when dealing with complex scenes, and it is difficult to balance rendering speed and quality, especially when using event cameras, where sparse reconstruction and slow rendering speeds are present.

Method used

By obtaining the event flow, the initial three-dimensional Gaussian splatter model is determined, and the dynamic spatiotemporal background is characterized by using time surface information, combining the differences between ground reality and rendered prediction values, the model is adaptively adjusted to realize the optimization of the three-dimensional Gaussian splatter model.

Benefits of technology

Generate accurate and detailed three-dimensional reconstruction in a dynamic environment, avoiding information loss in traditional methods, improving the timeliness and accuracy of reconstruction, especially when dealing with dynamic objects and complex lighting conditions, the rendering effect is more realistic.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942004B_ABST
    Figure CN119942004B_ABST
Patent Text Reader

Abstract

The present disclosure provides a three-dimensional reconstruction method based on event cameras using three-dimensional Gaussian splatting, belonging to the field of computer technology, aiming to solve the problem of poor performance of three-dimensional reconstruction methods in related technologies. The method includes: obtaining an event stream for a target scene; determining an initial three-dimensional Gaussian splatting model of the target scene according to the event stream; determining the temporal surface information of each pixel according to the event stream; the temporal surface information characterizing the dynamic spatio-temporal background information of the event; determining the ground truth of the target scene according to the temporal surface information of each pixel; the ground truth characterizing the actual brightness change of the target scene; determining a rendering prediction value according to the initial three-dimensional Gaussian splatting model, the rendering prediction value characterizing the rendering intensity difference of the initial three-dimensional Gaussian splatting model at two viewpoints; and adjusting the initial three-dimensional Gaussian splatting model based on the difference between the ground truth and the rendering prediction value to obtain a three-dimensional Gaussian splatting model corresponding to the target scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular to a method for implementing three-dimensional reconstruction based on three-dimensional Gaussian splattering of an event camera. Background Art

[0002] Creating detailed, realistic, real-time renderings of 3D scenes is a persistent challenge in computer vision and graphics. The event camera is a new biologically inspired visual sensor that uses asynchronous dynamic visual perception technology. Unlike traditional RGB cameras, event cameras record asynchronous event streams triggered by changes in pixel brightness rather than capturing absolute brightness values at a fixed frame rate. Each pixel responds independently to brightness changes, and a timestamp event is generated when the change exceeds a preset threshold. This unique capture method enables event cameras to perform well in high-speed motion and high dynamic range scenes, making them ideal for applications requiring high temporal resolution and low latency, such as simultaneous localization and mapping, target recognition and tracking, and optical flow estimation.

[0003] Currently, there are two main methods for reconstructing three-dimensional scenes using mobile event cameras: the first is based on visual odometry and SLAM (Simultaneous Localization and Mapping), and the second is based on neural radiance field (Nerf) rendering.

[0004] The first method is mainly based on explicit feature matching between event-accumulated images to reconstruct sparse or semi-dense scene structures. However, the first method may be limited by its sparse reconstruction characteristics when dealing with complex scenes and cannot meet the needs of high-precision and high-density scene reconstruction. In recent years, the second method has gradually been applied to scene reconstruction of event cameras. These methods supervise the training of neural radiance fields through event flow integration to render realistic images. However, the second method relies on the sampling strategy of ray marching during training and rendering. Excessive sampling may lead to slow speed, while too little sampling will affect the rendering quality.

[0005] How to strike a balance between rendering speed and quality becomes a major challenge, which leads to the study of how to achieve dense, accurate and fast scene reconstruction using only event streams. Summary of the Invention

[0006] To overcome the problems in related technologies, the present disclosure provides a method for 3D reconstruction based on 3D Gaussian splattering of an event camera. The technical solution of the present disclosure is as follows:

[0007] According to a first aspect of an embodiment of the present disclosure, a method for implementing 3D reconstruction using 3D Gaussian splatting based on an event camera is provided, comprising:

[0008] Obtain the event stream for the target scenario;

[0009] Determine the initial three-dimensional Gaussian splash model for the target scenario according to the event stream;

[0010] Determine the temporal surface information of each pixel according to the event stream; the temporal surface information characterizes the dynamic spatio-temporal background information of the event;

[0011] Determine the ground truth of the target scenario according to the temporal surface information of each pixel; the ground truth characterizes the actual brightness change of the target scenario;

[0012] Determine the rendering prediction value according to the initial three-dimensional Gaussian splash model, where the rendering prediction value characterizes the rendering intensity difference of the initial three-dimensional Gaussian splash model at two perspectives;

[0013] Adjust the initial three-dimensional Gaussian splash model based on the difference between the ground truth and the rendering prediction value to obtain the three-dimensional Gaussian splash model corresponding to the target scenario.

[0014] Optionally, it further includes:

[0015] Determine the event threshold corresponding to this round of iteration within a reasonable range; the reasonable range is determined by the event threshold configured by the event camera;

[0016] Determine the event threshold loss corresponding to this round of iteration according to the event threshold; the event threshold loss characterizes the suitability of the event threshold selected in this round of iteration;

[0017] Adjusting the initial three-dimensional Gaussian splash model based on the difference between the ground truth and the rendering prediction value to obtain the three-dimensional Gaussian splash model corresponding to the target scenario includes:

[0018] Determine the event rendering loss according to the ground truth and the rendering prediction value;

[0019] Adjust the initial three-dimensional Gaussian splash model according to the event rendering loss and the event threshold loss to obtain the three-dimensional Gaussian splash model corresponding to the target scenario.

[0020] Optionally, determining the event threshold loss corresponding to this round of iteration according to the event threshold includes:

[0021] Determine the selection interval of the event threshold; the threshold interval includes an upper threshold and a lower threshold;

[0022] Determine the event threshold loss according to the unit step function based on the upper threshold, the lower threshold, and the event threshold.

[0023] Optionally, based on the difference between the ground truth and the rendering prediction value, adjust the initial three-dimensional Gaussian splash model to obtain a three-dimensional Gaussian splash model corresponding to the target scene, including:

[0024] Determine the first difference corresponding to the pixel through the ground truth and the rendering prediction value corresponding to the same pixel;

[0025] Determine the difference between the ground truth and the rendering prediction value through the first differences corresponding to each pixel;

[0026] Based on the difference between the ground truth and the rendering prediction value, adjust the initial three-dimensional Gaussian splash model to obtain a three-dimensional Gaussian splash model corresponding to the target scene.

[0027] Optionally, determine the ground truth of the target scene according to the temporal surface information of each pixel, including:

[0028] Determine the training time period corresponding to this round of iteration;

[0029] Determine the temporal surface information of each pixel included in the training time period at each moment;

[0030] Integrate the temporal surface information of each pixel respectively to determine the brightness change of each pixel during the training time period, and generate each model supervision signal;

[0031] Determine the model supervision signal as the ground truth of the corresponding pixel during the training time period.

[0032] Optionally, determine the temporal surface information of each event of each pixel according to the event stream, including:

[0033] Determine the target pixel and the target timestamp corresponding to each event;

[0034] Determine each neighboring pixel included in the neighborhood range of the target pixel;

[0035] Determine each neighboring event of each neighboring pixel;

[0036] Select the target neighboring event that occurs before the target timestamp and is closest to the occurrence time of the target timestamp from each neighboring event;

[0037] Apply the decay rate to determine the temporal surface corresponding to the target pixel at the target timestamp through the neighboring event and the event.

[0038] Optionally, the training time period includes a first moment at the start of the training time period and a second moment at the end of the training time period; determining the rendering prediction value according to the initial three-dimensional Gaussian splash model includes:

[0039] Determining the first pose of the event camera at the first moment and determining the second pose of the event camera at the second moment;

[0040] Based on the first pose, projecting the initial three-dimensional Gaussian splash model onto a two-dimensional image to obtain a first image;

[0041] Based on the second pose, projecting the initial three-dimensional Gaussian splash model onto a two-dimensional image to obtain a second image;

[0042] Determining the first rendering intensity of each pixel in the first image and determining the second rendering intensity of each pixel in the second image;

[0043] According to the first rendering intensity and the second rendering intensity corresponding to the same pixel, determining the rendering intensity difference of each pixel;

[0044] Determining the rendering intensity difference as the rendering prediction value of the corresponding pixel.

[0045] Optionally, projecting the initial three-dimensional Gaussian splash model onto a two-dimensional image includes:

[0046] Determining each three-dimensional Gaussian function included in the three-dimensional Gaussian splash model;

[0047] Projecting each of the three-dimensional Gaussian functions onto a two-dimensional plane to determine the two-dimensional Gaussian corresponding to each of the three-dimensional Gaussian functions; the two-dimensional Gaussian represents the color of the three-dimensional Gaussian on the two-dimensional plane;

[0048] For any pixel in the two-dimensional plane, determining the two-dimensional Gaussians corresponding to the pixel, sorting the projected two-dimensional Gaussians based on a tile rasterizer, and determining the color corresponding to the pixel by means of alpha blending.

[0049] Optionally, determining the initial three-dimensional Gaussian splash model of the target scene according to the event stream includes:

[0050] Performing shallow training on the initial target model based on the event stream to obtain a target model;

[0051] Inputting the event stream into the target model, and recovering the depth information of each pixel from the event stream through the target model to obtain a rough depth map;

[0052] Converting the rough depth map into a three-dimensional point cloud through back-projection;

[0053] Process the three-dimensional point cloud through a time surface map constituted by the time surface information to obtain an initial three-dimensional Gaussian splash model; the processing at least includes: coloring the three-dimensional point cloud through the time surface map.

[0054] Optionally, obtain an event stream for a target scene, including:

[0055] Determine the brightness change value of each pixel at a preset time interval;

[0056] Generate an event stream when the brightness change value exceeds the event threshold corresponding to the event camera; each event in the event stream includes pixel coordinates, a timestamp, and an event polarity.

[0057] According to a second aspect of the embodiments of the present disclosure, there is provided a three-dimensional Gaussian splash-based three-dimensional reconstruction device for an event camera, including:

[0058] An acquisition module, configured to acquire an event stream for a target scene;

[0059] A model determination module, configured to determine an initial three-dimensional Gaussian splash model of a target scene according to the event stream;

[0060] A time surface determination module, configured to determine the time surface information of each pixel according to the event stream; the time surface information represents the dynamic spatio-temporal background information of the event;

[0061] A ground truth determination module, configured to determine the ground truth of the target scene according to the time surface information of each pixel; the ground truth represents the actual brightness change of the target scene;

[0062] A predicted value determination module, configured to determine a rendering predicted value according to the initial three-dimensional Gaussian splash model, where the rendering predicted value represents the rendering intensity difference of the initial three-dimensional Gaussian splash model from two perspectives;

[0063] An adjustment module, configured to adjust the initial three-dimensional Gaussian splash model based on the difference between the ground truth and the rendering predicted value to obtain a three-dimensional Gaussian splash model corresponding to the target scene.

[0064] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where when the computer program is executed by the processor, the steps of the three-dimensional reconstruction method described in the first aspect are implemented.

[0065] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the three-dimensional reconstruction method described in the first aspect are implemented.

[0066] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the three-dimensional reconstruction method described in the first aspect are implemented.

[0067] By combining the high temporal accuracy of the event stream and the meticulous construction of the Gaussian model, the three-dimensional reconstruction result will have high accuracy and real-time performance, and is suitable for the reconstruction of dynamic environments. The accurate description of the dynamic spatio-temporal background through temporal surface information can help the system better capture and reconstruct the three-dimensional shape of fast-moving objects, avoiding information loss caused by frame rate limitations in traditional methods. By the difference between the ground truth and the rendering prediction value, the initial three-dimensional Gaussian splash model can be adaptively adjusted to achieve a more refined and natural rendering effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0068] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings required for the description of the embodiments of the present disclosure will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0069] Figure 1 FIG. is a schematic diagram of the steps of a three-dimensional reconstruction method implemented by three-dimensional Gaussian splash based on an event camera shown in the embodiments of the present disclosure;

[0070] Figure 2 FIG. is a block diagram of a three-dimensional reconstruction device implemented by three-dimensional Gaussian splash based on an event camera shown in the embodiments of the present disclosure;

[0071] Figure 3 FIG. is a schematic diagram of an electronic device shown in the embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0072] The technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present disclosure without creative efforts shall fall within the protection scope of the present disclosure.

[0073] The terms "first", "second", etc. in the description and claims of the present disclosure are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same type, and do not limit the number of objects. For example, the first object can be one or multiple. In addition, "and / or" in the description and claims means at least one of the connected objects, and the character " / ", generally represents an "or" relationship between the associated objects before and after.

[0074] The latest progress of the 3D Gaussian Splatting (3DGS) technology provides a promising solution to the problems existing in the related technologies. Compared with NeRF, 3DGS has obvious advantages in training and rendering speed, which benefits from its explicit scene representation, highly parallelized workflow, and efficient computing strategy that combines a differentiable rendering pipeline with point-based rendering technology. In addition, the ability of 3DGS to control scene dynamics is particularly important in scenes with complex geometries.

[0075] However, the method of 3DGS is mainly oriented to RGB (Red, Green, Blue) frames, and scalability problems will occur when applied to event streams. For example, 3DGS requires point cloud initialization, while traditional Structure from Motion (SfM) - based methods are not applicable to event streams, and the point clouds generated by event - based point cloud reconstruction strategies are too sparse to meet the initialization requirements of 3DGS.

[0076] Therefore, in view of the above - mentioned technical problems, the present disclosure proposes a 3D reconstruction method based on 3D Gaussian Splatting with an event camera. The method can well combine the event stream and 3D Gaussian Splatting for 3D scene reconstruction.

[0077] Figure 1 It is a schematic diagram of the steps of the 3D reconstruction method based on 3D Gaussian Splatting with an event camera shown in the embodiments of the present disclosure. According to Figure 1 as shown, the method may specifically include the following steps:

[0078] Step S11: Obtain an event stream for the target scene.

[0079] Collect an event stream of a target scene through a moving event camera. An event camera is a bionic sensor different from a standard camera. Instead of directly outputting pixel values, an event camera asynchronously measures the changes of each pixel and outputs information encoding the time, position, and sign of these changes. The pixel is the pixel corresponding to the viewfinder of the event camera. For example, if the viewfinder of the event camera consists of 4*4 pixels, then the event camera will capture the brightness changes of each pixel within the 4*4 pixel range.

[0080] Step S12: Determine an initial three-dimensional Gaussian splash model of the target scene according to the event stream.

[0081] Different from the method in the related art of determining the initial three-dimensional Gaussian splash model corresponding to 3DGS through RGB, the initial three-dimensional Gaussian splash model corresponding to 3DGS can be determined through the event stream.

[0082] Using the time, position, and sign information in the event stream, the three-dimensional shape and position of the object in the target scene can be inferred. The initial three-dimensional Gaussian splash model is obtained by representing these inference results in the form of a Gaussian distribution.

[0083] Step S13: Determine the time surface information of each pixel according to the event stream; the time surface information characterizes the dynamic spatio-temporal background information of the event.

[0084] The time surface information characterizes the dynamic spatio-temporal background information of the event. By analyzing the timestamp information in the event stream, the time surface information of each pixel can be obtained. The dynamic spatio-temporal background information represents the influence of the past events of neighboring pixels on the current event of the target pixel.

[0085] Step S14: Determine the ground truth of the target scene according to the time surface information of each pixel; the ground truth characterizes the actual brightness change of the target scene.

[0086] The ground truth refers to the actual brightness change within a period of observation time. The ground truth of the target scene can reflect the actual brightness change of each pixel within the observation time.

[0087] Step S15: Determine a rendering prediction value according to the initial three-dimensional Gaussian splash model, and the rendering prediction value characterizes the rendering intensity difference of the initial three-dimensional Gaussian splash model at two viewpoints.

[0088] By projecting the initial three-dimensional Gaussian splash model onto two different viewpoints and calculating the rendering intensity difference between these two viewpoints, the rendering prediction value can be obtained.

[0089] The selection of the projection view is determined by the ground truth. By the observation time corresponding to the ground truth, two views for projecting the initial three-dimensional Gaussian splash model are determined. Usually, the starting and ending moments of the observation time are selected to determine the corresponding views respectively.

[0090] Determine the difference in rendering intensity of the pixels at the same position in the images corresponding to the two views, and determine the difference in rendering intensity of the initial three-dimensional Gaussian splash model in the two views. The difference in rendering intensity is a predicted value used to evaluate the accuracy of the initial three-dimensional Gaussian splash model.

[0091] Step S16: Based on the difference between the ground truth and the rendering prediction value, adjust the initial three-dimensional Gaussian splash model to obtain the three-dimensional Gaussian splash model corresponding to the target scene.

[0092] By comparing the difference between the ground truth and the rendering prediction value, the accuracy of the initial three-dimensional Gaussian splash model can be evaluated. If the difference is large, it indicates that there is a certain deviation between the initial three-dimensional Gaussian splash model and the real scene. At this time, the initial three-dimensional Gaussian splash model needs to be adjusted to reduce the difference and improve the accuracy of the model.

[0093] The adjustment method can be to iteratively optimize the model parameters based on an optimization algorithm until the difference between the ground truth and the rendering prediction value reaches the minimum.

[0094] Adopting the embodiments of the present disclosure, through the combination of the event stream and the three-dimensional Gaussian splash model, the method can generate accurate and detailed three-dimensional reconstructions in a dynamic environment, avoiding the accuracy loss caused by motion blur or insufficient sampling frequency in traditional methods. The temporal surface information enables the model to better reflect the dynamic characteristics of objects in a rapidly changing scene, solving the problem that traditional methods are difficult to handle high-speed motion and light changes, and improving the timeliness and accuracy of three-dimensional reconstruction. The comparison and adjustment of the rendering prediction and the actual brightness difference can achieve a more refined rendering effect. Especially when dealing with dynamic objects and complex lighting conditions, the optimized three-dimensional model is more realistic in terms of visual effects and detail processing. Based on the difference between the rendering prediction value and the ground truth, the initial three-dimensional Gaussian splash model is adaptively adjusted, and the adaptive adjustment mechanism improves the accuracy and stability of three-dimensional reconstruction.

[0095] Among them, in an optional embodiment, obtaining the event stream for the target scene includes: determining the brightness change value of each pixel at a preset time interval; generating an event stream when the brightness change value exceeds the event threshold corresponding to the event camera; each event in the event stream includes pixel coordinates, a timestamp, and an event polarity.

[0096] Each pixel refers to the pixel corresponding to the viewfinder of the event camera. When the event camera moves, the pixels in the viewfinder of the event camera will change.

[0097] For each pixel, the event camera continuously monitors its brightness. The brightness change can be obtained by calculating the brightness difference within consecutive time intervals.

[0098] Calculate the brightness change of pixel x at timestamp t , which can be determined by the following formula:

[0099]

[0100] where represents a very short time interval, , represents the brightness of pixel x at timestamp t, represents the brightness of pixel x at timestamp .

[0101] When the brightness change of a certain pixel exceeds the set event threshold , the event camera considers that a significant change has occurred to this pixel and generates an event stream. The event threshold of the event camera is usually 10% - 50% of the brightness value. The generated event stream can be represented by the following formula:

[0102]

[0103] where represents the pixel coordinates of the event; represents the timestamp when the event occurs; represents the polarity of the event.

[0104] The pixel coordinates represent the specific location where the event change occurs, indicating the change location in the image. The timestamp records the exact time when the event occurs and can be used to reconstruct the dynamic changes of the scene. The event polarity represents the direction of the brightness change, including two polarities: the positive polarity indicating an increase in brightness and the negative polarity indicating a decrease in brightness.

[0105] When generating the event of pixel at timestamp , the event does not carry the event threshold set when the event camera collects the event stream.

[0106] After obtaining the event, model the event of pixel as follows: as follows:

[0107]

[0108] where is the unit impulse function of time and represents the polarity of the event. Using the embodiments of the present disclosure, the event stream generates events based on the brightness change value, avoiding the generation of redundant data and improving the efficiency of data storage and calculation. Through the pixel coordinates, timestamps, and event polarities, the dynamic changes of the scene can be accurately analyzed. It can precisely capture the rapid dynamic changes in the scene and is applicable to scenes with high-speed motion and light changes.

[0109] Among them, in an optional embodiment, determining the initial three-dimensional Gaussian splash model of the target scene according to the event stream includes: performing shallow training on the initial target model based on the event stream to obtain the target model; inputting the event stream into the target model, and through the target model, recovering the depth information of each pixel from the event stream to obtain a rough depth map; through back-projection, converting the rough depth map into a three-dimensional point cloud; processing the three-dimensional point cloud through the time surface map composed of the time surface information to obtain the initial three-dimensional Gaussian splash model; the processing at least includes: coloring the three-dimensional point cloud through the time surface map.

[0110] Shallow training means performing preliminary training on the preliminary target model to quickly construct a preliminary three-dimensional model. The preliminary target model can be the EventNeRF model. In the shallow training stage, the target model only needs to perform preliminary adjustment and learning on the model according to the input information of the event stream.

[0111] After preliminary training, the target model receives the data input from the event stream. Through the analysis of the event stream, the target model uses the patterns obtained from the training to deduce the depth value of each pixel and obtains a rough depth map. The rough depth map is an image, where the gray value of each pixel represents the depth information of the corresponding scene of the pixel.

[0112] After obtaining the rough depth map, the rough depth map is converted into a three-dimensional point cloud through back-projection. Back-projection is to map each pixel on the image plane back to the three-dimensional space through the known depth information and the internal parameters of the camera, so as to obtain the corresponding three-dimensional points, and the three-dimensional points will form a three-dimensional point cloud.

[0113] The three-dimensional point cloud can be obtained through back-projection by the following formula:

[0114]

[0115]

[0116] where u and v are pixel coordinates; d is the depth value; K is the internal parameter matrix; M is the external parameter matrix; represents the transpose of the matrix, represents the three-dimensional point cloud.​

[0117] A time surface diagram takes time as a dimension of space and is used to represent the dynamic process of a physical quantity changing over time. The time surface diagram consists of various time surface information.

[0118] By further processing the three-dimensional point cloud with the time surface diagram, the information in the time surface diagram can be utilized to color each point in the point cloud, thereby reflecting features such as the lighting and motion state of the target scene. The information in the time surface diagram includes the timestamp of each point in the event stream, light intensity changes, and other information.

[0119] After further processing the three-dimensional point cloud with the time surface diagram, a three-dimensional point cloud for initializing the three-dimensional Gaussian splash model is obtained. Based on this three-dimensional point cloud for initializing the three-dimensional Gaussian splash model, a set of anisotropic three-dimensional Gaussian functions is used to perform a display representation on the three-dimensional point cloud, resulting in an initial three-dimensional Gaussian splash model. The initial three-dimensional Gaussian splash model represents the spatial distribution, object shape, and time evolution of the target scene by fusing the event stream, depth information, point cloud, and time surface information.

[0120] Adopting the embodiments of the present disclosure, event-driven shallow training can achieve generating a three-dimensional point cloud for initializing the three-dimensional Gaussian splash model through the event stream. Due to the high temporal resolution information provided by the event camera, depth can be quickly recovered from the event stream. Especially under fast motion or complex lighting conditions, it can provide more accurate depth information than traditional frame cameras. Through the time surface diagram, the dynamic features changing over time in the scene can be accurately represented, and coloring the three-dimensional point cloud with the time surface diagram can enhance the authenticity of the visual effect.

[0121] Among them, in an optional embodiment, based on the difference between the ground truth and the rendering prediction value, adjusting the initial three-dimensional Gaussian splash model to obtain a three-dimensional Gaussian splash model corresponding to the target scene includes: determining the first difference corresponding to the pixel through the ground truth and the rendering prediction value corresponding to the same pixel; determining the difference between the ground truth and the rendering prediction value through the first differences corresponding to each pixel; and adjusting the initial three-dimensional Gaussian splash model based on the difference between the ground truth and the rendering prediction value to obtain a three-dimensional Gaussian splash model corresponding to the target scene.

[0122] The ground truth is real-world data obtained through sensors, reflecting the actual situation of the target scene.

[0123] The rendering prediction value is a scene image predicted according to the current three-dimensional Gaussian splash model and the rendering algorithm, which is a calculation result with a certain error.

[0124] For each pixel, by comparing the ground truth value of the pixel with the corresponding rendered prediction value, a first difference can be obtained. The first difference reflects the error at each pixel level, indicating the deviation between the pixel at that location and the actual scene. The first difference is the event rendering loss corresponding to the pixel. The calculation formula for the event rendering loss of each pixel is as follows:

[0125]

[0126] where, represents the rendered prediction value of the pixel, represents the ground truth of the pixel.

[0127] After calculating the first difference of each pixel, next, the differences of all pixels need to be aggregated to obtain the global difference of the entire scene. The global difference can be obtained by performing a weighted sum of the first differences of each pixel.

[0128] According to the calculated global difference, the initial three-dimensional Gaussian splash model is optimized.

[0129] By adopting the embodiments of the present disclosure, by calculating the differences of each pixel and performing global optimization, the accuracy of the three-dimensional model can be significantly improved, reducing the error between the model and the actual scene. Automatically adjusting the model according to the difference between the ground truth and the rendered prediction value can adapt to different scene characteristics. Since only the pixel differences need to be calculated and globally aggregated each time, and then the three-dimensional model is adjusted through an optimization algorithm, the overall calculation process has high efficiency and is suitable for real-time or near-real-time three-dimensional modeling and rendering tasks.

[0130] Among them, in an optional embodiment, according to the temporal surface information of each pixel, determining the ground truth of the target scene includes: determining the training time period corresponding to the current iteration; determining the temporal surface information of each pixel included in the training time period at each moment; respectively integrating the temporal surface information of each pixel to determine the brightness change of each pixel during the training time period, and generating each model supervision signal; and determining the model supervision signal as the ground truth of the corresponding pixel during the training time period.

[0131] The temporal surface information can be used to capture the change process of each pixel in the target scene. The temporal surface information refers to the characteristic performance of the pixel at different time points, such as brightness, color, depth, etc. The characteristic performance can help the model identify and understand the dynamic changes of the scene.

[0132] Set a time range for the entire optimization process to focus on processing the data within the training time period. The data within the training time period will be used to capture and analyze the temporal characteristics of each pixel. For example, if a 1-minute event stream is collected, then the complete temporal surface information is one minute. The time period from the 20th second to the 30th second can be selected as the training time period, and the respective temporal surface information included in this training time period is determined.

[0133] For each pixel point in the target scene, extract the temporal surface information of this pixel throughout the training time period. The temporal surface information can help the model identify the dynamic evolution of the pixel.

[0134] Integrate the temporal surface information of each pixel to quantify the brightness change of this pixel throughout the training time period. Integration represents the accumulation of brightness change over time and reflects the overall brightness change of this pixel during this time period. The integration process can use simple time accumulation or weighted integration. The calculation of brightness change is used to quantify the brightness change trend of the pixel from start to end, such as changing from bright to dark or from dark to bright.

[0135] Through the integration calculation, the model supervision signal obtained is determined as the ground truth of this pixel during the training time period. That is, the true performance or change of this pixel throughout the training time period. The ground truth refers to the true data performance of the scene, so the supervision signal provides labeled data for subsequent model training and helps the model more accurately restore the true situation of the scene.

[0136] During the training time period In the case of The formula for determining the model supervision signal of pixel

[0137]

[0138] where represents the event corresponding to pixel represents the temporal surface information of pixel at each moment during the training time period, represents the event threshold selected in this round of iteration, and dt represents the minimum time step for integrating the event supervision signal.

[0139] ​Among them, in an optional embodiment, determining the temporal surface information of each event of each pixel according to the event stream includes: determining the target pixel and target timestamp corresponding to each event; determining each neighboring pixel included in the neighborhood range of the target pixel; determining each neighboring event of each neighboring pixel; selecting, from each neighboring event, a target neighboring event that occurs before the target timestamp and is closest to the occurrence time of the target timestamp; applying an attenuation rate, and determining the temporal surface information corresponding to the target pixel at the target timestamp through the neighboring event and the event.

[0140] The temporal surface information is determined through the acquired event stream. There is a one-to-one relationship between an event and the temporal surface. Each event in the event stream contains spatio-temporal information. By analyzing this information, the temporal surface information corresponding to each event can be determined. The temporal surface information reflects the influence of neighboring events on this event and can truly reflect the rapid changes in brightness and objects in the target scene.

[0141] For any event in the event stream, its temporal surface information can be determined in the following way.

[0142] Determine the target pixel and target timestamp corresponding to the event.

[0143] Determine the neighborhood range of the target pixel. The neighborhood range represents the pixel area around the target pixel. Each neighboring pixel within the neighborhood range provides information about the changes in the surrounding environment. The neighborhood range can be a fixed pixel matrix range, such as 3x3, 5x5, etc., or a range that is dynamically adjusted according to specific applications.

[0144] For each neighboring pixel, filter out the event information related to these neighboring pixels from the event stream, that is, the neighboring events. The neighboring events refer to all events related to the neighboring pixels. The neighboring events occur at different timestamps and reflect the states of the neighboring pixels at different time points.

[0145] According to the principle of occurring before the target timestamp and being closest to the occurrence time of the target timestamp, select a target neighboring event from all neighboring events. Select the neighboring events that are before the target timestamp from all neighboring events. The neighboring events that are before the target timestamp are events that have a certain influence on the change of the target pixel at the target timestamp. By selecting the event closest to the target timestamp, it can be ensured that the target neighboring event is the most relevant to the change of the target pixel at the target timestamp.

[0146] Target neighboring event The timestamp of can be determined by the following formula:

[0147]

[0148] Among them, , , represents the nearest event timestamp within the spatial neighborhood, represents the pixels within the u neighborhood, and k represents the k-th event.

[0149] Apply the decay rate to handle the relationship between the selected target neighborhood events and the target event. The decay rate refers to the process in which the influence of an event gradually weakens over time.

[0150] By combining the decayed neighborhood events with the event information of the target pixel, the temporal surface information of the target pixel at the target timestamp is finally calculated.

[0151] Pixel 's corresponding temporal surface can be determined by the following formula:

[0152]

[0153] Among them, provides the dynamic spatio-temporal background of the event, and the exponential decay extends the influence of past target neighborhood events on the event and provides information on the activity history within the neighborhood. represents the nearest event timestamp within the spatial neighborhood.

[0154] Adopting the embodiments of the present disclosure can utilize timestamps and event streams to capture the temporal surface information of each pixel in a dynamic scene. By considering the events of neighboring pixels, the information of the surrounding environment of the target pixel can be integrated, thereby more accurately reflecting the changes of the target pixel. By introducing the decay rate, the influence of past events on the current state can be reasonably modeled, avoiding unnecessary interference of outdated information on the modeling results.

[0155] Adopting the embodiments of the present disclosure, by analyzing the temporal surface information of each pixel during the training period, the brightness changes in the time series data can be effectively captured, which can help the model understand the law of the scene evolving over time. By integrating the brightness changes of each pixel and generating a model supervision signal, the model can be effectively guided to learn the temporal change features, thereby improving the fitting accuracy and stability of the model to the scene. By integrating the temporal surface information of each pixel and generating a supervision signal, this method can efficiently provide high-quality labels for model training.

[0156] Among them, in an optional embodiment, the training time period includes a first moment at the start of the training time period and a second moment at the end of the training time period; determining the rendering prediction value according to the initial three-dimensional Gaussian splash model includes: determining the first pose of the event camera at the first moment and determining the second pose of the event camera at the second moment; projecting the initial three-dimensional Gaussian splash model onto a two-dimensional image based on the first pose to obtain a first image; projecting the initial three-dimensional Gaussian splash model onto a two-dimensional image based on the second pose to obtain a second image; determining the first rendering intensity of each pixel in the first image and determining the second rendering intensity of each pixel in the second image; determining the rendering intensity difference of each pixel according to the first rendering intensity and the second rendering intensity corresponding to the same pixel; and determining the rendering intensity difference as the rendering prediction value of the corresponding pixel.

[0157] The training time period includes two moments, the first moment and the second moment. The first moment is the start of the training time period, and the second moment is the end of the training time period. For example, when the training time period is , the first moment is , and the second moment is .

[0158] Pose refers to the position and orientation of the camera. The pose includes the position and orientation of the camera in three-dimensional space.

[0159] The first pose of the event camera at the first moment and the second pose of the event camera at the second moment are the poses of the time camera relative to the reference object when collecting the event stream.

[0160] Through the pose of the event camera at the target moment, the projection direction of the three-dimensional Gaussian splash model can be determined. For example, when the event camera collects the event stream of the target scene, within the event segment, the event camera moves from the front of the target scene to the side of the target scene. Then at the moment, the event camera is in the pose of collecting the front of the target scene, and at the moment, the event camera is in the pose of collecting the side of the target scene. Then when projecting the initial three-dimensional Gaussian splash model, the first pose should also correspond to the front of the initial three-dimensional Gaussian splash model, and the second pose should also correspond to the side of the initial three-dimensional Gaussian splash model.

[0161] According to the poses of the event camera at the first moment and the second moment respectively, project the three-dimensional Gaussian splash model into a two-dimensional image to obtain two images at different time points, namely the first image and the second image.

[0162] Rendering intensity refers to the brightness or intensity of each pixel in an image. For each pixel, the rendering intensities at two moments are calculated separately. Specifically, in the first image, the first rendering intensity of each pixel is determined, which reflects the brightness or intensity of the pixel at the first moment; in the second image, the second rendering intensity of each pixel is determined, which reflects the brightness or intensity of the pixel at the second moment.

[0163] The scales of the first image and the second image are consistent, and any pixel point in the first image can find another pixel point with the same position in the second image. Therefore, for each pixel, the difference in rendering intensity between the first image and the second image is calculated.

[0164] The rendering prediction value can be determined according to the following formula:

[0165]

[0166]

[0167]

[0168] where represents the rendering prediction value, is the gamma correction value, and represent the rendering intensities of the initial three-dimensional Gaussian splash model at time stamps and respectively.

[0169] By adopting the embodiments of the present disclosure, by defining a specific training time period and performing projections at two moments, the dynamic changes in the time series can be efficiently modeled. By determining the projection direction of the initial three-dimensional Gaussian splash model through the pose of the event camera, it can be ensured that the two-dimensional images generated at each moment can fully reflect the dynamic changes in the scene. By calculating the difference in rendering intensity of the same pixel at two time points, factors such as light changes, object movements, and occlusion relationships in the scene can be reflected.

[0170] The initial three-dimensional Gaussian splash model can be projected onto a two-dimensional plane according to the steps provided in the following embodiments to obtain a two-dimensional image.

[0171] Among them, in an optional embodiment, projecting the initial three-dimensional Gaussian splash model onto a two-dimensional image includes: determining each three-dimensional Gaussian function included in the three-dimensional Gaussian splash model; projecting each of the three-dimensional Gaussian functions onto a two-dimensional plane to determine the corresponding two-dimensional Gaussian for each of the three-dimensional Gaussian functions; the two-dimensional Gaussian representing the color of the three-dimensional Gaussian in the two-dimensional plane; for any pixel in the two-dimensional plane, determining the corresponding two-dimensional Gaussians for the pixel, sorting the projected two-dimensional Gaussians based on a tile rasterizer, and determining the color corresponding to the pixel by means of alpha blending.

[0172] The initial three-dimensional Gaussian splash model is a model composed of multiple three-dimensional Gaussian functions, and the three-dimensional Gaussian function represents the density or weight of a certain point in a three-dimensional space. Each Gaussian function represents a different part or feature in the model.

[0173] The three-dimensional Gaussian function of any three-dimensional point in the three-dimensional point cloud can be represented by the following formula:

[0174]

[0175] Among them, represents the center of the 3DGS; represents the opacity; represents the three-dimensional covariance matrix. The three-dimensional covariance matrix is as follows:

[0176]

[0177] Among them, represents the rotation matrix, represents the scaling matrix, represents the transpose of the scaling matrix.

[0178] Project each three-dimensional Gaussian function from three-dimensional space onto a two-dimensional plane. Projecting onto a two-dimensional plane means mapping the Gaussian distribution in the three-dimensional coordinate system to the two-dimensional image plane. After the three-dimensional Gaussian function is projected onto the two-dimensional plane, it becomes a two-dimensional Gaussian function. Each three-dimensional Gaussian function corresponds to a two-dimensional Gaussian function, and the two-dimensional Gaussian function can be used to represent the color of the pixel after projection.

[0179] After projecting all the three-dimensional Gaussian functions onto the two-dimensional plane, calculate the corresponding two-dimensional Gaussian function for each pixel in the two-dimensional image according to the position of the pixel.

[0180] The color of each pixel is determined by the combination of multiple two-dimensional Gaussian functions associated with the pixel's position. This is because when different three-dimensional Gaussian functions are projected onto a two-dimensional image, they may affect multiple pixels. Therefore, the color of each pixel requires a comprehensive calculation of the contributions of multiple two-dimensional Gaussians. Each two-dimensional Gaussian function affects the color of a region in the image, and the range of influence is determined by its standard deviation. The smaller the standard deviation, the smaller the affected region; the larger the standard deviation, the wider the affected region.

[0181] Rasterization refers to the process of converting graphical objects into pixels. For the calculation of two-dimensional Gaussian projection, rasterization techniques are used to process each pixel in the image one by one, and the color of the pixel is determined according to its corresponding two-dimensional Gaussian function. The rasterizer of tiles is a technique to accelerate the rendering process. The rasterizer of tiles can divide the image into multiple small regions, i.e., tiles, and perform calculations on each tile separately. Tiling techniques help improve rendering efficiency.

[0182] In multi-level rendering, alpha blending is used to combine the color information of different layers. For the result of two-dimensional Gaussian projection, multiple two-dimensional Gaussian functions may overlap at a pixel position. Therefore, for a pixel, in order to obtain the final color of this pixel, the overlapping Gaussian functions need to be blended, and the result of the blending will generate the final color value.

[0183] For the determination of the final color value, it can be determined by the following formula:

[0184]

[0185] where, represents the number of sorted two-dimensional Gaussians related to the query pixel; represents the product of the opacity and density of the projected two-dimensional Gaussian distribution, represents the calculated color, i represents the index of the i-th two-dimensional Gaussian distribution being processed currently, and j is the index used for cumulative multiplication, representing the index of all Gaussian distribution terms before the i-th term being processed currently.

[0186] The final color of each pixel in the two-dimensional image can be used to determine the rendering intensity of the pixel in the two-dimensional image.

[0187] Adopting the embodiments of the present disclosure, each three-dimensional Gaussian function is mapped onto the two-dimensional image plane and a two-dimensional Gaussian distribution is generated through transformation, which can effectively generate a two-dimensional image from complex three-dimensional data. Through the rasterizer, the influence from different two-dimensional Gaussians can be calculated and synthesized for each pixel, thereby generating a high-quality image. Using alpha blending technology to synthesize different two-dimensional Gaussian functions ensures smooth transitions between multiple levels in the image and generates natural color variations.

[0188] Among them, in an optional embodiment, it further includes: determining an event threshold corresponding to the current iteration within a reasonable range; the reasonable range is determined by the event threshold configured by the event camera; determining an event threshold loss corresponding to the current iteration according to the event threshold; the event threshold loss represents the suitability of the event threshold selected in the current iteration; based on the difference between the ground truth and the rendering prediction value, adjusting the initial three-dimensional Gaussian splash model to obtain a three-dimensional Gaussian splash model corresponding to the target scene, including: determining an event rendering loss according to the ground truth and the rendering prediction value; adjusting the initial three-dimensional Gaussian splash model according to the event rendering loss and the event threshold loss to obtain a three-dimensional Gaussian splash model corresponding to the target scene.

[0189] In the related art, it is usually assumed that the event trigger threshold is constant; however, in fact, the event trigger threshold may vary between different perspectives, and setting it to a constant value may reduce the reconstruction quality. Therefore, the present disclosure proposes that the loss of the event threshold also needs to be considered.

[0190] In each iteration, an event threshold is re-determined. The event threshold is a simulated value, different from the real event threshold corresponding to the event camera. By simulating different event thresholds, a more accurate three-dimensional Gaussian splash model can be obtained. The event threshold selected in each iteration needs to be within a reasonable range, which is determined by the real event thresholds configured in each event camera. The determination of the simulated event threshold cannot deviate from the real event threshold.

[0191] The event threshold determines which brightness changes of pixels are regarded as valid events. The event threshold is usually set based on the magnitude of the brightness change. Only when the brightness change of a pixel exceeds a certain predetermined threshold will it be considered a valid event. For example, in a dynamic scene, if the brightness change of a certain pixel point exceeds the set threshold, the event is recorded. Conversely, if the brightness change is small, the event is ignored.

[0192] The event threshold loss reflects whether the event threshold selected in the current iteration is appropriate. The event threshold loss can measure how the selected event threshold affects the validity of the event stream and indirectly affects the accuracy of the three-dimensional model. For example, if the selected event threshold is too low, it may lead to too many redundant events, which may not represent important dynamic changes in the scene, thus having a negative impact on the accuracy of the model. If the event threshold is set too high, some small but meaningful events may be missed, resulting in the model being unable to accurately capture the details in the scene.

[0193] The goal of the event threshold loss is to optimize the selection of the threshold by evaluating the impact of the selected threshold on the effect of the three-dimensional reconstruction model.

[0194] The ground truth and the rendered prediction value can be used as the event rendering loss.

[0195] Considering both the event rendering loss and the event threshold loss, an adaptive threshold is determined based on the event rendering loss and the event threshold loss. The formula for the adaptive threshold is as follows:

[0196]

[0197] Wherein, represents the event rendering loss, represents the event threshold loss, and represent the weight coefficients.

[0198] The initial three-dimensional Gaussian splash model is adjusted through the adaptive threshold to obtain the three-dimensional Gaussian splash model corresponding to the target scene.

[0199] By adopting the embodiments of the present disclosure, by considering the influence of the event threshold loss on three-dimensional reconstruction, the quality of three-dimensional reconstruction can be improved. By minimizing the event threshold loss, the threshold selection can be made more appropriate, thereby reducing redundant information and noise, and making the result of three-dimensional reconstruction more accurate. The optimization of the event threshold loss and the event rendering loss enables the algorithm to dynamically adjust the three-dimensional model to adapt to rapid changes in the scene. Compared with traditional inter-frame image processing methods, data processing based on event streams can better handle rapid motion and sudden illumination changes, improving the update speed and accuracy of the model.

[0200] Wherein, in an alternative embodiment, determining the event threshold loss corresponding to the current iteration according to the event threshold includes: determining the selection interval of the event threshold; the threshold interval includes a threshold upper limit and a threshold lower limit; according to the unit step function, based on the threshold upper limit, the threshold lower limit, and the event threshold, determining the event threshold loss.

[0201] The threshold upper limit and the threshold lower limit constitute a reasonable threshold interval. The threshold interval is used to judge the rationality of the event threshold selected in each iteration.

[0202] The calculation of the event threshold loss is realized by introducing the unit step function. The unit step function is usually used to represent "switch" or "threshold" type behaviors. If the selected event threshold is less than the threshold lower limit, it means that too many valid events are obtained, increasing the computational burden. If the selected event threshold is greater than the threshold upper limit, it may cause too many valid events to be discarded, thus affecting the accuracy of the three-dimensional model, and the loss is also large.

[0203] Event threshold loss The calculation formula is as follows:

[0204]

[0205] Among them, represents the unit step function. and respectively represent the upper threshold and the lower threshold of the event trigger threshold, represents the event threshold selected in this round of iteration.

[0206] In each round of iteration, according to the current event threshold loss, adjusting the position of the event threshold can minimize the loss function, thereby ensuring that it is more adaptable to the current scenario.

[0207] Adopting the embodiments of the present disclosure, by introducing the calculation of the event threshold loss, combining the selection interval of the upper threshold and the lower threshold, and the unit step function, it is possible to avoid model errors caused by too high or too low event thresholds.

[0208] Based on the above three-dimensional reconstruction method based on event cameras provided by the present disclosure, a specific embodiment is used here to introduce the method proposed by the present disclosure in detail, and the specific steps are as follows:

[0209] Step 1: Construct a dataset for training the three-dimensional reconstruction method

[0210] Adopt the commonly used evaluation datasets for three-dimensional reconstruction based on event cameras. The evaluation datasets include synthetic datasets and real-world datasets, and the algorithm is evaluated through the synthetic datasets and real-world datasets.

[0211] The synthetic dataset includes seven objects with challenging features, such as fine structures and viewpoint-dependent colors. By rotating the event camera 360 degrees around each three-dimensional object, event streams for synthesizing the corresponding scenes of each object are captured from 1000 viewpoints.

[0212] The real-world dataset contains 10 event sequences. An event sequence represents a scene, and each sequence is recorded using a DAVIS 346C event camera under low light conditions, specifically a 5W light source, and the object rotates at a speed of 45 revolutions per minute. Based on the movement of the object, the external camera parameters of the sequence can be calculated.

[0213] Step 2: Based on EventNeRF, construct a rough 3D point cloud for each scene as the initialization of event-based 3DGS.

[0214] For each scene, the EventNeRF is trained for 30,000 iterations using the event stream corresponding to each scene to obtain a shallow training model. The event stream corresponding to the scene is input into the EventNeRF after shallow training, and the rough depth of the scene is calculated by the shallow-trained EventNeRF model. The rough depth map is converted into a rough 3D point cloud through back-projection:

[0215]

[0216] where u and v are pixel coordinates; d is the depth value; K is the intrinsic matrix; M is the extrinsic matrix; represents the transpose of the matrix, represents the 3D point cloud.

[0217] A temporal surface map encoding spatio-temporal information is obtained through the cumulative event polarity and scaled to the range [0, 255], where [0, 255] refers to the color range. Subsequently, the temporal surface map is used to color the rough 3D point cloud. Random downsampling is then employed to reduce the density of the rough 3D point cloud with a downsampling rate of 10% to obtain a 3D point cloud for 3DGS initialization.

[0218] Step 3: Determine the initial 3D Gaussian splash model based on the initialized point cloud

[0219] Based on the initialized point cloud, a set of anisotropic 3D Gaussians are used for explicit representation. Each 3D Gaussian is represented by the following formula for view-dependent appearance rendering:

[0220]

[0221] where, represents the center of the 3DGS; represents the opacity; represents the 3D covariance matrix.

[0222] Step 4: Process the event stream using a preprocessing method based on the spatio-temporal surface map to form a differentiable event supervision signal.

[0223] First, the spatio-temporal context around the event is extracted to construct a temporal surface map. Specifically, the temporal context around the event is extracted and denoted as : :

[0224]

[0225] where, , . represents the nearest event timestamp within the spatial neighborhood, Pixels in the u domain are represented, and k represents the k-th event.

[0226] After obtaining the temporal context, for the pixels Apply an exponential decay kernel with a decay rate of to obtain the temporal surface at timestamp t , the temporal surface represents the brightness change of the pixel at timestamp t under the influence of the temporal context on the event , as follows:

[0227]

[0228] where, provides a dynamic spatio-temporal background for the event, and the exponential decay extends the influence of past events and provides information about the activity history in the neighborhood, represents the timestamp of the nearest event in the spatial neighborhood.

[0229] Integrate the brightness changes of the events within the time periods and in the form of a temporal surface map as a differentiable supervision signal for training the 3DGS model:

[0230]

[0231] where, represents the event corresponding to the pixel , represents the temporal surface information of the pixel at each moment during the training time period, represents the event threshold selected in this iteration, and dt represents the minimum time step for integrating the event supervision signal.

[0232] Step Five: The 3D Gaussian splash model effectively presents the view through point splashes

[0233] Determine moment and moment, and the corresponding scanning poses of the event camera respectively. The scanning pose refers to the pose of the event camera when collecting the event stream.

[0234] According to the corresponding scanning poses of the event camera at moment and moment respectively, project the 3D Gaussian in the scene onto the 2D image plane to obtain each 2D Gaussian. The color of each 2D Gaussian is calculated through spherical harmonic parameters. Then, sort the 2D Gaussians through a tile-based rasterizer.

[0235] Based on the opacity and color of the two-dimensional Gaussian, the final color of each pixel is calculated by alpha blending as follows:

[0236]

[0237] where, represents the number of sorted two-dimensional Gaussians related to the query pixel; represents the product of the opacity and density of the projected two-dimensional Gaussian distribution, represents the calculated color, i represents the index of the i-th two-dimensional Gaussian distribution being processed currently, j is the index used for the cumulative multiplication operation, representing the index of all Gaussian distribution terms before the i-th term being processed currently.

[0238] In the case of projecting the three-dimensional Gaussian in the scene onto the two-dimensional image plane based on the scanning pose, since different scanning poses correspond to different times at time and time respectively, two two-dimensional images corresponding to the two times will be obtained.

[0239] Determine the final colors of the pixels at the same position on the two two-dimensional images respectively.

[0240] Step Six: Based on the event supervision signal, use the adaptive threshold loss to supervise the training of E-3DGS, so as to generate an accurate scene representation

[0241] E-3DGS is a three-dimensional Gaussian splash model determined by the three-dimensional reconstruction method provided by the present disclosure.

[0242] The adaptive threshold includes the event rendering loss and the time threshold loss.

[0243] Determine the pixel through Step Four at and The integral brightness change of the event polarity is as follows:

[0244]

[0245] where, represents the event corresponding to the pixel ; represents the time surface information of the pixel at each moment during the training time period, represents the event threshold selected in this round of iteration, dt represents the minimum time step for integrating the event supervision signal, represents the training time period.

[0246] Determine the pixel through Step Five at The grayscale difference between the rendered images of the 3DGS model between and Specifically, for the timestamps and the rendering intensity of the pixel on the two-dimensional images corresponding to the two timestamps is determined using the formula provided in Step Five for determining the final color of the pixel. The logarithmic luminance images for training are denoted as and respectively, where g is the gamma correction value set to 2.2 in the experiment.

[0247] The logarithmic luminance difference of the pixel between and is expressed as:

[0248]

[0249] By minimizing the difference between the two and calculating the mean squared error loss through MSE() as the event rendering loss, the event rendering loss is obtained:

[0250]

[0251] where represents the rendering prediction value of the pixel, and

[0252] represents the ground truth of the pixel. In each iteration, the event threshold is dynamically modeled, and through the event threshold loss, the threshold for event triggering is limited within a reasonable range. The event threshold loss term can be determined by the following formula:

[0253]

[0254] where represents the unit step function, and represent the upper and lower limits of the event triggering threshold respectively. In the experiment, the upper threshold can be set to = 0.2, and the lower threshold to = -0.2.

[0255] The event rendering loss and the event threshold loss jointly supervise the optimization of the 3DGS function. The weight coefficients of the event rendering loss and the event threshold loss are set to 0.9 and 0.1 respectively to obtain the adaptive loss.

[0256]

[0257] The E-3DGS of this iteration is adjusted through the adaptive loss.

[0258] Step 7: Model training was carried out 30,000 times on the NVIDIA RTX 4090 GPU, which took about 6 minutes to complete. In the experimental implementation, it was assumed that the pixel background value was known and set to 156, where 156 represents the color of the background. Rendering a new view image with a resolution of 246*260 through the three-dimensional Gaussian splash model requires approximately 73 FPS.

[0259] Compared with EventNeRF, which also only performs dense reconstruction from the event stream, and E2VID-3DGS, which first uses E2VID to recover grayscale frames from the event stream and then uses these frames to train the original frame-based 3DGS for three-dimensional reconstruction, the three-dimensional reconstruction method provided by the present disclosure, i.e., E-3DGS in Tables 1 and 2, has reached the state-of-the-art level in rendering quality.

[0260] Table 1: Quantitative comparison of synthetic scenes

[0261]

[0262] As shown in Table 1, compared with EventNeRF on the synthetic dataset, the three-dimensional reconstruction method provided by the present disclosure has an average increase in PSNR (Peak Signal-to-Noise Ratio) of 5.06 dB, an increase in SSIM (Structural Similarity Index) of 6.2%, and a decrease in LPIPS (Learned Perceptual Image Patch Similarity) of 71%.

[0263] Table 2: Quantitative comparison of real scenes

[0264]

[0265] As shown in Table 2, compared with EventNeRF on the real dataset, the PSNR of the three-dimensional reconstruction method provided by the present disclosure has increased by 1.95 dB.

[0266] In addition, our method only requires 1% of the training time of EventNeRF.

[0267] Moreover, in the real-world dataset, E2VID+3DGS performs significantly poorly, while the rendering quality of the three-dimensional reconstruction method provided by the present disclosure is significantly better than that of E2VID+3DGS, because the present disclosure directly constructs a three-dimensional Gaussian splash model through events.

[0268] Based on the same technical concept, the present disclosure provides a three-dimensional reconstruction device based on three-dimensional Gaussian splashing using an event camera. Figure 2 It is a block diagram of a three-dimensional reconstruction device based on three-dimensional Gaussian splashing using an event camera shown in an embodiment of the present disclosure. According to Figure 2 as shown, the three-dimensional reconstruction device includes:

[0269] An acquisition module 210, configured to acquire an event stream for a target scene;

[0270] A model determination module 220, configured to determine an initial three-dimensional Gaussian splashing model of the target scene according to the event stream;

[0271] A temporal surface determination module 230, configured to determine temporal surface information of each pixel according to the event stream; the temporal surface information characterizes the dynamic spatio-temporal background information of the event;

[0272] A ground truth determination module 240, configured to determine the ground truth of the target scene according to the temporal surface information of each pixel; the ground truth characterizes the actual brightness change of the target scene;

[0273] A predicted value determination module 250, configured to determine a rendering predicted value according to the initial three-dimensional Gaussian splashing model, where the rendering predicted value characterizes the rendering intensity difference of the initial three-dimensional Gaussian splashing model at two viewpoints;

[0274] An adjustment module 260, configured to adjust the initial three-dimensional Gaussian splashing model based on the difference between the ground truth and the rendering predicted value to obtain a three-dimensional Gaussian splashing model corresponding to the target scene.

[0275] An embodiment of the present disclosure also provides an electronic device. Referring to Figure 3 , Figure 3 it is a schematic diagram of an electronic device shown in an embodiment of the present disclosure. As Figure 3 shown, the electronic device 300 includes: a memory 310 and a processor 320. The memory 310 is communicatively connected to the processor 320 through a bus. A computer program is stored in the memory 310, and the computer program can run on the processor 320 to implement the steps in the three-dimensional reconstruction method disclosed in the embodiment of the present disclosure.

[0276] An embodiment of the present disclosure also provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the steps in the three-dimensional reconstruction method disclosed in the embodiment of the present disclosure are implemented.

[0277] An embodiment of the present disclosure also provides a computer program product, including a computer program, which when executed by a processor, implements the steps in the three-dimensional reconstruction method disclosed in the embodiments of the present disclosure.

[0278] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other.

[0279] Those skilled in the art should understand that the embodiments of the present disclosure can be provided as a method, a device, or a computer program product. Therefore, the embodiments of the present disclosure can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present disclosure can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0280] The embodiments of the present disclosure are described with reference to the flowcharts and / or block diagrams of methods, devices, electronic devices, and computer program products according to the embodiments of the present disclosure. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing terminal devices generate a device for implementing the specified functions in one Figure 1 flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0281] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the specified functions in one Figure 1 flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0282] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, such that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable terminal device provide steps for implementing the specified functions in one Figure 1 flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0283] Although some embodiments of the present disclosure have been described, those skilled in the art can make additional changes and modifications to these embodiments once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the embodiments of the present disclosure.

[0284] The above has introduced in detail the method for three-dimensional reconstruction based on an event camera using three-dimensional Gaussian splashing provided by the present disclosure. Specific examples are used herein to elaborate on the principle and implementation manner of the present disclosure. The description of the above embodiments is only used to help understand the method and its core idea of the present disclosure; at the same time, for those of ordinary skill in the art, according to the idea of the present disclosure, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present disclosure.

Claims

1. A method for three-dimensional reconstruction based on three-dimensional Gaussian splashing using an event camera, characterized in that, Including: Obtain an event stream for a target scene, including: determining the brightness change value of each pixel at a preset time interval; generating an event stream when the brightness change value exceeds the event threshold corresponding to the event camera; each event in the event stream includes a pixel coordinate, a timestamp, and an event polarity; Determine an initial three-dimensional Gaussian splash model for the target scene according to the event stream; Determine the time surface information of each pixel according to the event stream; the time surface information characterizes the dynamic spatio-temporal background information of the event; Determine the ground truth of the target scene according to the time surface information of each pixel; the ground truth characterizes the actual brightness change of the target scene; Determine a rendering prediction value according to the initial three-dimensional Gaussian splash model, where the rendering prediction value characterizes the rendering intensity difference of the initial three-dimensional Gaussian splash model at two viewpoints; Adjust the initial three-dimensional Gaussian splash model based on the difference between the ground truth and the rendering prediction value to obtain a three-dimensional Gaussian splash model corresponding to the target scene.

2. The method according to claim 1, wherein Also including: Determine the event threshold corresponding to the current iteration within a reasonable range; The reasonable range is determined by the event threshold configured by the event camera; Determine the event threshold loss corresponding to the current iteration according to the event threshold; The event threshold loss characterizes the suitability of the event threshold selected in the current iteration; Adjust the initial three-dimensional Gaussian splash model based on the difference between the ground truth and the rendering prediction value to obtain a three-dimensional Gaussian splash model corresponding to the target scene, including: Determine an event rendering loss according to the ground truth and the rendering prediction value; Adjust the initial three-dimensional Gaussian splash model according to the event rendering loss and the event threshold loss to obtain a three-dimensional Gaussian splash model corresponding to the target scene.

3. The method according to claim 2, characterized in that, Determine the event threshold loss corresponding to the current iteration according to the event threshold, including: Determine the selection interval of the event threshold; the threshold interval includes a threshold upper limit and a threshold lower limit; Determine the event threshold loss according to the unit step function based on the threshold upper limit, the threshold lower limit, and the event threshold.

4. The method according to claim 1, wherein Adjust the initial three-dimensional Gaussian splash model based on the difference between the ground truth and the rendering prediction value to obtain a three-dimensional Gaussian splash model corresponding to the target scene, including: Determine a first difference corresponding to the pixel through the ground truth and the rendering prediction value corresponding to the same pixel; Determine the difference between the ground truth and the rendering prediction value through the first differences corresponding to each pixel; Adjust the initial three-dimensional Gaussian splash model based on the difference between the ground truth and the rendering prediction value to obtain a three-dimensional Gaussian splash model corresponding to the target scene.

5. The method according to claim 1, characterized in that Determine the ground truth of the target scene according to the time surface information of each pixel, including: Determine the training time period corresponding to the current iteration; Determine the time surface information of each pixel at each moment included in the training time period; Integrate the time surface information of each pixel respectively to determine the brightness change of each pixel during the training time period and generate each model supervision signal; Determine the model supervision signal as the ground truth of the corresponding pixel during the training period.

6. The method according to claim 1, wherein Determine the temporal surface information of each event of each pixel according to the event stream, including: Determine the target pixel and the target timestamp corresponding to each event; Determine each neighboring pixel included in the neighborhood range of the target pixel; Determine each neighboring event of each neighboring pixel; Select, from each neighboring event, the target neighboring event that occurs before the target timestamp and is closest to the occurrence time of the target timestamp; Apply an attenuation rate, and determine the temporal surface information corresponding to the target pixel at the target timestamp through the neighboring event and the event.

7. The method according to claim 5, wherein The training period includes a first moment at the start of the training period and a second moment at the end of the training period; determine the rendering prediction value according to the initial three-dimensional Gaussian splash model, including: Determine the first pose of the event camera at the first moment and determine the second pose of the event camera at the second moment; Project the initial three-dimensional Gaussian splash model onto a two-dimensional image based on the first pose to obtain a first image; Project the initial three-dimensional Gaussian splash model onto a two-dimensional image based on the second pose to obtain a second image; Determine the first rendering intensity of each pixel in the first image and determine the second rendering intensity of each pixel in the second image; Determine the rendering intensity difference of each pixel according to the first rendering intensity and the second rendering intensity corresponding to the same pixel; Determine the rendering intensity difference as the rendering prediction value of the corresponding pixel.

8. The method according to claim 7, characterized in that, Project the initial three-dimensional Gaussian splash model onto a two-dimensional image, including: Determine each three-dimensional Gaussian function included in the three-dimensional Gaussian splash model; Project each of the three-dimensional Gaussian functions onto a two-dimensional plane to determine the two-dimensional Gaussian corresponding to each of the three-dimensional Gaussian functions; the two-dimensional Gaussian represents the color of the three-dimensional Gaussian in the two-dimensional plane; For any pixel in the two-dimensional plane, determine the two-dimensional Gaussians corresponding to the pixel, sort the projected two-dimensional Gaussians based on a tile rasterizer, and determine the color corresponding to the pixel by means of alpha blending.

9. The method according to claim 1, characterized in that, Determine the initial three-dimensional Gaussian splash model of the target scene according to the event stream, including: Perform shallow training on the initial target model based on the event stream to obtain a target model; Input the event stream into the target model, and recover the depth information of each pixel from the event stream through the target model to obtain a rough depth map; Convert the rough depth map into a three-dimensional point cloud through back-projection; Process the three-dimensional point cloud through a temporal surface map composed of the temporal surface information to obtain an initial three-dimensional Gaussian splash model; the processing at least includes: coloring the three-dimensional point cloud through the temporal surface map.

Citation Information

Patent Citations

  • Nerve radiation field artifact suppression method based on event camera three-dimensional reconstruction

    CN118887334A

  • High-fidelity large-scale scene rendering method, system and equipment based on Gaussian splashing and medium

    CN119295621A