Method for realizing three-dimensional reconstruction based on three-dimensional Gaussian splashing of event camera
Through the event camera-based three-dimensional Gaussian splattering method, combining event stream and temporal surface information, the three-dimensional Gaussian splattering model is adjusted to match the ground real-life and rendering prediction values, and the problem of difficult balance between rendering speed and quality in the prior art is solved, and a high-precision and high-density three-dimensional reconstruction effect is achieved.
Patent Information
- Application Number
- CN202510429811.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-08
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-08
AI Technical Summary
The prior art is difficult to balance rendering speed and quality when using mobile event cameras for three-dimensional reconstruction, especially when dealing with complex scenes, and cannot meet the needs of high-precision and high-density scene reconstruction.
Three-dimensional Gaussian splattering method based on event camera is used to achieve three-dimensional reconstruction by obtaining event streams, determining the initial 3-dimensional Gaussian splattering model, analyzing temporal surface information, and adjusting the model to match the difference between ground real-life and rendered predicted values.
It realizes the generation of accurate and detailed three-dimensional reconstruction in a dynamic environment, avoids the loss of information caused by frame rate limitation in traditional methods, and improves the timeliness and accuracy of three-dimensional reconstruction.
Smart Images

Figure CN119942004A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a method for implementing three-dimensional reconstruction based on three-dimensional Gaussian splashing of an event camera. Background Art
[0002] Creating detailed, realistic, real-time 3D scene renderings is a persistent challenge in computer vision and graphics. Event cameras are a new type of biologically inspired visual sensor that uses asynchronous dynamic visual perception technology. Unlike traditional RGB cameras, event cameras record asynchronous event streams triggered by pixel brightness changes rather than capturing absolute brightness values at a fixed frame rate. Each pixel responds to brightness changes independently, and generates an event with a timestamp when the change exceeds a preset threshold. This unique capture method enables event cameras to perform well in high-speed motion and high dynamic range scenes, and is very suitable for applications that require high temporal resolution and low latency, such as simultaneous positioning and mapping, target recognition and tracking, optical flow estimation, etc.
[0003] Currently, there are two main methods for using mobile event cameras to reconstruct three-dimensional scenes: the first is based on visual odometry and SLAM (Simultaneous Localization and Mapping), and the second is a rendering method based on neural radiance field (Nerf).
[0004] The first method mainly reconstructs sparse or semi-dense scene structures based on explicit feature matching between event accumulation images. However, the first method may be limited by its sparse reconstruction characteristics when dealing with complex scenes and cannot meet the needs of high-precision and high-density scene reconstruction. In recent years, the second method has gradually been applied to scene reconstruction of event cameras. These methods supervise the training of neural radiance fields through event flow integration to render realistic images. However, the second method relies on the sampling strategy of ray stepping during training and rendering. Excessive sampling may lead to slower speed, while too little sampling will affect the rendering quality.
[0005] How to strike a balance between rendering speed and quality becomes a major challenge, which leads to the study of how to achieve dense, accurate and fast scene reconstruction using only event streams. Summary of the invention
[0006] In order to overcome the problems existing in the related art, the present disclosure provides a method for implementing 3D reconstruction based on 3D Gaussian splashing of an event camera. The technical solution of the present disclosure is as follows: According to a first aspect of an embodiment of the present disclosure, a method for implementing three-dimensional reconstruction based on three-dimensional Gaussian splashing of an event camera is provided, comprising: Get the event stream for the target scenario; Determining an initial three-dimensional Gaussian splash model of a target scene according to the event stream; Determine the time surface information of each pixel according to the event stream; the time surface information represents the dynamic spatiotemporal background information of the event; Determining a ground truth of a target scene based on temporal surface information of each pixel; the ground truth characterizing actual brightness changes of the target scene; Determining a rendering prediction value according to the initial three-dimensional Gaussian splash model, wherein the rendering prediction value represents a rendering intensity difference of the initial three-dimensional Gaussian splash model at two viewing angles; Based on the difference between the ground truth and the rendering prediction value, the initial three-dimensional Gaussian splash model is adjusted to obtain a three-dimensional Gaussian splash model corresponding to the target scene.
[0007] Optionally, it also includes: Determine the event threshold corresponding to the current iteration from a reasonable range; the reasonable range is determined by the event threshold configured by the event camera; According to the event threshold, determining the event threshold loss corresponding to the current iteration; the event threshold loss represents: the suitability of the event threshold selected in the current iteration; Based on the difference between the ground truth and the rendering prediction value, the initial three-dimensional Gaussian splash model is adjusted to obtain a three-dimensional Gaussian splash model corresponding to the target scene, including: determining an event rendering loss based on the ground truth and the rendering prediction value; The initial three-dimensional Gaussian splash model is adjusted according to the event rendering loss and the event threshold loss to obtain a three-dimensional Gaussian splash model corresponding to the target scene.
[0008] Optionally, determining the event threshold loss corresponding to the current iteration according to the event threshold includes: Determine a selection interval of the event threshold; the threshold interval includes an upper threshold limit and a lower threshold limit; The event threshold loss is determined based on the upper threshold, the lower threshold, and the event threshold according to a unit step function.
[0009] Optionally, based on the difference between the ground truth and the rendering prediction value, adjusting the initial three-dimensional Gaussian splash model to obtain a three-dimensional Gaussian splash model corresponding to the target scene includes: Determining a first difference corresponding to the pixel by using the ground truth and the rendering prediction value corresponding to the same pixel; Determine the difference between the ground truth and the rendering prediction value by using the first difference corresponding to each pixel; Based on the difference between the ground truth and the rendering prediction value, the initial three-dimensional Gaussian splash model is adjusted to obtain a three-dimensional Gaussian splash model corresponding to the target scene.
[0010] Optionally, determining the ground truth of the target scene based on the temporal surface information of each pixel includes: Determine the training time period corresponding to this round of iteration; Determining the temporal surface information of each pixel at each moment in the training time period; Integrate the temporal surface information of each pixel respectively, determine the brightness change of each pixel in the training time period, and generate each model supervision signal; The model supervision signal is determined as the ground truth of the corresponding pixel during the training time period.
[0011] Optionally, determining the time surface information of each event of each pixel according to the event stream includes: Determine the target pixel and target timestamp corresponding to each event; Determine each neighborhood pixel included in the neighborhood range of the target pixel; Determine each neighborhood event of each neighborhood pixel; Select a target neighborhood event that occurs before a target timestamp and is closest to the target timestamp from each neighborhood event; A decay rate is applied to determine a time surface of the target pixel corresponding to a target timestamp through the neighborhood events and the events.
[0012] Optionally, the training time period includes a first moment at the beginning of the training time period and a second moment at the end of the training time period; and determining the rendering prediction value according to the initial three-dimensional Gaussian splash model includes: Determine a first pose of the event camera at a first moment, and determine a second pose of the event camera at a second moment; Based on the first posture, projecting the initial three-dimensional Gaussian splash model onto a two-dimensional image to obtain a first image; Based on the second posture, projecting the initial three-dimensional Gaussian splash model onto a two-dimensional image to obtain a second image; Determine a first rendering intensity of each pixel in the first image, and determine a second rendering intensity of each pixel in the second image; Determining a rendering intensity difference of each pixel according to the first rendering intensity and the second rendering intensity corresponding to the same pixel; The rendering intensity difference is determined as the rendering prediction value of the corresponding pixel.
[0013] Optionally, projecting the initial three-dimensional Gaussian splash model onto a two-dimensional image comprises: Determine each three-dimensional Gaussian function included in the three-dimensional Gaussian splash model; Projecting each of the three-dimensional Gaussian functions onto a two-dimensional plane, and determining a two-dimensional Gaussian corresponding to each of the three-dimensional Gaussian functions; the two-dimensional Gaussian represents the color of the three-dimensional Gaussian on the two-dimensional plane; For any pixel in a two-dimensional plane, each two-dimensional Gaussian corresponding to the pixel is determined, and the projected two-dimensional Gaussians are sorted based on a tile-based rasterizer, and the color corresponding to the pixel is determined by alpha blending.
[0014] Optionally, determining an initial three-dimensional Gaussian splash model of a target scene according to the event stream includes: Performing shallow training on the initial target model based on the event stream to obtain a target model; Inputting the event stream into the target model, and recovering depth information of each pixel from the event stream through the target model to obtain a rough depth map; Converting the rough depth map into a three-dimensional point cloud by back-projection; The three-dimensional point cloud is processed by a time surface graph composed of the time surface information to obtain an initial three-dimensional Gaussian splash model; the processing at least includes: coloring the three-dimensional point cloud by the time surface graph.
[0015] Optionally, obtain the event stream for the target scenario, including: Determine the brightness change value of each pixel at a preset time interval; When the brightness change value exceeds the event threshold corresponding to the event camera, an event stream is generated; each event in the event stream includes pixel coordinates, a timestamp and an event polarity.
[0016] According to a second aspect of an embodiment of the present disclosure, a device for implementing three-dimensional reconstruction based on three-dimensional Gaussian splashing of an event camera is provided, comprising: An acquisition module is used to obtain the event stream for the target scenario; A model determination module, used to determine an initial three-dimensional Gaussian splash model of a target scene according to the event stream; A time surface determination module, used to determine the time surface information of each pixel according to the event stream; the time surface information represents the dynamic spatiotemporal background information of the event; A ground truth determination module, for determining the ground truth of the target scene based on the temporal surface information of each pixel; the ground truth represents the actual brightness change of the target scene; A prediction value determination module, used to determine a rendering prediction value according to the initial three-dimensional Gaussian splash model, wherein the rendering prediction value represents a rendering intensity difference of the initial three-dimensional Gaussian splash model at two viewing angles; An adjustment module is used to adjust the initial three-dimensional Gaussian splash model based on the difference between the ground truth and the rendering prediction value to obtain a three-dimensional Gaussian splash model corresponding to the target scene.
[0017] According to a third aspect of an embodiment of the present disclosure, there is provided an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the computer program is executed by the processor, the steps of the three-dimensional reconstruction method described in the first aspect are implemented.
[0018] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the three-dimensional reconstruction method described in the first aspect are implemented.
[0019] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, comprising a computer program, and when the computer program is executed by a processor, the steps of the three-dimensional reconstruction method as described in the first aspect are implemented.
[0020] The present disclosure combines the high temporal precision of the event stream with the careful construction of the Gaussian model, so that the three-dimensional reconstruction results will have high precision and real-time performance, and are suitable for the reconstruction of dynamic environments. The accurate description of the dynamic spatiotemporal background through the time surface information can help the system better capture and reconstruct the three-dimensional form of fast-moving objects, avoiding the information loss caused by the frame rate limitation of the traditional method. Through the difference between the ground truth and the rendering prediction value, the initial three-dimensional Gaussian splash model is adaptively adjusted to achieve a more refined and natural rendering effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings required for use in the description of the embodiments of the present disclosure will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0022] Figure 1 is a schematic diagram of the steps of a method for implementing three-dimensional reconstruction based on three-dimensional Gaussian splashing of an event camera shown in an embodiment of the present disclosure; Figure 2 is a block diagram of a device for implementing three-dimensional reconstruction by three-dimensional Gaussian splashing based on an event camera shown in an embodiment of the present disclosure; Figure 3 It is a schematic diagram of an electronic device shown in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0023] The following will be combined with the drawings in the embodiments of the present disclosure to clearly and completely describe the technical solutions in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.
[0024] The terms "first", "second", etc. in the specification and claims of the present disclosure are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable when appropriate, so that the embodiments of the present disclosure can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.
[0025] Recent advances in 3D Gaussian Splatting (3DGS) technology provide a promising solution to the problems of related technologies. Compared with NeRF, 3DGS has obvious advantages in training and rendering speed, thanks to its explicit scene representation, highly parallelized workflow, and efficient computing strategy that combines differentiable rendering pipeline with point-based rendering technology. In addition, 3DGS's ability to control scene dynamics is particularly important in scenes with complex geometric structures.
[0026] However, the 3DGS method is mainly oriented towards RGB (Red, Green, Blue) frames, and there will be scalability issues when applied to event streams. For example, 3DGS requires point cloud initialization, while the traditional Structure from Motion (SfM)-based method is not applicable to event streams, and the point cloud generated by the event-based point cloud reconstruction strategy is too sparse to meet the requirements of 3DGS initialization.
[0027] Therefore, in order to solve the above technical problems, the present disclosure proposes a method for implementing 3D reconstruction based on 3D Gaussian splashing of an event camera. The method can well combine event stream and 3D Gaussian splashing to reconstruct a 3D scene.
[0028] Figure 13D Gaussian splashing based on event camera is a schematic diagram of the steps of implementing 3D reconstruction method according to an embodiment of the present disclosure. Figure 1 As shown, the method may specifically include the following steps: Step S11: Obtain the event stream for the target scene.
[0029] The event stream of the target scene is collected through a moving event camera. An event camera is a bionic sensor different from a standard camera. The event camera does not output direct pixel values, but asynchronously measures the changes in each pixel and outputs information encoding the time, position and sign of these changes. The pixels are the pixels corresponding to the event camera's viewfinder. For example, if the event camera's viewfinder is composed of 4*4 pixels, then the event camera will capture the brightness changes of each pixel within the 4*4 pixel range.
[0030] Step S12: Determine an initial three-dimensional Gaussian splash model of the target scene according to the event stream.
[0031] Different from the method of determining the initial three-dimensional Gaussian splash model corresponding to 3DGS through RGB in the related art, the initial three-dimensional Gaussian splash model corresponding to 3DGS can be determined through the event stream.
[0032] Using the time, location and symbol information in the event stream, the 3D shape and location of the object in the target scene can be inferred. The initial 3D Gaussian splash model is obtained by representing these inference results in the form of Gaussian distribution.
[0033] Step S13: determining the time surface information of each pixel according to the event stream; the time surface information represents the dynamic spatiotemporal background information of the event.
[0034] The temporal surface information characterizes the dynamic spatiotemporal background information of the event. By analyzing the timestamp information in the event stream, the temporal surface information of each pixel can be obtained. The dynamic spatiotemporal background information represents the impact of the past events of the neighboring pixels on the current events of the target pixel.
[0035] Step S14: determining the ground truth of the target scene according to the temporal surface information of each pixel; the ground truth represents the actual brightness change of the target scene.
[0036] The ground truth refers to the actual brightness change during a period of observation time. The ground truth of the target scene can reflect the actual brightness change of each pixel during the observation time.
[0037] Step S15: determining a rendering prediction value according to the initial three-dimensional Gaussian splash model, wherein the rendering prediction value represents a rendering intensity difference of the initial three-dimensional Gaussian splash model at two viewing angles.
[0038] The rendering prediction value can be obtained by projecting the initial 3D Gaussian splash model onto two different viewpoints and calculating the rendering intensity difference between the two viewpoints.
[0039] The selection of the projection viewing angle is determined by the ground truth. The two viewing angles of the initial three-dimensional Gaussian splash model projection are determined by the observation time corresponding to the ground truth. Usually, the starting time and the ending time of the observation time are selected to determine the corresponding viewing angles respectively.
[0040] Determine the rendering intensity difference of the pixel at the same position in the images corresponding to the two perspectives, and determine the rendering intensity difference of the initial three-dimensional Gaussian splash model at the two perspectives. The rendering intensity difference is a prediction value used to evaluate the accuracy of the initial three-dimensional Gaussian splash model.
[0041] Step S16: Based on the difference between the ground truth and the rendering prediction value, the initial three-dimensional Gaussian splash model is adjusted to obtain a three-dimensional Gaussian splash model corresponding to the target scene.
[0042] By comparing the difference between the ground truth and the rendering prediction value, the accuracy of the initial 3D Gaussian splash model can be evaluated. If the difference is large, it means that there is a certain deviation between the initial 3D Gaussian splash model and the real scene. At this time, the initial 3D Gaussian splash model needs to be adjusted to reduce the difference and improve the accuracy of the model.
[0043] The adjustment method can be to iteratively optimize the model parameters based on an optimization algorithm until the difference between the ground truth and the rendered prediction value is minimized.
[0044] By adopting the embodiments of the present disclosure, through the combination of event stream and three-dimensional Gaussian splash model, the method can generate accurate and detailed three-dimensional reconstruction in a dynamic environment, avoiding the loss of accuracy caused by motion blur or insufficient sampling frequency in traditional methods. The temporal surface information enables the model to better reflect the dynamic characteristics of objects in rapidly changing scenes, solving the problem that traditional methods are difficult to cope with high-speed motion and lighting changes, and improving the timeliness and accuracy of three-dimensional reconstruction. The contrast adjustment of the rendering prediction and the actual brightness difference can achieve a more refined rendering effect, especially when dealing with dynamic objects and complex lighting conditions. The optimized three-dimensional model is more realistic in visual effects and detail processing. Based on the difference between the rendering prediction value and the ground truth, the initial three-dimensional Gaussian splash model is adaptively adjusted, and the adaptive adjustment mechanism improves the accuracy and stability of three-dimensional reconstruction.
[0045] Among them, in an optional embodiment, obtaining an event stream for a target scene includes: determining the brightness change value of each pixel in a preset time interval; generating an event stream when the brightness change value exceeds an event threshold corresponding to an event camera; each event in the event stream includes pixel coordinates, timestamp and event polarity.
[0046] Each pixel refers to a pixel corresponding to the event camera's viewfinder. When the event camera moves, each pixel in the event camera's viewfinder will change.
[0047] For each pixel, the event camera continuously monitors its brightness. The brightness change can be obtained by calculating the brightness difference between consecutive time intervals.
[0048] Calculate the brightness change of pixel x at timestamp t , can be determined by the following formula:
[0049] in, Represents a short time interval, , represents the brightness of pixel x at timestamp t, Represents pixel x at timestamp brightness.
[0050] When the brightness of a pixel changes Exceeding the set event threshold When the event camera considers that the pixel has changed significantly, it generates an event stream, where the event threshold of the event camera is usually 10%-50% of the brightness value. The generated event stream can be expressed as follows:
[0051] in, The pixel coordinates representing the event; The timestamp indicating when the event occurred; Indicates the polarity of the event.
[0052] The pixel coordinates indicate the specific location where the event change occurred, indicating the location of the change in the image. The timestamp records the exact time when the event occurred, which can be used to reconstruct the dynamic changes of the scene. The event polarity indicates the direction of the brightness change, including two polarities: positive polarity indicating an increase in brightness and negative polarity indicating a decrease in brightness.
[0053] Generating pixels In timestamp In the case of an event, the event does not carry the event threshold set when the event camera collects the event stream.
[0054] After getting the event, the pixel Events Modeling is performed as follows:
[0055] in, It's time The unit pulse function, Indicates the polarity of the event.
[0056] With the embodiments of the present disclosure, the event stream generates events according to the brightness change value, avoiding the generation of redundant data and improving the efficiency of data storage and calculation. The dynamic changes of the scene can be accurately analyzed through pixel coordinates, timestamps and event polarity. It can accurately capture the rapid dynamic changes in the scene and is suitable for scenes with high-speed motion and lighting changes.
[0057] Among them, in an optional embodiment, an initial three-dimensional Gaussian splash model of the target scene is determined according to the event stream, including: shallow training of the initial target model based on the event stream to obtain a target model; inputting the event stream into the target model, and recovering the depth information of each pixel from the event stream through the target model to obtain a rough depth map; converting the rough depth map into a three-dimensional point cloud through back projection; processing the three-dimensional point cloud through a time surface map composed of the time surface information to obtain an initial three-dimensional Gaussian splash model; the processing at least includes: coloring the three-dimensional point cloud through the time surface map.
[0058] Shallow training means preliminary training of the preliminary target model to quickly build a preliminary 3D model. The preliminary target model can be an EventNeRF model. In the shallow training stage, the target model only needs to make preliminary adjustments and learn the model according to the input information of the event stream.
[0059] After initial training, the target model receives data input from the event stream. By analyzing the event stream, the target model uses the patterns obtained from training to infer the depth value of each pixel and obtain a coarse depth map. A coarse depth map is an image in which the grayscale value of each pixel represents the depth information of the scene corresponding to the pixel.
[0060] After obtaining the rough depth map, the rough depth map is converted into a 3D point cloud through back projection. Back projection is to map each pixel on the image plane back to the 3D space through the known depth information and the internal parameters of the camera, so as to obtain the corresponding 3D points, which will form a 3D point cloud.
[0061] The three-dimensional point cloud can be obtained by reverse projection using the following formula:
[0062] Among them, u and v are pixel coordinates; d is the depth value; K is the intrinsic parameter matrix; M is the extrinsic parameter matrix; represents the transpose of a matrix, Represents a 3D point cloud.
[0063] The time surface diagram uses time as a dimension of space to represent the dynamic process of a physical quantity changing over time. The time surface diagram is composed of various time surface information.
[0064] By further processing the 3D point cloud through the time surface graph, each point in the point cloud can be colored using the information in the time surface graph, thereby reflecting the characteristics of the target scene such as illumination and motion status. The information in the time surface graph includes the timestamp of each point in the event stream, the change in light intensity, and other information.
[0065] After further processing the 3D point cloud through the time surface graph, a 3D point cloud for initializing the 3D Gaussian splash model is obtained. Based on the 3D point cloud for initializing the 3D Gaussian splash model, a set of anisotropic 3D Gaussian functions are used to display the 3D point cloud to obtain an initial 3D Gaussian splash model. The initial 3D Gaussian splash model represents the spatial distribution, object morphology and time evolution of the target scene by fusing event flow, depth information, point cloud and time surface information.
[0066] By using the embodiments of the present disclosure, event stream driven shallow training can realize the generation of a three-dimensional point cloud for initializing a three-dimensional Gaussian splash model through an event stream. Due to the high temporal resolution information provided by the event camera, depth can be quickly recovered from the event stream, especially under fast motion or complex lighting conditions, and more accurate depth information can be provided than traditional frame cameras. Through the time surface map, the dynamic features of the scene that change over time can be accurately represented, and coloring the three-dimensional point cloud through the time surface map can improve the authenticity of the visual effect.
[0067] Among them, in an optional embodiment, based on the difference between the ground truth and the rendering prediction value, the initial three-dimensional Gaussian splash model is adjusted to obtain a three-dimensional Gaussian splash model corresponding to the target scene, including: determining the first difference corresponding to the pixel through the ground truth and the rendering prediction value corresponding to the same pixel; determining the difference between the ground truth and the rendering prediction value through the first difference corresponding to each pixel; based on the difference between the ground truth and the rendering prediction value, the initial three-dimensional Gaussian splash model is adjusted to obtain a three-dimensional Gaussian splash model corresponding to the target scene.
[0068] The ground truth is the real-world data acquired through sensors, reflecting the actual situation of the target scene.
[0069] The rendering prediction value is a scene image predicted based on the current 3D Gaussian splash model and rendering algorithm. It is a calculation result with a certain error.
[0070] For each pixel, by comparing the ground truth value of the pixel with the corresponding rendering prediction value, a first difference can be obtained. The first difference reflects the error at each pixel level, indicating the deviation between the pixel at this position and the actual scene. The first difference is the event rendering loss corresponding to the pixel. The calculation formula for the event rendering loss of each pixel is as follows:
[0071] in, represents the rendering prediction value of the pixel, Represents the ground truth of a pixel.
[0072] After calculating the first difference of each pixel, the differences of all pixels need to be summarized to obtain the global difference of the entire scene. The global difference can be obtained by weighted summing of the first differences of each pixel.
[0073] The initial three-dimensional Gaussian splash model is optimized based on the calculated global difference.
[0074] By using the embodiments of the present disclosure, the accuracy of the three-dimensional model can be significantly improved by calculating the difference of each pixel and performing global optimization, so that the error between the model and the actual scene is reduced. The model can be automatically adjusted according to the difference between the ground truth and the rendering prediction value to adapt to different scene characteristics. Since only the pixel difference needs to be calculated and globally aggregated each time, and then the three-dimensional model is adjusted by the optimization algorithm, the overall calculation process has a high efficiency and is suitable for real-time or near real-time three-dimensional modeling and rendering tasks.
[0075] Among them, in an optional embodiment, the ground truth of the target scene is determined based on the time surface information of each pixel, including: determining the training time period corresponding to the current iteration; determining the time surface information of each pixel at each moment contained in the training time period; integrating each of the time surface information of each pixel respectively, determining the brightness change of each pixel in the training time period, and generating each model supervision signal; determining the model supervision signal as the ground truth of the corresponding pixel in the training time period.
[0076] Temporal surface information can be used to capture the changing process of each pixel in the target scene. Temporal surface information refers to the characteristic performance of pixels at different time points, such as brightness, color, depth, etc. The characteristic performance can help the model recognize and understand the dynamic changes of the scene.
[0077] Set a time range for the entire optimization process so that the data within the training time period can be processed intensively. The data within the training time period will be used to capture and analyze the temporal characteristics of each pixel. For example, if a 1-minute event stream is collected, the complete time surface information will be one minute. You can select the time period from the 20th second to the 30th second as the training time period, and determine the various time surface information included in the training time period.
[0078] For each pixel in the target scene, the temporal surface information of the pixel during the entire training period is extracted. The temporal surface information can help the model recognize the dynamic evolution of the pixel.
[0079] The temporal surface information of each pixel is integrated to quantify the brightness change of the pixel during the entire training time period. Integration represents the accumulation of brightness changes over time, reflecting the overall brightness change of the pixel during this time period. The integration process can be a simple time accumulation or a weighted integration. The calculation of brightness change is used to quantify the brightness change trend of the pixel from the beginning to the end, such as the change from bright to dark, or from dark to bright.
[0080] Through integral calculation, the obtained model supervision signal is determined as the ground truth of the pixel in the training time period. That is, the actual performance or change of the pixel during the entire training time period. The ground truth refers to the real data performance of the scene, so the supervision signal provides labeled data for subsequent model training, helping the model to more accurately restore the real situation of the scene.
[0081] During the training period In the case of pixels The formula for determining the model supervision signal is as follows:
[0082] in, Represents pixels The corresponding event, Represents pixels Time surface information at each moment in the training period, represents the event threshold selected in this round of iteration, and dt represents the minimum time step for integrating the event supervision signal.
[0083] Among them, in an optional embodiment, the time surface information of each event of each pixel is determined according to the event stream, including: determining the target pixel and target timestamp corresponding to each event; determining each neighborhood pixel included in the neighborhood range of the target pixel; determining each neighborhood event of each neighborhood pixel; selecting a target neighborhood event that occurs before the target timestamp and is closest to the target timestamp from each neighborhood event; and applying the attenuation rate to determine the time surface information corresponding to the target pixel at the target timestamp through the neighborhood events and the events.
[0084] The time surface information is determined by the collected event stream. There is a one-to-one relationship between events and time surfaces. Each event in the event stream contains spatiotemporal information. By analyzing this information, the time surface information corresponding to each event can be determined. The time surface information reflects the impact of neighboring events on the event and can truly reflect the rapid changes in brightness and objects in the target scene.
[0085] For any event in the event stream, its time surface information can be determined in the following way.
[0086] Determine the target pixel and target timestamp corresponding to the event.
[0087] Determine the neighborhood range of the target pixel. The neighborhood range represents the pixel area around the target pixel. Each neighboring pixel in the neighborhood range provides information about changes in the surrounding environment. The neighborhood range can be a fixed pixel matrix range, such as 3x3, 5x5, etc., or a range that is dynamically adjusted according to specific applications.
[0088] For each neighborhood pixel, the event information related to these neighborhood pixels is filtered out from the event stream, which is also called neighborhood events. Neighborhood events refer to all events related to neighborhood pixels. Neighborhood events occur at different timestamps, reflecting the status of neighborhood pixels at different time points.
[0089] According to the principle of occurring before the target timestamp and being closest to the target timestamp, a target neighborhood event is selected from all neighborhood events. A neighborhood event located before the target timestamp is selected from all neighborhood events. The neighborhood event located before the target timestamp is an event that has a certain influence on the change of the target pixel at the target timestamp. By selecting the event closest to the target timestamp, it can be ensured that the target neighborhood event is the most relevant to the change of the target pixel at the target timestamp.
[0090] Target neighborhood events The timestamp can be determined by the following formula:
[0091] in, , , Indicated in The timestamp of the most recent event in the spatial neighborhood, represents the pixel in the u domain, and k represents the kth event.
[0092] The decay rate is applied to process the relationship between the selected target domain events and the target events. The decay rate refers to the process in which the impact of an event gradually weakens over time.
[0093] By combining the attenuated neighborhood events with the event information of the target pixel, the temporal surface information of the target pixel at the target timestamp is finally calculated.
[0094] Pixel Events The corresponding time surface can be determined by the following formula:
[0095] in, Provides a dynamic spatiotemporal context for events, exponential decay Extended past target neighborhood events to events and provides information on the history of activities in the neighborhood. Indicated in The timestamp of the most recent event in the spatial neighborhood.
[0096] By using the embodiments of the present disclosure, the time surface information of each pixel in the dynamic scene can be captured by using the timestamp and event stream. By considering the events of the neighboring pixels, the information of the surrounding environment of the target pixel can be integrated to more accurately reflect the changes of the target pixel. By introducing the decay rate, the impact of past events on the current state can be reasonably modeled to avoid unnecessary interference of outdated information on the modeling results.
[0097] By adopting the embodiments of the present disclosure, by analyzing the temporal surface information of each pixel during the training time period, the brightness changes in the time series data can be effectively captured, which can help the model understand the law of scene evolution over time. By integrating the brightness changes of each pixel and generating a model supervision signal, the model can be effectively guided to learn the temporal change characteristics, thereby improving the model's fitting accuracy and stability to the scene. By integrating the temporal surface information of each pixel and generating a supervision signal, the method can efficiently provide high-quality labels for model training.
[0098] Among them, in an optional embodiment, the training time period includes a first moment at the beginning of the training time period and a second moment at the end of the training time period; determining a rendering prediction value according to the initial three-dimensional Gaussian splash model, including: determining a first pose of the event camera at the first moment, and determining a second pose of the event camera at the second moment; based on the first pose, projecting the initial three-dimensional Gaussian splash model to a two-dimensional image to obtain a first image; based on the second pose, projecting the initial three-dimensional Gaussian splash model to a two-dimensional image to obtain a second image; determining a first rendering intensity of each pixel in the first image, and determining a second rendering intensity of each pixel in the second image; determining a rendering intensity difference of each pixel according to the first rendering intensity and the second rendering intensity corresponding to the same pixel; determining the rendering intensity difference as the rendering prediction value of the corresponding pixel.
[0099] The training time period includes two moments, the first moment and the second moment. The first moment is the beginning of the training time period, and the second moment is the end of the training time period. In the case of , the second moment is .
[0100] Pose refers to the position and orientation of the camera. Pose includes the position and orientation of the camera in three-dimensional space.
[0101] The first pose of the event camera at the first moment and the second pose of the event camera at the second moment are the poses of the time camera relative to the reference object when collecting the event stream.
[0102] The projection direction of the 3D Gaussian splash model can be determined by the position of the event camera at the target time. In the event segment, the event camera moves from the front of the target scene to the side of the target scene. At time , the event camera is in the position of collecting the front of the target scene. At this moment, the event camera is in a posture of collecting the side of the target scene. Then, when the initial three-dimensional Gaussian splash model is projected, the first posture should also correspond to the front of the initial three-dimensional Gaussian splash model, and the second posture should also correspond to the side of the initial three-dimensional Gaussian splash model.
[0103] According to the positions of the event camera at the first moment and the second moment, the three-dimensional Gaussian splash model is projected into the two-dimensional image to obtain two images at different time points, namely the first image and the second image.
[0104] Rendering intensity refers to the brightness or intensity of each pixel in the image. For each pixel, its rendering intensity at two moments is calculated respectively. Specifically, in the first image, the first rendering intensity of each pixel is determined, and the first rendering intensity reflects the brightness or intensity of the pixel at the first moment; in the second image, the second rendering intensity of each pixel is determined, and the second rendering intensity reflects the brightness or intensity of the pixel at the second moment.
[0105] The scales of the first image and the second image are consistent, and any pixel in the first image can find another pixel with the same position in the second image. Therefore, for each pixel, the difference in its rendering intensity in the first image and the second image is calculated.
[0106] The rendering prediction value can be determined according to the following formula:
[0107]
[0108]
[0109] in, represents the rendering prediction value, is the gamma correction value, and Indicates timestamp and The rendering intensity of the initial 3D Gaussian splatter model.
[0110] By adopting the embodiments of the present disclosure, by defining a specific training time period and performing projection at two moments, dynamic changes in the time series can be efficiently modeled. By determining the projection direction of the initial three-dimensional Gaussian splash model through the event camera pose, it can be ensured that the two-dimensional image generated at each moment can fully reflect the dynamic changes in the scene. By calculating the rendering intensity difference of the same pixel at two time points, factors such as illumination changes, object motion, and occlusion relationships in the scene can be reflected.
[0111] The initial three-dimensional Gaussian splash model can be projected onto a two-dimensional plane according to the steps provided in the following embodiment to obtain a two-dimensional image.
[0112] Among them, in an optional embodiment, the initial three-dimensional Gaussian splash model is projected onto a two-dimensional image, including: determining each three-dimensional Gaussian function included in the three-dimensional Gaussian splash model; projecting each of the three-dimensional Gaussian functions onto a two-dimensional plane, and determining the two-dimensional Gaussian corresponding to each of the three-dimensional Gaussian functions; the two-dimensional Gaussian represents the color of the three-dimensional Gaussian in the two-dimensional plane; for any pixel in the two-dimensional plane, determining each two-dimensional Gaussian corresponding to the pixel, sorting the projected two-dimensional Gaussian based on a tile-based rasterizer, and determining the color corresponding to the pixel through alpha blending.
[0113] The initial 3D Gaussian splash model is a model composed of multiple 3D Gaussian functions, which represent the density or weight of a point in a 3D space. Each Gaussian function represents a different part or feature in the model.
[0114] The 3D Gaussian function of any 3D point in the 3D point cloud can be expressed by the following formula:
[0115] in, represents the center of 3DGS; Indicates opacity; Represents the three-dimensional covariance matrix. The three-dimensional covariance matrix is as follows:
[0116] in, represents the rotation matrix, represents the scaling matrix, Represents the transpose of the scaling matrix.
[0117] Project each 3D Gaussian function from 3D space to 2D plane. Projecting to 2D plane means mapping Gaussian distribution in 3D coordinate system to 2D image plane. After projecting to 2D plane, 3D Gaussian function will become a 2D Gaussian function. Each 3D Gaussian function corresponds to a 2D Gaussian function, and 2D Gaussian function can be used to represent the color of the projected pixel.
[0118] After projecting all three-dimensional Gaussian functions onto a two-dimensional plane, the two-dimensional Gaussian function corresponding to each pixel is calculated according to the position of the pixel in the two-dimensional image.
[0119] The color of each pixel is determined by the combination of multiple two-dimensional Gaussian functions associated with the pixel position. This is because different three-dimensional Gaussian functions may affect multiple pixels when projected onto a two-dimensional image, so the color of each pixel needs to be calculated comprehensively based on the contributions of multiple two-dimensional Gaussians. Each two-dimensional Gaussian function has an impact on the color of an area in the image, and the range of the impact is determined by its standard deviation. The smaller the standard deviation, the smaller the affected area; the larger the standard deviation, the wider the affected area.
[0120] Rasterization refers to the process of converting graphic objects into pixels. For the calculation of two-dimensional Gaussian projection, rasterization technology is used to process each pixel in the image one by one, and the color of the pixel is determined according to its corresponding two-dimensional Gaussian function. Tile rasterizer is a technology to accelerate the rendering process. Tile rasterizer can divide the image into multiple small areas, namely tiles, and calculate each tile separately. Tiling technology helps to improve rendering efficiency.
[0121] In multi-layer rendering, alpha blending is used to combine the color information of different layers. For the result of two-dimensional Gaussian projection, there may be multiple two-dimensional Gaussian functions overlapping at a pixel position. Therefore, for a pixel, in order to get the final color of the pixel, the overlapping Gaussian functions need to be blended, and the result of the blending will generate the final color value.
[0122] The final color value can be determined by the following formula:
[0123] in, represents the number of sorted 2D Gaussians associated with the query pixel; represents the product of the opacity and density of the projected two-dimensional Gaussian distribution, Represents the calculated color, i represents the index of the i-th two-dimensional Gaussian distribution currently being processed, and j is the index used for cumulative multiplication operations, which represents the index of all Gaussian distribution items before the i-th item currently being processed.
[0124] The final color of each pixel in the two-dimensional image can be used to determine the rendering intensity of the pixel in the two-dimensional image.
[0125] With the embodiments of the present disclosure, each three-dimensional Gaussian function is mapped to a two-dimensional image plane, and a two-dimensional Gaussian distribution is generated by conversion, which can effectively generate a two-dimensional image from complex three-dimensional data. Through the rasterizer, the influence from different two-dimensional Gaussians can be calculated and synthesized for each pixel to generate a high-quality image. Different two-dimensional Gaussian functions are synthesized using alpha blending technology to ensure smooth transitions between multiple levels in the image and generate natural color changes.
[0126] Among them, in an optional embodiment, it also includes: determining the event threshold corresponding to this round of iteration from a reasonable range; the reasonable range is determined by the event threshold configured by the event camera; according to the event threshold, determining the event threshold loss corresponding to this round of iteration; the event threshold loss characterizes: the suitability of the event threshold selected in this round of iteration; based on the difference between the ground truth and the rendering prediction value, adjusting the initial three-dimensional Gaussian splash model to obtain a three-dimensional Gaussian splash model corresponding to the target scene, including: determining the event rendering loss according to the ground truth and the rendering prediction value; adjusting the initial three-dimensional Gaussian splash model according to the event rendering loss and the event threshold loss to obtain a three-dimensional Gaussian splash model corresponding to the target scene.
[0127] In the related art, it is usually assumed that the event trigger threshold is constant; however, the event trigger threshold actually varies between different perspectives, and setting it to a constant value may reduce the reconstruction quality. Therefore, the present disclosure proposes that the loss of the event threshold also needs to be considered.
[0128] In each iteration, an event threshold is re-determined. The event threshold is a simulated value, which is different from the real event threshold corresponding to the event camera. By simulating different event thresholds, a more accurate three-dimensional Gaussian splash model can be obtained. The event threshold selected in each iteration needs to be within a reasonable range. This is determined by the real event threshold configured in each event camera. The determination of the simulated event threshold cannot be separated from the real event threshold.
[0129] The event threshold determines which brightness changes of pixels are considered valid events. The event threshold is usually set based on the size of the brightness change. Only when the brightness change of a pixel exceeds a certain predetermined threshold will it be considered a valid event. For example, in a dynamic scene, if the brightness change of a pixel exceeds the set threshold, the event is recorded. On the contrary, if the brightness change is small, the event is ignored.
[0130] The event threshold loss reflects whether the event threshold selected in this iteration is appropriate. The event threshold loss can measure how the selected event threshold affects the validity of the event stream and indirectly affects the accuracy of the 3D model. For example, if the selected event threshold is too low, it may lead to too many redundant events, which may not represent important dynamic changes in the scene, and thus have a negative impact on the accuracy of the model. If the event threshold is set too high, some small but meaningful events may be missed, resulting in the model being unable to accurately capture the details in the scene.
[0131] The goal of the event threshold loss is to optimize the choice of threshold by evaluating the impact of the selected threshold on the quality of the 3D reconstructed model.
[0132] The ground truth and the rendering prediction value can be used as the event rendering loss.
[0133] Considering both event rendering loss and event threshold loss, the adaptive threshold is determined by the event rendering loss and event threshold loss. The formula of the adaptive threshold is as follows:
[0134] in, Indicates event rendering loss, represents the event threshold loss, and Represents the weight coefficient.
[0135] The initial three-dimensional Gaussian splash model is adjusted through an adaptive threshold to obtain a three-dimensional Gaussian splash model corresponding to the target scene.
[0136] By adopting the embodiments of the present disclosure, the quality of three-dimensional reconstruction can be improved by considering the impact of event threshold loss on three-dimensional reconstruction. By minimizing the event threshold loss, the threshold selection can be made more appropriate, thereby reducing redundant information and noise, making the three-dimensional reconstruction result more accurate. The optimization of event threshold loss and event rendering loss enables the algorithm to dynamically adjust the three-dimensional model to adapt to rapid changes in the scene. Compared with traditional inter-frame image processing methods, data processing based on event streams can better cope with rapid motion and sudden changes in illumination, and improve the update speed and accuracy of the model.
[0137] Among them, in an optional embodiment, according to the event threshold, the event threshold loss corresponding to this round of iteration is determined, including: determining a selection interval of the event threshold; the threshold interval includes an upper threshold limit and a lower threshold limit; according to a unit step function, based on the upper threshold limit, the lower threshold limit and the event threshold, determining the event threshold loss.
[0138] The upper and lower thresholds constitute a reasonable threshold interval, which is used to determine the rationality of the event threshold selected in each iteration.
[0139] The calculation of event threshold loss is achieved by introducing a unit step function. Unit step functions are usually used to represent "switch" or "threshold" type behaviors. If the selected event threshold is less than the lower threshold, it means that too many valid events are obtained, which increases the calculation burden. If the selected event threshold is greater than the upper threshold, too many valid events may be discarded, thereby affecting the accuracy of the 3D model, and the loss is also large.
[0140] Event Threshold Loss The calculation formula is as follows:
[0141] in, represents the unit step function. and Respectively represent the upper and lower limits of the event trigger threshold. Indicates the event threshold selected in this iteration.
[0142] In each round of iteration, the position of the event threshold is adjusted according to the current event threshold loss to minimize the loss function, thereby ensuring that it is more adapted to the current scenario.
[0143] By adopting the embodiments of the present disclosure, by introducing the calculation of event threshold loss, combining the selection interval of the upper threshold limit and the lower threshold limit and the unit step function, the model error caused by too high or too low event threshold can be avoided.
[0144] Based on the above-mentioned method for implementing 3D reconstruction based on 3D Gaussian splashing of event camera provided by the present disclosure, the method proposed by the present disclosure is introduced in detail through a specific embodiment, which specifically includes the following steps: Step 1: Build a dataset for training 3D reconstruction methods A commonly used evaluation dataset for event-based camera 3D reconstruction is used. The evaluation dataset includes a synthetic dataset and a real-world dataset. The algorithm is evaluated using the synthetic dataset and the real-world dataset.
[0145] The synthetic dataset includes seven objects with challenging features such as subtle structure and viewpoint-dependent color. The event stream of the scene corresponding to each object is generated by letting the event camera rotate 360 degrees around each 3D object and capture data from 1000 viewpoints.
[0146] The real-world dataset contains 10 event sequences. One event sequence represents one scene, and each sequence is recorded using a DAVIS 346C event camera under low light conditions, specifically a 5W light source, with the object rotating at 45 revolutions per minute. Based on the motion of the object, the camera extrinsics of the sequence can be calculated.
[0147] Step 2: Construct a rough 3D point cloud for each scene based on EventNeRF as the initialization of the event-based 3DGS.
[0148] For each scene, EventNeRF is trained for 30,000 iterations using the event stream corresponding to each scene to obtain a shallowly trained model. The event stream corresponding to the scene is input into the shallowly trained EventNeRF, and the rough depth of the scene is calculated by the shallowly trained EventNeRF model. The rough depth map is converted into a rough 3D point cloud by backprojection:
[0149] Among them, u and v are pixel coordinates; d is the depth value; K is the intrinsic parameter matrix; M is the extrinsic parameter matrix; represents the transpose of a matrix, Represents a 3D point cloud.
[0150] The temporal surface map encoding spatiotemporal information is obtained by accumulating event polarities and scaled to the range of [0, 255], where [0, 255] refers to the color range. Subsequently, the rough 3D point cloud is colored using the temporal surface map. Random downsampling is then used to reduce the density of the rough 3D point cloud with a downsampling rate of 10% to obtain the 3D point cloud used for 3DGS initialization.
[0151] Step 3: Determine the initial 3D Gaussian splash model based on the initial point cloud Based on the initialization point cloud, a set of anisotropic 3D Gaussians are used for explicit representation. Each 3D Gaussian is represented by the following formula for view-dependent appearance rendering:
[0152] in, represents the center of 3DGS; Indicates opacity; Represents the three-dimensional covariance matrix.
[0153] Step 4: Process the event stream based on the preprocessing method of the spatiotemporal surface graph to form a differentiable event supervision signal.
[0154] First, we extract the spatiotemporal context around the event and construct a temporal surface graph. Specifically, we extract the event The surrounding temporal context, denoted as :
[0155] in, , . Indicated in The timestamp of the most recent event in the spatial neighborhood, represents the pixel in the u domain, and k represents the kth event.
[0156] After obtaining the temporal context, the pixels The applied attenuation rate is The exponential decay kernel is used to obtain the time surface at time stamp t , the time surface Represents the event in the temporal context The pixels that affect The brightness change at timestamp t is as follows:
[0157] in, Provides a dynamic spatiotemporal background for events, exponential decay Extends the impact of past events and provides information about the history of activity in the neighborhood, Indicated in The timestamp of the most recent event in the spatial neighborhood.
[0158] Set time period and The brightness changes of the internal events are integrated by means of a time surface map as a differentiable supervision signal for training the 3DGS model:
[0159] in, Represents pixels The corresponding event, Represents pixels Time surface information at each moment in the training period, represents the event threshold selected in this round of iteration, and dt represents the minimum time step for integrating the event supervision signal.
[0160] Step 5: 3D Gaussian splatter model effectively renders the view through point splatter Sure Moment and At each moment, the event camera corresponds to the scanning posture. The scanning posture refers to the posture of the event camera when collecting the event stream.
[0161] According to the event camera Moment and The scanning pose corresponding to each moment is used to project the 3D Gaussian in the scene onto the 2D image plane to obtain each 2D Gaussian. The color of each 2D Gaussian is calculated by spherical harmonic parameters. Then, the 2D Gaussian is sorted by a tile-based rasterizer.
[0162] Based on the opacity and color of the two-dimensional Gaussian, the final color of each pixel is calculated by alpha blending as follows:
[0163] in, represents the number of sorted 2D Gaussians associated with the query pixel; represents the product of the opacity and density of the projected two-dimensional Gaussian distribution, Represents the calculated color, i represents the index of the i-th two-dimensional Gaussian distribution currently being processed, and j is the index used for cumulative multiplication operations, which represents the index of all Gaussian distribution items before the i-th item currently being processed.
[0164] In the case of projecting the 3D Gaussian in the scene onto the 2D image plane based on the scanning pose, due to Moment and The moments correspond to different scanning postures, so two-dimensional images corresponding to the two moments are obtained.
[0165] Determine the final colors of pixels at the same position in two two-dimensional images.
[0166] Step 6: Based on the event supervision signal, use the adaptive threshold loss to supervise the training of E-3DGS to generate accurate scene representation E-3DGS is a three-dimensional Gaussian splash model determined by the three-dimensional reconstruction method provided by the present disclosure.
[0167] The adaptive threshold includes event rendering loss and time threshold loss.
[0168] Determine the pixel through step 4 exist and The event polarity integrated brightness change between is as follows:
[0169] in, Represents pixels The corresponding event, Represents pixels Time surface information at each moment in the training period, represents the event threshold selected in this round of iteration, dt represents the minimum time step for integrating the event supervision signal, Indicates the training period.
[0170] Determine the pixel through step 5 exist and The grayscale difference between the 3DGS model rendering images. Specifically, for the timestamp and , using the formula provided in step 5 to determine the final color of the pixel, determine the pixel Rendering intensity on a 2D image corresponding to two timestamps and The logarithmic brightness images used for training are denoted as and , where g is the gamma correction value set to 2.2 in the experiment.
[0171] Pixel exist and The logarithmic brightness difference between is expressed as:
[0172] By minimizing the difference between the two and calculating the mean square error loss through MSE(), we get the event rendering loss:
[0173] in, represents the rendering prediction value of the pixel, Represents the ground truth of a pixel.
[0174] In each iteration, the event threshold is dynamically modeled and the event triggering threshold is limited to a reasonable range through the event threshold loss. It can be determined by the following formula:
[0175] in, represents the unit step function, and Respectively represent the upper and lower limits of the event triggering threshold. In the experiment, the upper threshold can be set to = 0.2, the lower threshold is = -0.2.
[0176] The event rendering loss and the event threshold loss jointly supervise the optimization of the 3DGS function. The weight coefficients of the event rendering loss and the event threshold loss are set to 0.9 and 0.1, respectively, to obtain an adaptive loss.
[0177]
[0178] The E-3DGS of this iteration is adjusted through adaptive loss.
[0179] Step 7: The model training was performed on the NVIDIA RTX 4090 GPU for 30,000 iterations, which took about 6 minutes to complete. In the experimental implementation, it was assumed that the pixel background value was known and set to 156, where 156 represents the color of the background. It took about 73 FPS to render a new perspective image with a resolution of 246*260 using the 3D Gaussian splatter model.
[0180] Compared with EventNeRF, which also performs dense reconstruction only from event streams, and E2VID-3DGS, which first uses E2VID to recover grayscale frames from event streams and then uses these frames to train the original frame-based 3DGS for 3D reconstruction, the 3D reconstruction method provided by the present disclosure, namely E-3DGS in Tables 1 and 2, achieves the most advanced level in rendering quality.
[0181] Table 1: Quantitative comparison of synthetic scenes
[0182] As shown in Table 1, compared with EventNeRF, the 3D reconstruction method provided by the present invention has an average PSNR (Peak Signal-to-Noise Ratio) improvement of 5.06 dB, a SSIM (Structural Similarity Index) improvement of 6.2%, and a LPIPS (Learned Perceptual Image Patch Similarity) reduction of 71% on the synthetic dataset.
[0183] Table 2: Quantitative comparison of real scenarios
[0184] As shown in Table 2, the 3D reconstruction method provided by the present disclosure improves the PSNR by 1.95 dB compared with EventNeRF on the real data set.
[0185] Furthermore, our method only requires 1% of the training time of EventNeRF.
[0186] Moreover, in real-world datasets, E2VID+3DGS performs significantly poorly, while the rendering quality of the 3D reconstruction method provided by the present disclosure is significantly better than that of E2VID+3DGS. This is because the present disclosure directly constructs a 3D Gaussian splash model through events.
[0187] Based on the same technical concept, the present disclosure provides a three-dimensional reconstruction device based on three-dimensional Gaussian splashing of an event camera. Figure 2 3D Gaussian splashing based on an event camera is a block diagram of a 3D reconstruction device according to an embodiment of the present disclosure. Figure 2 As shown, the three-dimensional reconstruction device comprises: An acquisition module 210 is used to acquire an event stream for a target scene; A model determination module 220, configured to determine an initial three-dimensional Gaussian splash model of a target scene according to the event stream; A time surface determination module 230 is used to determine the time surface information of each pixel according to the event stream; the time surface information represents the dynamic spatiotemporal background information of the event; A ground truth determination module 240 is used to determine the ground truth of the target scene based on the temporal surface information of each pixel; the ground truth represents the actual brightness change of the target scene; A prediction value determination module 250 is used to determine a rendering prediction value according to the initial three-dimensional Gaussian splash model, wherein the rendering prediction value represents a rendering intensity difference of the initial three-dimensional Gaussian splash model at two viewing angles; The adjustment module 260 is used to adjust the initial three-dimensional Gaussian splash model based on the difference between the ground truth and the rendering prediction value to obtain a three-dimensional Gaussian splash model corresponding to the target scene.
[0188] The present disclosure also provides an electronic device, referring to Figure 3 , Figure 3 is a schematic diagram of an electronic device shown in an embodiment of the present disclosure. Figure 3 As shown, the electronic device 300 includes: a memory 310 and a processor 320. The memory 310 and the processor 320 are connected via a bus communication. A computer program is stored in the memory 310. The computer program can be run on the processor 320 to implement the steps in the three-dimensional reconstruction method disclosed in the embodiment of the present disclosure.
[0189] The embodiment of the present disclosure further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the three-dimensional reconstruction method disclosed in the embodiment of the present disclosure are implemented.
[0190] The embodiment of the present disclosure further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps in the three-dimensional reconstruction method disclosed in the embodiment of the present disclosure are implemented.
[0191] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0192] It should be understood by those skilled in the art that the embodiments of the present disclosure may be provided as methods, devices or computer program products. Therefore, the embodiments of the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the embodiments of the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0193] The embodiments of the present disclosure are described with reference to the flowcharts and / or block diagrams of the methods, apparatuses, electronic devices, and computer program products according to the embodiments of the present disclosure. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0194] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0195] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0196] Although some embodiments of the present disclosure have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiment and all changes and modifications that fall within the scope of the present disclosure.
[0197] The above is a detailed introduction to the three-dimensional reconstruction method based on event camera three-dimensional Gaussian splashing provided by the present disclosure. This article uses specific examples to illustrate the principles and implementation methods of the present disclosure. The description of the above embodiments is only used to help understand the method of the present disclosure and its core idea. At the same time, for general technical personnel in this field, according to the ideas of the present disclosure, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as a limitation on the present disclosure.
Claims
1. A three-dimensional reconstruction method based on three-dimensional Gaussian splashing of an event camera, characterized in that: include: Get the event stream for the target scenario; Determining an initial three-dimensional Gaussian splash model of a target scene according to the event stream; Determine the time surface information of each pixel according to the event stream; the time surface information represents the dynamic spatiotemporal background information of the event; Determining a ground truth of a target scene based on temporal surface information of each pixel; the ground truth characterizing actual brightness changes of the target scene; Determining a rendering prediction value according to the initial three-dimensional Gaussian splash model, wherein the rendering prediction value represents a rendering intensity difference of the initial three-dimensional Gaussian splash model at two viewing angles; Based on the difference between the ground truth and the rendering prediction value, the initial three-dimensional Gaussian splash model is adjusted to obtain a three-dimensional Gaussian splash model corresponding to the target scene.
2. The method according to claim 1, characterized in that Also includes: Determine the event threshold corresponding to this round of iteration within a reasonable range; The reasonable range is determined by the event threshold configured by the event camera; According to the event threshold, determining the event threshold loss corresponding to the current iteration; The event threshold loss represents: the suitability of the event threshold selected in this round of iteration; Based on the difference between the ground truth and the rendering prediction value, the initial three-dimensional Gaussian splash model is adjusted to obtain a three-dimensional Gaussian splash model corresponding to the target scene, including: determining an event rendering loss based on the ground truth and the rendering prediction value; The initial three-dimensional Gaussian splash model is adjusted according to the event rendering loss and the event threshold loss to obtain a three-dimensional Gaussian splash model corresponding to the target scene.
3. The method according to claim 2, characterized in that According to the event threshold, the event threshold loss corresponding to the current iteration is determined, including: Determine a selection interval of the event threshold; the threshold interval includes an upper threshold limit and a lower threshold limit; The event threshold loss is determined based on the upper threshold, the lower threshold, and the event threshold according to a unit step function.
4. The method according to claim 1, characterized in that: Based on the difference between the ground truth and the rendering prediction value, the initial three-dimensional Gaussian splash model is adjusted to obtain a three-dimensional Gaussian splash model corresponding to the target scene, including: Determining a first difference corresponding to the pixel by using the ground truth and the rendering prediction value corresponding to the same pixel; Determine the difference between the ground truth and the rendering prediction value by using the first difference corresponding to each pixel; Based on the difference between the ground truth and the rendering prediction value, the initial three-dimensional Gaussian splash model is adjusted to obtain a three-dimensional Gaussian splash model corresponding to the target scene.
5. The method according to claim 1, characterized in that Based on the temporal surface information of each pixel, the ground truth of the target scene is determined, including: Determine the training time period corresponding to this round of iteration; Determining the temporal surface information of each pixel at each moment in the training time period; Integrate the temporal surface information of each pixel respectively, determine the brightness change of each pixel in the training time period, and generate each model supervision signal; The model supervision signal is determined as the ground truth of the corresponding pixel during the training time period.
6. The method according to claim 1, characterized in that Determining temporal surface information of each event of each pixel according to the event stream includes: Determine the target pixel and target timestamp corresponding to each event; Determine each neighborhood pixel included in the neighborhood range of the target pixel; Determine each neighborhood event of each neighborhood pixel; Select a target neighborhood event that occurs before a target timestamp and is closest to the target timestamp from each neighborhood event; The decay rate is applied to determine the temporal surface information of the target pixel corresponding to the target timestamp through the neighborhood events and the events.
7. The method according to claim 5, characterized in that The training time period includes a first moment at the beginning of the training time period and a second moment at the end of the training time period; determining a rendering prediction value according to the initial three-dimensional Gaussian splash model, including: Determine a first pose of the event camera at a first moment, and determine a second pose of the event camera at a second moment; Based on the first posture, projecting the initial three-dimensional Gaussian splash model onto a two-dimensional image to obtain a first image; Based on the second posture, projecting the initial three-dimensional Gaussian splash model onto a two-dimensional image to obtain a second image; Determine a first rendering intensity of each pixel in the first image, and determine a second rendering intensity of each pixel in the second image; Determine a rendering intensity difference of each pixel according to the first rendering intensity and the second rendering intensity corresponding to the same pixel; The rendering intensity difference is determined as the rendering prediction value of the corresponding pixel.
8. The method according to claim 7, characterized in that Projecting the initial three-dimensional Gaussian splash model onto a two-dimensional image includes: Determine each three-dimensional Gaussian function included in the three-dimensional Gaussian splash model; Projecting each of the three-dimensional Gaussian functions onto a two-dimensional plane, and determining a two-dimensional Gaussian corresponding to each of the three-dimensional Gaussian functions; the two-dimensional Gaussian represents the color of the three-dimensional Gaussian on the two-dimensional plane; For any pixel in a two-dimensional plane, each two-dimensional Gaussian corresponding to the pixel is determined, and the projected two-dimensional Gaussians are sorted based on a tile-based rasterizer, and the color corresponding to the pixel is determined by alpha blending.
9. The method according to claim 1, characterized in that: Determining an initial three-dimensional Gaussian splash model of a target scene according to the event stream includes: Performing shallow training on the initial target model based on the event stream to obtain a target model; Inputting the event stream into the target model, and recovering depth information of each pixel from the event stream through the target model to obtain a rough depth map; Converting the rough depth map into a three-dimensional point cloud by back-projection; The three-dimensional point cloud is processed by a time surface graph composed of the time surface information to obtain an initial three-dimensional Gaussian splash model; the processing at least includes: coloring the three-dimensional point cloud by the time surface graph.
10. The method according to claim 1, characterized in that Get the event stream for the target scenario, including: Determine the brightness change value of each pixel at a preset time interval; When the brightness change value exceeds the event threshold corresponding to the event camera, an event stream is generated; each event in the event stream includes pixel coordinates, a timestamp and an event polarity.
Citation Information
Patent Citations
Nerve radiation field artifact suppression method based on event camera three-dimensional reconstruction
CN118887334A
Dynamic real-time rendering method for large assembly scene based on three-dimensional Gaussian splashing
CN119229031A
High-fidelity large-scale scene rendering method, system and equipment based on Gaussian splashing and medium
CN119295621A
Systems and methods for generating splat-based differentiable two-dimensional renderings
WO2022046113A1
Cited By
Gaussian splash dynamic three-dimensional reconstruction method and system based on spiking neurons
CN121010707A
GPU-accelerated 3D Gaussian splash three-dimensional coordinate high-precision real-time pickup method and system
CN121095466A
Three-dimensional scene reconstruction method and device, equipment, storage medium and program product
CN121280668A
Video viewpoint prediction method based on event-driven time domain modeling and multiple scales
CN121963057A
Event-driven time-domain modeling and multi-scale based video viewpoint prediction method
CN121963057B