Driving scene physical simulation method based on neural rendering
By using neural rendering technology and hybrid rendering pipelines to generate high-reality scenes in driving simulation scenarios, and introducing collision distances to simulate physical collisions, the problems of poor rendering reality and lack of collision sensing in the existing technology are solved, and more efficient driving simulation and learning of assisted driving systems are achieved.
Patent Information
- Application Number
- CN202510088878.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-01-21
AI Technical Summary
The prior art renders poorly realistic and lacks collision sensing in driving simulation scenarios, resulting in a degradation of the perceived performance of the model in real scenarios.
Using a physical simulation method of driving scenes based on neural rendering, a driving scene rendering network is constructed, including a motion structure recovery module and a hybrid rendering pipeline, a neural radiation field pipeline and a Gaussian splashing pipeline are used to generate a high-reality driving scene, and a collision distance is introduced to simulate physical collisions.
It realizes the generation of high-reality scenarios in the driving simulation environment, provides interactive collision sensing, helps the assisted driving system learn driving rules, improves robustness and solves the problem of scarcity of real environment data.
Smart Images

Figure CN120088384A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a neural rendering method in the field of driving scenario simulation, and particularly relates to a physical simulation method of a driving scenario based on neural rendering. Background Art
[0002] Simulation technology is a key component in the field of autonomous driving, which can significantly improve the iteration efficiency of the system, reduce the cost pressure of real vehicle testing, and shorten the time for business implementation.
[0003] Although data generated by common simulators such as CARLA, Baidu Apollo, and Tencent TAD Sim provide a test environment for driving assistance systems, the data generated by these traditional rendering pipelines have significant differences from real-world scenarios, especially in terms of texture and details, and cannot fully meet the requirements of real vehicle testing. Training a perception model or conducting vehicle simulation using such low-fidelity scenario data may lead to overfitting of the model to the simulation environment, resulting in a decline in perception performance in the real scenario.
[0004] Technologies based on neural radiance fields or Gaussian splatting methods have achieved certain results in multi-view synthesis of realistic scenes, but their applications are mainly concentrated in multi-view synthesis, unable to provide sensor collision and physical simulation at the driving scenario level, and it is also difficult to construct an interactive exploration environment for the assisted driving system to learn driving rules.
[0005] One of the current prior arts is the method of Alexey et al. in the paper "CARLA: An Open Urban Driving Simulator". Although the CARLA platform uses a physical rendering engine to achieve the creation of data assets and the modeling of physical scenes, due to the limitations of the rendering engine, there are still significant differences between the 3D simulation scenes generated in real time in this way and the real world, making it difficult to meet the actual requirements. Training a perception model or conducting vehicle simulation using such low-fidelity scenario data may lead to overfitting of the model to the simulation environment, resulting in a decline in perception performance in the real environment.
[0006] The second current prior art is the method of Matthew et al. in the paper "Block-NeRF: Scalable Large Scene Neural View Synthesis". Block-NeRF is an extension of neural radiance fields, aiming to represent large-scale environments. This method divides the urban scene into multiple modules and independently assigns neural radiance fields to each module for rendering, and then dynamically combines these radiance fields during prediction. This decomposition process effectively decouples the rendering time from the scene size, enabling rendering to scale to any size and supporting piecemeal updates of the environment. However, this method cannot effectively predict the position of geometric information, and the rendered results are blurred or there are floating objects in the air, unable to provide a realistic driving scene.
[0007] The third current prior art is the method of Zhou et al. in the paper "DrivingGaussian: Composite Gaussian Splatting for Surrounding Dynamic Autonomous Driving Scenes". In the face of the existence of moving objects in complex scenes, this method first adopts an incremental static Gaussian model to sequentially and progressively simulate the static background of the entire scene. Then, it uses a composite dynamic Gaussian map to process multiple moving objects, reconstructs them separately, and accurately restores their positions and occlusion relationships in the scene. Although this method provides high-fidelity and multi-camera consistent realistic surround view synthesis. However, this method only considers synthesizing realistic views and does not consider interacting with the environment, lacking a collision sensing design, which is not conducive to the auxiliary driving system learning rules. Summary of the Invention
[0008] In order to effectively solve the problems of poor rendering realism and lack of collision sensing in existing methods in driving simulation scenarios, the present invention proposes a physical simulation method for driving scenes based on neural rendering. The present invention can render a driving scene with high realism in real time, reducing the gap between the simulation system and the real world. And a collision distance is introduced in the present invention to simulate collisions in the physical world and help the auxiliary driving system learn driving rules.
[0009] The technical solution adopted by the present invention is as follows:
[0010] 1. A physical simulation method for driving scenes based on neural rendering
[0011] S1: Construct a driving scene rendering network, the driving scene rendering network includes a connected motion structure recovery module and a hybrid rendering pipeline, and the hybrid rendering pipeline includes a connected neural radiance field pipeline and a Gaussian splatting pipeline;
[0012] S2: The sequence images containing the same driving simulation scenario are used to form the training dataset corresponding to each driving scenario. After training the hybrid rendering pipeline in the driving scenario rendering network using the training dataset of the current driving scenario, the neural radiance field pipeline outputs the depth map of the current driving scenario and obtains the trained Gaussian splatter pipeline;
[0013] S3: After processing the depth map of the current driving scenario using the truncated sigmoid function, an interactive collision grid space is obtained;
[0014] S4: The target pose is input into the trained Gaussian splatter pipeline to obtain the scene color observed at the target pose. The reinforcement learning-based agent explores and provides feedback in the interactive collision grid space by combining the scene color observed at the target pose.
[0015] The described method for physical simulation of driving scenarios based on neural rendering further includes the following steps:
[0016] S5: Obtain the sequence images of different driving simulation scenarios, repeat S2 - S4, and complete the simulation exploration in different driving simulation scenarios.
[0017] In the above S1, the motion structure recovery module of the driving scenario rendering network is used to extract the pose of the input image. The pose of the input image and the position and orientation in the scene three-dimensional space are used as the input of the neural radiance field pipeline. The neural radiance field pipeline outputs the depth map, and the Gaussian splatter pipeline outputs the scene color.
[0018] The motion structure recovery module includes a feature extraction module, a feature matching module, and a pose estimation module connected in sequence.
[0019] In the above S2, training the hybrid rendering pipeline in the driving scenario rendering network using the training dataset of the current driving scenario specifically includes:
[0020] The pose estimated by the structure from motion recovery module for each input image serves as the input pose for the neural radiance field pipeline. The neural radiance field pipeline predicts the color and density of points in the three-dimensional space of the current scene based on the input pose and the position and direction in the three-dimensional space of the current scene, and outputs the voxel color rendering result, depth map, and scene geometry information. The Gaussian splash pipeline learns the Gaussian sphere attributes based on the scene geometry information, thereby outputting the scene color and the optimized scene geometry information. The position in the three-dimensional space of the current scene is updated using the optimized scene geometry information. The network parameters in the neural radiance field pipeline are optimized and adjusted according to the first loss function calculated from the voxel color rendering result and the current input image, and the parameters in the Gaussian splash pipeline are optimized and adjusted according to the second loss function calculated from the scene color and the current input image, completing one training session. Then, the estimated poses corresponding to other input images in the current training dataset are successively input into the neural radiance field pipeline, and the neural radiance field pipeline and the Gaussian splash pipeline are trained multiple times while updating the three-dimensional space of the scene, thereby obtaining the trained neural radiance field pipeline and Gaussian splash pipeline. The trained neural radiance field pipeline outputs the depth map.
[0021] The neural radiance field includes a color estimation module and a geometry estimation module. After spherical harmonic encoding of the direction in the three-dimensional space of the scene, it is input into the color estimation module, and the color estimation module outputs the voxel color rendering result. And after hash encoding of the position in the three-dimensional space of the scene, it is input into the geometry estimation module, and the geometry estimation module outputs the density information. Then, the density information is converted into a depth map and the depth map is subjected to point densification processing to obtain the scene geometry information.
[0022] In the neural radiance field pipeline, there is also:
[0023] The depth map output by the geometry estimation module is subjected to smooth regularization processing to obtain the processed depth map. After point densification processing of the processed depth map, the scene geometry information is obtained and used as the input for the Gaussian splash pipeline.
[0024] In the agent based on reinforcement learning, the final reward function of the agent satisfies the following formula:
[0025]
[0026] r basic (s t ,a t )=0.1*(lap(s t )-u)
[0027]
[0028] where r(s t ,a t) represents the final reward function value, r basic (s t , a t ) represents the basic reward function value, s t represents the state of the agent interacting and exploring with the environment at time t, a t represents the action of the agent, f(a, b) represents the collision distance, lap(s t ) represents using the Laplace operator to estimate the quality of the rendered image, u represents the threshold for whether the observed scene color is clear; x a , y a , z a are respectively the central coordinate values of the ego vehicle in the world coordinate system, x b , y b , z b are respectively the central coordinate values of the collided object mesh in the world coordinate system, a r is the radius of the ego vehicle, b r is the radius of the collided object mesh.
[0029] The color estimation module consists of 2 connected fully-connected layers; the geometric estimation module consists of 1 connected fully-connected layer.
[0030] The depth map is smoothed and regularized according to the following formula:
[0031]
[0032] where, represents the loss function of the depth map, θ represents the learning parameter, r ij represents the ray passing through the pixel (i, j), r i+1j+1 represents the ray passing through the pixel (i + 1, j + 1), r i+1j represents the ray passing through the pixel (i + 1, j), r ij+1 represents the ray passing through the pixel (i, j + 1), r ij represents the ray passing through the pixel (i, j), S patch represents the size of the rendered image block, and this image block is the area centered on the ray r; represents the set of rays sampled from the camera pose, and the purpose of smoothing regularization is to minimize the depth difference between adjacent pixels in a specific image sub-block.
[0033] In the above S3, the depth map of the current driving scene is processed using the following formula:
[0034]
[0035] tsdf(x) = max[-1, min(1, sdf(x) / λ)]
[0036] Among them, represents the depth map of the current driving scenario, cam z (x) represents the depth of voxel x relative to the camera, λ represents the threshold of the depth difference between voxel x and the object cross-section depth, and tsdf(x) represents the operation of the truncated sign function on voxel x.
[0037] In the reinforcement learning-based agent, the state s of the agent interacting and exploring with the environment at time t t satisfies the following formula:
[0038] s t =(C t , h t )
[0039] Among them, h t represents the color of the historical scene observed at time t, and C t represents the color of the scene observed by the agent at time t.
[0040] The action space of the agent is the driving behavior in the world coordinate system, specifically including four displacement actions and two rotation actions.
[0041] II. A computer device
[0042] The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of the method for physical simulation of a driving scenario based on neural rendering.
[0043] III. A computer-readable storage medium
[0044] The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the steps of the method for physical simulation of a driving scenario based on neural rendering.
[0045] IV. A computer program product
[0046] The computer program product includes a computer program / instructions, and when the computer program / instructions are executed by the processor, they implement the steps of the method for physical simulation of a driving scenario based on neural rendering.
[0047] The beneficial effects of the present invention are:
[0048] By adopting the above technical solutions, the method proposed by the present invention can synthesize a highly realistic driving simulation environment through a driving scenario rendering network and provide a rendered scene with clear texture details.
[0049] The proposed hybrid renderer in the present invention optimizes and learns the scene depth through mixing, and then constructs an interactive and collidable grid space, which can solve the problem of lack of physical sensing in the current rendering methods.
[0050] Through reinforcement learning technology, the assisted driving system can explore and collect a large amount of data in this simulation environment through trial and error, learn to simulate perception and decision-making in the real environment, and does not rely on traditional rule-based methods. The present invention can improve the problem of unrealistic texture details in existing driving simulation software on the basis of ensuring real-time interaction performance, and narrow the gap between simulation and the real world. At the same time, it provides the collision distance and learns the driving rules of a highly realistic simulation environment; it not only improves the robustness of the assisted driving system in complex driving environments, but also solves the problem of scarce real environment data. In addition, the real-time interaction performance can provide real-time feedback information for the agent to learn and correct errors. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 is the overall flowchart of a physical simulation method for a driving scene based on neural rendering according to an embodiment of the present invention.
[0052] Figure 2 is the framework of the physical simulation method for a driving scene based on neural rendering according to an embodiment of the present invention.
[0053] Figure 3 is the flowchart of the neural radiance field in the hybrid renderer according to an embodiment of the present invention.
[0054] Figure 4 is the comparison result when performing driving scene simulation according to an embodiment of the present invention, compared with a traditional driving simulation engine.
[0055] Figure 5 is a highly realistic visualization result diagram of the hybrid renderer rendering at different angles when performing driving scene simulation according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0056] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.
[0057] The present invention proposes a physical simulation method for a driving scene based on neural rendering, as Figure 1 shown, the method includes the following steps:
[0058] S1: Construct a driving scene rendering network, as Figure 2 shown, the driving scene rendering network includes a connected motion structure recovery module and a hybrid rendering pipeline; the hybrid rendering pipeline includes a connected neural radiance field pipeline and a Gaussian splash pipeline;
[0059] The motion structure recovery module of the driving scene rendering network is used to extract the pose of the input image. The pose of the input image and the position and orientation in the three-dimensional space of the scene are used as the input of the neural radiance field pipeline. The neural radiance field pipeline outputs a depth map, and the Gaussian splash pipeline outputs the scene color.
[0060] The motion structure recovery module includes a feature extraction module, a feature matching module, and a pose estimation module connected in sequence. The input image is input into the feature extraction module, and finally the pose estimation module outputs the initial pose corresponding to the input image. The pose represents the position and orientation of the object in the coordinate system.
[0061] S2: The sequence images containing the same driving simulation scene are used to form the training dataset corresponding to each driving scene. After end-to-end training of the hybrid rendering pipeline in the driving scene rendering network using the training dataset of the current driving scene, the neural radiance field pipeline outputs the depth map of the current driving scene and obtains a trained Gaussian splash pipeline;
[0062] In this embodiment, three general datasets, namely Waymo, Argoverse, and Kitti, are selected. These datasets contain images collected under good vision and include diverse driving scenes such as urban streets, alleys, and highways. The collected datasets are divided into multi-camera scenes and single-camera scenes according to the number of cameras. Then, the datasets are divided into a training set and a test set according to the number of frames. Every 10 frames are used as the test set, and the rest are used as the training set.
[0063] In S2, end-to-end training of the hybrid rendering pipeline in the driving scene rendering network using the training dataset of the current driving scene specifically includes:
[0064] The pose estimated by the structure from motion module for each input image serves as the input pose for the neural radiance field pipeline. The neural radiance field pipeline predicts the color and density of points in the three-dimensional space of the current scene based on the input pose and the position and direction in the three-dimensional space of the current scene, and outputs the voxel color rendering result, depth map, and scene geometry information (calculated from the density). The first position and direction in the three-dimensional space of the scene are initialized and generated. Among them, the position is the sampling point position passed by the ray caused by each pixel in the corresponding input image starting from the camera origin, and the direction refers to the direction of the ray. The Gaussian splash pipeline learns the Gaussian sphere attributes such as the shape, size, color, etc. of the Gaussian sphere based on the scene geometry information, so as to output the scene color and the optimized scene geometry information. The method for optimizing the scene geometry information in the Gaussian splash pipeline is to use differentiable point cloud splitting and cloning techniques. The optimized scene geometry information is used to update the position in the three-dimensional space of the current scene, thereby improving the accuracy of the reconstructed geometric space. The network parameters in the neural radiance field pipeline are optimized and adjusted according to the first loss function calculated from the voxel color rendering result and the current input image, and the parameters in the Gaussian splash pipeline are optimized and adjusted according to the second loss function calculated from the scene color and the current input image to complete one training. The specific training process of the Gaussian splash pipeline satisfies the following formula, c i represents the color of the Gaussian sphere, and α i is the opacity of Gaussian sphere i, and α j is the opacity of Gaussian sphere j, and N is the number of Gaussian spheres, and C t represents the output scene color; then, the estimated poses corresponding to other input images in the current training dataset are sequentially input into the neural radiance field pipeline, and the neural radiance field pipeline and the Gaussian splash pipeline are trained multiple times while updating the three-dimensional space of the scene, so as to obtain the trained neural radiance field pipeline and Gaussian splash pipeline. The trained neural radiance field pipeline outputs a depth map, which is used to calculate the subsequent collision grid space; the Gaussian splash pipeline outputs the color of the scene, which is used as the observation of the subsequent agent. Among them, the position in the updated three-dimensional space of the scene is hash-coded to obtain a new depth map, and then the new scene geometry information x is obtained by using the point densification method. The new scene geometry information x contains the new position and is used to fill the problematic area.
[0065] such as Figure 3As shown, the neural radiance field includes a color estimation module and a geometry estimation module. The directions in the three-dimensional space of the scene are spherically harmonically encoded and then input into the color estimation module, which outputs the voxel color rendering result. And the positions in the three-dimensional space of the scene are hash-encoded and then input into the geometry estimation module, which outputs the density information. Then the density information is converted into a depth map and then the depth map is point-densified to obtain the scene geometry information. Among them,
[0066] c θ , σ θ = f θ (γ(x), γ(d))
[0067] Among them, c θ is the color, σ θ is the density, f θ is the neural radiance field with parameters, x is the position, d is the direction, and γ represents the predefined spherical harmonic encoding or position encoding. Given the neural radiance field, the pixels are rendered by casting rays r(u) = o + ud along the direction d from the camera center o. The rendered pixels are used to represent the simulated driving scene, and at the same time, a depth map still needs to be learned to represent the grid space.
[0068] The depth map satisfies the following formula:
[0069]
[0070] Among them, represents the depth value corresponding to the ray r, u n and u f are the predefined near plane and far plane for rendering respectively, Q(u) represents the cumulative transmittance along the ray, and σ θ represents the opacity estimated by the neural radiance field; r(u) represents the ray of the sampling point, and u represents the position of the sampling point.
[0071] Specifically, the following formula is used to perform point densification on the depth map:
[0072]
[0073] Among them, x represents the scene geometry information, o is the ray origin, and d r is the ray direction.
[0074] The color estimation module consists of two connected fully-connected layers with 64 channels.
[0075] The geometry estimation module consists of one connected fully-connected layer with 64 channels.
[0076] In the neural radiance field pipeline, there is also:
[0077] The depth map output by the geometric estimation module is smoothed and regularized to obtain a processed depth map. After point densification processing on the processed depth map, scene geometry information is obtained and used as the input to the Gaussian splatter pipeline.
[0078] The optimization of driving scene depth consists of multi-view Figure 1 consistency and depth map smoothing regularization:
[0079] For the depth map obtained by hybrid optimization through the neural radiance field pipeline and the Gaussian splatter pipeline, to ensure 3D consistency, through multi-view Figure 1 consistency regularization: For each pose, two camera depth maps are selected from adjacent viewpoints, and the key points of the current viewpoint are transformed into the camera coordinate system. These points are projected back to the current view to measure the per-pixel difference.
[0080] The depth map is smoothed and regularized according to the following formula:
[0081]
[0082] where, represents the loss function of the depth map, θ represents the learned parameter, r ij represents the ray passing through pixel (i, j), r ij represents the ray passing through pixel (i, j), r i+1j+1 represents the ray passing through pixel (i + 1, j + 1), r i+1j represents the ray passing through pixel (i + 1, j), r ij+1 represents the ray passing through pixel (i, j + 1), r ij represents the ray passing through pixel (i, j), S patch represents the size of the rendered image patch, which is the region centered on ray r; represents the set of rays sampled from the camera pose. The purpose of smoothing regularization is to minimize the depth difference between adjacent pixels in a specific image sub-patch.
[0083] S3: After processing the depth map of the current driving scene using the truncated sign function, an interactive collision grid space is obtained for simulating physical collisions between vehicles or interactions between vehicles and the environment;
[0084] Specifically, the depth map of the current driving scene is processed using the following formula:
[0085]
[0086] tsdf(x) = max[-1, min(1, sdf(x) / λ)]
[0087] where, Depth map representing the current driving scenario, cam z (x) represents the depth of voxel x relative to the camera, λ represents the threshold of the depth difference between voxel x and the object cross-section depth, and tsdf(x) represents the operation of the truncated sign function on voxel x.
[0088] If tsdf(x) is close to 0, it means the voxel is close to the surface. Conversely, when tsdf(x) is close to 1 or -1, it means the voxel is far from the surface.
[0089] S4: Use an agent based on reinforcement learning to explore in the interactive collision grid space.
[0090] In the agent based on reinforcement learning, the state s of the agent's interaction and exploration with the environment at time t t satisfies the following formula:
[0091] s t =(C t , h t )
[0092] where h t represents the color of the historical scene observed at time t, and C t represents the color of the scene observed by the agent at time t.
[0093] As Figure 4 shown, the Gaussian splash pipeline in the hybrid rendering pipeline outputs the color of the scene, combining the historical image frames as the observation of the agent. Compared with traditional simulators, the rendering result of the present invention is more realistic in texture, reducing the gap with the real world. In addition, the hybrid rendering pipeline proposed by the present invention only requires the driving data of the forward view to render different perspective renderings for the agent to observe, as Figure 5 shown. Compared with the existing methods, it avoids the collection of a large amount of repeated data.
[0094] The action space of the agent is the driving behavior in the world coordinate system, specifically including four displacement actions and two rotation actions. The four displacement actions are forward, backward, left, and right; the two rotation actions are rotating 10° clockwise and rotating 10° counterclockwise.
[0095] The agent learns the value Q(s t , a t ) of the action through multiple explorations, selects the action value that maximizes in the current state, and thus learns the driving rules in the real world. Specifically, given the initial state s t , the agent determines the behavior a * = argmaxQ(s t , a t ) with the highest value in the current environment through the value function, and the agent executes the optimal policy a* After that, it transfers to the next state s t+1 , through a continuous sequence decision-making process (a t , a t+1 , a t+2 , a t+3 , a t+4 , a t+4… ), the agent finally learns the optimal driving strategy.
[0096] The final reward function of the agent satisfies the following formula:
[0097]
[0098] r basic (s t , a t ) = 0.1 * (lap(s t ) - u)
[0099]
[0100] where r(s t , a t ) represents the final reward function value, r basic (s t , a t ) represents the basic reward function value, s t represents the state of the interaction and exploration between the agent and the environment at time t, a t represents the action of the agent, f(a, b) represents the collision distance, lap(s t ) represents estimating the quality of the rendered image using the Laplace operator, u represents the threshold for whether the observed scene color is clear, set to 100; x a , y a , z a are the central coordinate values of the ego vehicle in the world coordinate system respectively, x b , y b , z b are the central coordinate values of the object grid of the collision in the world coordinate system respectively, a r is the radius of the ego vehicle, b r is the radius of the object grid of the collision.
[0101] The interaction and exploration between the agent and the environment are mainly based on the reinforcement learning technology of the value function. By interacting with the collision grid space to obtain the positions of objects such as vehicles and houses, it is judged whether the actions executed by the agent obtain positive or negative feedback; thus learning the driving rules of the assisted driving system in complex environments and enhancing the robustness in complex driving environments.
[0102] S5: Obtain the sequence images of different driving simulation scenarios, repeat S2 - S4, and complete the simulation exploration in different driving simulation scenarios.
[0103] The method proposed by the present invention can improve the problem of untrue texture details in existing driving simulation software on the basis of ensuring real - time interaction performance, and narrow the gap between simulation and the real world. In addition, the real - time interaction performance can provide real - time feedback information for the agent to learn and correct errors.
[0104] The present invention also proposes a computer device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, it implements the steps of a physical simulation method for driving scenarios based on neural rendering.
[0105] The present invention also proposes a computer - readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of a physical simulation method for driving scenarios based on neural rendering.
[0106] The present invention also proposes a computer program product, which includes computer programs / instructions. When the computer programs / instructions are executed by a processor, they implement the steps of a physical simulation method for driving scenarios based on neural rendering.
[0107] A physical simulation method for driving scenarios based on neural rendering proposed by an embodiment of the present invention can, through the design of a hybrid neural renderer, render a driving scenario with high authenticity in real time. The present invention optimizes the depth map through multi - view Figure 1 consistency regularization and smooth regularization, and then generates a high - quality mesh space, provides accurate collision sensing for the agent, and thus learns the rules based on the driving simulation scenario. The present invention will be more conducive to the implementation of the driving simulation method in the field of autonomous driving.
[0108] Finally, it should be noted that the above - mentioned embodiments and descriptions are only used to illustrate the technical solutions of the present invention and not to limit them. Those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced. Without departing from the spirit and scope of the disclosure of the technical solutions of the present invention, they should all be covered within the protection scope of the claims of the present invention.
Claims
1. A driving scene physical simulation method based on neural rendering, characterized in that: The following steps are involved: S1: constructing a driving scene rendering network, the driving scene rendering network comprising a connected motion structure recovery module and a hybrid rendering pipeline, the hybrid rendering pipeline comprising a connected neural radiation field pipeline and a Gaussian splash pipeline; S2: A training data set corresponding to each driving scene is formed by the acquired sequence images containing the same driving simulation scene. After the hybrid rendering pipeline in the driving scene rendering network is trained using the training data set of the current driving scene, the neural radiance field pipeline outputs the depth map of the current driving scene and obtains the trained Gaussian splash pipeline; S3: After processing the depth map of the current driving scene using the truncated sign function, the interactive collision grid space is obtained; S4: The target pose is input into the trained Gaussian splash pipeline to obtain the scene color observed at the target pose. The reinforcement learning-based agent combines the scene color observed at the target pose to perform exploration feedback in the interactive collision grid space.
2. The method for physical simulation of driving scenes based on neural rendering according to claim 1, characterized in that: The method further comprises the following steps: S5: Obtain sequence images of different driving simulation scenes, repeat S2-S4, and complete simulation exploration in different driving simulation scenes.
3. The method for physical simulation of driving scenes based on neural rendering according to claim 1, characterized in that: In S1, the motion structure recovery module of the driving scene rendering network is used to extract the pose of the input image, and the pose of the input image and the position and direction in the three-dimensional space of the scene are used as inputs of the neural radiation field pipeline, the neural radiation field pipeline outputs a depth map, and the Gaussian splash pipeline outputs the scene color.
4. The method for physical simulation of driving scenes based on neural rendering according to claim 1, characterized in that: The motion structure recovery module includes a feature extraction module, a feature matching module and a posture estimation module which are connected in sequence.
5. The method for physical simulation of driving scenes based on neural rendering according to claim 1, characterized in that: In S2, the hybrid rendering pipeline in the driving scene rendering network is trained using the training data set of the current driving scene, specifically including: The pose of each input image estimated by the motion structure recovery module is used as the input pose of the neural radiation field pipeline. The neural radiation field pipeline predicts the color and density of the midpoint in the three-dimensional space of the current scene based on the input pose and the position and direction in the three-dimensional space of the current scene, and outputs the voxel color rendering result, the depth map and the scene geometry information. The Gaussian splash pipeline learns the Gaussian sphere properties based on the scene geometry information, thereby outputting the scene color and the optimized scene geometry information. The optimized scene geometry information is used to update the position in the three-dimensional space of the current scene. The network parameters in the neural radiation field pipeline are optimized and adjusted according to the first loss function calculated by the voxel color rendering result and the current input image, and the parameters in the Gaussian splash pipeline are optimized and adjusted according to the second loss function calculated by the scene color and the current input image, and a training is completed; then, the estimated poses corresponding to other input images in the current training data set are input into the neural radiation field pipeline in turn, and the neural radiation field pipeline and the Gaussian splash pipeline are trained multiple times, and the three-dimensional space of the scene is updated at the same time, so as to obtain the trained neural radiation field pipeline and Gaussian splash pipeline, and the trained neural radiation field pipeline outputs a depth map.
6. The method for physical simulation of driving scenes based on neural rendering according to claim 5, characterized in that: The neural radiation field includes a color estimation module and a geometry estimation module. The direction in the three-dimensional space of the scene is spherically harmonic encoded and then input into the color estimation module. The color estimation module outputs the voxel color rendering result. The position in the three-dimensional space of the scene is hash-encoded and then input into the geometry estimation module. The geometry estimation module outputs density information, which is then converted into a depth map and then the depth map is densified to obtain the scene geometry information.
7. The method for physical simulation of driving scenes based on neural rendering according to claim 5, characterized in that: The neural radiation field pipeline further includes: The depth map output by the geometry estimation module is smoothed and regularized to obtain a processed depth map. After point densification, the processed depth map is subjected to scene geometry information and used as the input of the Gaussian splash pipeline.
8. The method for physical simulation of driving scenes based on neural rendering according to claim 1, characterized in that: In the reinforcement learning-based agent, the final reward function of the agent satisfies the following formula: r basic (s t ,a t )=0.1*(lap(s t )-u) Among them, r(s t , a t ) represents the final reward function value, r basic (s t ,a t ) represents the basic reward function value, s t represents the state of the agent's interactive exploration with the environment at time t, a t represents the action of the agent, f(a, b) represents the collision distance, lap(s t ) represents the quality of the rendered image estimated using the Laplacian operator, u represents the threshold value of whether the observed scene color is clear; x a ,y a ,z a are the center coordinates of the vehicle in the world coordinate system, x b ,y b , z b are the center coordinates of the colliding object mesh in the world coordinate system, a r is the radius of the vehicle, b r The mesh radius of the colliding object.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of a driving scene physical simulation method based on neural rendering as described in any one of claims 1 to 8 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of a driving scene physical simulation method based on neural rendering as described in any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Neural radiation field for vehicle
CN117523078A
Dynamic human body modeling method based on three-dimensional Gaussian
CN117671108A
Three-dimensional scene model construction method based on neural radiation field
CN117911618A
Nerve radiation field rendering method and device based on nerve base and tensor decomposition
CN119180898A
Automatic driving simulation method, storage medium and computer equipment
CN119203358A