A three-dimensional driving recorder system based on NeRF
Through the NeRF-based three-dimensional driving recording system, the improved NeRF network and block reconstruction technology are used to solve the problems of large amount of calculation and insufficient clarity of the existing three-dimensional reconstruction, and multi-angle observation and fine image reconstruction are realized.
Patent Information
- Application Number
- CN202310766307.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-26
- Publication Date
- 2025-08-29
- Estimated Expiration
- 2043-06-26
AI Technical Summary
The existing three-dimensional driving recording system has high calculation volume and requires a lot of images during three-dimensional reconstruction, making it difficult to achieve photo-level clarity.
The three-dimensional driving recording system based on NeRF is adopted to obtain the video data around the vehicle through the data acquisition module, and the three-dimensional reconstruction of static street scenes and dynamic objects is used to use the improved NeRF network module to carry out three-dimensional reconstruction of static street scenes and dynamic objects. Combined with Block-NeRF's block reconstruction idea and adversarial neural network, static street scenes and dynamic object models are integrated to output multi-angle observation images.
Multi-angle observation is realized, and fine reconstruction result images are generated, which reduces the amount of calculation and increases clarity, and can render viewing images that do not exist in the original data.
Smart Images

Figure CN116704643B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of vehicle driving recorders, and in particular to a three-dimensional vehicle driving recorder system based on NeRF. Background Art
[0002] A car dashcam, also known as a "black box," is a digital, fully automatic, and intelligent on-board monitoring device controlled by a microcomputer, adapting the design principles of aircraft black boxes to automobiles. This technology, transferred from military aviation technology to civilian use, is a digital, fully automatic, and intelligent on-board monitoring device controlled by a microcomputer. Integrating mechanics, electronics, and microcomputers, it is suitable for accident prevention, violation detection, and scientific management, providing a fair, accurate, and scientific basis for accident analysis.
[0003] Car dashcams have multiple functions for recording timely driving data, playing an irreplaceable role in road traffic management, even beyond the reach of traffic police. They record traffic violations and prevent accidents. Addressing major causes of traffic accidents, such as speeding and fatigue driving, car dashcams automatically, continuously, and accurately record a driver's speed and driving progress over hundreds of days. This effectively acts as a permanent traffic officer on board, providing comprehensive, long-term, and timely monitoring of the driver's safe driving. This data can be used to establish a driver safety profile, providing a scientific basis for assessments, evaluations, and rewards and penalties. It also helps cultivate good driving habits, such as safe and moderate speeds, to reduce violations and prevent accidents. They can also scientifically analyze and handle road traffic accidents. Car dashcams can record a vehicle's speed, time, deceleration and braking, and driving conditions before an accident, providing quantitative and accurate data for handling traffic accidents. This data is used to assist with on-site investigations and evidence collection, enabling fair determination of accident causes and responsibility. This provides a scientific basis for improving the efficiency of traffic police and effectively protecting the legitimate rights and interests of vehicle owners and drivers.
[0004] At present, the most popular form of car dashcams is to record the situation in front of the car in real time, which also means that it is difficult to take into account other directions. In order to facilitate the owners to drive in congested and narrow areas, most vehicles are now equipped with a car panoramic view system, which uses fisheye cameras in four directions to collect the situation around the car, and transform and splice it into a bird's-eye view to help the owner drive more precisely. However, this still limits the viewing angle to a certain extent.
[0005] If the driving scene can be reconstructed in three dimensions, users will be able to observe the situation in all directions while driving from a freer perspective. However, the currently popular three-dimensional reconstruction method requires a large number of images while requiring high computational effort, and it is difficult to obtain a reconstruction result with photo-level clarity. Summary of the Invention
[0006] The present invention aims to solve the technical problems existing in the prior art and provides a NeRF-based three-dimensional vehicle driving recorder system to solve the problem that the three-dimensional reconstruction method of the prior art three-dimensional vehicle driving recorder system has high computational complexity and requires a large number of images.
[0007] According to a first aspect of the present invention, there is provided a NeRF-based three-dimensional driving recorder system, comprising: a data acquisition module, an improved NeRF network module, and a reconstruction module;
[0008] The data acquisition module collects and stores video data of the vehicle's surroundings from various perspectives centered on the vehicle while the vehicle is traveling; each perspective includes at least four perspectives: front, back, left, and right;
[0009] The improved NeRF network module includes: a static street scene model construction unit and a dynamic object model construction unit;
[0010] The static street scene model construction unit trains the video data images based on the Block-NeRF block reconstruction concept to obtain a 3D reconstruction model of the static street scene; the dynamic object model construction unit pre-sets the types of dynamic objects based on the 3D reconstruction model of the static street scene obtained by the static street scene model construction unit, and trains the video data images through the NeRF network to obtain 3D reconstruction models of various types of dynamic objects; the improved NeRF network module fuses the dynamic objects with the static street scene and outputs a reconstructed model;
[0011] The reconstruction module reconstructs the corresponding image according to the viewing angle requirement information input by the user and the output of the improved NeRF network module.
[0012] On the basis of the above technical solution, the present invention can also make the following improvements.
[0013] Optionally, the data acquisition module includes: a data storage terminal and at least four fisheye cameras;
[0014] The fisheye camera collects videos from four viewing angles of the vehicle, front, back, left, and right, in real time and saves the videos to the data storage terminal.
[0015] Optionally, the processing of the static street view model building unit includes:
[0016] Acquire an image with a time difference less than a set threshold and remove dynamic objects, perform a traditional 3D reconstruction structure-from-motion operation on the image to obtain sparse depth data in the corresponding scene;
[0017] According to the BlockNeRF blocking strategy, the recorded data is divided into several different blocks in time. When processing each individual block, the corresponding blocks are fused according to needs to obtain an implicit 3D model of the static street scene.
[0018] Optionally, the types of dynamic objects pre-set by the dynamic object model building unit include: vehicles and pedestrians;
[0019] The dynamic object model construction unit obtains three-dimensional reconstruction models of various types of dynamic objects by training the images of the video data in combination with the NeRF network of the adversarial neural network.
[0020] Optionally, the viewing angle requirement information input by the user includes: time point, camera position, and direction angle information.
[0021] Optionally, the processing of the reconstruction module includes:
[0022] The camera position and direction angle information input by the user is converted into a three-dimensional affine transformation matrix and rotation value input into the NeRF network, and the target image is reconstructed using the reconstruction model output by the improved NeRF network module, and output to the user interface for display to the user. According to a second aspect of the present invention, a three-dimensional driving recorder system based on NeRF is provided, comprising:
[0023] According to a third aspect of the present invention, an electronic device is provided, comprising a memory and a processor, wherein the processor is configured to implement steps of a NeRF-based three-dimensional vehicle driving recorder system when executing a computer management program stored in the memory.
[0024] According to a fourth aspect of the present invention, a computer-readable storage medium is provided, on which a computer management program is stored. When the computer management program is executed by a processor, the steps of a NeRF-based three-dimensional vehicle driving recorder system are implemented.
[0025] The present invention provides a NeRF-based three-dimensional driving recorder system, system, electronic device and storage medium, which can first provide multi-angle observation. Existing driving recorders all use video recording as a method. This method reduces the equipment requirements while also limiting the observation angle. After completing the reconstruction calculation, this method can provide observation images from various angles, including angles that do not exist in the original data. This provides a brand new solution for current driving records. Secondly, it can produce fine result images. Traditional three-dimensional reconstruction methods usually use a certain area in space as the basic unit of three-dimensional reconstruction, but if this unit is divided into a small size, it will lead to a huge amount of reconstruction calculation; if this unit is divided into too large, it will lead to insufficient clarity of the reconstruction result. This method uses implicit three-dimensional reconstruction based on NeRF, which can maximize the use of the clarity of the original data. It can make the clarity of the rendering result close to the original data through continuous iteration; it can also terminate after a certain number of iterations, making a compromise between clarity and training time. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 A processing flow chart of an embodiment of a NeRF-based three-dimensional vehicle driving recorder system provided by the present invention;
[0027] Figure 2 A schematic diagram of the position distribution of a fisheye camera provided by an embodiment of the present invention;
[0028] Figure 3 A schematic diagram of the processing process of an embodiment of an improved NeRF network module provided by the present invention;
[0029] Figure 4 A schematic diagram of the hardware structure of a possible electronic device provided by the present invention;
[0030] Figure 5 A schematic diagram of the hardware structure of a possible computer-readable storage medium provided by the present invention. DETAILED DESCRIPTION
[0031] The principles and features of the present invention are described below with reference to the accompanying drawings. The examples given are only used to explain the present invention and are not used to limit the scope of the present invention.
[0032] Figure 1 A processing flow chart of an embodiment of a NeRF-based three-dimensional vehicle driving recorder system provided by the present invention is shown as follows: Figure 1 As shown, the system includes:
[0033] Data acquisition module, improved NeRF network module and reconstruction module.
[0034] The data acquisition module collects and saves video data around the vehicle from various perspectives centered on the vehicle while the vehicle is traveling; each perspective includes at least four perspectives: front, back, left, and right.
[0035] The improved NeRF network module includes: a static street scene model construction unit and a dynamic object model construction unit.
[0036] The static street scene model construction unit trains the images of video data based on the Block-NeRF block reconstruction idea to obtain a three-dimensional reconstruction model of the static street scene; the dynamic object model construction unit pre-sets the types of dynamic objects based on the three-dimensional reconstruction model of the static street scene obtained by the static street scene model construction unit, and trains the images of video data through the NeRF network to obtain three-dimensional reconstruction models of various types of dynamic objects; the improved NeRF network module fuses the dynamic objects with the static street scene and outputs the reconstructed model.
[0037] The reconstruction module reconstructs the corresponding image based on the viewing angle requirement information input by the user and the output of the improved NeRF network module.
[0038] NeRF (Neural Radiance Field) is an implicit 3D reconstruction method proposed in 2020. It uses images of a specific object at multiple angles as input and a neural network as a carrier, and combines it with the volume rendering method to reconstruct images observed from perspectives that do not exist in the original data. Because the implicitly represented 3D scene is continuous, it is easier to produce photo-realistic images than traditional 3D reconstruction.
[0039] However, NeRF also has problems such as requiring a large amount of image data and taking a long time to train and reconstruct. In addition, there is relatively little research on applying NeRF to inside-out scenarios (i.e., observing the surrounding scenes radially from the center). Therefore, NeRF needs to be improved to adapt to the scenarios required by the three-dimensional driving recorder system.
[0040] Existing dashcams all rely on video recording, which reduces equipment requirements but also limits the viewing angle. The present invention provides a NeRF-based 3D dashcam system that can provide multi-angle observations and, after completing reconstruction calculations, can provide observation images from various perspectives, including perspectives not present in the original data. The improved NeRF network uses less image data as input than the original NeRF network, and uses the collected data to train one or more implicit 3D reconstruction models. Ultimately, these implicit 3D reconstruction models are input into the reconstruction module, which reconstructs the image from the desired perspective based on the time point, observation direction, and angle information requested by the user.
[0041] Example 1
[0042] The embodiment 1 provided by the present invention is an embodiment of a NeRF-based three-dimensional driving recorder system provided by the present invention, which divides the target task into two subtasks. The first is a three-dimensional reconstruction of the road passed by the vehicle during driving, and the goal is to obtain an implicit three-dimensional reconstruction of the street scene that can generate a virtual perspective. This three-dimensional reconstruction model does not need to include dynamic objects such as vehicles and pedestrians; the second is an implicit three-dimensional scene picture at a specific moment during driving, and the goal is to obtain an implicit three-dimensional reconstruction of the scene around the vehicle at this moment. In this task, as many of the photographed objects as possible should be displayed one by one. Combined with Figure 1 It can be seen that the embodiment of the system includes:
[0043] Data acquisition module, improved NeRF network module and reconstruction module.
[0044] In a possible embodiment, the data acquisition module includes: a data storage terminal and at least four fisheye cameras.
[0045] The fisheye camera collects videos from four perspectives of the vehicle in real time and saves them to the data storage terminal.
[0046] like Figure 2 Figure 1 shows a schematic diagram of the fisheye camera layout provided by an embodiment of the present invention. This layout is similar to the cameras used in popular automotive surround-view systems, located at the front, back, left, and right sides of the vehicle. When the module is enabled, the cameras simultaneously turn on, capturing real-time video from all four perspectives of the vehicle and saving it to the data terminal. Each camera also records its activation time to address any unsynchronized activations.
[0047] The data acquisition module collects and saves video data around the vehicle from various perspectives centered on the vehicle while the vehicle is traveling; each perspective includes at least four perspectives: front, back, left, and right.
[0048] The improved NeRF network module includes: a static street scene model construction unit and a dynamic object model construction unit.
[0049] The static street scene model construction unit trains the images of video data based on the Block-NeRF block reconstruction idea to obtain a three-dimensional reconstruction model of the static street scene; the dynamic object model construction unit pre-sets the types of dynamic objects based on the three-dimensional reconstruction model of the static street scene obtained by the static street scene model construction unit, and trains the images of video data through the NeRF network to obtain three-dimensional reconstruction models of various types of dynamic objects; the improved NeRF network module fuses the dynamic objects with the static street scene and outputs the reconstructed model.
[0050] like Figure 3FIG. 1 is a schematic diagram of a processing process of an embodiment of an improved NeRF network module provided by the present invention. In one possible embodiment, the processing process of the static street scene model construction unit includes:
[0051] Images with a time difference less than a set threshold are acquired and dynamic objects are removed. The images are subjected to a traditional 3D reconstruction structure from motion (SfM) operation to obtain sparse depth data in the corresponding scene.
[0052] According to the BlockNeRF blocking strategy, the recorded data is divided into several different blocks in time. When processing each individual block, the corresponding blocks are fused according to needs to obtain an implicit 3D model of the static street scene.
[0053] Dynamic objects in the scene are separated from the static environment, and the entire driving route is treated as a large static object. Cost-volume processing is introduced in explicit 3D reconstruction. An "inside-out" scenario approach is used, using information such as the original image, dense depth map, and uncertainty map to guide NeRF's 3D reconstruction of static street scenes. Since the data is sampled within a single driving time, the data collected by the data acquisition module is continuous image data. In this case, the time interval between adjacent image sampling is not long, so there is no need to consider the impact of time on environmental lighting and other effects. For longer scenes, a block-by-block fusion method can also be used. However, due to the uncertainty of the scene shape, the block-by-block and fusion methods may require certain restrictions and improvements.
[0054] In a possible embodiment, the types of dynamic objects preset by the dynamic object model building unit include: vehicles and pedestrians.
[0055] The dynamic object model construction unit obtains three-dimensional reconstruction models of various types of dynamic objects by training the images of video data in combination with the NeRF network of the adversarial neural network.
[0056] Using the 3D scene constructed by the static street scene model building unit as prior knowledge, the dynamic object model building unit only needs to complete the 3D reconstruction of dynamic objects in the scene, such as cars and people. Since this involves objects that are relatively dynamic relative to the scene, fewer images are required to complete the reconstruction. Because the number of dynamic objects in this scene is relatively small, NeRF combined with an adversarial neural network can be used to perform 3D reconstruction of several types of dynamic objects in the environment.
[0057] Finally, the reconstructed dynamic objects are fused with the previously reconstructed static scene to achieve the desired effect. The improved NeRF network module takes less image data as input than the original NeRF and ultimately outputs a trained implicit 3D reconstruction model.
[0058] The reconstruction module reconstructs the corresponding image based on the viewing angle requirement information input by the user and the output of the improved NeRF network module.
[0059] In a possible embodiment, the viewing angle requirement information input by the user includes: time point, camera position, and direction angle information.
[0060] In a possible embodiment, the processing of the reconstruction module includes:
[0061] The camera position and direction angle information input by the user is converted into a three-dimensional affine transformation matrix and rotation value input into the NeRF network. The target image is reconstructed using the reconstruction model output by the improved NeRF network module and output to the user interface for display to the user.
[0062] During rendering, NeRF uses a three-dimensional affine transformation matrix and a rotation value to determine the camera's position and orientation. After rendering the current image, it saves the image directly to the corresponding directory. To make the program more user-friendly, the reconstruction module first guides the user to enter the camera position and orientation information corresponding to the desired image. It then uses the affine transformation formula to convert these two pieces of information into a three-dimensional affine transformation matrix and rotation value that can be input into NeRF. The target image is then reconstructed using the acquired implicit three-dimensional reconstruction model and output to the program's user interface, allowing the user to preview and save the image. The implicit three-dimensional reconstruction model trained by the improved NeRF network module is a neural network that takes perspective and position information as input and outputs color information of the corresponding area. The reconstruction module, on the other hand, integrates the previously trained implicit neural network to create a visual input and output interface for the user.
[0063] Example 2
[0064] The embodiment 2 provided by the present invention is an embodiment of a three-dimensional driving recording method based on NeRF provided by the present invention, combined with Figure 1-Figure 3 It can be seen that the embodiment of the method includes:
[0065] Step 1: Collect and save video data around the vehicle from various perspectives centered on the vehicle while the vehicle is traveling; each perspective includes at least four perspectives: front, back, left, and right.
[0066] In step 2, the video data images are trained based on the Block-NeRF block reconstruction concept to obtain a 3D reconstruction model of the static street scene. Based on the 3D reconstruction model of the static street scene, the types of dynamic objects are pre-set, and the video data images are trained through the NeRF network to obtain 3D reconstruction models of various types of dynamic objects. The dynamic objects are fused with the static street scene and the reconstructed model is output.
[0067] Step 3: Reconstruct the corresponding image based on the user's input viewing angle requirement information and the output of the improved NeRF network module.
[0068] It can be understood that the NeRF-based three-dimensional driving recording method provided by the present invention corresponds to the NeRF-based three-dimensional driving recording system provided in the aforementioned embodiments. The relevant technical features of the NeRF-based three-dimensional driving recording method can refer to the relevant technical features of the NeRF-based three-dimensional driving recording system, which will not be repeated here.
[0069] See also Figure 4 , Figure 4 Schematic diagram of an embodiment of an electronic device provided by an embodiment of the present invention. Figure 4 As shown, an embodiment of the present invention provides an electronic device, including a memory 1310, a processor 1320, and a computer program 1311 stored in the memory 1310 and executable on the processor 1320. When the processor 1320 executes the computer program 1311, the following steps are implemented: collecting and saving video data around the vehicle from various perspectives centered on the vehicle when the vehicle is traveling; each perspective includes at least four perspectives: front, back, left, and right; training the image of the video data based on the Block-NeRF block reconstruction idea to obtain a three-dimensional reconstruction model of a static street scene; based on the three-dimensional reconstruction model of the static street scene, pre-setting the type of dynamic object, and training the image of the video data through the NeRF network to obtain three-dimensional reconstruction models of various types of dynamic objects; fusing the dynamic object with the static street scene and outputting the reconstructed model; reconstructing the corresponding image based on the perspective requirement information input by the user and the output of the improved NeRF network module.
[0070] See also Figure 5 , Figure 5 Schematic diagram of an embodiment of a computer-readable storage medium provided by the present invention. Figure 5As shown, this embodiment provides a computer-readable storage medium 1400, on which a computer program 1411 is stored. When the computer program 1411 is executed by a processor, the following steps are implemented: collecting and saving video data around the vehicle from various perspectives centered on the vehicle when the vehicle is driving; each perspective includes at least four perspectives: front, back, left, and right; training the image of the video data based on the block reconstruction idea of Block-NeRF to obtain a three-dimensional reconstruction model of the static street scene; based on the three-dimensional reconstruction model of the static street scene, the type of dynamic object is pre-set, and the image of the video data is trained through the NeRF network to obtain three-dimensional reconstruction models of various types of dynamic objects; the dynamic object is fused with the static street scene and the reconstructed model is output; the corresponding image is reconstructed according to the perspective requirement information input by the user and the output of the improved NeRF network module.
[0071] The embodiments of the present invention provide a NeRF-based three-dimensional driving recorder system, system, electronic device, and storage medium. First, they can provide multi-angle observation. Existing driving recorders all use video recording as a method. While this method reduces equipment requirements, it also limits the observation angle. After completing the reconstruction calculation, this method can provide observation images from various angles, including angles that do not exist in the original data. This provides a brand new solution for current driving records. Second, it can produce detailed result images. Traditional three-dimensional reconstruction methods usually use a certain area in space as the basic unit of three-dimensional reconstruction. However, if this unit is divided into small units, the reconstruction calculation will be huge; if this unit is divided into too large units, the clarity of the reconstruction result will be insufficient. This method uses implicit three-dimensional reconstruction based on NeRF, which can maximize the clarity of the original data. It can make the clarity of the rendering result close to that of the original data through continuous iteration; it can also terminate after a certain number of iterations, making a compromise between clarity and training time.
[0072] It should be noted that, in the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant description of other embodiments.
[0073] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0074] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowcharts and / or block diagrams, as well as combinations of processes and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowcharts and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0075] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0076] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0077] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0078] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.
Claims
1. A three-dimensional driving recorder system based on NeRF, characterized in that: The system includes: a data acquisition module, an improved NeRF network module and a reconstruction module; The data acquisition module collects and stores video data of the vehicle's surroundings from various perspectives centered on the vehicle while the vehicle is traveling; each perspective includes at least four perspectives: front, back, left, and right; The improved NeRF network module includes: a static street scene model construction unit and a dynamic object model construction unit; The static street scene model construction unit trains the video data images based on the Block-NeRF block reconstruction concept to obtain a 3D reconstruction model of the static street scene; the dynamic object model construction unit pre-sets the types of dynamic objects based on the 3D reconstruction model of the static street scene obtained by the static street scene model construction unit, and trains the video data images through the NeRF network to obtain 3D reconstruction models of various types of dynamic objects; the improved NeRF network module fuses the dynamic objects with the static street scene and outputs a reconstructed model; The reconstruction module reconstructs the corresponding image according to the viewing angle requirement information input by the user and the output of the improved NeRF network module.
2. The system according to claim 1, wherein: The data acquisition module includes: a data storage terminal and at least four fisheye cameras; The fisheye camera collects videos from four viewing angles of the vehicle, front, back, left, and right, in real time and saves the videos to the data storage terminal.
3. The system according to claim 1, wherein: The processing process of the static street view model construction unit includes: Acquire an image with a time difference less than a set threshold and remove dynamic objects, perform a traditional 3D reconstruction structure-from-motion operation on the image to obtain sparse depth data in the corresponding scene; According to the BlockNeRF blocking strategy, the recorded data is divided into several different blocks in time. When processing each individual block, the corresponding blocks are fused according to needs to obtain an implicit 3D model of the static street scene.
4. The system according to claim 1, wherein: The types of dynamic objects preset by the dynamic object model building unit include: vehicles and pedestrians; The dynamic object model construction unit obtains three-dimensional reconstruction models of various types of dynamic objects by training the images of the video data in combination with the NeRF network of the adversarial neural network.
5. The system according to claim 1, wherein: The viewing angle requirement information input by the user includes: time point, camera position and direction angle information.
6. The system according to claim 5, characterized in that The processing of the reconstruction module includes: The camera position and direction angle information input by the user is converted into a three-dimensional affine transformation matrix and rotation value input into the NeRF network, and the target image is reconstructed using the reconstruction model output by the improved NeRF network module and output to the user interface for display to the user.
7. A three-dimensional driving recording method based on NeRF, characterized in that: include: Step 1: Collect and save video data of the vehicle's surroundings from various perspectives centered on the vehicle while the vehicle is traveling; Each perspective includes at least four perspectives: front, back, left, and right; Step 2: Constructing an improved NeRF network module including a static street scene model construction unit and a dynamic object model construction unit; the static street scene model construction unit trains the video data images based on the Block-NeRF block reconstruction concept to obtain a 3D reconstruction model of the static street scene; the dynamic object model construction unit pre-sets the types of dynamic objects based on the 3D reconstruction model of the static street scene, and trains the video data images through the NeRF network to obtain 3D reconstruction models of various types of dynamic objects; the improved NeRF network module fuses the dynamic objects with the static street scene and outputs a reconstructed model; Step 3: reconstruct the corresponding image according to the viewing angle requirement information input by the user and the output of the improved NeRF network module.
8. An electronic device, characterized in that: It includes a memory and a processor, and the processor is used to implement the steps of the NeRF-based three-dimensional driving recording method as claimed in claim 7 when executing the computer management program stored in the memory.
9. A computer-readable storage medium, characterized in that A computer management program is stored thereon, and when the computer management program is executed by the processor, the steps of the NeRF-based three-dimensional driving recording method as claimed in claim 7 are implemented.
Citation Information
Patent Citations
Method for generating three-dimensional dynamic scene based on multi-view video and dynamic nerve radiation field
CN115423924A
Image processing method, neural radiation field training method and neural network
CN115631418A