A simulation rendering method, device, equipment and storage medium
By acquiring and processing information such as the target image sample set and combining with the rendering engine to generate the target image, the problems of low rendering efficiency and low image authenticity of traditional simulation are solved, and efficient and realistic new perspective image generation is achieved.
Patent Information
- Application Number
- CN202410089630.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-22
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2044-01-22
AI Technical Summary
Traditional simulation rendering methods are inefficient and have low image authenticity, making it difficult to effectively generate a large number of high-quality images from new perspectives.
By obtaining the target image sample set, target pose information, target camera exposure time, 3D model to be inserted, and 3D model position information to be inserted, the background image and environment map are determined, combined with the rendering engine to generate the foreground image, and finally the target image is generated based on the background and foreground image.
It improves the authenticity and efficiency of simulation rendering, and can generate high-quality new perspective images more quickly, suitable for application scenarios such as autonomous driving.
Smart Images

Figure CN117893660B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of image processing technologies, and in particular, to a simulation rendering method, apparatus, device, and storage medium. Background Art
[0002] In various application scenarios such as autonomous driving, gaming, virtual reality, and augmented reality, it is often necessary to render images of new perspectives in a specific scene. For example, in the field of autonomous driving, simulation rendering technology is required to provide simulation data for the training of detection models, which is essential for training and validating autonomous driving algorithms. Through simulation, it is possible to effectively create rare or dangerous boundary cases in reality and provide comprehensive test scenarios for autonomous driving systems. In addition, simulation helps to reduce the cost of field data collection and ensure the safety of vehicles before they hit the road, which is crucial for the sustainable development of the entire autonomous driving industry.
[0003] Currently, the traditional simulation rendering method is to perform three-dimensional scene modeling based on computer graphics and render the three-dimensional scene model through a rendering engine to obtain image data from a specific perspective. In this solution, the image quality of the new perspective depends on the accuracy of the three-dimensional model and the capabilities of the relevant rendering engine. If a large number of images from new perspectives need to be generated, a large amount of resource costs will be incurred. It can be seen that the traditional simulation rendering process is relatively complex, has low efficiency, and the authenticity of the simulated images is also low. Summary of the Invention
[0004] Embodiments of the present invention provide a simulation rendering method, apparatus, device, and storage medium to achieve improved simulation authenticity and efficiency.
[0005] According to one aspect of the present invention, a simulation rendering method is provided, including:
[0006] Obtaining a target image sample set, target pose information, target camera exposure time, a 3D model to be inserted, and position information of the 3D model to be inserted;
[0007] Determining a background image corresponding to the target pose information according to the target pose information and the target camera exposure time;
[0008] Determining an environment map according to the target image sample set, the target camera exposure time, and the position information of the 3D model to be inserted;
[0009] Inputting the environment map and the 3D model to be inserted into a rendering engine to obtain a foreground image;
[0010] Generating a target image according to the background image and the foreground image.
[0011] According to another aspect of the present invention, a simulation rendering device is provided, and the simulation rendering device includes:
[0012] An acquisition module, configured to acquire a target image sample set, target pose information, a target camera exposure time, a 3D model to be inserted, and position information of the 3D model to be inserted;
[0013] A background image determination module, configured to determine a background image corresponding to the target pose information according to the target pose information and the target camera exposure time;
[0014] An environment map determination module, configured to determine an environment map according to the target image sample set, the target camera exposure time, and the position information of the 3D model to be inserted;
[0015] A foreground image determination module, configured to input the environment map and the 3D model to be inserted into a rendering engine to obtain a foreground image;
[0016] A target image generation module, configured to generate a target image according to the background image and the foreground image.
[0017] According to another aspect of the present invention, an electronic device is provided, and the electronic device includes:
[0018] At least one processor; and
[0019] A memory communicatively connected to the at least one processor; wherein,
[0020] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the simulation rendering method according to any embodiment of the present invention.
[0021] According to another aspect of the present invention, a computer-readable storage medium is provided, and the computer-readable storage medium stores computer instructions, and the computer instructions are used to implement the simulation rendering method according to any embodiment of the present invention when executed by a processor.
[0022] In the embodiments of the present invention, by acquiring a target image sample set, target pose information, a target camera exposure time, a 3D model to be inserted, and position information of the 3D model to be inserted; determining a background image corresponding to the target pose information according to the target pose information and the target camera exposure time; determining an environment map according to the target image sample set, the target camera exposure time, and the position information of the 3D model to be inserted; inputting the environment map and the 3D model to be inserted into a rendering engine to obtain a foreground image; and generating a target image according to the background image and the foreground image, the simulation authenticity and simulation efficiency can be improved.
[0023] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. Description of the Drawings
[0024] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.
[0025] Figure 1 is a flowchart of a simulation rendering method in an embodiment of the present invention;
[0026] Figure 2 is a flowchart of the second model training in an embodiment of the present invention;
[0027] Figure 3 is a flowchart of determining an environment map in an embodiment of the present invention;
[0028] Figure 4 is a schematic structural diagram of a simulation rendering device in an embodiment of the present invention;
[0029] Figure 5 is a schematic structural diagram of an electronic device in an embodiment of the present invention. Detailed Embodiments
[0030] In order to enable those skilled in the art of the present technology to better understand the solutions of the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of the present invention.
[0031] It should be noted that the terms "first", "second", etc. in the specification, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0032] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to users and the authorization of users should be obtained through appropriate means in accordance with relevant laws and regulations.
[0033] Embodiment 1
[0034] Figure 1 It is a flowchart of a simulation rendering method provided by an embodiment of the present invention. This embodiment is applicable to the situation of simulation rendering. This method can be executed by the simulation rendering device in the embodiment of the present invention. The device can be implemented in software and / or hardware, such as Figure 1 As shown, the method specifically includes the following steps:
[0035] S110, obtain a target image sample set, target pose information, target camera exposure time, a 3D model to be inserted, and the position information of the 3D model to be inserted.
[0036] Among them, the target image sample set includes: image samples, pose information corresponding to the image samples, and camera exposure time corresponding to the image samples. The target camera exposure time can be a preset camera exposure time, or the camera exposure time corresponding to any image sample in the target image sample set, or the average value of the camera exposure times corresponding to the image samples in the target image sample set. The embodiments of the present invention do not limit this.
[0037] Among them, the target pose information is the pose information of the camera, and the target pose information includes: the camera view direction and the camera coordinates.
[0038] Among them, the 3D model to be inserted can be a vehicle 3D model, a human body 3D model, or a 3D model of other objects. The embodiments of the present invention do not limit this. The position information of the 3D model to be inserted can be the insertion position coordinates of the 3D model to be inserted.
[0039] Specifically, the method for obtaining the target pose information may be: using MetaShape software to annotate the pose information corresponding to multiple frames of images collected by multiple cameras and unifying them into the same world coordinate system.
[0040] S120. Determine the background image corresponding to the target pose information according to the target pose information and the target camera exposure time.
[0041] Specifically, the method for determining the background image corresponding to the target pose information according to the target pose information and the target camera exposure time may be: determining the HDR intensity value corresponding to each pixel point according to the target pose information and the target camera exposure time, performing Gamma correction on the HDR intensity value corresponding to each pixel point to obtain the LDR color value corresponding to each pixel point, and generating a background image according to the LDR color value corresponding to each pixel point.
[0042] Optionally, determining the background image corresponding to the target pose information according to the target pose information and the target camera exposure time includes:
[0043] Inputting the target pose information and the target camera exposure time into the target neural radiance field network to obtain the background image corresponding to the target pose information, where the target neural radiance field network is obtained by iteratively training the neural radiance field network to be trained through a target image sample set.
[0044] Specifically, the method for inputting the target pose information and the target camera exposure time into the target neural radiance field network to obtain the background image corresponding to the target pose information may be: inputting the target pose information and the target camera exposure time into the target neural radiance field network to obtain the HDR intensity value corresponding to each pixel point, performing Gamma correction on the HDR intensity value corresponding to each pixel point to obtain the LDR color value corresponding to each pixel point, and generating a background image according to the LDR color value corresponding to each pixel point.
[0045] Optionally, the target image sample set includes: image samples, the pose information corresponding to the image samples, and the camera exposure time corresponding to the image samples;
[0046] Iteratively training the neural radiance field network to be trained through the target image sample set includes:
[0047] Determining the light origin and light direction corresponding to each pixel point in the image sample according to the pose information corresponding to the image sample;
[0048] Determining the sampling point set corresponding to each pixel point in the image sample according to the light origin and light direction corresponding to each pixel point in the image sample, where the sampling point set includes: at least two sampling points and the position coordinates and viewing coordinates corresponding to each sampling point;
[0049] Input the position coordinates and viewing coordinates corresponding to each sampling point in the set of sampling points corresponding to each pixel point in the image sample into the neural radiance field network to be trained, and obtain the scene radiance value and density corresponding to each sampling point in the set of sampling points corresponding to each pixel point in the image sample;
[0050] Perform volume rendering based on the scene radiance value, density, and attribute information of each sampling point corresponding to each sampling point in the set of sampling points corresponding to each pixel point to obtain the scene radiance value corresponding to each pixel point;
[0051] Determine the HDR intensity value corresponding to each pixel point according to the camera exposure time corresponding to the image sample and the scene radiance value corresponding to each pixel point;
[0052] Determine the predicted background image according to the HDR intensity value corresponding to each pixel point;
[0053] Train the neural radiance field network to be trained according to the objective function formed by the predicted background image and the image sample to obtain the target neural radiance field network.
[0054] Among them, the image sample can be image samples collected by multiple cameras.
[0055] Specifically, the method for determining the set of sampling points corresponding to each pixel point in the image sample according to the light origin and light direction corresponding to each pixel point in the image sample can be: sampling the light according to the light origin and light direction corresponding to each pixel point to obtain the set of sampling points corresponding to each pixel point. The set of sampling points includes several sampling points. Among them, the attribute information of each sampling point in the set of sampling points can be the sampling interval of each sampling point in the set of sampling points.
[0056] Specifically, the method for determining the HDR intensity value corresponding to each pixel point according to the camera exposure time corresponding to the image sample and the scene radiance value corresponding to each pixel point can be: normalizing the camera exposure time corresponding to the image sample, and multiplying the normalized camera exposure time corresponding to the image sample by the output of volume rendering to obtain the HDR intensity value corresponding to each pixel point.
[0057] Specifically, the method for determining the predicted background image according to the HDR intensity value corresponding to each pixel point can be: performing Gamma correction on the HDR intensity value corresponding to each pixel point to obtain the LDR color value corresponding to each pixel point, and generating the predicted background image according to the LDR color value corresponding to each pixel point.
[0058] Specifically, the method of training the neural radiance field network to be trained with the objective function formed by the predicted background image and the image sample to obtain the target neural radiance field network can be as follows: training the parameters of the neural radiance field network to be trained with the objective function formed by the predicted background image and the image sample, and returning to execute the operation of inputting the position coordinates and viewing angle coordinates corresponding to each sampling point in the sampling point set corresponding to each pixel point in the image sample into the neural radiance field network to be trained to obtain the scene radiance value and density corresponding to each sampling point in the sampling point set corresponding to each pixel point in the image sample, until the target neural radiance field network is obtained.
[0059] S130. Determine the environment map according to the target image sample set, the target camera exposure time, and the position information of the 3D model to be inserted.
[0060] Specifically, the method of determining the environment map according to the target image sample set, the target camera exposure time, and the position information of the 3D model to be inserted can be as follows: obtain the peak direction vector, peak intensity vector, and sky content vector corresponding to each target image sample group in the target image sample set, generate the first HDR sky panoramic image according to the peak direction vector, peak intensity vector, and sky content vector corresponding to each target image sample group, determine the surrounding light panoramic image according to the position information of the 3D model to be inserted and the target camera exposure time, and determine the environment map according to the surrounding light panoramic image and the first HDR sky panoramic image.
[0061] Optionally, determining the environment map according to the target image sample set, the target camera exposure time, and the position information of the 3D model to be inserted includes:
[0062] Input the target image sample group in the target image sample set into the first encoder to obtain the peak direction vector, peak intensity vector, and sky content vector corresponding to the target image sample group, where the first encoder is obtained by iteratively training the first model with the first image sample set;
[0063] Input the peak direction vector, peak intensity vector, and sky content vector corresponding to the target image sample group into the decoding network to obtain the first HDR sky panoramic image;
[0064] Input the position information of the 3D model to be inserted and the target camera exposure time into the target neural radiance field network to obtain the surrounding light panoramic image;
[0065] Overlay the surrounding light panoramic image and the first HDR sky panoramic image to obtain the environment map.
[0066] The target image sample group includes: a front view sample, a left front view sample, and a right front view sample, and the front view sample, the left front view sample, and the right front view sample are respectively acquired by different cameras.
[0067] Specifically, the position information of the 3D model to be inserted and the exposure time of the target camera are input into the target neural radiation field network, and the method for obtaining the panoramic view of the surrounding light can be: emitting sampling light to the upper hemisphere from the position information of the 3D model to be inserted, setting the exposure time to the exposure time of the target camera, and obtaining the HDR intensity value through the high dynamic range multi-camera neural radiation field network, and then determining the panoramic view of the surrounding light according to the obtained HDR intensity value.
[0068] Specifically, the ambient lighting panorama and the first HDR sky panoramic image are superimposed to obtain an environment map by weightedly mixing the ambient lighting panorama and the first HDR sky panoramic image using the transmittance of the last sampling point of the target neural radiation field network to obtain an environment map.
[0069] Optionally, the decoding network includes: a target multi-layer perceptron and a target decoder;
[0070] Before inputting the peak direction vector, the peak intensity vector and the sky content vector corresponding to the target image sample group into a decoding network to obtain a first HDR sky panoramic image, the method further includes:
[0071] Acquire a second image sample set, wherein the second image sample set includes: LDR sky panorama samples and HDR sky panorama samples corresponding to the LDR sky panorama samples;
[0072] Inputting the LDR sky panorama samples in the second image sample set into the second model to obtain a predicted HDR sky panorama;
[0073] The parameters of the second model are trained according to a second function formed by the predicted HDR sky panorama and the HDR sky panorama corresponding to the LDR sky panorama sample to obtain a target image processing model, wherein the target image processing model includes: a target sky dome decoder and a decoding network.
[0074] The second image sample set includes: a pair of LDR and HDR sky panorama samples, and the pair of LDR and HDR sky panorama samples includes: an LDR sky panorama sample and an HDR sky panorama corresponding to the LDR sky panorama sample.
[0075] Specifically, the parameters of the second model may be trained according to a second function formed by the predicted HDR sky panorama and the HDR sky panorama corresponding to the LDR sky panorama samples, and the method for obtaining the target image processing model may be: the parameters of the second model may be trained according to a second function formed by the predicted HDR sky panorama and the HDR sky panorama corresponding to the LDR sky panorama samples, and the operation of inputting the LDR sky panorama samples in the second image sample set into the second model to obtain the predicted HDR sky panorama may be returned to execute, until the target image processing model is obtained.
[0076] Optionally, the second model includes: a skydome encoder to be trained, a multi-layer perceptron to be trained, and a decoder to be trained;
[0077] Inputting the LDR sky panorama sample in the second image sample set into the second model to obtain a predicted HDR sky panorama, including:
[0078] Inputting the LDR sky panorama sample in the second image sample set into a sky encoder to be trained, and obtaining a peak direction vector, a peak intensity vector, and a sky content vector corresponding to the LDR sky panorama sample;
[0079] Converting the peak direction vector corresponding to the LDR sky panorama sample into a peak direction map;
[0080] Converting the peak intensity vector corresponding to the LDR sky panorama sample into a peak brightness map;
[0081] Determine a position coding map according to a direction vector corresponding to each pixel point in the LDR sky panorama sample;
[0082] The peak direction map, the peak brightness map and the position coding map are concatenated to obtain a first feature map;
[0083] Inputting the sky content vector corresponding to the LDR sky panorama sample into the multi-layer perceptron to be trained to obtain a second feature map;
[0084] Inputting the first feature map and the second feature map into a decoder to be trained to obtain a first HDR panoramic image;
[0085] Multiply the peak directional map and the peak brightness map to obtain the spherical Gaussian lobe coded attenuation;
[0086] determining a second HDR panoramic image according to the spherical Gaussian lobe coding attenuation;
[0087] The first HDR panorama and the second HDR panorama are superimposed to obtain a predicted HDR sky panorama.
[0088] Among them, the decoder to be trained can be a UNet network structure.
[0089] Specifically, the method of converting the peak direction vector corresponding to the LDR sky panorama sample into a peak direction map can be: using equirectangular projection to convert the peak direction vector corresponding to the LDR sky panorama sample into a peak direction map encoded by spherical Gaussian lobes.
[0090] Specifically, the method of converting the peak intensity vector corresponding to the LDR sky panorama sample into a peak luminance map can be: using equirectangular projection to convert the peak intensity vector corresponding to the LDR sky panorama sample into a peak luminance map encoded by spherical Gaussian lobes.
[0091] Specifically, the method of determining the position encoding map according to the direction vector corresponding to each pixel point in the LDR sky panorama sample can be: encoding the direction vector corresponding to each pixel point in the LDR sky panorama sample into a position encoding map.
[0092] Among them, the decoder to be trained includes: an encoding layer, a bottleneck layer, and a decoding layer. Specifically, the method of inputting the first feature map and the second feature map into the decoder to be trained to obtain the first HDR panorama can be: inputting the first feature map into the encoding layer and the bottleneck layer of the decoder to be trained; inputting the second feature map and the output of the bottleneck layer of the decoder to be trained after superposition into the decoding layer of the decoder to be trained to obtain the first HDR panorama.
[0093] It should be noted that in order to eliminate the overly smooth output of the neural network, consider a spherical Gaussian lobe encoding attenuation, multiply the peak direction map and the peak luminance map to obtain the spherical Gaussian lobe encoding attenuation, and determine the second HDR panorama according to the spherical Gaussian lobe encoding attenuation; superimpose the first HDR panorama and the second HDR panorama to obtain the predicted HDR sky panorama. This can restore extremely high intensity at the peak position.
[0094] In a specific example, such as Figure 2As shown in the figure, the second model includes: a to-be-trained sky encoder, a to-be-trained multi-layer perceptron, and a to-be-trained decoder. The to-be-trained decoding network includes: a to-be-trained multi-layer perceptron and a to-be-trained decoder. The to-be-trained decoder includes: an encoding layer, a bottleneck layer, and a decoding layer. Input the LDR sky panoramic image sample into the to-be-trained sky encoder to obtain the peak direction vector, peak intensity vector, and sky content vector corresponding to the LDR sky panoramic image sample; convert the peak direction vector corresponding to the LDR sky panoramic image sample into a peak direction map; convert the peak intensity vector corresponding to the LDR sky panoramic image sample into a peak brightness map; determine the position encoding map according to the direction vector corresponding to each pixel point in the LDR sky panoramic image sample; splice the peak direction map, the peak brightness map, and the position encoding map to obtain a first feature map; input the sky content vector corresponding to the LDR sky panoramic image sample into the to-be-trained multi-layer perceptron to obtain a second feature map; input the first feature map into the encoding layer and bottleneck layer of the to-be-trained decoder; stack the second feature map and the output of the bottleneck layer of the to-be-trained decoder and input them into the decoding layer of the to-be-trained decoder to obtain a first HDR panoramic image; multiply the peak direction map and the peak brightness map to obtain the spherical Gaussian lobe encoding attenuation; determine the second HDR panoramic image according to the spherical Gaussian lobe encoding attenuation; stack the first HDR panoramic image and the second HDR panoramic image to obtain the predicted HDR sky panoramic image.
[0095] Optionally, iteratively training the first model through the first image sample set includes:
[0096] Obtain the first image sample set, where the first image sample set includes: HDR sky panoramic image samples;
[0097] Crop the HDR sky panoramic image samples in the first image sample set to obtain a first image sample group;
[0098] Input the first image sample group into the first model to obtain the peak direction vector, peak intensity vector, and sky content vector corresponding to the first image sample group;
[0099] Input the peak direction vector, peak intensity vector, and sky content vector corresponding to the first image sample group into the decoding network to obtain the predicted HDR sky panoramic image;
[0100] Train the parameters of the first model according to the first function formed by the predicted HDR sky panoramic image and the HDR sky panoramic image sample to obtain a first encoder.
[0101] Among them, the first image sample group includes: a front view sample, a left front view sample, and a right front view sample, and the front view sample, the left front view sample, and the right front view sample are respectively collected by different cameras.
[0102] It should be noted that after training the parameters of the second model and obtaining the decoding network, the parameters of the first model are trained.
[0103] In a specific example, as Figure 3 shown, the target image sample group is input into the first encoder to obtain the peak direction vector, peak intensity vector, and sky content vector corresponding to the target image sample group; the peak direction vector, peak intensity vector, and sky content vector corresponding to the target image sample group are input into the decoding network to obtain the first HDR sky panoramic image; the position information of the 3D model to be inserted and the exposure time of the target camera are input into the target neural radiance field network to obtain the surrounding illumination panoramic image; the surrounding illumination panoramic image and the first HDR sky panoramic image are mixed to obtain the environment map.
[0104] S140, input the environment map and the 3D model to be inserted into a rendering engine to obtain a foreground image.
[0105] Specifically, the way of inputting the environment map and the 3D model to be inserted into a rendering engine to obtain a foreground image can be: inputting the environment map and the 3D model to be inserted into an existing rendering engine for foreground rendering to obtain a foreground image. Among them, the rendering results include: an RGB image, physically based depth information, and mask information of the foreground object.
[0106] S150, generate a target image according to the background image and the foreground image.
[0107] Among them, the target image is the output of the simulation camera.
[0108] It should be noted that by using the target neural radiance field network, background images and depth information from arbitrary perspectives can be obtained; by using light estimation and traditional rendering engines, foreground images, depth information, and mask information from arbitrary perspectives can be obtained. According to the depth information and mask information, the foreground image is superimposed on the background image and synthesized into the output of the simulation camera from an arbitrary perspective, that is, the target image.
[0109] The technical solution provided by the embodiments of the present invention has the following technical effects:
[0110] Improve simulation authenticity: Through neural radiance field technology, highly realistic three-dimensional scene reconstruction can be achieved. This means that the autonomous driving system can be trained and tested in an environment extremely close to the real world, thereby improving its performance in the real world.
[0111] Enhance the diversity and richness of data: Combining neural radiance fields and traditional rendering engines can create various complex and variable traffic scenarios, including extreme situations that are difficult to encounter in reality. This can not only provide more comprehensive training data but also test and enhance the adaptability of autonomous driving systems under various conditions.
[0112] Improve the generalization ability of the model: Since the autonomous driving system can be exposed to a wider range of scenarios and situations, its generalization ability is enhanced. This means that the system can better handle various unknown situations and improve its stability and reliability in unknown environments.
[0113] Reduce costs and improve safety: By reducing the dependence on field tests, the costs of data collection and physical testing can be reduced. At the same time, potential safety issues can be identified and corrected in a virtual environment, improving the safety of the autonomous driving system before it hits the road.
[0114] Accelerate the development of autonomous driving technology: This efficient and cost-effective simulation technology significantly shortens the development and testing cycle of autonomous driving systems, accelerating the transformation of autonomous driving technology from the laboratory to the market.
[0115] The technical solution of this embodiment, by obtaining a target image sample set, target pose information, target camera exposure time, a 3D model to be inserted, and the position information of the 3D model to be inserted; determining a background image corresponding to the target pose information according to the target pose information and the target camera exposure time; determining an environment map according to the target image sample set, the target camera exposure time, and the position information of the 3D model to be inserted; inputting the environment map and the 3D model to be inserted into a rendering engine to obtain a foreground image; and generating a target image according to the background image and the foreground image, can improve the simulation authenticity and simulation efficiency.
[0116] Embodiment 2
[0117] Figure 4 It is a schematic structural diagram of a simulation rendering device provided by an embodiment of the present invention. This embodiment is applicable to the situation of simulation rendering. The device can be implemented in software and / or hardware, and the device can be integrated into any device providing simulation rendering functions, such as Figure 4 As shown, the simulation rendering device specifically includes: an acquisition module 410, a background image determination module 420, an environment map determination module 430, a foreground image determination module 440, and a target image generation module 450.
[0118] Among them, the acquisition module is used to obtain a target image sample set, target pose information, target camera exposure time, a 3D model to be inserted, and the position information of the 3D model to be inserted;
[0119] A background image determination module, configured to determine a background image corresponding to the target pose information according to the target pose information and the target camera exposure time;
[0120] An environment map determination module, configured to determine an environment map according to the target image sample set, the target camera exposure time, and the position information of the 3D model to be inserted;
[0121] A foreground image determination module, configured to input the environment map and the 3D model to be inserted into a rendering engine to obtain a foreground image;
[0122] A target image generation module, configured to generate a target image according to the background image and the foreground image.
[0123] The above product can execute the method provided in any embodiment of the present invention, and has function modules and beneficial effects corresponding to the execution of the method.
[0124] Embodiment III
[0125] Figure 5 FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device (such as a helmet, glasses, a watch, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0126] As Figure 5 shown, the electronic device 10 includes at least one processor 11, and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. Among them, the memory stores a computer program executable by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.
[0127] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0128] The processor 11 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the simulation rendering method.
[0129] In some embodiments, the simulation rendering method can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the simulation rendering method described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the simulation rendering method in any other suitable manner (e.g., by means of firmware).
[0130] The various embodiments of the systems and technologies described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs, the one or more computer programs can be executed and / or interpreted on a programmable system including at least one programmable processor, the programmable processor can be a dedicated or general-purpose programmable processor, can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0131] A computer program for implementing the method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer programs are executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer programs can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0132] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0133] In order to provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0134] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected with each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0135] A computing system can include a client and a server. The client and the server are generally far from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0136] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.
[0137] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A simulation rendering method, characterized in that: include: Obtain target image sample set, target pose information, target camera exposure time, 3D model to be inserted, and position information of the 3D model to be inserted; Determine a background image corresponding to the target posture information according to the target posture information and the target camera exposure time; Determine an environment map according to the target image sample set, the target camera exposure time, and the position information of the 3D model to be inserted; Inputting the environment map and the 3D model to be inserted into a rendering engine to obtain a foreground image; Generate a target image according to the background image and the foreground image; Determining an environment map according to the target image sample set, the target camera exposure time, and the position information of the 3D model to be inserted includes: Inputting a target image sample group in the target image sample set into a first encoder to obtain a peak direction vector, a peak intensity vector, and a sky content vector corresponding to the target image sample group, wherein the first encoder is obtained by iteratively training a first model with a first image sample set; Inputting the peak direction vector, the peak intensity vector and the sky content vector corresponding to the target image sample group into a decoding network to obtain a first HDR sky panoramic image; Inputting the position information of the 3D model to be inserted and the exposure time of the target camera into the target neural radiation field network to obtain a panoramic view of the surrounding lighting; The ambient lighting panorama and the first HDR sky panorama image are superimposed to obtain an environment map.
2. The method according to claim 1, characterized in that Determining a background image corresponding to the target posture information according to the target posture information and the target camera exposure time includes: The target pose information and the target camera exposure time are input into a target neural radiation field network to obtain a background image corresponding to the target pose information, wherein the target neural radiation field network is obtained by iteratively training the neural radiation field network to be trained using a target image sample set.
3. The method according to claim 2, characterized in that The target image sample set includes: image samples, posture information corresponding to the image samples, and camera exposure time corresponding to the image samples; The neural radiation field network to be trained is iteratively trained through the target image sample set, including: Determine the light origin and light direction corresponding to each pixel in the image sample according to the posture information corresponding to the image sample; Determine a sampling point set corresponding to each pixel point in the image sample according to the light origin and the light direction corresponding to each pixel point in the image sample, wherein the sampling point set includes: at least two sampling points and position coordinates and viewing angle coordinates corresponding to each sampling point; The position coordinates and the viewing angle coordinates of each sampling point in the sampling point set corresponding to each pixel in the image sample are input into the neural radiation field network to be trained, and the scene radiation brightness value and density of each sampling point in the sampling point set corresponding to each pixel in the image sample are obtained; Volume rendering is performed according to the scene radiation brightness value and density corresponding to each sampling point in the sampling point set corresponding to each pixel point and the attribute information of each sampling point in the sampling point set to obtain the scene radiation brightness value corresponding to each pixel point; Determine the HDR intensity value corresponding to each pixel point according to the camera exposure time corresponding to the image sample and the scene radiance value corresponding to each pixel point; Determine the predicted background image according to the HDR intensity value corresponding to each pixel; The neural radiation field network to be trained is trained according to the objective function formed by the predicted background image and the image samples to obtain a target neural radiation field network.
4. The method according to claim 1, characterized in that: The decoding network includes: a target multi-layer perceptron and a target decoder; Before inputting the peak direction vector, the peak intensity vector and the sky content vector corresponding to the target image sample group into a decoding network to obtain a first HDR sky panoramic image, the method further includes: Acquire a second image sample set, wherein the second image sample set includes: LDR sky panorama samples and HDR sky panorama samples corresponding to the LDR sky panorama samples; Inputting the LDR sky panorama samples in the second image sample set into the second model to obtain a predicted HDR sky panorama; The parameters of the second model are trained according to a second function formed by the predicted HDR sky panorama and the HDR sky panorama corresponding to the LDR sky panorama sample to obtain a target image processing model, wherein the target image processing model includes: a target sky dome decoder and a decoding network.
5. The method according to claim 4, characterized in that The second model includes: a skydome encoder to be trained, a multi-layer perceptron to be trained, and a decoder to be trained; Inputting the LDR sky panorama sample in the second image sample set into the second model to obtain a predicted HDR sky panorama, including: Inputting the LDR sky panorama sample in the second image sample set into a sky encoder to be trained, and obtaining a peak direction vector, a peak intensity vector, and a sky content vector corresponding to the LDR sky panorama sample; Converting the peak direction vector corresponding to the LDR sky panorama sample into a peak direction map; Converting the peak intensity vector corresponding to the LDR sky panorama sample into a peak brightness map; Determine a position coding map according to a direction vector corresponding to each pixel point in the LDR sky panorama sample; The peak direction map, the peak brightness map and the position coding map are concatenated to obtain a first feature map; Inputting the sky content vector corresponding to the LDR sky panorama sample into the multi-layer perceptron to be trained to obtain a second feature map; Inputting the first feature map and the second feature map into a decoder to be trained to obtain a first HDR panoramic image; Multiply the peak directional map and the peak brightness map to obtain the spherical Gaussian lobe coded attenuation; determining a second HDR panoramic image according to the spherical Gaussian lobe coding attenuation; The first HDR panorama and the second HDR panorama are superimposed to obtain a predicted HDR sky panorama.
6. The method according to claim 5, characterized in that Iteratively training a first model through a first image sample set includes: Acquire a first image sample set, wherein the first image sample set includes: HDR sky panoramic image samples; Cropping the HDR sky panoramic image samples in the first image sample set to obtain a first image sample group; Inputting the first image sample group into a first model to obtain a peak direction vector, a peak intensity vector, and a sky content vector corresponding to the first image sample group; Inputting the peak direction vector, the peak intensity vector and the sky content vector corresponding to the first image sample group into a decoding network to obtain a predicted HDR sky panoramic image; The parameters of the first model are trained according to a first function formed by the predicted HDR sky panoramic image and the HDR sky panoramic image sample to obtain a first encoder.
7. A simulation rendering device, characterized in that: include: An acquisition module is used to acquire a target image sample set, target posture information, target camera exposure time, a 3D model to be inserted, and position information of the 3D model to be inserted; A background image determination module, used to determine the background image corresponding to the target posture information according to the target posture information and the target camera exposure time; An environment map determination module is used to determine the environment map according to the target image sample set, the target camera exposure time and the position information of the 3D model to be inserted; A foreground image determination module, used for inputting the environment map and the 3D model to be inserted into a rendering engine to obtain a foreground image; A target image generation module, used to generate a target image according to the background image and the foreground image; The environment map determination module is specifically used for: Inputting a target image sample group in the target image sample set into a first encoder to obtain a peak direction vector, a peak intensity vector, and a sky content vector corresponding to the target image sample group, wherein the first encoder is obtained by iteratively training a first model with a first image sample set; Inputting the peak direction vector, the peak intensity vector and the sky content vector corresponding to the target image sample group into a decoding network to obtain a first HDR sky panoramic image; Inputting the position information of the 3D model to be inserted and the exposure time of the target camera into the target neural radiation field network to obtain a panoramic view of the surrounding lighting; The ambient lighting panorama and the first HDR sky panorama image are superimposed to obtain an environment map.
8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the simulation rendering method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the simulation rendering method according to any one of claims 1 to 6 when executed.
Citation Information
Patent Citations
Method for rendering virtual object based on brightness estimation, method for training neural network and related product
CN115039137A
Text-driven immersive open scene neural rendering and hybrid enhancement method
CN116563459A