Image acquisition method and device, computer equipment, storage medium and program product
Through acquisition, rendering and synthesis technology, high-precision multi-view images are generated, which solves the problem of low image acquisition efficiency in autonomous driving tests of new energy vehicles and improves the perception ability of the autonomous driving system.
Patent Information
- Application Number
- CN202510312933.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, the testing method of the autonomous driving function of new energy vehicles is single, resulting in low image acquisition efficiency and inability to effectively simulate environmental data from different perspectives.
Point cloud data is collected through vehicle sensors, converted into Gaussian distribution form, geometric and dynamic surface features are extracted, and a three-dimensional Gaussian sputtering technology is used to render it into a three-dimensional model. The diffusion model is input to generate new perspective synthetic data and render it to obtain high-quality new perspective synthetic images.
It improves image acquisition efficiency, breaks through the limitations of traditional single-view environment data, generates high-precision multi-view synthesized images, and enhances the perception ability of the autonomous driving system.
Smart Images

Figure CN120259507A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the technical field of vision change of intelligent driving sensors, and in particular, to an image acquisition method, apparatus, computer device, storage medium, and program product. Background Art
[0002] With the rapid development of the new energy vehicle industry, the autonomous driving function of new energy vehicles is deeply favored by users due to its strong perception and accurate and safe driving performance. Therefore, before the product is launched, the testing work for the autonomous driving function is particularly important.
[0003] In the related art, vehicle manufacturers can conduct on-site tests on vehicles with different model configurations, drive for a long distance (for example, hundreds of thousands of kilometers), and collect test data of the vehicles in different environments in real time to verify the test effect of the vehicles.
[0004] However, the above method for vehicles to obtain test data is relatively single, and the efficiency of obtaining images from different perspectives of the vehicle is relatively low. Summary of the Invention
[0005] Embodiments of the present application provide an image acquisition method, apparatus, computer device, storage medium, and program product, which can improve the efficiency of image acquisition. The technical solution is as follows:
[0006] On the one hand, an image acquisition method is provided, and the method includes:
[0007] Collect point cloud data around the vehicle at a first perspective through a sensor of the vehicle;
[0008] Convert the point cloud data into a Gaussian distribution form;
[0009] Extract a first geometric feature and a first dynamic surface feature from the point cloud data in the Gaussian distribution form;
[0010] Render the first geometric feature and the first dynamic surface feature through a three-dimensional Gaussian sputtering technique to obtain rendered three-dimensional model data;
[0011] Extract feature data of the three-dimensional model data;
[0012] Input the feature data of the three-dimensional model data into a diffusion model to obtain new perspective synthesis data output by the diffusion model; the new perspective synthesis data is environmental data around the vehicle at a second perspective; the second perspective is different from the first perspective;
[0013] Render the new perspective synthesis data to obtain a rendered new perspective synthesis image; the new perspective synthesis image is an image corresponding to the second perspective.
[0014] On the other hand, an image acquisition device is provided, and the device includes:
[0015] A point cloud data acquisition module, configured to acquire point cloud data around the vehicle from a first perspective through a sensor of the vehicle;
[0016] A point cloud data conversion module, configured to convert the point cloud data into a Gaussian distribution form;
[0017] A feature extraction module, configured to extract a first geometric feature and a first dynamic surface feature from the point cloud data in the Gaussian distribution form;
[0018] A feature rendering model, configured to render the first geometric feature and the first dynamic surface feature through a three-dimensional Gaussian sputtering technique to obtain rendered three-dimensional model data;
[0019] A feature data extraction module, configured to extract feature data of the three-dimensional model data;
[0020] A new perspective synthesis data acquisition module, configured to input the feature data of the three-dimensional model data into a diffusion model to obtain new perspective synthesis data output by the diffusion model; the new perspective synthesis data is environmental data around the vehicle from a second perspective; the second perspective is different from the first perspective;
[0021] A new perspective synthesis image acquisition module, configured to render the new perspective synthesis data to obtain a rendered new perspective synthesis image; the new perspective synthesis image is an image corresponding to the second perspective.
[0022] In some embodiments, the device further includes:
[0023] A video data acquisition module, configured to acquire video data of the vehicle through the sensor of the vehicle before converting the point cloud data into a Gaussian distribution form;
[0024] A preprocessing operation execution module, configured to perform a preprocessing operation on the point cloud data based on the video data before converting the point cloud data into a Gaussian distribution form.
[0025] In some embodiments, the preprocessing operation execution module is configured to perform a time synchronization operation and a spatial alignment operation on the video data and the point cloud data.
[0026] In some embodiments, the new perspective synthesis image acquisition module is configured to extract a second geometric feature and a second dynamic surface feature from the new perspective synthesis data; and perform high-precision lighting calculation, texture optimization, and dynamic object synthesis based on the second geometric feature and the second dynamic surface feature to obtain a rendered new perspective synthesis image.
[0027] In some embodiments, the geometric feature includes the shape, size, and position information of the point cloud; and the dynamic surface feature includes the movement trajectory and appearance change of the objects around the vehicle.
[0028] In some embodiments, the apparatus further includes:
[0029] An image application module, configured to perform perception enhancement training and backpropagation simulation experiments on the vehicle based on the new perspective synthesis image; the perception enhancement training includes object detection, object tracking, and driving scene understanding of the vehicle's autonomous driving system; and the backpropagation simulation experiment includes performance testing and performance optimization of the vehicle's autonomous driving algorithm in a virtual environment.
[0030] In another aspect, a computer device is provided, which includes a processor and a memory. At least one instruction, at least one program, a code set, or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the image acquisition method as described above.
[0031] In another aspect, a computer-readable storage medium is provided, in which at least one instruction, at least one program, a code set, or an instruction set is stored, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the image acquisition method as described above.
[0032] In still another aspect, a computer program product is provided, which includes a computer program stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the image acquisition method provided in the above various optional implementation manners.
[0033] The technical solution provided by this application may include the following beneficial effects:
[0034] The autonomous driving system provides high-precision spatial information through the point cloud data captured by the vehicle's sensors. By using three-dimensional Gaussian sputtering technology to render the first geometric features and the first dynamic surface features obtained from the point cloud data, the feature data of the generated three-dimensional model data can simulate complex appearance characteristics while retaining details, enhancing the realism of the three-dimensional model. The feature data of the three-dimensional model data is input into the diffusion model. Through the generation ability of the diffusion model, environmental data from different perspectives is synthesized, breaking through the limitations of traditional single-perspective environmental data. Further rendering the synthesized data of the new perspective to generate new perspective synthetic images of other perspectives improves the acquisition efficiency of the new perspective synthetic images.
[0035] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit this application. Brief Description of the Drawings
[0036] The drawings herein are incorporated into the specification and form a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application.
[0037] Figure 1 It is a system configuration diagram of an image acquisition method according to an embodiment of this application;
[0038] Figure 2 It is a flowchart of an image acquisition method provided by an embodiment of this application;
[0039] Figure 3 It is a flowchart of an image acquisition method provided by an embodiment of this application;
[0040] Figure 4 It is a flowchart of an image acquisition method provided by an embodiment of this application;
[0041] Figure 5 It is a method flowchart of an autonomous driving simulation three-dimensional reconstruction technology of this application;
[0042] Figure 6 It is a schematic flowchart of an autonomous driving simulation three-dimensional reconstruction technology of this application;
[0043] Figure 7 It is a block diagram of an image acquisition device provided by an exemplary embodiment of this application;
[0044] Figure 8 It is a schematic structural diagram of a computer device provided by an exemplary embodiment of this application. Detailed Description of the Embodiments
[0045] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application.
[0046] On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0047] Figure 1 is a system configuration diagram of an image acquisition method according to an embodiment of the present application. As Figure 1 shown, Figure 1 it includes a vehicle 100. The vehicle 100 is provided with a plurality of sensors and an autonomous driving system 100a. The autonomous driving system 100a collects point cloud data around the vehicle 100 from the first perspective through the sensors of the vehicle 100; converts the point cloud data into a Gaussian distribution form; extracts the first geometric feature and the first dynamic surface feature from the point cloud data in the Gaussian distribution form; renders the first geometric feature and the first dynamic surface feature through a three-dimensional Gaussian sputtering technique to obtain the rendered three-dimensional model data; extracts the feature data of the three-dimensional model data; inputs the feature data of the three-dimensional model data into a diffusion model to obtain the environmental data around the vehicle from the second perspective output by the diffusion model, that is, the new perspective synthesis data, where the second perspective is different from the first perspective; renders the new perspective synthesis data to obtain the image corresponding to the second perspective after rendering, that is, the new perspective synthesis image.
[0048] In the embodiment of the present application, the above solution takes the execution of the autonomous driving system of the vehicle as an example. In some embodiments, the data collected by the autonomous driving system of the vehicle can also be uploaded to the server for the server to execute.
[0049] Exemplarily, please refer to Figure 2 , Figure 2 is a flowchart of an image acquisition method provided by an embodiment of the present application. Figure 2 The image acquisition method shown can be executed by the autonomous driving system of the vehicle. For example, the vehicle can be the vehicle 100 shown above Figure 1 , and the autonomous driving system can be the autonomous driving system 100a shown Figure 1 .
[0050] As Figure 2 shown, the above image acquisition method may include step 210, step 220, step 230, step 240, step 250, step 260, and step 270, and the specific implementation is as follows.
[0051] Step 210: Collect the point cloud data around the vehicle from the first perspective through the sensors of the vehicle.
[0052] Among them, the sensors of the above vehicle are electronic devices used to collect data in and around the vehicle. The sensors of the vehicle include, but are not limited to, lidar, cameras, and millimeter-wave radars. The lidar obtains the distance information between the object and the vehicle by emitting laser beams and measuring the reflection time of the laser beams; the camera captures two-dimensional images of the vehicle environment; the millimeter-wave radar uses radio waves to detect the distance, speed, and angle of the object.
[0053] Among them, the above point cloud data refers to a data set of a series of discrete points in physical space collected by lidar. Each point contains the coordinate information of the target object in three-dimensional space and other attributes of the target object, such as color information and reflection intensity.
[0054] In some embodiments, the autonomous driving system collects the point cloud data around the vehicle from the first perspective through the lidar of the vehicle.
[0055] In the embodiments of the present application, the lidar of the vehicle can quickly and accurately collect the point cloud data around the vehicle from the first perspective, improving the collection efficiency of the point cloud data.
[0056] In some embodiments, the autonomous driving system respectively collects the video data around the vehicle and the point cloud data around the vehicle from the first perspective through at least two sensors of the vehicle; performs fusion processing on the video data and the point cloud data to obtain target point cloud data; the fusion processing includes, but is not limited to, data registration, feature matching, and error correction, and the target point cloud data is used to perform subsequent new perspective synthetic image generation operations.
[0057] In the embodiments of the present application, the autonomous driving system can respectively input the video data and the point cloud data into the fusion model, and obtain the fused point cloud data output by the fusion model as the target point cloud data.
[0058] The training method of the above fusion model is as follows: the autonomous driving system obtains the pre-prepared video data samples, point cloud data samples, and labeled point cloud data. The labeled point cloud data is complete point cloud data, and the point cloud data sample is incomplete point cloud data obtained by removing part of the data from the labeled point cloud data; the autonomous driving system inputs the video data samples and the point cloud data samples into the fusion model, obtains the predicted point cloud data output by the fusion model, calculates the loss function value through the difference between the predicted point cloud data and the labeled point cloud data, and updates the parameters of the fusion model through the loss function value until the parameters of the fusion model converge.
[0059] In the embodiments of the present application, by performing fusion processing on video data and point cloud data collected by different sensors, data collected by other sensors can be used as auxiliary data to further process and correct the content of the point cloud data, so as to improve the quality of the point cloud data.
[0060] Step 220: Convert the point cloud data into a Gaussian distribution form.
[0061] Among them, the above Gaussian distribution is a continuous probability distribution model used to describe the probability density function of a random variable.
[0062] In the embodiments of the present application, the autonomous driving system can apply KDE (Kernel Density Estimation) or other statistical methods to convert the point cloud data into a Gaussian distribution form.
[0063] Step 230: Extract the first geometric feature and the first dynamic surface feature from the point cloud data in Gaussian distribution form.
[0064] Among them, the above first geometric feature refers to the basic geometric attributes of an object, such as edges, corners, and planes.
[0065] Among them, the above first dynamic surface feature describes the characteristics of the object surface changing over time, such as the boundary contour of a moving object.
[0066] In some embodiments, the autonomous driving system can extract the first geometric feature and the first dynamic surface feature from the point cloud data in Gaussian distribution form through a feature extraction model. The feature extraction model is a machine learning model capable of extracting the first geometric feature and the first dynamic surface feature from the point cloud data in Gaussian distribution form.
[0067] In the embodiments of the present application, the training method of the above feature extraction model is as follows: The autonomous driving system obtains the pre-prepared point cloud data samples and labeled features. The labeled features include the labeled first geometric feature and the labeled first dynamic surface feature; input the point cloud data samples into the feature extraction model to obtain the predicted features output by the feature extraction model. The predicted features include the predicted first geometric feature and the predicted first dynamic surface feature; calculate the loss function value through the difference between the predicted features and the labeled features, and perform parameter update on the feature extraction model through the loss function value until the parameters of the feature extraction model converge.
[0068] Step 240: Render the first geometric feature and the first dynamic surface feature through three-dimensional Gaussian sputtering technology to obtain the rendered three-dimensional model data.
[0069] Among them, the above three-dimensional Gaussian sputtering technology is a rendering technology used to convert geometric features and dynamic surface features into realistic three-dimensional model data.
[0070] Step 250: Extract the feature data of the 3D model data.
[0071] In the embodiment of the present application, the autonomous driving system can extract the feature data from the 3D model data through a feature extraction model; the feature extraction network is a machine learning model with the ability to extract feature data from the 3D model data.
[0072] Step 260: Input the feature data of the 3D model data into the diffusion model to obtain the synthesized data of the new perspective output by the diffusion model; the synthesized data of the new perspective is the environmental data around the vehicle from the second perspective; the second perspective is different from the first perspective.
[0073] Among them, the above diffusion model is a machine learning model for processing and transforming the features of 3D model data. In the embodiment of the present application, the above diffusion model is used to generate the synthesized data from a new perspective by simulating the diffusion process of the feature data of the 3D model data in space.
[0074] Step 270: Render the synthesized data of the new perspective to obtain the rendered synthesized image of the new perspective; the synthesized image of the new perspective is the image corresponding to the second perspective.
[0075] Among them, the image corresponding to the above second perspective is an image with a perspective different from the acquisition image of the vehicle's sensor.
[0076] In the embodiment of the present application, the autonomous driving system provides high-precision spatial information through the point cloud data from the first perspective collected by the vehicle's sensor. By using the three-dimensional Gaussian sputtering technology to render the first geometric feature and the first dynamic surface feature obtained from the point cloud data, the feature data of the generated 3D model data can simulate complex appearance characteristics while retaining details, enhancing the realism of the 3D model. Input the feature data of the 3D model data into the diffusion model, and through the generation ability of the diffusion model, synthesize the environmental data of different perspectives, breaking through the limitations of traditional single-perspective environmental data. Further render the synthesized data of the new perspective to generate the synthesized image of the new perspective from other perspectives, improving the acquisition efficiency of the synthesized image of the new perspective.
[0077] Based on the solutions shown in any one or more of the above embodiments of the present application, please refer to Figure 3 , Figure 3 is a flowchart of an image acquisition method provided by an embodiment of the present application. Before the above step 220, it includes step 214 and step 218, which are specifically as follows.
[0078] Step 214: Collect the video data of the vehicle through the vehicle's sensor.
[0079] Among them, the above video data is the video data of the surrounding environment of the vehicle or the vehicle itself collected by sensors during the operation of the vehicle.
[0080] In the embodiments of the present application, multiple cameras can be installed at different positions of the vehicle, such as the front windshield, the rearview mirror, and the roof, to ensure that as large a field of view as possible is covered to collect the video data of the vehicle.
[0081] Step 218: Perform a preprocessing operation on the point cloud data based on the video data.
[0082] In the embodiments of the present application, the above preprocessing operation is a preliminary processing of the point cloud data to remove noise and fill in missing data.
[0083] In the embodiments of the present application, the above video data provides rich visual information. Performing a preprocessing operation on the point cloud data based on the video data can supplement the texture details of the point cloud data through the content of the video data, and improve the quality of the preprocessed point cloud data.
[0084] Based on the solutions shown in any one or more of the above embodiments of the present application, the above step 218 can be implemented as: performing a time synchronization operation and a spatial alignment operation on the video data and the point cloud data.
[0085] Among them, the above time synchronization operation is to ensure that the data collected by different sensors in the video data and the point cloud data are consistent on the time axis.
[0086] In the embodiments of the present application, each piece of data collected by a sensor is attached with an accurate time stamp for recording the data collection moment.
[0087] In the embodiments of the present application, the autonomous driving system can identify and correct the time deviation caused by the internal processing delay of the sensor through a time stamp matching algorithm, and implement a time synchronization operation on the video data and the point cloud data. For example, the autonomous driving system can use interpolation or translation methods to compensate for the time-misaligned data, so that each frame of video data can be accurately matched with the corresponding point cloud data.
[0088] Among them, the above spatial alignment is to convert the data from different sensors into the same coordinate system to ensure that their spatial correspondence is accurate.
[0089] In the embodiments of the present application, the autonomous driving system can perform an initial calibration. By using a standard geometric target or a reference object with a known size, determine the position and orientation of each sensor relative to the vehicle coordinate system, and use the iterative closest point algorithm or other registration methods to match the point cloud data with the feature points extracted from the video frame, and gradually optimize the transformation matrix to make the geometric structures of the video data and the point cloud data as aligned as possible.
[0090] In the embodiments of the present application, the time synchronization operation ensures the consistency of video data and point cloud data on the time axis, eliminates the time deviation that may be caused by different sensor sampling rates, enables each frame of video data to be accurately matched with the corresponding point cloud data, and ensures the position consistency of moving objects in different data sources; the spatial alignment operation unifies the video data and the point cloud data into the same coordinate system through a coordinate transformation matrix, solving the problems of position and attitude differences between sensors; through the above preprocessing operations, it can effectively ensure that the preprocessed point cloud data is consistent with the video data in content and ensure the accuracy of the point cloud data.
[0091] Based on the solutions shown in any one or more of the above embodiments of the present application, step 270 above can be implemented as: extracting a second geometric feature and a second dynamic surface feature from the newly-viewpoint synthesized data; based on the second geometric feature and the second dynamic surface feature, performing high-precision lighting calculation, texture optimization, and dynamic object synthesis to obtain a rendered newly-viewpoint synthesized image.
[0092] Among them, the above second geometric feature refers to the geometric attributes of an object recognized and extracted from a new viewpoint, such as shape, size, and position.
[0093] Among them, the above second dynamic surface feature is the characteristic of the object surface changing over time observed from a new viewpoint, such as the contour of a moving object.
[0094] In the embodiments of the present application, the autonomous driving control system can extract a second geometric feature and a second dynamic surface feature from the newly-viewpoint synthesized data through a feature extraction network; the feature extraction network is a neural network capable of extracting a second geometric feature and a second dynamic surface feature from the newly-viewpoint synthesized data.
[0095] Among them, the above high-precision lighting calculation is a calculation technology used to simulate the propagation and interaction of light in a three-dimensional scene, and is used to generate realistic shadows, reflections, refractions, and other optical effects. High-precision lighting calculation provides high-quality visual effects and enhances the realism of 3D models by accurately calculating the influence of each light source on each object in the scene and the multiple reflections and scattering of light between these objects.
[0096] In the embodiments of the present application, the autonomous driving system constructs a three-dimensional scene representation based on the extracted second geometric feature and second dynamic surface feature, ensures that the position, shape, and material attributes of each object are accurately described, and applies a path tracing algorithm to simulate the propagation path of light in the three-dimensional scene, thereby generating high-quality shadow and reflection effects.
[0097] Among them, the above texture optimization is a technology for enhancing the surface details and realism of 3D models. Texture optimization ensures the best visual effects under different viewing angles and lighting conditions by adjusting and improving the quality of texture images.
[0098] In the embodiments of the present application, the autonomous driving system analyzes the original texture, identifies the defects of the original texture, uses image processing technology to optimize the defects of the original texture, and increases the detail clarity of the texture, ensuring high-quality visual effects even in enlarged scenes.
[0099] Among them, the above dynamic object synthesis is a technology for generating and simulating dynamic objects in 3D scenes. Dynamic object synthesis combines geometric features, dynamic surface features, and physical behavior models to generate realistic motion effects and interaction behaviors, enabling objects in the virtual environment to move, deform, and interact naturally like in the real world.
[0100] In the embodiments of the present application, the autonomous driving system creates a skeleton structure for deformable objects (such as humans or animals) based on the second geometric feature and the second dynamic surface feature using skeletal animation technology, binds the skeleton structure to the 3D model, and drives the model to produce natural and smooth motion effects by manipulating the position and rotation parameters of the skeleton nodes.
[0101] In some embodiments, the autonomous driving system extracts the second geometric feature, the second dynamic surface feature, and the environmental feature from the synthesized data of the new view; based on the second geometric feature, the second dynamic surface feature, and the environmental feature, performs high-precision lighting calculation, texture optimization, and dynamic object synthesis to obtain the synthesized image of the new view after rendering.
[0102] In the embodiments of the present application, the above environmental feature is a feature indicating the environment where the vehicle is located in the synthesized data of the new view, and the environmental feature includes but is not limited to weather features (such as cloudy days, rainy days) and time features (such as morning, late at night).
[0103] In the embodiments of the present application, during the lighting calculation stage, the autonomous driving system adjusts the direction, intensity, and color temperature of the light source according to the environmental feature to simulate the lighting effects under different weather and time periods; during the texture optimization stage, uses the environmental feature to enhance the realism of the material under different environmental conditions. For example, in a rainy environment, increases the reflection effect of the wet road surface on light; during the dynamic object synthesis stage, on the basis of considering the geometric shape and motion trajectory of the object, incorporates environmental factors. For example, in a rainy environment, the effect of water splashing when the vehicle is driving. Finally, renders the above features through rendering technology to obtain the synthesized image of the new view.
[0104] In the embodiments of the present application, in addition to geometric features and dynamic surface features, the autonomous driving system combines environmental features to make the synthesized new perspective image after rendering consider the display environmental factors, making it more realistic and improving the quality of the synthesized new perspective image.
[0105] In the embodiments of the present application, high-precision lighting calculation can enhance the visual realism of the synthesized new perspective image. Texture optimization ensures a more realistic material performance of the synthesized new perspective image by finely adjusting the surface details of objects. Dynamic object synthesis can make the motion effects more smooth; the second geometric features and the second dynamic surface features provide rich environmental descriptions. Performing the above rendering operation based on the second geometric features and the second dynamic surface features enables the obtained synthesized new perspective image after rendering to display more fine detail features on the basis of ensuring the correctness of the image content, greatly improving the generation quality of the synthesized new perspective image.
[0106] Based on the solutions shown in any one or more of the above embodiments of the present application, in some embodiments, the geometric features include the shape, size, and position information of the point cloud; the dynamic surface features include the motion trajectories and appearance changes of the objects around the vehicle.
[0107] Among them, the above shape is the three-dimensional form of the object, for example, a sphere or a cube.
[0108] Among them, the above size is the scale of the object in each dimension, for example, length, width, and height.
[0109] Among them, the above position information is the coordinate of the object in space, usually the position of the object relative to a certain reference point or coordinate system.
[0110] Among them, the above motion trajectory is the path that the object moves within a period of time, and the motion trajectory can be composed of continuous timestamps and position coordinates.
[0111] Among them, the above appearance change is the change that occurs on the surface of the object over time, for example, color change, texture change, or shape deformation.
[0112] In the embodiments of the present application, a variety of specific geometric features and dynamic surface features are extended. Through more refined geometric features, it is convenient for the autonomous driving system to accurately understand the static structure in the environment. The dynamic surface features include not only the motion trajectories of objects but also the appearance changes of objects, enabling the autonomous driving system to track and predict the behavior of objects in real time in a complex dynamic environment and enhancing the understanding ability of the scene; combining more detailed geometric features and dynamic surface features can obtain more detailed features from more angles. These detailed features are beneficial for the subsequent autonomous driving system to generate more realistic and detailed three-dimensional models and perform high-precision synthesis and rendering from different perspectives, improving the generation quality of the synthesized new perspective image.
[0113] Referring to the solutions shown in any one or more of the above embodiments of the present application, please refer to Figure 4 , Figure 4 which is a flowchart of an image acquisition method provided by an embodiment of the present application. Based on the Figure 2 image acquisition method shown, the method further includes step 280, which is specifically as follows.
[0114] Step 280: Synthesize an image based on a new perspective, perform perception enhancement training on the vehicle, and conduct a backfill simulation experiment; the perception enhancement training includes object detection, object tracking, and driving scenario understanding of the vehicle's autonomous driving system; the backfill simulation experiment includes performance testing and performance optimization of the vehicle's autonomous driving algorithm in a virtual environment.
[0115] Among them, the above object detection is used to identify and locate specific objects in the image, such as pedestrians and vehicles.
[0116] In the embodiment of the present application, the autonomous driving system can use the synthesized image from the new perspective as training data, input it into the object detection model, and implement the training of the object detection model, so that the trained object detection model can perform object detection from various perspectives to achieve the object detection training of the autonomous driving system.
[0117] Among them, the above object tracking is used to continuously track the detected objects to understand their movement trajectories and behaviors.
[0118] In the embodiment of the present application, the autonomous driving system uses a deep learning model to identify and locate the target objects (such as pedestrians) in the synthesized image from the new perspective, calculates the displacement and speed of the target objects through motion estimation technology, and combines a Kalman filter for trajectory smoothing and prediction. The multi-sensor fusion technology combines point cloud data and video data to correct the target position in the two-dimensional image to ensure high-precision tracking, so as to achieve the object tracking training of the autonomous driving system.
[0119] Among them, the above driving scenario understanding is a high-level understanding of the entire driving environment by the autonomous driving system, and the understanding content includes but is not limited to road structures, traffic signs, and other dynamic elements.
[0120] In the embodiment of the present application, the autonomous driving system can use a semantic segmentation algorithm to perform per-pixel classification on the synthesized image from the new perspective, identify environmental elements such as roads and buildings, introduce a graph neural network to analyze the interactions between different target objects, provide in-depth scene understanding, combine time series data, and use a recurrent neural network to process consecutive frames, capture the changes in the scene over time, and predict future states, so as to achieve the driving scenario understanding training of the autonomous driving system.
[0121] Among them, the above performance test is a test used to evaluate the performance of the autonomous driving algorithm in a virtual environment, and the test content includes but is not limited to the accuracy, response time, and stability of the autonomous driving algorithm.
[0122] In the embodiment of the present application, the autonomous driving system can construct a highly simulated virtual environment, which includes various road types, traffic conditions, and weather conditions. Taking the newly synthesized perspective image as input data, it generates a virtual scene with rich details to ensure it is as consistent as possible with the actual driving environment. Running the autonomous driving algorithm in this virtual environment, simulating the vehicle's behavior in different scenarios, automatically recording and analyzing various performance indicators of the autonomous driving algorithm, so as to realize the performance test of the vehicle's autonomous driving algorithm in the virtual environment.
[0123] Among them, the above performance optimization is an operation of adjusting algorithm parameters according to the test results to improve the performance of the autonomous driving algorithm.
[0124] In the embodiment of the present application, based on the results of the performance test, the autonomous driving system identifies the deficiencies in the autonomous driving algorithm, adopts reinforcement learning to optimize the parameters of the autonomous driving algorithm, so that the improved autonomous driving algorithm with optimized parameters performs better in the virtual environment, in order to achieve the performance optimization of the autonomous driving algorithm in the virtual environment.
[0125] In the embodiment of the present application, the newly synthesized perspective image is an image of the vehicle from different perspectives. Through the newly synthesized perspective image, the characteristic information of the vehicle in different environments can be obtained. Based on the characteristic information of the vehicle in different environments, performing perception enhancement training and backpropagation simulation experiments can improve the ability of the vehicle's autonomous driving system in complex environments and further enhance the training effect of the vehicle's autonomous driving system.
[0126] In some embodiments, the autonomous driving system includes a perception model. After obtaining the newly synthesized perspective image, the autonomous driving system can perform enhancement training on the perception model in the autonomous driving system through the newly synthesized perspective image to improve the perception function of the autonomous driving system; any vehicle installed with this autonomous driving system can obtain this newly synthesized perspective image and realize road perception in different scenarios based on this newly synthesized perspective image, so as to achieve the full-platform sharing of autonomous driving experience.
[0127] Based on the above Figures 2 to 4 steps in the embodiments, the embodiment of the present application shows an autonomous driving simulation three-dimensional reconstruction technology.
[0128] At present, when new models are launched by automobile enterprises, hundreds of thousands of kilometers of testing are required, the testing and development costs reach more than tens of millions, and dozens of vehicles are required for testing and verification at the same time. The verification cycle reaches several months, resulting in relatively high costs and time consumption. The purpose of this autonomous driving simulation 3D reconstruction technology is to reuse the collected real vehicle data, generate high-fidelity sensor data for different vehicle configurations, support the perception enhancement training and back-injection simulation verification of the autonomous driving system, reduce the sampling cost, reduce the number of samplings, and accelerate the development and verification process.
[0129] The technical solution shown in the embodiments of this application is based on AIGC (AI-Generated Content) and 3D (Three–Dimensional) Gaussian sputtering technology. Through million-level image fine-tuning training, high-precision static scene reconstruction and dynamic object synthesis are achieved, and high-quality visual and lidar fusion data are generated for perception enhancement training and back-injection simulation verification. First, for a certain vehicle model, based on the video and point cloud data collected in real time by the sensors installed on the vehicle body, the collected data is processed. The 3D points are represented as Gaussian distributions, and geometric features, dynamic surface features, etc. are extracted for rendering. After rendering, the data feature data is output. The feature data is processed in combination with the diffusion model, and the AIGC technology is used to generate enhanced new perspective synthesis data. The synthesis data is further processed, features are extracted, and higher-quality scene reconstruction and new perspective synthesis are rendered and output.
[0130] Exemplarily, please refer to Figure 5 , Figure 5 which is a flowchart of the method for an autonomous driving simulation 3D reconstruction technology in this application. This technology includes steps S1, S2, S3, S4, S5, S6, S7, and S8. This method is executed by the autonomous driving system of the vehicle, and is specifically as follows.
[0131] S1: Data collection
[0132] In the embodiments of this application, multiple sensors are installed on the vehicle of the target vehicle model, and vehicle video data and the point cloud data around the vehicle are obtained through real vehicle road sampling by the sensors; the sensors include but are not limited to cameras and lidars.
[0133] S2: Data preprocessing
[0134] In the embodiments of this application, the autonomous driving system preprocesses the collected video data and point cloud data. The preprocessing includes but is not limited to denoising, alignment, and format conversion to ensure the accuracy of the point cloud data and the consistency between the content of the point cloud data and the video content.
[0135] S3: Data processing
[0136] In the embodiments of the present application, the autonomous driving system represents 3D point cloud data as a Gaussian distribution, and extracts geometric features and dynamic surface features from the point cloud data under the Gaussian distribution; the geometric features include the shape, size, and position information of the point cloud, and the dynamic surface features include the movement trajectory and surface appearance changes of the object.
[0137] S4: Rendering
[0138] In the embodiments of the present application, the autonomous driving system uses 3D Gaussian sputtering technology to render the processed point cloud data to generate preliminary 3D model data; the rendering process includes steps such as lighting calculation, texture mapping, and surface smoothing.
[0139] S5: Feature Extraction
[0140] In the embodiments of the present application, the autonomous driving system outputs the rendered feature data as the basis for subsequent processing.
[0141] S6: AIGC Enhancement Processing
[0142] In the embodiments of the present application, the autonomous driving system inputs the rendered feature data into a diffusion model and uses AIGC technology to generate enhanced new perspective synthesis data; the diffusion model enhances the data through deep learning algorithms to generate new perspective synthesis data with higher resolution and more details.
[0143] S7: Feature Extraction and Re-rendering
[0144] In the embodiments of the present application, the autonomous driving system performs further processing on the generated new perspective synthesis data to extract finer geometric features and dynamic surface features from the new perspective synthesis data, and uses the extracted geometric features and dynamic surface features for re-rendering to generate a new perspective synthesis image of higher quality; the re-rendering process includes high-precision lighting calculation, texture optimization, and dynamic object synthesis.
[0145] S8: Data Output and Application
[0146] In the embodiments of the present application, the autonomous driving system outputs the finally generated new perspective synthesis image for perception enhancement training and backfilling simulation verification of the vehicle's autonomous driving system; the perception enhancement training includes tasks such as object detection, tracking, and scene understanding of the autonomous driving system, and the backfilling simulation verification includes performance testing and optimization of the autonomous driving algorithm in a virtual environment.
[0147] Exemplarily, based on Figure 5 , please refer to Figure 6 , Figure 6 is a schematic flow diagram of an autonomous driving simulation three-dimensional reconstruction technology of the present application. As Figure 6 shown, there are two stages.
[0148] Stage 1 includes steps S1, S2, S3, and S4:
[0149] In an embodiment of the present application, an image acquisition device and a lidar point can be installed on a vehicle. The vehicle's autonomous driving system acquires information data of the vehicle from the first perspective through the vehicle's image acquisition device and lidar point, inputs the information into an encoder, and obtains geometric features and appearance features in the information data through the encoder, and renders the geometric features and appearance features to generate preliminary 3D model data.
[0150] Stage 2 includes steps S5, S6, S7, and S8:
[0151] In an embodiment of the present application, the autonomous driving system can extract feature data from the preliminary 3D model data, input the extracted features into a diffusion model, obtain more refined and higher-quality new perspective synthesis data output by the diffusion model, and perform re-selection on the new perspective synthesis data to obtain multiple new perspective synthesis images from different perspectives.
[0152] In an embodiment of the present application, the above new perspective synthesis images can be uploaded to the perception model in the autonomous driving system for enhanced training to improve the perception function of the autonomous driving system; any vehicle equipped with this autonomous driving system can obtain the new perspective synthesis images, and based on these new perspective synthesis images, road perception can be achieved in different scenarios to realize the full-platform sharing of autonomous driving experience.
[0153] Exemplarily, Vehicle A obtains new perspective synthesis images through the above autonomous driving simulation three-dimensional reconstruction technology, uploads the new perspective synthesis images to the perception model in the autonomous driving system for enhanced training to improve the perception function of the autonomous driving system; through data migration, when Vehicle B is installed with the same above autonomous driving system as Vehicle A, Vehicle B can obtain the new perspective synthesis images uploaded by Vehicle A in this autonomous driving system.
[0154] Based on Figure 5 and Figure 6 the autonomous driving simulation three-dimensional reconstruction technology shown, the above solutions can achieve the following beneficial effects:
[0155] 1) Joint reconstruction of multi-modal data, empowered by high-quality lidar simulation, improving the perception accuracy of the autonomous driving system.
[0156] 2) Using AIGC to improve the data synthesis quality of 3D Gaussian sputtering, with key technical indicators leading the world.
[0157] 3) Reusing data assets, saving the cost of acquisition and annotation, and improving the development efficiency of autonomous driving.
[0158] The above solution saves the vehicle acquisition and annotation costs through data asset reuse, and improves the efficiency of autonomous driving development.
[0159] Please refer to Figure 7 , which shows a block diagram of an image acquisition device provided by an exemplary embodiment of the present application. The image acquisition device can be implemented as all or part of a computer device in a hardware or a combination of hardware and software manner to implement all or part of the steps in the embodiment as described above Figures 2 to 4 shown. As Figure 7 shown, the image acquisition device includes:
[0160] A point cloud data acquisition module 701, configured to acquire point cloud data around the vehicle from a first perspective through a sensor of the vehicle;
[0161] A point cloud data conversion module 702, configured to convert the point cloud data into a Gaussian distribution form;
[0162] A feature extraction module 703, configured to extract a first geometric feature and a first dynamic surface feature from the point cloud data in the Gaussian distribution form;
[0163] A feature rendering model 704, configured to render the first geometric feature and the first dynamic surface feature through a three-dimensional Gaussian sputtering technique to obtain rendered three-dimensional model data;
[0164] A feature data extraction module 705, configured to extract feature data of the three-dimensional model data;
[0165] A new perspective synthesis data acquisition module 706, configured to input the feature data of the three-dimensional model data into a diffusion model to obtain new perspective synthesis data output by the diffusion model; the new perspective synthesis data is environmental data around the vehicle from a second perspective; the second perspective is different from the first perspective;
[0166] A new perspective synthesis image acquisition module 707, configured to render the new perspective synthesis data to obtain a rendered new perspective synthesis image; the new perspective synthesis image is an image corresponding to the second perspective.
[0167] In some embodiments, the device further includes:
[0168] A video data acquisition module, configured to acquire video data of the vehicle through a sensor of the vehicle before converting the point cloud data into a Gaussian distribution form;
[0169] A preprocessing operation execution module, configured to perform a preprocessing operation on the point cloud data based on the video data before converting the point cloud data into a Gaussian distribution form.
[0170] In some embodiments, the preprocessing operation execution module is configured to perform a time synchronization operation and a spatial alignment operation on the video data and the point cloud data.
[0171] In some embodiments, the new perspective synthesis image acquisition module 707 is configured to extract a second geometric feature and a second dynamic surface feature from the new perspective synthesis data; and perform high-precision illumination calculation, texture optimization, and dynamic object synthesis based on the second geometric feature and the second dynamic surface feature to obtain a rendered new perspective synthesis image.
[0172] In some embodiments, the geometric feature includes the shape, size, and position information of the point cloud; and the dynamic surface feature includes the movement trajectory and appearance change of the objects around the vehicle.
[0173] In some embodiments, the apparatus further includes:
[0174] An image application module, configured to perform perception enhancement training and backfill simulation experiments on the vehicle based on the new perspective synthesis image; the perception enhancement training includes object detection, object tracking, and driving scene understanding of the vehicle's autonomous driving system; and the backfill simulation experiment includes performance testing and performance optimization of the vehicle's autonomous driving algorithm in a virtual environment.
[0175] Please refer to Figure 8 , Figure 8 , which is a schematic structural diagram of a computer device provided by an exemplary embodiment of the present application. The computer device 800 includes a central processing unit (CPU) 801, a system memory 804 including a random access memory (RAM) 802 and a read-only memory (ROM) 803, and a system bus 805 connecting the system memory 804 and the central processing unit 801. The computer device 800 further includes a basic input / output system (I / O system) 806 for facilitating information transmission between various components within the computer, and a mass storage device 807 for storing an operating system 813, application programs 814, and other program modules 815.
[0176] The basic input / output system 806 includes a display 808 for displaying information and input devices 809 such as a mouse and a keyboard for user input of information. The display 808 and the input devices 809 are both connected to the central processing unit 801 through an input / output controller 810 connected to the system bus 805. The basic input / output system 806 may further include an input / output controller 810 for receiving and processing inputs from multiple other devices such as a keyboard, a mouse, or an electronic stylus. Similarly, the input / output controller 810 also provides output to a display screen, a printer, or other types of output devices.
[0177] The mass storage device 807 is connected to the central processing unit 801 through a mass storage controller (not shown) connected to the system bus 805. The mass storage device 807 and its associated computer-readable medium provide non-volatile storage for the computer device 800. That is, the mass storage device 807 may include a computer-readable medium (not shown) such as a hard disk or a CD-ROM (Compact Disc Read-Only Memory) drive.
[0178] Without loss of generality, the computer-readable medium may include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes RAM (Random Access Memory), ROM (Read-Only Memory), EPROM (Erasable Programmable Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), flash memory or other solid-state storage technologies, CD-ROM, DVD (Digital Video Disc) or other optical storage, magnetic tape cartridges, tapes, disk storage or other magnetic storage devices. Of course, those skilled in the art will know that computer storage media is not limited to the above several. The above-mentioned system memory 804 and mass storage device 807 may be collectively referred to as memory.
[0179] The computer device 800 may be connected to the Internet or other network devices through a network interface unit 811 connected to the system bus 805.
[0180] The memory further includes one or more programs. The one or more programs are stored in the memory, and the central processing unit 801 implements Figures 2 to 4 all or part of the steps in the method shown.
[0181] In an exemplary embodiment, a chip is further provided. The chip includes programmable logic circuits and / or program instructions, which are used to implement all or part of the steps of the methods shown in the above various embodiments of the present application when the chip runs on a computer device.
[0182] In an exemplary embodiment, a computer program product is further provided. The computer program product includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor reads and executes the computer instructions to implement all or part of the steps of the methods shown in the foregoing various embodiments of the present application.
[0183] In an exemplary embodiment, a computer-readable storage medium is further provided. A computer program is stored in the computer-readable storage medium, and the computer program is loaded and executed by a processor to implement all or part of the steps of the methods shown in the foregoing various embodiments of the present application.
[0184] Those of ordinary skill in the art can understand that all or part of the steps of implementing the foregoing embodiments can be completed by hardware, or can be completed by a program instructing related hardware. The foregoing program can be stored in a computer-readable storage medium, and the storage medium mentioned above can be a read-only memory, a magnetic disk, an optical disk, or the like.
[0185] Those skilled in the art should be able to realize that in the above one or more examples, the functions described in the embodiments of the present application can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. The computer-readable medium includes a computer storage medium and a communication medium, where the communication medium includes any medium that facilitates the transfer of a computer program from one place to another. The storage medium can be any available medium accessible by a general-purpose or special-purpose computer.
[0186] The foregoing are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. An image acquisition method, characterized in that, The method includes: Collecting point cloud data around the vehicle from a first perspective through sensors of the vehicle; Converting the point cloud data into a Gaussian distribution form; Extracting a first geometric feature and a first dynamic surface feature from the point cloud data in the Gaussian distribution form; Rendering the first geometric feature and the first dynamic surface feature through a three-dimensional Gaussian sputtering technique to obtain rendered three-dimensional model data; Extracting feature data of the three-dimensional model data; Inputting the feature data of the three-dimensional model data into a diffusion model to obtain new perspective synthesis data output by the diffusion model; the new perspective synthesis data is environmental data around the vehicle from a second perspective; the second perspective is different from the first perspective; Rendering the new perspective synthesis data to obtain a rendered new perspective synthesis image; the new perspective synthesis image is an image corresponding to the second perspective.
2. The method according to claim 1, characterized in that, Before converting the point cloud data into a Gaussian distribution form, the method further includes: Collecting video data of the vehicle through sensors of the vehicle; Performing a preprocessing operation on the point cloud data based on the video data.
3. The method according to claim 2, wherein The performing a preprocessing operation on the point cloud data based on the video data includes: Performing a time synchronization operation and a spatial alignment operation on the video data and the point cloud data.
4. The method according to claim 1, wherein The rendering the new perspective synthesis data to obtain a rendered new perspective synthesis image includes: Extracting a second geometric feature and a second dynamic surface feature from the new perspective synthesis data; Performing high-precision lighting calculation, texture optimization, and dynamic object synthesis based on the second geometric feature and the second dynamic surface feature to obtain a rendered new perspective synthesis image.
5. The method according to claim 1, wherein The geometric feature includes the shape, size, and position information of the point cloud; the dynamic surface feature includes the movement trajectory and appearance change of objects around the vehicle.
6. The method according to claim 1, characterized in that, The method further includes: Performing perception enhancement training and backpropagation simulation experiments on the vehicle based on the new perspective synthesis image; The perception enhancement training includes object detection, object tracking, and driving scene understanding of the vehicle's autonomous driving system; the backpropagation simulation experiment includes performance testing and performance optimization of the vehicle's autonomous driving algorithm in a virtual environment.
7. An image acquisition device, characterized in that, The device includes: A point cloud data collection module for collecting point cloud data around the vehicle from a first perspective through sensors of the vehicle; A point cloud data conversion module for converting the point cloud data into a Gaussian distribution form; A feature extraction module for extracting a first geometric feature and a first dynamic surface feature from the point cloud data in the Gaussian distribution form; A feature rendering model for rendering the first geometric feature and the first dynamic surface feature through a three-dimensional Gaussian sputtering technique to obtain rendered three-dimensional model data; A feature data extraction module for extracting feature data of the three-dimensional model data; A new perspective synthetic data acquisition module for inputting the feature data of the 3D model data into a diffusion model to obtain new perspective synthetic data output by the diffusion model; the new perspective synthetic data is environmental data around the vehicle from a second perspective; the second perspective is different from the first perspective. A new perspective synthetic image acquisition module for rendering the new perspective synthetic data to obtain a rendered new perspective synthetic image; the new perspective synthetic image is an image corresponding to the second perspective.
8. A computer device, characterized in that, The computer device includes a processor and a memory, and instructions are stored in the memory, and the instructions are executed by the processor to implement the image acquisition method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, Instructions are stored in the storage medium, and the instructions are executed by the processor of the computer device to implement the image acquisition method according to any one of claims 1 to 6.
10. A computer program product, characterized in that, The computer program product includes computer instructions, and the computer instructions are stored in a computer-readable storage medium; the computer instructions are read and executed by the processor of the computer device to implement the image acquisition method according to any one of claims 1 to 6.