Method and device for generating synthesized scene data and model training method
Through three-dimensional reconstruction technology, multi-view image data and object point cloud data are combined to generate synthetic scene data that is adapted to different target scenarios, solving the problems of insufficient comprehensive data synthesis and legal restrictions in the existing technology, and achieving high-quality data generation and optimization of artificial intelligence models.
Patent Information
- Application Number
- CN202311819482.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-27
- Publication Date
- 2025-06-27
AI Technical Summary
It is difficult for the prior art to generate data that can fully reflect the application scenarios of the model, especially when the data collection process is subject to legal and regulatory restrictions, and GAN-based data synthesis methods have limitations that rely heavily on input data and single-modal processing.
By performing three-dimensional reconstruction based on multi-view image data and object point cloud data, the reconstructed three-dimensional object model is generated and combined with the target background model to generate synthetic scene data that is suitable for different target scenarios.
The generated synthetic scene data can accurately and comprehensively represent the target scene, provide color and position/geometric information, be suitable for training and optimization of artificial intelligence models, and be able to desensitize processing to avoid the emergence of sensitive information.
Smart Images

Figure CN120219604A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the field of computer vision, and more particularly to methods and apparatuses for generating synthetic scene data, and model training methods. Background Art
[0002] In the field of computer vision, data synthesis technology is developing rapidly. Data synthesis refers to the process of generating data through computer algorithms. The synthesized data can be used in various application scenarios, among which the application of the synthesized data in the field of artificial intelligence has received extensive attention. The performance of an artificial intelligence model highly depends on the data used for training and optimizing the model. For this reason, data synthesis technology has been proposed to obtain more training data in order to improve the performance of the artificial intelligence model. Summary of the Invention
[0003] The present disclosure provides an improved mechanism for generating synthetic scene data. The proposed synthetic scene data generation mechanism can generate synthetic scene data adapted to different target scenes based on real image data and point cloud data obtained from real scenes.
[0004] According to one aspect of the present disclosure, there is provided a method for generating synthetic scene data, including: obtaining multi-view image data for a real scene; object point cloud data for an object in the real scene; performing 3D reconstruction based on the multi-view image data and the object point cloud data to generate a reconstructed 3D object model for the object; and generating the synthetic scene data based on the reconstructed 3D object model.
[0005] According to another aspect of the present disclosure, there is provided an apparatus for generating synthetic scene data, including: a memory and a processor. The processor is coupled to the memory and is configured to execute the method according to any one of the various embodiments of the present disclosure.
[0006] According to still another aspect of the present disclosure, there is provided a computer-readable medium storing a computer program including instructions, the instructions, when executed by a processor, causing the processor to be configured to execute the method according to any one of the various embodiments of the present disclosure.
[0007] According to still another aspect of the present disclosure, there is provided a computer program product including computer-executable instructions, the instructions, when executed, causing one or more processors to execute the method according to any one of the various embodiments of the present disclosure.
[0008] According to another aspect of the present disclosure, there is provided a model training method, including: obtaining training data; training an artificial intelligence model based on the training data to obtain a trained artificial intelligence model; wherein, the training data includes synthetic scene data generated according to the method of any one of the various embodiments of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Various embodiments of the claimed subject matter will now be described by way of example with reference to the accompanying drawings. In the different drawings, the same reference numerals are used to denote the same or similar components.
[0010] Figure 1 A schematic diagram showing the overall architecture of a synthetic scene data generation mechanism according to an example embodiment of the present disclosure is shown.
[0011] Figure 2 An example flowchart showing the operations of obtaining multi-view image data according to an example embodiment of the present disclosure is shown.
[0012] Figure 3 An example flowchart showing the operations of obtaining object point cloud data according to an example embodiment of the present disclosure is shown.
[0013] Figure 4 A schematic diagram showing a neural network architecture for reconstructing a three-dimensional object model according to an example embodiment of the present disclosure is shown.
[0014] Figure 5 An example flowchart showing the operations of generating synthetic scene data based on the reconstructed three-dimensional object according to an example embodiment of the present disclosure is shown.
[0015] Figure 6 An example flowchart showing the operations of cyclically training an artificial intelligence model based on synthetic scene data according to an example embodiment of the present disclosure is shown.
[0016] Figure 7 A schematic diagram showing a rendered scene image based on synthetic scene data according to an example embodiment of the present disclosure is shown.
[0017] Figure 8 An example flowchart showing a method for generating synthetic scene data according to an example embodiment of the present disclosure is shown.
[0018] Figure 9 A block diagram showing a device according to an example embodiment of the present disclosure, which can implement the method for generating synthetic scene data, is shown. DETAILED DESCRIPTION
[0019] In the following description, numerous specific details are set forth to provide a thorough understanding of embodiments of the present disclosure. However, those skilled in the relevant art will recognize that the present disclosure may be practiced without one or more of the specific details, or may be practiced using alternative methods, components, etc. In some instances, well-known structures, operations are not shown or described in detail so as not to unnecessarily obscure the present disclosure.
[0020] In the field of artificial intelligence, people have gradually realized the importance of data for the training and optimization of artificial intelligence models. Generally, it is required that the data should be able to reflect the data characteristics of the model application scenario as accurately and comprehensively as possible in order to improve the accuracy and generalization ability of the model.
[0021] Conventionally, real data can be collected from real scenarios as training data. However, the real data collected from real scenarios may contain sensitive information stipulated by relevant laws, regulations, etc. of a specific country or region, resulting in such real data being unable to be directly used or processed.
[0022] An alternative is that if the data collection process in a specific country or region is restricted by relevant laws, regulations, etc., then directly use the existing (e.g., open-source) data sets collected from other countries or regions as training data. However, such training data may not accurately reflect the characteristics of the data in that specific country or region, which may lead to the trained artificial intelligence model being unable to provide good performance.
[0023] To address the above difficulties faced in the data collection process, data synthesis techniques can be used to synthesize data. Commonly used data synthesis techniques mainly include simulation-based data synthesis, deep learning-based data synthesis, and so on. Simulation-based data synthesis can use a virtual engine to generate simulation data. Although simulation technology has been widely applied in fields such as games and movies, the data generated in this way cannot fully conform to real data and cannot accurately reflect the data characteristics of the model application scenario. In other words, there is a large domain gap between simulation data and real data. Deep learning-based data synthesis techniques use deep learning models to generate synthetic data. For example, the deep learning model can be a generative adversarial network (GAN), which can generate new images based on the input images. However, the synthetic data generated by the GAN-based data synthesis method is usually strongly dependent on the input data. For example, the generated new images have the same background as the input images. This results in the generated synthetic data being unable to comprehensively reflect the data characteristics of the model application scenario, which may affect the generalization ability of the artificial intelligence model. In addition, the GAN-based data synthesis method is unimodal. For example, it can usually only process image data.
[0024] The present disclosure provides an improved synthetic scene data generation mechanism. A three-dimensional object model of an object can be reconstructed based on multi-view image data of a real scene and object point cloud data of the object in the real scene. The multi-view image data provides rich color information for the object, and the object point cloud data provides accurate position / geometry information for the object. This enables the reconstructed three-dimensional object model to accurately represent the object with more details. Then, the reconstructed three-dimensional object model can be combined with any target background model to generate synthetic scene data, and the generated synthetic scene data can accurately and comprehensively represent the target scene. The improved synthetic scene data generation mechanism of the present disclosure is multi-modal, at least because it can process point cloud data and image data simultaneously, and the generated scene data contains both color information and position / geometry information. The mechanism is multi-view, at least because the point cloud data and image data it processes are obtained from different perspectives, and the generated scene data can represent the details of the target scene from different perspectives of observing the target scene.
[0025] The synthetic scene data generated using the mechanism of the present disclosure can be desensitized, that is, it does not contain sensitive image information stipulated by relevant laws, regulations, etc. Such desensitized synthetic scene data can be used to train and optimize artificial intelligence models without being restricted by relevant laws, regulations, etc.
[0026] The synthetic scene data generated using the mechanism of the present disclosure is accurate because the reconstructed three-dimensional object model can provide accurate and real detailed data about the scene object. The synthetic scene data generated using the mechanism of the present disclosure is comprehensive, and a large amount of available data for the target scene can be generated by flexibly combining the reconstructed three-dimensional object model with the target background model. Thus, using such synthetic scene data to train and optimize artificial intelligence models helps to improve the model performance.
[0027] The improved synthetic scene data generation mechanism proposed in the present disclosure will be further discussed in detail below with reference to the accompanying drawings. It should be noted that although an example application scenario of the mechanism of the present disclosure is described by taking autonomous driving as an example for simplicity and clarity in the following discussion, the present disclosure is not limited thereto, but can be applied to any computer vision application scenario that requires synthetic data.
[0028] Figure 1 A schematic diagram showing the overall architecture 100 of a synthetic scene data generation mechanism according to an example embodiment of the present disclosure is shown.
[0029] The synthetic scene data generation mechanism of the present disclosure can partially generate synthetic scene data for a target scene based on data from a real scene.
[0030] Taking the autonomous driving application scenario as an example, the real scenario can be the traffic scenario of a specific traffic section or intersection. The traffic scenario may include a background and traffic participants. The background may include various static targets in the real scenario, including roads, buildings, traffic lights, trees, and so on. The traffic participants may include various moving targets in the real scenario, which may also be referred to as foreground objects and are uniformly referred to as objects in this disclosure. The objects may include any type of traffic participants, including for example pedestrians, motor vehicles, non-motor vehicles, and so on.
[0031] The data from the real scenario includes multi-view image data 102 for the real scenario. The multi-view image data 102 can represent the color information of the real scenario from different perspectives. Generally, the information provided by a single-view image for the targets in the real scenario may be limited. For example, some targets may be occluded and lost, some targets may be truncated and incomplete, and the effective pixel information of some small-sized targets may be less, and so on. The multi-view image data 102 can avoid the above problems as much as possible, so as to more accurately represent the real scenario. The multi-view image data 102 also helps to supplement the missing depth information in the single-view image, thus contributing to the generation of a three-dimensional object model.
[0032] The data from the real scenario also includes object point cloud data 104 for the objects in the real scenario. The object point cloud data 104 is also multi-view to represent the position / geometry information of the object from different perspectives. There may be data missing in the single-view point cloud data. For example, the position / geometry information of the back of the detected object is missing. The multi-view object point cloud data avoids potential data missing problems by detecting the object from different perspectives. It can be understood that such multi-view object point cloud data is three-dimensional point cloud data in the true sense. The position / geometry information provided by the object point cloud data 104 helps to improve the spatial accuracy of the reconstructed three-dimensional object model.
[0033] Three-dimensional reconstruction 106 can be performed based on the multi-view image data 102 and the object point cloud data 104. Three-dimensional reconstruction refers to the process of establishing a data model suitable for computer representation and processing for the real scenario objects. Through the three-dimensional reconstruction 106, a reconstructed three-dimensional object model 108 for the object can be generated. The reconstructed three-dimensional object model 108 contains dense data points, and each data point can include both position / geometry information and color information.
[0034] The reconstructed three-dimensional object model 108 can be combined with the target background model to generate a large amount of available data for the target scene, that is, the synthesized scene data 110. The generated synthesized scene data 110 can be used to train or optimize an artificial intelligence model. In one example, the artificial intelligence model can be a model for performing autonomous driving perception tasks, which is used to sense, identify, and track the driving environment and traffic objects therein based on image, radar, and lidar sensor data, so as to help the backend decision-making module make accurate planning decisions.
[0035] Figure 2 FIG. 200 is an example flowchart showing operations for obtaining multi-view image data (e.g., the multi-view image data 102 discussed above in connection with Figure 1 ).
[0036] At S202, multiple images of the real scene can be obtained from multiple viewpoints. For example, the multiple images can be obtained from multiple image sensors, which can be arranged at different orientations to observe the real scene from multiple different viewpoints. In the example where the real scene discussed above is a traffic scene, the multiple image sensors can be multiple traffic cameras arranged at different orientations at a specific road section or intersection.
[0037] At step S204, the multiple images can be aligned to obtain aligned multiple images. The image alignment operation aims to fuse the image data from different image sensors. Any known image alignment technique in the prior art can be used to complete the operation of step S204. In one example, image alignment can include time synchronization and spatial calibration. Time synchronization aims to synchronize multiple frames of images from multiple image sensors to the same moment. Spatial calibration aims to calibrate the respective coordinate systems of multiple image sensors to a unified coordinate system.
[0038] At step S206, sensitive image information in the aligned multiple images can be removed. For a traffic scene, the images captured for the traffic scene may include sensitive information related to the following example aspects: sensitive information associated with the road, such as the height and longitude / latitude of the road surface, etc. Sensitive information associated with traffic participants, such as face information, license plate information, etc., and any other type of sensitive information. Any known image recognition technique in the prior art can be used to detect whether the aligned multiple images contain sensitive image information, such as deep learning-based image recognition, etc. In the case where sensitive image information is identified, the sensitive image information can be removed. For example, the sensitive image information can be removed by applying a mosaic to the area containing the sensitive image information in each of the aligned multiple images.
[0039] By combining Figure 2 the operations described, multi-view image data for a real scene can be obtained, and such multi-view image data can be used in subsequent processes to perform three-dimensional reconstruction of objects in the real scene.
[0040] Figure 3 FIG. 300 is an example flowchart showing operations for obtaining object point cloud data according to an example embodiment of the present disclosure.
[0041] In step S302, multiple sets of point cloud data for the real scene from multiple viewpoints can be obtained. For example, the above multiple sets of point cloud data can be obtained from multiple lidars, and these lidars can be arranged in different orientations to detect the real scene from multiple different viewpoints. In the example where the real scene discussed above is a traffic scene, the multiple lidars can be multiple roadside lidars arranged in different orientations at a specific road section or intersection.
[0042] In step S304, alignment can be performed on the multiple sets of point cloud data to obtain multi-viewpoint cloud data. The point cloud alignment operation aims to fuse image data from different lidars. Any known point cloud alignment technique in the prior art can be used to complete the operation in step S304. In one example, point cloud alignment can include time synchronization and spatial calibration. Time synchronization aims to synchronize multiple sets of point cloud data from multiple lidars to the same moment. Spatial calibration aims to calibrate the respective coordinate systems of multiple lidars to a unified coordinate system.
[0043] In step S306, background point cloud data for the background in the real scene can be obtained. As mentioned above, the background of the real scene can include a series of stationary targets, excluding any moving targets. In the example where the real scene discussed above is a traffic scene, the background does not include any traffic participants. The background point cloud data can be obtained by collecting point clouds of a stable stationary background for the real scene over a period of time.
[0044] In step S308, the multi-viewpoint cloud data obtained in step S304 can be segmented based on the background point cloud data obtained in step S306. The multi-viewpoint cloud data can be compared with the background point cloud data, and the data in the multi-viewpoint cloud data associated with the background can be deleted to obtain the segmented point cloud data. Through the above segmentation operation, the segmented point cloud data only contains information associated with objects in the real scene to obtain annotations for objects in the real scene.
[0045] In one example, it can be by combining Figure 2The multi-view image data obtained by the operations discussed is fused with the multi-view point cloud data obtained at step 304 to obtain fused scene data. Then, the segmentation operation at step S308 is performed on the fused scene data. This helps to further improve the accuracy of background segmentation. The fusion between the multi-view image data and the multi-view point cloud data can be achieved by using sensor fusion techniques for cameras and lidar that are known in the prior art. For example, the point cloud data can be projected onto the reference coordinate system of the camera based on the parameter matrix of the camera, thereby realizing the mapping between the points in the three-dimensional point cloud data and the pixel points in the two-dimensional image. In one exemplary aspect, the fused scene data can be segmented based on the background point cloud data in a similar manner as described above in step S308. For example, since the background point cloud data provides information about the position of the background in the scene, it can be determined based on the background point cloud data which data in the fused scene data is associated with the background, and such data can be deleted, leaving only the fused object data associated with the objects in the scene. Then, the point cloud data associated with the objects can be separated from the fused object data associated with the objects, thereby obtaining the segmented point cloud data. In another exemplary aspect, the fused scene data can also be segmented based on the fused background data. For example, first, the background point cloud data and the background image data obtained for the background can be fused to obtain the fused background data. Then, the fused scene data can be compared with the fused background data to delete the data in the fused scene data that is associated with the background, leaving only the fused object data associated with the objects in the scene. Next, the point cloud data associated with the objects can be separated from the fused object data associated with the objects, thereby obtaining the segmented point cloud data.
[0046] At step S310, the segmented point cloud data obtained at step S308 can be voxelized to obtain voxelized point cloud data. Point clouds typically include a large amount of data. For example, a point cloud can include a large number of sparsely arranged points, where each point can be represented by a multi-dimensional vector including information such as three-dimensional coordinates, intensity values, time, etc. This makes the processing of point cloud data usually consume a large amount of computing resources. Point cloud voxelization is used to map the large number of discretely distributed points in the point cloud to a regular voxel grid, which significantly reduces the amount of data contained in the voxelized point cloud, thereby facilitating subsequent data processing.
[0047] At step S312, the voxelized point cloud data can be filtered to filter out some abnormal point cloud data that may affect subsequent processing, including, for example, noise points, outliers, and the like. After step S312 is completed, object point cloud data can be obtained. Such object point cloud data can be used in subsequent processes to perform three-dimensional reconstruction of objects in the real scene.
[0048] Figure 4 FIG. shows a schematic diagram of a neural network architecture 400 for reconstructing a three-dimensional object model according to an example embodiment of the present disclosure.
[0049] It has been proposed to use the Neural Radiance Field (NeRF) algorithm to perform three-dimensional reconstruction. The NeRF algorithm pioneered the use of neural radiance fields to represent spatial information, which can establish a three-dimensional representation of a scene based on a series of images from known viewpoints, and can render the established scene representation to generate images from new viewpoints.
[0050] The synthetic scene data generation mechanism of the present disclosure can adopt a NeRF-based three-dimensional reconstruction algorithm to perform three-dimensional reconstruction to generate a reconstructed three-dimensional object model. Different from the conventional NeRF algorithm that only takes a series of images from known viewpoints as input, the input of this NeRF-based three-dimensional reconstruction algorithm includes both of the following: 1) multi-view image data for a real scene obtained by combining the Figure 2 operations described, and 2) object point cloud data for the objects in the real scene obtained by combining the Figure 3 operations described. In other words, the NeRF-based three-dimensional reconstruction algorithm of the present disclosure further introduces object point cloud data into the input of the algorithm. The object point cloud data is a multi-view three-dimensional point cloud directly belonging to the object to be reconstructed itself. Thus, the object point cloud data provides information describing the accurate three-dimensional positions of the objects. Introducing object point cloud data into the input of the algorithm is equivalent to adding three-dimensional position supervision data on the basis of two-dimensional image data in the three-dimensional reconstruction process. Thus, through the NeRF-based three-dimensional reconstruction algorithm of the present disclosure, it helps to establish a more accurate spatial representation for real-scene objects.
[0051] The output of the NeRF-based three-dimensional reconstruction algorithm is the reconstructed three-dimensional object model for real-scene objects. Different from explicit spatial representation methods such as meshes, point clouds, and voxels, the reconstructed three-dimensional object model adopts an implicit spatial representation method called neural radiance field. The NeRF-based three-dimensional reconstruction algorithm can be implemented by a neural network. In this case, the three-dimensional reconstruction process can be understood as a process of modeling real-scene objects into the parameters (e.g., weights) of a neural network, and the reconstructed three-dimensional object model generated through the three-dimensional reconstruction process can be a neural network including the above parameters.
[0052] In the three-dimensional reconstruction process, the neural point cloud can be determined first. The neural point cloud abstracts the image data and the point cloud data into a large number of neural points representing neural features. The neural point cloud P can be represented as
[0053] P = {(pi , f i , γ i ), i = 1, ..., N}
[0054] where N is the number of neural points; p i is the color position information of the i-th neural point; f i is the feature vector of the i-th neural point; γ i is the confidence score of the i-th neural point.
[0055] The color position information p i can be determined according to the multi-view image data described above in combination with Figure 2 and the object point cloud data described above in combination with Figure 3 and can be represented as (x, y, z, r, g, b).
[0056] The feature vector f i can be obtained through feature extraction. Feature extraction can be implemented using a feature extraction backbone dedicated to extracting high-level feature representations of data. Figure 4 An exemplary feature extraction backbone 402 is shown. In one example, the feature extraction backbone 402 can include a 3D deep learning backbone, for example, 3D CNN models of various architectures, etc. The feature extraction backbone 402 can take multi-view image data and object point cloud data as inputs to extract local scene features.
[0057] The confidence score γ i is used to represent the likelihood that the neural point is near the surface of a real-world object and can be determined in any manner known in the prior art, for example, by performing trilinear interpolation on the probability space.
[0058] Then, a neural radiance field, that is, a reconstructed three-dimensional object model, can be generated based on the neural point cloud. This can be performed by Figure 4 the neural network 404 shown, which includes multiple multi-layer perceptrons (MLPs).
[0059] For a given three-dimensional RGB point x, K adjacent neural points around it can be queried within a certain radius R. These K adjacent neural points can be represented as (p1, f1, γ1, ..., p K , f K , γ K)。The neural radiance field can aggregate information from these adjacent neural points to regress the volume density σ of the given 3D RGB point x and the viewing direction-dependent radiance r along any viewing direction d. In other words, the neural radiance field establishes a mapping between the volume density σ, the viewing direction-dependent radiance r, the 3D RGB point x, the viewing direction d, and K adjacent neural points. The above process can be described as
[0060] (σ,r) = RGB Point-NeRF(x,d,p1,f1,γ1,...,p K ,f K ,γ K )
[0061] where RGB Point-NeRF represents the generated neural radiance field.
[0062] The reconstructed 3D object model can include dense 3D RGB points. In other words, each point contains both position / geometry information and color information (or RGB information).
[0063] Figure 5 FIG. 500 is an example flowchart showing operations for generating synthetic scene data based on a reconstructed 3D object according to an example embodiment of the present disclosure.
[0064] In step S502, a target background model can be obtained. The target background model can be used to characterize the background in the target scene. The background can be the background of another real scene different from the real background in the real scene. In the above-discussed autonomous driving example, the background in the target scene can be the background of another road section or intersection different from the road section or intersection in the real scene. The background in the target scene can also be any simulated background generated by data simulation techniques, or any fixed test background. The target background model can be a 3D model including color and position / geometry information similar to the reconstructed 3D object model, so as to facilitate combination with the reconstructed 3D object model. The combination method is, for example, to overlay the reconstructed 3D object model on the target background model to generate synthetic scene data.
[0065] In step S504, the reconstructed 3D object model and the target background model can be combined to generate synthetic scene data. The synthetic scene data can be used to characterize the target scene.
[0066] The synthesized scene data can be further post - processed. As discussed above, the reconstructed three - dimensional object model generated through 3D reconstruction operations includes dense data points. In one example, the reconstructed three - dimensional object model can be downsampled to reduce the density of the data points. LiDARs used to generate point cloud data have different numbers of threads, for example, 64 - thread, 32 - thread, 16 - thread, etc. The higher the number of threads of the LiDAR, the denser the data points in the generated point cloud data. Thus, in one example, the reconstructed three - dimensional object model including dense data points can be downsampled based on the attributes of the LiDAR associated with the target scene, such as the number of threads. The above downsampling operation helps to customize the synthesized scene data to fit the attributes of the target scene and can improve the efficiency of data processing.
[0067] The synthesized scene data generated by combining Figure 5 the operations described above can be used as training data for training an artificial intelligence model for the target scene.
[0068] Figure 6 FIG. 600 shows an example flowchart of operations for cyclically training an artificial intelligence model based on synthesized scene data according to an example embodiment of the present disclosure. In one example, at least one step in flowchart 600 can be automatically executed by corresponding software functional modules.
[0069] In step S602, a specific target background model can be selected from a background database storing multiple target background models, and one or more reconstructed three - dimensional object models can be selected from an object database storing multiple reconstructed three - dimensional object models. Among them, the reconstructed three - dimensional object models stored in the object database are constructed based on real - world scene data by combining Figures 2 to 4 the operations described above.
[0070] In step S604, the selected target background model and one or more reconstructed three - dimensional object models can be combined to generate synthesized scene data.
[0071] In step S606, the artificial intelligence model can be trained using the synthesized scene data.
[0072] In step S608, the trained artificial intelligence model can be model - verified. For example, a corner - case analysis can be performed to determine whether there are corner cases for the trained artificial intelligence model. Corner cases generally refer to some extreme situations that rarely occur in real - world scenes. These corner cases may cause the trained artificial intelligence model to output incorrect results when encountering similar extreme situations in actual applications.
[0073] In step S610, the synthetic scene data can be regenerated based on the results of model validation. For example, in the case where an edge case is determined, the information associated with the edge case can be fed back to the software functional module that executes step S602. The software functional module can repeat the operation of step S602 based on the feedback to select a new target background model from the background database and / or select one or more new reconstructed three-dimensional object models from the object database to regenerate the synthetic scene data. Then, other operations can be repeated Figure 6 to iteratively train the artificial intelligence model to improve the performance of the trained artificial intelligence model.
[0074] Figure 7 FIG. shows a schematic diagram of a rendered scene image obtained based on synthetic scene data according to an exemplary embodiment of the present disclosure.
[0075] The synthetic scene data can be stored in a computer as a three-dimensional data model as a large amount of data for characterizing the color and position / geometry information of the target scene. The synthetic scene data can be rendered to achieve two-dimensional visualization of the synthetic scene data. Any three-dimensional rendering algorithm known in the prior art can be used to perform the rendering of the synthetic scene data. In one example, a volume rendering algorithm can be employed to implement the rendering process based on the volume density σ and the line-of-sight-dependent radiance r discussed above in connection with Figure 4 the discussion.
[0076] Since the synthetic scene data generated by using the mechanism of the present disclosure can provide accurate and realistic detailed data about the scene objects, including color information and position / geometry information, the rendered scene image can include more object detail information.
[0077] The synthetic scene data can be rendered into multiple rendered scene images for different perspectives to simulate the scene images observed when observing the target scene from different perspectives. Figure 7 FIG. shows four rendered scene images, corresponding to the images of the target scene observed from the perspectives of obliquely above, directly above, left, and right, respectively. It should be noted that the above four perspectives are only exemplary, and scene images for any observation perspective can be rendered based on the synthetic scene data. Each rendered scene image not only includes RGB information but also can include some data points representing the position / geometry information of the object.
[0078] In one example, instead of performing the rendering on the synthetic scene data as discussed above, the target background model and the reconstructed three-dimensional object model discussed above can be rendered separately, and the rendering results can be combined. For example, the rendered object can be superimposed on the rendered background to generate the rendered scene image.
[0079] In one example, instead of directly using the synthetic scene data, a scene image rendered based on the synthetic scene data can be used to train the artificial intelligence model. Figure 6 The loop training process discussed above.
[0080] Figure 8 FIG. 800 is an example flowchart showing a method for generating synthetic scene data according to an example embodiment of the present disclosure.
[0081] In step S802, multi-view image data for a real scene can be obtained. The multi-view image data can be obtained by combining the operations described above. Figure 2 The operations described above.
[0082] In step S804, object point cloud data for an object in the real scene can be obtained. The object point cloud data can be obtained by combining the operations described above. Figure 3 The operations described above.
[0083] In step S806, 3D reconstruction can be performed based on the multi-view image data and the object point cloud data to generate a reconstructed 3D object model for the object. A 3D reconstruction algorithm based on Neural Radiance Field (NeRF) can be used, and the 3D object model can be generated by adopting the neural network architecture described above. Figure 4 The operations described above.
[0084] In step S808, the synthetic scene data can be generated based on the reconstructed 3D object model. The synthetic scene data can be generated by combining the operations described above. Figure 5 The operations described above.
[0085] Figure 9 FIG. 900 is a block diagram showing a device 900 according to an example embodiment of the present disclosure, which can implement the method for generating synthetic scene data.
[0086] The exemplary device 900 includes a processor 904 connected to an internal communication bus 902. The processor 904 is configured to execute instructions in a memory 906 to implement the method for generating synthetic scene data described in detail above. Examples of the processor 904 may include a central processing unit (CPU), a microcontroller, and the like. Examples of the processor 904 may also include a graphics processing unit (GPU) dedicated to performing graphics processing operations, and the like. The memory 906 suitable for tangibly embodying computer program instructions and data includes various forms of memory, such as EPROM, EEPROM, and flash memory devices, and the like. The device 900 may also include an input interface 908 and an output interface 910. The input interface 908 is configured to receive input signals and data. The output interface 910 is configured to send output signals and data.
[0087] A computer program may include instructions executable by a computer for causing a processor 904 of an apparatus 900 to execute the method of the present disclosure for generating synthetic scene data. The program may be recorded on any data storage medium including a memory. For example, the program may be implemented in digital electronic circuitry, or in computer hardware, firmware, software, or in combinations thereof. The process / method steps described in the present disclosure may be executed by a programmable processor executing program instructions to perform the method, steps, operations by operating on input data and generating output.
[0088] Embodiments of the present disclosure may be implemented in a computer-readable medium. The computer-readable medium may store a computer program including instructions. In one exemplary aspect, when executed, the instructions may cause at least one processor to: obtain multi-view image data for a real scene; obtain object point cloud data for an object in the real scene; perform three-dimensional reconstruction based on the multi-view image data and the object point cloud data to generate a reconstructed three-dimensional object model for the object; and generate the synthetic scene data based on the reconstructed three-dimensional object model.
[0089] Embodiments of the present disclosure may be implemented in a computer program product. The computer program product may include instructions. In one exemplary aspect, when executed, the instructions may cause a processor of a computing device to: obtain multi-view image data for a real scene; obtain object point cloud data for an object in the real scene; perform three-dimensional reconstruction based on the multi-view image data and the object point cloud data to generate a reconstructed three-dimensional object model for the object; and generate the synthetic scene data based on the reconstructed three-dimensional object model.
[0090] In addition to what is described herein, various modifications may be made to the disclosed embodiments and implementations of the present invention without departing from the scope thereof. Therefore, the description and examples herein should be construed as illustrative rather than limiting in nature. The scope of the present invention should be measured only by reference to the claims.
Claims
1. A method for generating synthetic scene data, comprising: Obtaining multi-view image data for a real scene; Obtaining object point cloud data for an object in the real scene; Performing 3D reconstruction based on the multi-view image data and the object point cloud data to generate a reconstructed 3D object model for the object; and Generating the synthetic scene data based on the reconstructed 3D object model.
2. The method according to claim 1, wherein Including at least one of the following: Sensitive image information in the multi-view image data is removed; Each data point in the reconstructed 3D object model includes position information and color information.
3. The method according to claim 1, wherein The obtaining of the multi-view image data includes: Obtaining multiple images for the real scene from multiple perspectives; and Performing alignment on the multiple images to obtain the multi-view image data.
4. The method according to claim 1, wherein, The obtaining of the object point cloud data includes: Obtaining multi-view point cloud data for the real scene; Obtaining background point cloud data for the background in the real scene; and Performing segmentation on the multi-view point cloud data based on the background point cloud data.
5. The method according to claim 4, wherein, The obtaining of the multi-view point cloud data includes: Obtaining multiple sets of point cloud data for the real scene from multiple perspectives; and Performing alignment on the multiple sets of point cloud data to obtain the multi-view point cloud data.
6. The method according to claim 4, further comprising: Fusing the multi-view image data with the multi-view point cloud data to obtain fused scene data; and And Wherein, the segmentation is performed on the fused scene data.
7. The method according to claim 4, further comprising: Voxelizing the segmented point cloud data obtained by the segmentation to obtain voxelized point cloud data.
8. The method according to claim 7, further comprising: Filtering the voxelized point cloud data to obtain the object point cloud data.
9. The method according to claim 1, wherein The 3D reconstruction is performed using a 3D reconstruction algorithm based on Neural Radiance Field (NeRF).
10. The method according to claim 9, wherein, The NeRF-based 3D reconstruction algorithm is implemented using multiple Multi-Layer Perceptrons (MLP).
11. The method according to claim 10, wherein, A feature extraction backbone network is attached before the multiple MLPs to generate feature vectors based on the multi-view image data and the object point cloud data.
12. The method according to claim 1, wherein, The generating of the synthetic scene data based on the reconstructed 3D object model includes: Obtaining a target background model, wherein the target background model is used to characterize the background in the target scene; and Combining the reconstructed 3D object model and the target background model to generate the synthetic scene data, wherein the synthetic scene data characterizes the target scene.
13. The method according to claim 12, further comprising at least one of the following: Downsampling the reconstructed 3D object model based on the attributes of a lidar associated with the target scene; Rendering the synthetic scene data to visualize the synthetic scene data as multiple rendered scene images for specific perspectives respectively.
14. The method according to claim 1, wherein, The synthetic scene data is used as training data for training an artificial intelligence model.
15. The method according to claim 14, further comprising: regenerating the synthetic scene data for the training based on the result of validating the trained artificial intelligence model; and / or the artificial intelligence model includes an autonomous driving perception model.
16. An apparatus for generating synthetic scene data, comprising: a memory; a processor coupled to the memory, the processor being configured to execute the method according to any one of claims 1-15.
17. A computer-readable medium storing a computer program including instructions that, when executed by a processor, cause the processor to be configured to execute the method according to any one of claims 1-15.
18. A computer program product comprising computer-executable instructions that, when executed, cause one or more processors to execute the method according to any one of claims 1-15.
19. A model training method, comprising: obtaining training data; training an artificial intelligence model based on the training data to obtain a trained artificial intelligence model; wherein the training data includes synthetic scene data generated by the method according to any one of claims 1 to 15.
20. The model training method according to claim 19, further comprising: validating the trained artificial intelligence model; regenerating the synthetic scene data based on the result of the model validation; and training the artificial intelligence model based on the regenerated synthetic scene data.