Method and device for constructing multi-view multi-modal automatic driving data set
By detecting dynamic and static driving scenarios, combining camera sensor dataset construction and preprocessing, a multi-view multi-modal autonomous driving dataset is generated, which solves the shortcomings of the existing dataset in angle diversity and perspective evaluation, and achieves a more comprehensive scenario reconstruction evaluation.
Patent Information
- Application Number
- CN202411972942.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-06-06
AI Technical Summary
The existing autonomous driving datasets have significant limitations in the diversity of angles in scene reconstruction, and it is impossible to effectively evaluate the perspectives that lack high-quality truth values, and the perspectives are similar, making it difficult to fully evaluate the reconstruction quality of the model in multi-view scenarios.
By detecting dynamic and static driving scenarios, combining the pre-generated camera sensor dataset, a multi-view multi-modal initial autonomous driving dataset is constructed and pre-processed to generate a multi-view multi-modal autonomous driving dataset that meets the preset format conditions.
The coverage of the data set in the scene reconstruction angle diversity is improved, and the reconstruction quality of the model can be effectively evaluated in multi-view scenarios, solving the evaluation difficulties caused by similar perspectives.
Smart Images

Figure CN120107712A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of autonomous driving technology, and in particular to a method and device for constructing a multi-view multi-modal autonomous driving dataset. Background Art
[0002] In the field of autonomous driving, due to limitations in technical means or collection costs, the data returned by on-board cameras cannot satisfy the requirement of finding corresponding images at any angle.
[0003] In the related technology, the left-view image and the right-view image of the same scene can be segmented into superpixels to obtain corresponding superpixel blocks, and the effective superpixel blocks that are effective for scene depth estimation are selected for matching, and then the depth of field information of the left-view image and the right-view image is calculated, so as to realize fogging processing of the left-view image and / or the right-view image, thereby constructing a fogged image data set; it is also possible to detect at least one target human image from the sequence of images to be detected based on a preset human body tracking method according to N sequences of images to be detected captured by N cameras, and construct a three-dimensional posture estimation data set of the target human image in combination with the camera parameter information.
[0004] However, in related technologies, the camera installation position is relatively fixed and the camera's viewing angle is limited, which leads to significant limitations in the angle diversity of scene reconstruction in these datasets, and the lack of high-quality true values cannot be effectively evaluated. In addition, the perspectives in these datasets are relatively similar, making it difficult to fully evaluate the reconstruction quality of the model in multi-view scenes, which urgently needs to be improved. Summary of the invention
[0005] The present application provides a method and device for constructing a multi-view multi-modal autonomous driving dataset to solve the problem that the camera installation position is relatively fixed and the camera's viewing angle is limited in the related technology, resulting in significant limitations on the angle diversity of scene reconstruction in these datasets, and the view angles that lack high-quality true values cannot be effectively evaluated. In addition, the view angles in these datasets are relatively similar, making it difficult to fully evaluate the reconstruction quality of the model in multi-view scenes.
[0006] The first aspect of the present application provides a method for constructing a multi-view multi-modal autonomous driving dataset, comprising the following steps: detecting dynamic driving scenes and static driving scenes in autonomous driving; based on the dynamic driving scenes and the static driving scenes, combined with a pre-generated camera sensor dataset, to construct a multi-view multi-modal initial autonomous driving dataset for the autonomous driving; preprocessing the multi-view multi-modal initial autonomous driving dataset to obtain a preprocessed multi-view multi-modal initial autonomous driving dataset, and constructing a multi-view multi-modal autonomous driving dataset that meets preset format conditions based on the preprocessed multi-view multi-modal initial autonomous driving dataset.
[0007] Optionally, in one embodiment of the present application, before constructing a multi-view multi-modal initial autonomous driving dataset for the autonomous driving based on the dynamic driving scene and the static driving scene in combination with a pre-generated camera sensor dataset, it also includes: based on the dynamic image information of the dynamic driving scene and the static image information of the static driving scene, to obtain an image camera sensor dataset in the camera sensor dataset; based on the dynamic depth information of the dynamic driving scene and the static depth information of the static driving scene, to obtain a depth camera sensor dataset in the camera sensor dataset; based on the dynamic semantic segmentation information of the dynamic driving scene and the static semantic segmentation information of the static driving scene, to obtain a semantic segmentation camera sensor dataset in the camera sensor dataset; based on the dynamic point cloud information of the dynamic driving scene and the static point cloud information of the static driving scene, to obtain a radar camera sensor dataset in the camera sensor dataset; based on the image camera sensor dataset, the depth camera sensor dataset, the semantic segmentation camera sensor dataset and the radar camera sensor dataset, to obtain the camera sensor dataset.
[0008] Optionally, in one embodiment of the present application, before constructing a multi-view multi-modal initial autonomous driving dataset for the autonomous driving based on the dynamic driving scene and the static driving scene in combination with a pre-generated camera sensor dataset, it also includes: determining the image perspective position information of the image camera sensor corresponding to the image camera sensor in the image camera sensor dataset based on the image camera sensor dataset; determining the depth perspective position information of the depth camera sensor corresponding to the depth camera sensor in the depth camera sensor dataset based on the depth camera sensor dataset; determining the semantic segmentation perspective position information of the semantic segmentation camera sensor corresponding to the semantic segmentation camera sensor dataset based on the semantic segmentation camera sensor dataset; determining the radar perspective position information of the radar camera sensor corresponding to the radar camera sensor in the radar camera sensor dataset based on the radar camera sensor dataset in the camera sensor dataset; generating the camera sensor dataset based on the image perspective position information, the depth perspective position information, the semantic segmentation perspective position information and the radar perspective position information.
[0009] Optionally, in one embodiment of the present application, before constructing a multi-view multi-modal initial autonomous driving dataset for the autonomous driving based on the dynamic driving scene and the static driving scene in combination with a pre-generated camera sensor dataset, it also includes: based on the image camera sensor dataset in the camera sensor dataset, obtaining the image operating frequency of the corresponding image camera sensor in the image camera sensor dataset; based on the depth camera sensor dataset in the camera sensor dataset, obtaining the depth operating frequency of the corresponding depth camera sensor in the depth camera sensor dataset; based on the semantic segmentation camera sensor dataset in the camera sensor dataset, obtaining the semantic segmentation operating frequency of the corresponding semantic segmentation camera sensor in the semantic segmentation camera sensor dataset; based on the radar camera sensor dataset in the camera sensor dataset, obtaining the radar camera the radar operating frequency of the corresponding radar camera sensor in the sensor data set; judging whether the image operating frequency, the depth operating frequency, the semantic segmentation operating frequency and the radar operating frequency are all equal to each other; if the image operating frequency, the depth operating frequency, the semantic segmentation operating frequency and the radar operating frequency are all equal to each other, then based on the image operating frequency, the depth operating frequency, the semantic segmentation operating frequency and the radar operating frequency, it is allowed to construct the multi-view multi-modal initial autonomous driving data set; if the image operating frequency, the depth operating frequency, the semantic segmentation operating frequency and the radar operating frequency are not equal to each other, then adjusting at least one of the image operating frequency, the depth operating frequency, the semantic segmentation operating frequency and the radar operating frequency until the image operating frequency, the depth operating frequency, the semantic segmentation operating frequency and the radar operating frequency are all equal to each other.
[0010] Optionally, in one embodiment of the present application, the multi-view multi-modal initial autonomous driving dataset is preprocessed to obtain a preprocessed multi-view multi-modal initial autonomous driving dataset, and a multi-view multi-modal autonomous driving dataset that meets preset format conditions is constructed based on the preprocessed multi-view multi-modal initial autonomous driving dataset, including: obtaining a final image dataset of the initial image dataset based on an initial image dataset of the multi-view multi-modal initial autonomous driving dataset; obtaining a final point cloud dataset of the initial point cloud dataset based on an initial point cloud dataset of the multi-view multi-modal initial autonomous driving dataset; and obtaining the multi-view multi-modal autonomous driving dataset based on the final image dataset and the final point cloud dataset.
[0011] According to a second aspect of the present application, there is provided a device for constructing a multi-view multi-modal autonomous driving dataset, comprising: a detection module for detecting dynamic driving scenes and static driving scenes in autonomous driving; a first construction module for constructing a multi-view multi-modal initial autonomous driving dataset for the autonomous driving based on the dynamic driving scenes and the static driving scenes in combination with a pre-generated camera sensor dataset; and a second construction module for preprocessing the multi-view multi-modal initial autonomous driving dataset to obtain a preprocessed multi-view multi-modal initial autonomous driving dataset, and constructing a multi-view multi-modal autonomous driving dataset that meets preset format conditions based on the preprocessed multi-view multi-modal initial autonomous driving dataset.
[0012] Optionally, in one embodiment of the present application, it also includes: a first generation module, which is used to obtain an image camera sensor dataset in the camera sensor dataset based on the dynamic image information of the dynamic driving scene and the static image information of the static driving scene before constructing a multi-view multi-modal initial autonomous driving dataset for the autonomous driving based on the dynamic driving scene and the static driving scene in combination with a pre-generated camera sensor dataset; a second generation module, which is used to obtain a depth camera sensor dataset in the camera sensor dataset based on the dynamic depth information of the dynamic driving scene and the static depth information of the static driving scene; a third generation module, which is used to obtain a semantic segmentation camera sensor dataset in the camera sensor dataset based on the dynamic semantic segmentation information of the dynamic driving scene and the static semantic segmentation information of the static driving scene; a fourth generation module, which is used to obtain a radar camera sensor dataset in the camera sensor dataset based on the dynamic point cloud information of the dynamic driving scene and the static point cloud information of the static driving scene; a fifth generation module, which is used to obtain the camera sensor dataset based on the image camera sensor dataset, the depth camera sensor dataset, the semantic segmentation camera sensor dataset and the radar camera sensor dataset.
[0013] Optionally, in one embodiment of the present application, it also includes: a first determination module, which is used to determine the image perspective position information of the corresponding image camera sensor in the image camera sensor data set based on the image camera sensor data set in the camera sensor data set before constructing the multi-view multi-modal initial autonomous driving data set for the autonomous driving based on the dynamic driving scene and the static driving scene in combination with the pre-generated camera sensor data set; a second determination module, which is used to determine the depth perspective position information of the corresponding depth camera sensor in the depth camera sensor data set based on the depth camera sensor data set in the camera sensor data set; a third determination module, which is used to determine the semantic segmentation perspective position information of the corresponding semantic segmentation camera sensor in the semantic segmentation camera sensor data set based on the semantic segmentation camera sensor data set in the camera sensor data set; a fourth determination module, which is used to determine the radar perspective position information of the corresponding radar camera sensor in the radar camera sensor data set based on the radar camera sensor data set in the camera sensor data set; a fifth determination module, which is used to generate the camera sensor data set based on the image perspective position information, the depth perspective position information, the semantic segmentation perspective position information and the radar perspective position information.
[0014] Optionally, in one embodiment of the present application, it also includes: a sixth generation module, which is used to obtain the image operating frequency of the corresponding image camera sensor in the image camera sensor data set based on the image camera sensor data set in the camera sensor data set before constructing a multi-view multi-modal initial autonomous driving data set for the autonomous driving based on the dynamic driving scene and the static driving scene in combination with a pre-generated camera sensor data set; a seventh generation module, which is used to obtain the depth operating frequency of the corresponding depth camera sensor in the depth camera sensor data set based on the depth camera sensor data set in the camera sensor data set; an eighth generation module, which is used to obtain the semantic segmentation operating frequency of the corresponding semantic segmentation camera sensor in the semantic segmentation camera sensor data set based on the semantic segmentation camera sensor data set in the camera sensor data set; and a ninth generation module, which is used to obtain the radar camera sensor data set in the camera sensor data set. the radar operating frequency of the corresponding radar camera sensor in the radar camera sensor data set; a judgment module, used to judge whether the image operating frequency, the depth operating frequency, the semantic segmentation operating frequency and the radar operating frequency are all equal to each other; a third construction module, used to allow the construction of the multi-view multi-modal initial autonomous driving data set based on the image operating frequency, the depth operating frequency, the semantic segmentation operating frequency and the radar operating frequency when the image operating frequency, the depth operating frequency, the semantic segmentation operating frequency and the radar operating frequency are all equal to each other; an adjustment module, used to adjust at least one of the image operating frequency, the depth operating frequency, the semantic segmentation operating frequency and the radar operating frequency when the image operating frequency, the depth operating frequency, the semantic segmentation operating frequency and the radar operating frequency are not equal to each other, until the image operating frequency, the depth operating frequency, the semantic segmentation operating frequency and the radar operating frequency are all equal to each other.
[0015] Optionally, in one embodiment of the present application, the second construction module includes: a first generation unit, used to obtain a final image dataset of the initial image dataset based on the initial image dataset of the multi-view multi-modal initial autonomous driving dataset; a second generation unit, used to obtain a final point cloud dataset of the initial point cloud dataset based on the initial point cloud dataset of the multi-view multi-modal initial autonomous driving dataset; and a third generation unit, used to obtain the multi-view multi-modal autonomous driving dataset based on the final image dataset and the final point cloud dataset.
[0016] The third aspect of the present application provides an electronic device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method for constructing a multi-view multi-modal autonomous driving dataset as described in the above embodiment.
[0017] The fourth aspect embodiment of the present application provides a computer-readable storage medium, which stores a computer program, which, when executed by a processor, implements the above method for constructing a multi-view multi-modal autonomous driving dataset.
[0018] The fifth aspect embodiment of the present application provides a computer program product, including a computer program, which, when executed, implements the method for constructing a multi-view multi-modal autonomous driving dataset as described above.
[0019] The embodiments of the present application can construct a multi-view multi-modal initial autonomous driving dataset for autonomous driving based on dynamic driving scenes and static driving scenes in autonomous driving, combined with a pre-generated camera sensor dataset, and pre-process the multi-view multi-modal initial autonomous driving dataset to obtain a multi-view multi-modal autonomous driving dataset that meets certain format conditions, which can meet a variety of computer vision tasks for autonomous driving scenarios. The attached additional perspectives can be used to detect the reconstruction quality under new perspectives in work such as scene reconstruction. This solves the problem in the related art that the installation position of the camera is relatively fixed and the perspective of the camera is limited, resulting in significant limitations on the angle diversity of scene reconstruction in these datasets, and the perspective that lacks high-quality true values cannot be effectively evaluated. In addition, the perspectives in these data sets are relatively similar, making it difficult to fully evaluate the reconstruction quality of the model in multi-view scenarios.
[0020] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0022] Figure 1 A flowchart of a method for constructing a multi-view multi-modal autonomous driving dataset according to an embodiment of the present application;
[0023] Figure 2 A block diagram of a camera sensor configuration provided according to an embodiment of the present application;
[0024] Figure 3(a)-Figure 3(d)A schematic block diagram of a partial display of an RGB (RedGreenBlue) image according to an embodiment of the present application;
[0025] Figure 4(a)-Figure 4(d) A schematic block diagram showing a portion of a semantic segmentation image provided according to an embodiment of the present application;
[0026] Figure 5(a)-Figure 5(d) A schematic block diagram of a partial depth image provided according to an embodiment of the present application;
[0027] like Figure 6(a)-Figure 6(b) A schematic block diagram showing a portion of a radar point cloud provided according to an embodiment of the present application;
[0028] Figure 7 A flowchart of the working principle of a method for constructing a multi-view multi-modal autonomous driving dataset according to an embodiment of the present application;
[0029] Figure 8 A block diagram of a device for constructing a multi-view multi-modal autonomous driving dataset according to an embodiment of the present application;
[0030] Fig. 9 It is a schematic diagram of the structure of an electronic device provided according to an embodiment of the present application. DETAILED DESCRIPTION
[0031] Embodiments of the present application are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0032] The following describes the method and device for constructing a multi-view multi-modal autonomous driving dataset of the embodiment of the present application with reference to the accompanying drawings. In view of the fact that the installation position of the camera mentioned in the above background technology is relatively fixed and the viewing angle of the camera is limited, these datasets have significant limitations on the angle diversity of scene reconstruction, and the viewing angles that lack high-quality true values cannot be effectively evaluated. In addition, the viewing angles in these datasets are relatively similar, making it difficult to fully evaluate the reconstruction quality of the model in multi-view scenarios. The present application provides a method for constructing a multi-view multi-modal autonomous driving dataset, in which a multi-view multi-modal initial autonomous driving dataset for autonomous driving can be constructed based on dynamic driving scenes and static driving scenes in autonomous driving, combined with a pre-generated camera sensor dataset, and the multi-view multi-modal initial autonomous driving dataset is pre-processed to obtain a multi-view multi-modal autonomous driving dataset that meets certain format conditions, which can meet a variety of computer vision tasks for autonomous driving scenarios, and the attached additional viewing angles can be used to detect the reconstruction quality under new viewing angles in work such as scene reconstruction. Thus, the problem that the installation position of the camera is relatively fixed and the viewing angle of the camera is limited in the related technology, resulting in significant limitations on the angle diversity of scene reconstruction in these datasets, and the viewing angles that lack high-quality true values cannot be effectively evaluated. In addition, the perspectives in these datasets are relatively similar, making it difficult to fully evaluate issues such as the reconstruction quality of the model in multi-perspective scenarios.
[0033] Specifically, Figure 1 This is a flowchart of a method for constructing a multi-view multi-modal autonomous driving dataset according to an embodiment of the present application.
[0034] like Figure 1 As shown, the method for constructing the multi-view multi-modal autonomous driving dataset includes the following steps:
[0035] In step S101 , dynamic driving scenes and static driving scenes in autonomous driving are detected.
[0036] It is understood that in the embodiments of the present application, the driving scenes of the autonomous driving may include, but are not limited to, city streets, rural roads, and highways, etc., and the present application does not impose specific restrictions. In addition, the driving scenes of the embodiments of the present application may also include vehicles, pedestrians, bicycles, roads, external environments, road signs, traffic lights, etc. to simulate complex traffic conditions, and the present application does not impose specific restrictions.
[0037] Furthermore, in the embodiments of the present application, driving scenarios can be divided into dynamic driving scenarios and static driving scenarios. Among them, dynamic driving scenarios can be understood as the process of interaction between vehicles, pedestrians and other vehicles, infrastructure, weather, lighting, obstacles, etc. in the driving environment within a certain time and space range; static driving scenarios can be understood as only containing parked vehicles and fixed infrastructure (such as road signs, trees and buildings). These scenarios are used to evaluate the quality of the constructed autonomous driving data set under static conditions.
[0038] In addition, it should be noted that all data of dynamic driving scenes and static driving scenes in the embodiments of the present application are collected under sunny and cloudy weather conditions to enhance the diversity of the constructed autonomous driving data set.
[0039] In some embodiments, the embodiments of the present application can detect dynamic driving scenarios and static driving scenarios in autonomous driving, and then construct autonomous driving data sets in different driving scenarios.
[0040] Optionally, in one embodiment of the present application, before constructing a multi-view multi-modal initial autonomous driving dataset for autonomous driving based on dynamic driving scenes and static driving scenes in combination with a pre-generated camera sensor dataset, it also includes: based on the dynamic image information of the dynamic driving scene and the static image information of the static driving scene, to obtain an image camera sensor dataset in the camera sensor dataset; based on the dynamic depth information of the dynamic driving scene and the static depth information of the static driving scene, to obtain a depth camera sensor dataset in the camera sensor dataset; based on the dynamic semantic segmentation information of the dynamic driving scene and the static semantic segmentation information of the static driving scene, to obtain a semantic segmentation camera sensor dataset in the camera sensor dataset; based on the dynamic point cloud information of the dynamic driving scene and the static point cloud information of the static driving scene, to obtain a radar camera sensor dataset in the camera sensor dataset; based on the image camera sensor dataset, the depth camera sensor dataset, the semantic segmentation camera sensor dataset and the radar camera sensor dataset, to obtain a camera sensor dataset.
[0041] It can be understood that in the embodiments of the present application, the camera sensor may include but is not limited to image camera sensors, depth camera sensors, semantic segmentation camera sensors, radar camera sensors, etc., and the present application does not impose any specific restrictions.
[0042] In some embodiments, the image camera sensor data set of the embodiment of the present application may include, but is not limited to, dynamic image information of dynamic driving scenes and static image information of static driving scenes, etc., and the present application does not impose any specific limitations.
[0043] Among them, the image camera sensor data set can be obtained through an RGB camera sensor to capture a color image of the environment, and its resolution can be 1920x1080, which is not specifically limited in this application.
[0044] In some embodiments, the depth camera sensor data set of the embodiment of the present application may include, but is not limited to, dynamic depth information of dynamic driving scenes and static depth information of static driving scenes, etc., and the present application does not impose any specific limitations.
[0045] Among them, the depth camera sensor data set can be obtained by the depth camera sensor to provide depth information between the sensor and the object in the scene, and its resolution can be 1920x1080, which is not specifically limited in this application.
[0046] In some embodiments, the semantic segmentation camera sensor data set of the embodiments of the present application may include, but is not limited to, dynamic semantic segmentation information of dynamic driving scenes and static semantic segmentation information of static driving scenes, etc., and the present application does not impose any specific limitations.
[0047] Among them, the semantic segmentation camera sensor dataset can be obtained through the semantic segmentation camera sensor to generate a semantic label for each pixel in the scene, and its resolution can be 1920x1080, which is not specifically limited in this application.
[0048] In some embodiments, the radar camera sensor data set of the embodiment of the present application may include, but is not limited to, dynamic point cloud information of dynamic driving scenes and static point cloud information of static driving scenes, etc., and the present application does not impose any specific limitations.
[0049] Among them, the radar camera sensor data set can be obtained through the LiDAR sensor to provide a 360-degree LiDAR sensor to capture three-dimensional point clouds. Its maximum detection range can be 200 meters, generating 3 million points per second, and has 128 scanning channels. The specific settings can be made by technicians in this field according to actual conditions, and this application does not impose any specific restrictions.
[0050] Optionally, in one embodiment of the present application, before constructing a multi-view multi-modal initial autonomous driving dataset for autonomous driving based on dynamic driving scenarios and static driving scenarios in combination with a pre-generated camera sensor dataset, it also includes: determining the image perspective position information of the image camera sensor corresponding to the image camera sensor in the image camera sensor dataset based on the image camera sensor dataset; determining the depth perspective position information of the depth camera sensor corresponding to the depth camera sensor in the depth camera sensor dataset based on the depth camera sensor dataset in the camera sensor dataset; determining the semantic segmentation perspective position information of the semantic segmentation camera sensor corresponding to the semantic segmentation camera sensor in the semantic segmentation camera sensor dataset based on the semantic segmentation camera sensor dataset in the camera sensor dataset; determining the radar perspective position information of the radar camera sensor corresponding to the radar camera sensor in the radar camera sensor dataset based on the radar camera sensor dataset in the camera sensor dataset; generating a camera sensor dataset based on the image perspective position information, the depth perspective position information, the semantic segmentation perspective position information and the radar perspective position information.
[0051] It is understandable that in order to improve the quality of the reconstructed dataset in multi-view scenarios, the embodiment of the present application not only uses the CARLA simulator to generate it to cover dynamic driving scenarios and static driving scenarios, but also combines a variety of sensor data, which may include but is not limited to image information, depth information, semantic segmentation information, and point cloud information. In addition, the autonomous driving dataset constructed by the embodiment of the present application also provides 12 uniform viewpoints around the vehicle body. Compared with traditional datasets, the autonomous driving dataset constructed by the embodiment of the present application can find the corresponding picture truth value in almost every perspective, which makes it more suitable for verifying the performance of the model in new perspective synthesis.
[0052] As a possible implementation method, the embodiment of the present application can obtain a camera sensor data set based on the image perspective position information of the image camera sensor, the depth perspective position information of the depth camera sensor, the semantic segmentation perspective position information of the semantic segmentation camera sensor, and the radar perspective position information of the radar camera sensor.
[0053] Exemplary, combined Figure 2As shown, in order to improve the coverage of the data set, the embodiment of the present application can add camera sensors at 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11 and 12, which can be but not limited to image camera sensors, depth camera sensors and semantic segmentation camera sensors, etc., and the present application does not make specific restrictions. Further, in the embodiment of the present application, the field of view of each camera sensor can be 90 degrees, evenly distributed at intervals of 30 degrees. In addition, a radar camera sensor is installed at the top center of the vehicle to enhance three-dimensional environmental mapping. In general, the camera sensor set in the embodiment of the present application provides a 360-degree coverage range and can be accurately evaluated from a perspective never seen before. It is worth noting that the data set constructed by the embodiment of the present application can not only be used for the new perspective evaluation of 3DGS (3D Gaussian Splashing, three-dimensional Gaussian sputtering technology), but also can be used for multiple autonomous driving tasks, such as bird's-eye view perception and occupancy detection, etc., and the present application does not make specific restrictions.
[0054] Optionally, in one embodiment of the present application, before constructing a multi-view multi-modal initial autonomous driving dataset for autonomous driving based on a dynamic driving scenario and a static driving scenario in combination with a pre-generated camera sensor dataset, it also includes: based on an image camera sensor dataset in the camera sensor dataset, obtaining an image operating frequency of a corresponding image camera sensor in the image camera sensor dataset; based on a depth camera sensor dataset in the camera sensor dataset, obtaining a depth operating frequency of a corresponding depth camera sensor in the depth camera sensor dataset; based on a semantic segmentation camera sensor dataset in the camera sensor dataset, obtaining a semantic segmentation operating frequency of a corresponding semantic segmentation camera sensor in the semantic segmentation camera sensor dataset; based on a radar camera in the camera sensor dataset A sensor data set is used to obtain the radar operating frequency of the corresponding radar camera sensor in the radar camera sensor data set; determine whether the image operating frequency, the depth operating frequency, the semantic segmentation operating frequency and the radar operating frequency are all equal to each other; if the image operating frequency, the depth operating frequency, the semantic segmentation operating frequency and the radar operating frequency are all equal to each other, then based on the image operating frequency, the depth operating frequency, the semantic segmentation operating frequency and the radar operating frequency, it is allowed to construct a multi-view multi-modal initial autonomous driving data set; if the image operating frequency, the depth operating frequency, the semantic segmentation operating frequency and the radar operating frequency are not equal to each other, then adjust at least one of the image operating frequency, the depth operating frequency, the semantic segmentation operating frequency and the radar operating frequency until the image operating frequency, the depth operating frequency, the semantic segmentation operating frequency and the radar operating frequency are all equal to each other.
[0055] During the actual implementation process, before combining the pre-generated camera sensor data set to construct a multi-view multi-modal initial autonomous driving data set for autonomous driving, the embodiments of the present application can also determine whether the image operating frequency of the image camera sensor, the depth operating frequency of the depth camera sensor, the semantic segmentation operating frequency of the semantic segmentation camera sensor, and the radar operating frequency of the radar camera sensor are all equal to each other. If they are equal to each other, it is allowed to construct a multi-view multi-modal initial autonomous driving data set; if they are not equal to each other, at least one of the image operating frequency, depth operating frequency, semantic segmentation operating frequency, and radar operating frequency is adjusted until the image operating frequency, depth operating frequency, semantic segmentation operating frequency, and radar operating frequency are all equal to each other.
[0056] Exemplarily, the embodiments of the present application can set the image operating frequency of the image camera sensor, the depth operating frequency of the depth camera sensor, the semantic segmentation operating frequency of the semantic segmentation camera sensor, and the radar operating frequency of the radar camera sensor to 10 Hz, thereby constructing an autonomous driving data set.
[0057] In step S102, based on the dynamic driving scene and the static driving scene, a pre-generated camera sensor data set is combined to construct a multi-view multi-modal initial autonomous driving data set for autonomous driving.
[0058] As a possible implementation method, the embodiment of the present application can construct a multi-view and multi-modal initial autonomous driving data set for autonomous driving based on dynamic driving scenarios and static driving scenarios through different camera sensors, position information, operating frequencies, etc. in the camera sensor data set.
[0059] For example, the embodiments of the present application can be combined with Figure 2 Add camera sensors in the manner shown, adjust the working frequency, such as setting the image working frequency of the image camera sensor, the depth working frequency of the depth camera sensor, the semantic segmentation working frequency of the semantic segmentation camera sensor, and the radar working frequency of the radar camera sensor to 10Hz, and select 20 dynamic driving scenes and 20 static driving scenes. In each driving scene, it lasts for 10 seconds (about 100 meters of street driving), and each camera sensor generates 100 frames of data, thereby obtaining a multi-view multi-modal initial autonomous driving data set for autonomous driving. .
[0060] In addition, it should be noted that in each driving scene, the generated data may include but is not limited to 1200 RGB images, as shown in the schematic diagram Figure 3(a)-Figure 3(d) As shown; 1200 semantic segmentation images, the schematic diagram is as follows Figure 4(a)-Figure 4(d) As shown; 1200 depth images, the schematic diagram is as follows Figure 5(a)-Figure 5(d)As shown; 30 million radar point clouds, etc., the schematic diagram is as follows Figure 6(a)-Figure 6(b) shown.
[0061] In step S103, the multi-view multi-modal initial autonomous driving dataset is preprocessed to obtain a preprocessed multi-view multi-modal initial autonomous driving dataset, and a multi-view multi-modal autonomous driving dataset that meets preset format conditions is constructed based on the preprocessed multi-view multi-modal initial autonomous driving dataset.
[0062] It can be understood that in order to ensure the uniformity and availability of the autonomous driving dataset, the embodiments of the present application may preprocess the multi-view multi-modal initial autonomous driving dataset to obtain a preprocessed multi-view multi-modal initial autonomous driving dataset, and construct a multi-view multi-modal autonomous driving dataset that meets certain format conditions based on the preprocessed multi-view multi-modal initial autonomous driving dataset.
[0063] Among them, certain format conditions can be set by technicians in this field according to actual conditions, and this application does not impose specific restrictions.
[0064] Optionally, in one embodiment of the present application, the multi-view multi-modal initial autonomous driving dataset is preprocessed to obtain a preprocessed multi-view multi-modal initial autonomous driving dataset, and a multi-view multi-modal autonomous driving dataset that meets preset format conditions is constructed based on the preprocessed multi-view multi-modal initial autonomous driving dataset, including: obtaining a final image dataset of the initial image dataset based on an initial image dataset of the multi-view multi-modal initial autonomous driving dataset; obtaining a final point cloud dataset of the initial point cloud dataset based on an initial point cloud dataset of the multi-view multi-modal initial autonomous driving dataset; and obtaining a multi-view multi-modal autonomous driving dataset based on the final image dataset and the final point cloud dataset.
[0065] It can be understood that the storage formats of the initial image dataset and the initial point cloud dataset in the multi-view multi-modal initial autonomous driving dataset of the embodiment of the present application are different. Therefore, the embodiment of the present application can process the initial image dataset and the initial point cloud dataset separately to obtain the final image dataset of the initial image dataset and the final point cloud dataset of the initial point cloud dataset, and then obtain the multi-view multi-modal autonomous driving dataset based on the final image dataset and the final point cloud dataset.
[0066] For example, in the embodiments of the present application, when the initial image data sets are stored in PNG format and the initial point cloud data sets are stored in PCD format, each frame of data is timestamped to ensure accurate alignment in subsequent analysis.
[0067] The following is an introduction to the method for constructing a multi-view and multi-modal autonomous driving dataset proposed in the embodiment of the present application with reference to a specific example.
[0068] in, Figure 7 The present invention is a flowchart of the working principle of a method for constructing a multi-view multi-modal autonomous driving dataset according to an embodiment of the present application.
[0069] Step S701: Add a camera sensor.
[0070] Among them, the embodiments of the present application can be combined with Figure 2 Add a camera sensor.
[0071] Step S702: Determine the operating frequency of the camera sensor.
[0072] Among them, the camera sensor in the embodiment of the present application may include but is not limited to image camera sensors, depth camera sensors, semantic segmentation camera sensors and radar camera sensors, etc., and the present application does not make specific restrictions.
[0073] Step S703: Detect dynamic driving scenes and static driving scenes in autonomous driving.
[0074] Step S704: Construct a multi-view multi-modal initial autonomous driving dataset for autonomous driving.
[0075] Step S705: Construct a multi-view multi-modal autonomous driving dataset that meets certain format conditions.
[0076] According to the method for constructing a multi-view multi-modal autonomous driving dataset proposed in the embodiment of the present application, a multi-view multi-modal initial autonomous driving dataset for autonomous driving can be constructed based on dynamic driving scenes and static driving scenes in autonomous driving, combined with a pre-generated camera sensor dataset, and the multi-view multi-modal initial autonomous driving dataset is pre-processed to obtain a multi-view multi-modal autonomous driving dataset that meets certain format conditions, which can meet a variety of computer vision tasks for autonomous driving scenarios. The attached additional perspectives can be used to detect the reconstruction quality under new perspectives in work such as scene reconstruction. Thus, the problem in the related art that the camera installation position is relatively fixed and the camera's perspective is limited, resulting in significant limitations on the angle diversity of scene reconstruction in these datasets, and the perspective that lacks high-quality true values cannot be effectively evaluated. In addition, the perspectives in these datasets are relatively similar, making it difficult to fully evaluate the reconstruction quality of the model in multi-view scenarios.
[0077] Next, a device for constructing a multi-view multi-modal autonomous driving dataset proposed in an embodiment of the present application is described with reference to the accompanying drawings.
[0078] Figure 8 A block diagram of a device for constructing a multi-view multi-modal autonomous driving dataset according to an embodiment of the present application.
[0079] like Figure 8 As shown, the device 10 for constructing the multi-view multi-modal autonomous driving dataset includes: a detection module 100, a first construction module 200 and a second construction module 300.
[0080] Among them, the detection module 100 is used to detect dynamic driving scenes and static driving scenes in automatic driving.
[0081] The first building module 200 is used to build a multi-view multi-modal initial autonomous driving dataset for autonomous driving based on dynamic driving scenarios and static driving scenarios in combination with a pre-generated camera sensor dataset.
[0082] The second construction module 300 is used to preprocess the multi-view multi-modal initial autonomous driving dataset to obtain a preprocessed multi-view multi-modal initial autonomous driving dataset, and construct a multi-view multi-modal autonomous driving dataset that meets preset format conditions based on the preprocessed multi-view multi-modal initial autonomous driving dataset.
[0083] Optionally, in one embodiment of the present application, it further includes: a first generation module, a second generation module, a third generation module, a fourth generation module and a fifth generation module.
[0084] Among them, the first generation module is used to obtain the image camera sensor dataset in the camera sensor dataset based on the dynamic image information of the dynamic driving scene and the static image information of the static driving scene before building a multi-view multi-modal initial autonomous driving dataset for autonomous driving based on the dynamic driving scene and the static image information of the static driving scene in combination with the pre-generated camera sensor dataset.
[0085] The second generating module is used to obtain a depth camera sensor data set in the camera sensor data set based on the dynamic depth information of the dynamic driving scene and the static depth information of the static driving scene.
[0086] The third generation module is used to obtain a semantically segmented camera sensor data set in the camera sensor data set based on the dynamic semantic segmentation information of the dynamic driving scene and the static semantic segmentation information of the static driving scene.
[0087] The fourth generating module is used to obtain a radar camera sensor data set in the camera sensor data set based on the dynamic point cloud information of the dynamic driving scene and the static point cloud information of the static driving scene.
[0088] The fifth generation module is used to obtain a camera sensor dataset based on the image camera sensor dataset, the depth camera sensor dataset, the semantic segmentation camera sensor dataset and the radar camera sensor dataset.
[0089] Optionally, in an embodiment of the present application, it further includes: a first determination module, a second determination module, a third determination module, a fourth determination module and a fifth determination module.
[0090] Among them, the first determination module is used to determine the image perspective position information of the corresponding image camera sensor in the image camera sensor dataset based on the image camera sensor dataset in the camera sensor dataset before building a multi-view multi-modal initial autonomous driving dataset for autonomous driving based on dynamic driving scenarios and static driving scenarios in combination with a pre-generated camera sensor dataset.
[0091] The second determining module is used to determine the depth viewing angle position information of the corresponding depth camera sensor in the depth camera sensor data set based on the depth camera sensor data set in the camera sensor data set.
[0092] The third determination module is used to determine the semantic segmentation view position information of the corresponding semantic segmentation camera sensor in the semantic segmentation camera sensor data set based on the semantic segmentation camera sensor data set in the camera sensor data set.
[0093] The fourth determination module is used to determine the radar viewing angle position information of the corresponding radar camera sensor in the radar camera sensor data set based on the radar camera sensor data set in the camera sensor data set.
[0094] The fifth determination module is used to generate a camera sensor data set based on the image perspective position information, the depth perspective position information, the semantic segmentation perspective position information and the radar perspective position information.
[0095] Optionally, in one embodiment of the present application, it also includes: a sixth generation module, a seventh generation module, an eighth generation module, a ninth generation module, a judgment module, a third construction module and an adjustment module.
[0096] Among them, the sixth generation module is used to obtain the image working frequency of the corresponding image camera sensor in the image camera sensor data set based on the image camera sensor data set in the camera sensor data set before constructing a multi-view multi-modal initial autonomous driving data set for autonomous driving based on dynamic driving scenarios and static driving scenarios in combination with a pre-generated camera sensor data set.
[0097] The seventh generating module is used to obtain the depth operating frequency of the corresponding depth camera sensor in the depth camera sensor data set based on the depth camera sensor data set in the camera sensor data set.
[0098] The eighth generating module is used to obtain the semantic segmentation working frequency of the corresponding semantic segmentation camera sensor in the semantic segmentation camera sensor data set based on the semantic segmentation camera sensor data set in the camera sensor data set.
[0099] The ninth generating module is used to obtain the radar operating frequency of the corresponding radar camera sensor in the radar camera sensor data set based on the radar camera sensor data set in the camera sensor data set.
[0100] The judgment module is used to judge whether the image working frequency, the depth working frequency, the semantic segmentation working frequency and the radar working frequency are all equal to each other.
[0101] The third building module is used to allow the construction of a multi-view multi-modal initial autonomous driving dataset based on the image working frequency, the depth working frequency, the semantic segmentation working frequency and the radar working frequency when the image working frequency, the depth working frequency, the semantic segmentation working frequency and the radar working frequency are equal to each other.
[0102] An adjustment module is used to adjust at least one of the image operating frequency, the depth operating frequency, the semantic segmentation operating frequency and the radar operating frequency when the image operating frequency, the depth operating frequency, the semantic segmentation operating frequency and the radar operating frequency are not equal to each other, until the image operating frequency, the depth operating frequency, the semantic segmentation operating frequency and the radar operating frequency are all equal to each other.
[0103] Optionally, in one embodiment of the present application, the second building module 300 includes: a first generating unit, a second generating unit and a third generating unit.
[0104] A first generating unit is configured to obtain a final image dataset of the initial image dataset based on an initial image dataset of a multi-view multi-modal initial autonomous driving dataset;
[0105] A second generating unit is used to obtain a final point cloud dataset of the initial point cloud dataset based on the initial point cloud dataset of the multi-view multi-modal initial autonomous driving dataset;
[0106] The third generating unit is used to obtain a multi-view multi-modal autonomous driving dataset based on the final image dataset and the final point cloud dataset.
[0107] It should be noted that the aforementioned explanation of the embodiment of the method for constructing a multi-view multi-modal autonomous driving dataset is also applicable to the device for constructing a multi-view multi-modal autonomous driving dataset of this embodiment, and will not be repeated here.
[0108] According to the device for constructing a multi-view multi-modal autonomous driving dataset proposed in the embodiment of the present application, a multi-view multi-modal initial autonomous driving dataset for autonomous driving can be constructed based on dynamic driving scenes and static driving scenes in autonomous driving, combined with a pre-generated camera sensor dataset, and the multi-view multi-modal initial autonomous driving dataset is pre-processed to obtain a multi-view multi-modal autonomous driving dataset that meets certain format conditions, which can meet a variety of computer vision tasks for autonomous driving scenarios. The attached additional perspectives can be used to detect the reconstruction quality under new perspectives in work such as scene reconstruction. Thus, the problem that the camera installation position is relatively fixed and the camera perspective is limited in the related technology, resulting in significant limitations on the angle diversity of scene reconstruction in these datasets, and the perspective that lacks high-quality true values cannot be effectively evaluated. In addition, the perspectives in these datasets are relatively similar, making it difficult to fully evaluate the reconstruction quality of the model in multi-view scenarios.
[0109] Fig. 9 This is a schematic diagram of the structure of an electronic device provided according to an embodiment of the present application. The electronic device may include:
[0110] A memory 901 , a processor 902 , and a computer program stored in the memory 901 and executable on the processor 902 .
[0111] When the processor 902 executes the program, the method for constructing the multi-view multi-modal autonomous driving dataset provided in the above embodiment is implemented.
[0112] Furthermore, the electronic device further comprises:
[0113] The communication interface 903 is used for communication between the memory 901 and the processor 902 .
[0114] The memory 901 is used to store computer programs that can be executed on the processor 902 .
[0115] The memory 901 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0116] If the memory 901, the processor 902 and the communication interface 903 are implemented independently, the communication interface 903, the memory 901 and the processor 902 can be connected to each other through a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig. 9 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0117] Optionally, in a specific implementation, if the memory 901, the processor 902 and the communication interface 903 are integrated on a chip, the memory 901, the processor 902 and the communication interface 903 can communicate with each other through an internal interface.
[0118] The processor 902 may be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.
[0119] An embodiment of the present application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for constructing a multi-view multi-modal autonomous driving dataset as described above.
[0120] An embodiment of the present application also provides a computer program product, including a computer program, which, when executed, implements the above-mentioned method for constructing a multi-view multi-modal autonomous driving dataset.
[0121] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, without contradiction.
[0122] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of the features. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise clearly and specifically defined.
[0123] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or N executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in reverse order depending on the functions involved, which should be understood by technicians in the technical field to which the embodiments of the present application belong.
[0124] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in combination with these instruction execution systems, devices or apparatuses. For the purpose of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in combination with these instruction execution systems, devices or apparatuses. More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or N wirings (electronic devices), a portable computer disk box (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), a fiber optic device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically by optically scanning the paper or other medium and then editing, interpreting or processing in other suitable ways as necessary and then storing it in a computer memory.
[0125] It should be understood that the various parts of the present application can be implemented by hardware, software, firmware or a combination thereof. In the above embodiment, the N steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. If implemented by hardware, as in another embodiment, it can be implemented by any one or a combination of multiple of the following technologies known in the art: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0126] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.
[0127] In addition, each functional unit in each embodiment of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0128] The storage medium mentioned above may be a read-only memory, a magnetic disk or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limiting the present application. A person of ordinary skill in the art may change, modify, replace and modify the above embodiments within the scope of the present application.
Claims
1. A method for constructing a multi-view multi-modal autonomous driving dataset, characterized in that: The following steps are involved: Detect dynamic and static driving scenarios in autonomous driving; Based on the dynamic driving scene and the static driving scene, combined with a pre-generated camera sensor data set, a multi-view multi-modal initial autonomous driving data set for the autonomous driving is constructed; The multi-view multi-modal initial autonomous driving dataset is preprocessed to obtain a preprocessed multi-view multi-modal initial autonomous driving dataset, and a multi-view multi-modal autonomous driving dataset that meets preset format conditions is constructed based on the preprocessed multi-view multi-modal initial autonomous driving dataset.
2. The method according to claim 1, characterized in that Before constructing a multi-view multi-modal initial autonomous driving dataset for the autonomous driving based on the dynamic driving scenario and the static driving scenario in combination with a pre-generated camera sensor dataset, the method further includes: Based on the dynamic image information of the dynamic driving scene and the static image information of the static driving scene, an image camera sensor data set in the camera sensor data set is obtained; Based on the dynamic depth information of the dynamic driving scene and the static depth information of the static driving scene, a depth camera sensor data set in the camera sensor data set is obtained; Based on the dynamic semantic segmentation information of the dynamic driving scene and the static semantic segmentation information of the static driving scene, a semantically segmented camera sensor data set in the camera sensor data set is obtained; Based on the dynamic point cloud information of the dynamic driving scene and the static point cloud information of the static driving scene, a radar camera sensor data set in the camera sensor data set is obtained; The camera sensor dataset is obtained based on the image camera sensor dataset, the depth camera sensor dataset, the semantic segmentation camera sensor dataset and the radar camera sensor dataset.
3. The method according to claim 2, characterized in that Before constructing a multi-view multi-modal initial autonomous driving dataset for the autonomous driving based on the dynamic driving scenario and the static driving scenario in combination with a pre-generated camera sensor dataset, the method further includes: Based on the image camera sensor data set in the camera sensor data set, determining the image viewing angle position information of the corresponding image camera sensor in the image camera sensor data set; Based on a depth camera sensor data set in the camera sensor data set, determining depth viewing angle position information of a corresponding depth camera sensor in the depth camera sensor data set; Based on the semantic segmentation camera sensor data set in the camera sensor data set, determining the semantic segmentation view position information of the semantic segmentation camera sensor corresponding to the semantic segmentation camera sensor data set; Based on the radar camera sensor data set in the camera sensor data set, determining the radar viewing angle position information of the corresponding radar camera sensor in the radar camera sensor data set; The camera sensor data set is generated based on the image perspective position information, the depth perspective position information, the semantic segmentation perspective position information and the radar perspective position information.
4. The method according to claim 2, characterized in that: Before constructing a multi-view multi-modal initial autonomous driving dataset for the autonomous driving based on the dynamic driving scenario and the static driving scenario in combination with a pre-generated camera sensor dataset, the method further includes: Based on the image camera sensor data set in the camera sensor data set, obtaining the image operating frequency of the corresponding image camera sensor in the image camera sensor data set; Based on a depth camera sensor data set in the camera sensor data set, obtaining a depth operating frequency of a corresponding depth camera sensor in the depth camera sensor data set; Based on the semantic segmentation camera sensor data set in the camera sensor data set, obtaining the semantic segmentation working frequency of the semantic segmentation camera sensor corresponding to the semantic segmentation camera sensor data set; Based on the radar camera sensor data set in the camera sensor data set, obtaining the radar operating frequency of the corresponding radar camera sensor in the radar camera sensor data set; Determining whether the image operating frequency, the depth operating frequency, the semantic segmentation operating frequency, and the radar operating frequency are all equal to each other; If the image operating frequency, the depth operating frequency, the semantic segmentation operating frequency, and the radar operating frequency are all equal to each other, then the multi-view multi-modal initial autonomous driving dataset is allowed to be constructed based on the image operating frequency, the depth operating frequency, the semantic segmentation operating frequency, and the radar operating frequency; If the image operating frequency, the depth operating frequency, the semantic segmentation operating frequency and the radar operating frequency are not equal to each other, adjust at least one of the image operating frequency, the depth operating frequency, the semantic segmentation operating frequency and the radar operating frequency until the image operating frequency, the depth operating frequency, the semantic segmentation operating frequency and the radar operating frequency are all equal to each other.
5. The method according to claim 1, characterized in that The preprocessing of the multi-view multi-modal initial autonomous driving dataset to obtain a preprocessed multi-view multi-modal initial autonomous driving dataset, and constructing a multi-view multi-modal autonomous driving dataset that meets a preset format condition according to the preprocessed multi-view multi-modal initial autonomous driving dataset, includes: Based on the initial image dataset of the multi-view multi-modal initial autonomous driving dataset, obtaining a final image dataset of the initial image dataset; Based on the initial point cloud dataset of the multi-view multi-modal initial autonomous driving dataset, obtaining a final point cloud dataset of the initial point cloud dataset; The multi-view multi-modal autonomous driving dataset is obtained based on the final image dataset and the final point cloud dataset.
6. A device for constructing a multi-view multi-modal autonomous driving dataset, characterized in that: include: A detection module, used to detect dynamic driving scenes and static driving scenes in autonomous driving; A first building module is used to build a multi-view multi-modal initial autonomous driving dataset for the autonomous driving based on the dynamic driving scenario and the static driving scenario in combination with a pre-generated camera sensor dataset; The second construction module is used to preprocess the multi-view multi-modal initial autonomous driving dataset to obtain a preprocessed multi-view multi-modal initial autonomous driving dataset, and construct a multi-view multi-modal autonomous driving dataset that meets preset format conditions based on the preprocessed multi-view multi-modal initial autonomous driving dataset.
7. The device according to claim 6, characterized in that Also includes: A first generating module is configured to obtain an image camera sensor dataset in the camera sensor dataset based on the dynamic image information of the dynamic driving scene and the static image information of the static driving scene before constructing a multi-view multi-modal initial autonomous driving dataset for the autonomous driving in combination with a pre-generated camera sensor dataset based on the dynamic driving scene and the static driving scene; A second generating module is used to obtain a depth camera sensor data set in the camera sensor data set based on the dynamic depth information of the dynamic driving scene and the static depth information of the static driving scene; A third generating module is used to obtain a semantically segmented camera sensor data set in the camera sensor data set based on the dynamic semantic segmentation information of the dynamic driving scene and the static semantic segmentation information of the static driving scene; a fourth generating module, configured to obtain a radar camera sensor data set in the camera sensor data set based on the dynamic point cloud information of the dynamic driving scene and the static point cloud information of the static driving scene; A fifth generating module is used to obtain the camera sensor dataset based on the image camera sensor dataset, the depth camera sensor dataset, the semantic segmentation camera sensor dataset and the radar camera sensor dataset.
8. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method for constructing a multi-view multi-modal autonomous driving dataset as described in any one of claims 1 to 5.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: The program is executed by a processor to implement the method for constructing a multi-view multi-modal autonomous driving dataset as described in any one of claims 1 to 5.
10. A computer program product, characterized in that It includes a computer program, which, when executed, is used to implement the method for constructing a multi-view multi-modal autonomous driving dataset as described in any one of claims 1 to 5.