Multi-modal data adaptive acquisition and labeling method, system and equipment and storage medium
By extracting multi-scale features through a deep belief network and building a data labeling model, the problem of difficulty in integrating multimodal data in autonomous driving data collection and labeling is solved, and more efficient and accurate target object recognition and labeling are achieved.
Patent Information
- Application Number
- CN202510719156.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-09-16
AI Technical Summary
Existing technologies have problems in the process of autonomous driving data collection and labeling, such as difficulty in integrating heterogeneous multimodal data, inconsistent data quality, waste of resources and low efficiency, especially the insufficient accuracy of recognition and labeling in complex scenarios.
A multimodal data adaptive collection and annotation method is adopted to extract multi-scale feature representations through a deep belief network, and the network structure of the data annotation model is constructed. The network parameters are optimized to ultimately achieve automatic annotation of driving scene data.
It improves the accuracy of target object recognition and labeling results in autonomous driving scenarios, enhances the efficiency of data collection and labeling, reduces resource waste, and improves recognition capabilities in complex scenarios.
Smart Images

Figure CN120655972A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of machine learning, and in particular to a method, system, device and storage medium for adaptively collecting and annotating multimodal data. Background Art
[0002] With the rapid development of artificial intelligence (AI), data labeling has become a core, early-stage step in machine learning model training. However, existing data collection still faces numerous challenges. These challenges primarily stem from the traditional manual data collection model, which relies on a point-to-point collaborative architecture between the demand side and the labeling team, and relies on manual labor for data source screening, data collection equipment deployment, and sample collection.
[0003] In the field of autonomous driving, this challenge is particularly prominent. Autonomous driving systems need to rely on high-precision multimodal data (such as lidar point clouds, camera images, millimeter-wave radar signals, etc.) to achieve environmental perception and decision-making. Existing technologies rely on manual or decentralized automation tools, which makes it difficult to effectively integrate heterogeneous sensor data, resulting in inconsistent semantic information (such as spatial alignment errors between point clouds and images, time series asynchrony, etc.). The collection process has a low degree of standardization and uneven data quality, which directly affects the subsequent annotation efficiency and model training effect, and seriously restricts the reliability and safety of autonomous driving vehicles in complex scenarios (such as urban roads and bad weather). In addition, the lack of full-link collaborative management of data collection, annotation, and cleaning leads to resource waste and efficiency bottlenecks, which are specifically manifested as follows:
[0004] 1. Data quality and cost conflicts: Traditional point cloud annotation relies on manual labor, resulting in poor data quality and high costs. Differences in understanding of standards among annotation personnel can lead to data deviations. Faced with large data volumes and complex formats, tool limitations result in an acquisition rate of only 65%.
[0005] 2. Difficulty aligning multimodal heterogeneous data: The semantic information among collected modal data, such as images, point clouds, and text, is inconsistent. For example, the timing of images and audio in a video is misaligned (e.g., audio is delayed by two seconds compared to video). Existing methods have an adaptation rate of less than 85%.
[0006] 3. Difficulty in completing missing data: Up to 80% of data sources contain errors or incomplete data. Existing technologies generate synthetic data through generative adversarial networks (GANs) or self-supervised learning, but there is a risk of model bias introducing annotation noise. The authenticity of the generated data remains questionable, especially in long-tail distributions or small sample sizes. Summary of the Invention
[0007] The present application aims to at least solve the technical problems existing in the prior art and provide a method, system, device and storage medium for adaptive collection and annotation of multimodal data.
[0008] In a first aspect, the present invention provides a method for adaptively collecting and annotating multimodal data, comprising:
[0009] Collect driving scene data in autonomous driving scenarios; driving scene data includes lidar point cloud, RGB image sequence and inertial navigation data;
[0010] Use deep belief networks to extract features from driving scene data and obtain multi-scale feature representations of driving scene data;
[0011] Construct the network structure of the data annotation model and optimize the network parameters of the data annotation model based on the multi-scale feature representation to obtain the final data annotation model;
[0012] The driving scene data is input into the final data annotation model, which extracts the data features of the driving scene data to obtain the three-dimensional point cloud feature information of the driving scene, and determines the target objects in the driving scene and the labels corresponding to the target objects based on the three-dimensional point cloud feature information to obtain the annotation results.
[0013] In a second aspect, the present invention provides a multimodal data adaptive collection and annotation system, the system comprising:
[0014] The acquisition module is used to collect driving scene data in autonomous driving scenarios; the driving scene data includes lidar point cloud, RGB image sequence and inertial navigation data;
[0015] A feature extraction module is used to extract features from driving scene data using a deep belief network to obtain a multi-scale feature representation of the driving scene data;
[0016] The data annotation model construction module is used to build the network structure of the data annotation model and optimize the network parameters of the data annotation model based on the multi-scale feature representation to obtain the final data annotation model;
[0017] The annotation module is used to input the driving scene data into the final data annotation model. The final data annotation model extracts the data features of the driving scene data to obtain the three-dimensional point cloud feature information of the driving scene, and determines the target objects in the driving scene and the labels corresponding to the target objects based on the three-dimensional point cloud feature information to obtain the annotation results.
[0018] In a third aspect, the present invention provides an electronic device, comprising:
[0019] at least one processor; and,
[0020] a memory communicatively connected to the at least one processor; wherein,
[0021] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can perform the multimodal data adaptive acquisition and annotation method described above.
[0022] In a fourth aspect, the present invention also provides a computer-readable storage medium, in which at least one computer program is stored. The at least one computer program is executed by a processor in an electronic device to implement the above-mentioned multimodal data adaptive acquisition and annotation method.
[0023] In summary, this application has the following beneficial technical effects:
[0024] This application first extracts multi-scale feature representations that integrate multimodal information through a deep belief network, then uses the multi-scale feature representations to optimize the network parameters of the data annotation model to obtain the final data annotation model; finally, the final data annotation model is used to automatically predict the target objects and labels of the target objects in the driving scene to improve the annotation efficiency;
[0025] This application can improve the accuracy of complex scene recognition by fusing multimodal features corresponding to lidar point cloud, RGB image sequence and inertial navigation data, thereby improving the accuracy of target object recognition results and labeling results. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 A schematic diagram of a process for adaptively collecting and annotating multimodal data according to an embodiment of the present invention;
[0027] Figure 2 A diagram of a VAE-GAN hybrid architecture provided by one embodiment of the present invention;
[0028] Figure 3 A schematic diagram of the structure of an electronic device for implementing the multimodal data adaptive collection and annotation method provided by one embodiment of the present invention.
[0029] Reference numerals: 10, processor; 11, memory; 12, communication bus; 13, communication interface.
[0030] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION
[0031] The following describes embodiments of the present invention in detail. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended only to explain the present invention and are not to be construed as limiting the present invention.
[0032] In the description of the present invention, it should be understood that the terms "longitudinal", "transverse", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating the orientation or position relationship, are based on the orientation or position relationship shown in the accompanying drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as limiting the present invention.
[0033] In the description of the present invention, unless otherwise specified and limited, it should be noted that the terms "installed", "connected" and "connected" should be understood in a broad sense. For example, it can be a mechanical connection or an electrical connection, or it can be the internal communication between two components. It can be a direct connection or an indirect connection through an intermediate medium. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to the specific circumstances.
[0034] Reference Figure 1 FIG. 1 is a flow chart of a method for adaptively collecting and annotating multimodal data according to an embodiment of the present invention. In this embodiment, the method for adaptively collecting and annotating multimodal data includes:
[0035] S1. Collect driving scene data in an autonomous driving scenario.
[0036] The driving scene data includes a lidar point cloud, an RGB image sequence, and inertial navigation data (such as speed, acceleration, steering angle, etc.). Specifically, the lidar point cloud of the driving scene is collected by a lidar, the RGB image sequence is collected by a camera installed on the driving vehicle, and the inertial navigation data is collected using an inertial navigation device installed on the driving vehicle. In a preferred implementation of this embodiment, the lidar point cloud is a multi-perspective radar point cloud, and the RGB image sequence is a road scene image from different perspectives. The RGB image is a color image format based on the superposition and synthesis of the three primary color channels of red (Red), green (Green), and blue (Blue). The road scene image includes at least one of vehicles, pedestrians, roads, traffic signs, sky, green belts, and obstacles.
[0037] S2. Use the deep belief network to extract features from the driving scene data and obtain a multi-scale feature representation of the driving scene data.
[0038] The full name of deep belief network in English is Deep Bel ief Network, abbreviated as DBN model. Deep belief network is a deep neural network constructed using RBM (Restricted Boltzmann Machines). Restricted Boltzmann machine (RBM) contains visible layer and hidden layer. The two layers are connected by weights, which can perform probabilistic reasoning and learning.
[0039] A deep belief network consists of multiple layers of restricted Boltzmann machines (RBMs), stacked sequentially to form a deep structure. Lower-level RBMs can learn lower-level features (such as image edges and textures), while higher-level RBMs can learn higher-level features (such as object parts and semantic patterns) based on these lower-level features.
[0040] A deep belief network, comprised of multiple layers of restricted Boltzmann machines, is greedily trained layer by layer to obtain multi-scale feature representations of driving scene data. Specifically, each RBM is trained sequentially to learn the data distribution, initializing the network's parameters. Each RBM layer undergoes independent, unsupervised training to learn the data distribution. After pre-training, supervised learning is used to further adjust the network's parameters to suit specific tasks. The deep belief network automatically learns meaningful feature representations from driving scene data, which are critical for subsequent classification, regression, and other tasks.
[0041] The multi-scale feature representation is a multi-modal feature that integrates the contents of lidar point cloud, RGB image sequence and inertial navigation data; specifically, the multi-scale feature representation includes the content of multi-modal information such as point cloud features (such as geometric features and texture features) corresponding to the lidar point cloud, RGB color information and point cloud normal vector information; specifically, feature extraction is performed on the lidar point cloud, and then the extracted features are subjected to voxel grid downsampling and radius filtering denoising to obtain point cloud features; feature extraction is performed on the RGB image to obtain RGB color information, and feature extraction is performed on the inertial navigation data to obtain point cloud normal vector information; the point cloud features, RGB color information and point cloud normal vector information extracted by each layer of restricted Boltzmann machine are integrated to obtain a multi-scale feature representation that integrates multi-modal information; the above method can significantly improve the recognition accuracy and effectively reduce the consumption of computing resources.
[0042] S3. Build the network structure of the data labeling model and optimize the network parameters of the data labeling model based on the multi-scale feature representation to obtain the final data labeling model.
[0043] In this embodiment, the network structure of the data labeling model is a convolutional neural network (CNN). In this embodiment, the robust multi-scale feature representation learned by the deep belief network is passed as input to the convolutional neural network CNN, and the network parameters of the convolutional neural network are trained to obtain the data labeling model.
[0044] When training the data annotation model, we use a three-stage training method (DBN pre-training - CNN initialization training - end-to-end fine-tuning). The low-level parameters of the DBN are frozen for feature extraction, and the high-level parameters are fine-tuned jointly with the CNN. The fusion weights of geometric features and texture features are automatically adjusted through the gated network. The gated network is a learnable module used to dynamically adjust the weights of multimodal features.
[0045] This application improves the accuracy of complex scene recognition by more than 15% through the fusion of multimodal features; at the same time, through clear stage division and quantitative indicator design, it ensures both academic innovation and the feasibility of project implementation.
[0046] S4. Input the driving scene data into the final data annotation model. The final data annotation model extracts the data features of the driving scene data to obtain three-dimensional point cloud feature information of the driving scene, and determines the target object in the driving scene and the label corresponding to the target object based on the three-dimensional point cloud feature information to obtain the annotation result.
[0047] In this embodiment, the target object is an object that the autonomous driving model needs to identify. For example, in a certain autonomous driving scenario, objects such as green belts, signboards, vehicles, and obstacles need to be identified from a picture. Then the green belts, signboards, vehicles, and obstacles that need to be identified are the target objects, and the green belts, signboards, vehicles, and obstacles are the labels corresponding to the target objects.
[0048] After the lidar point cloud is input into the final data annotation model, the final data annotation model processes the driving scene data in the following steps:
[0049] S41, extracting features of the laser radar point cloud to obtain point cloud feature information;
[0050] S42, extracting features of the RGB image sequence in sequence to obtain color feature information;
[0051] S43, extracting features of the inertial navigation data to obtain camera pose feature information;
[0052] S44, determining geometric shape features of a three-dimensional model corresponding to the driving scene based on the point cloud feature information and the camera pose feature information;
[0053] S45, determining texture distribution characteristics of the three-dimensional model corresponding to the driving scene based on the point cloud feature information, the color feature information, and the camera pose feature information;
[0054] S46. Determine the target object in the driving scene based on the geometric shape features and texture distribution features, and predict the label corresponding to the target object to obtain a labeling result.
[0055] In a preferred embodiment of this embodiment, after collecting driving scene data, before inputting the driving scene data into a data annotation model, extracting data features of the driving scene data, and predicting category labels of the driving scene data based on the data features, and obtaining data annotation results, the multimodal data adaptive collection and annotation method further includes:
[0056] S5. Perform spatiotemporal alignment of the multi-perspective lidar point clouds based on the inertial navigation data, and reconstruct a three-dimensional model of the driving scene based on the aligned lidar point clouds to obtain an optimized three-dimensional model; determine the target objects and the labels corresponding to the target objects in the optimized three-dimensional model based on the feature information of the three-dimensional point clouds to obtain the labeling results.
[0057] The three-dimensional model of the driving scene generated by the data annotation model based on the driving scene data may have problems with three-dimensional point cloud defects and occlusions. Using inertial navigation data to temporally and spatially align the multi-perspective lidar point clouds and reconstruct the three-dimensional model of the driving scene based on the aligned lidar point clouds can further improve the accuracy of target object recognition results and the accuracy of the annotation results.
[0058] Specifically, the steps of performing spatiotemporal alignment of multi-view LiDAR point clouds based on inertial navigation data, and reconstructing a three-dimensional model of the driving scene based on the aligned LiDAR point clouds to obtain an optimized three-dimensional model include:
[0059] S51, performing spatiotemporal calibration on the multi-view LiDAR point clouds according to the inertial navigation data to obtain aligned LiDAR point clouds;
[0060] Spatiotemporal alignment includes time alignment and space alignment. Spatiotemporal alignment is a key technology to ensure that the generated or repaired images maintain precise matching in spatial position (geometric consistency) and time sequence (dynamic coherence).
[0061] Spatial alignment ensures that objects under different viewpoints, occlusions, or deformations remain geometrically consistent (for example, the position, shape, and edges of the same object in different images are aligned).
[0062] Temporal alignment can maintain the continuity and temporal consistency of object motion (for example, the motion trajectory of an object in adjacent frames transitions smoothly), thereby improving the accuracy of subsequent point cloud inpainting results.
[0063] S52: stitch the aligned lidar point clouds to construct a globally consistent point cloud scene;
[0064] S53, extracting local geometric features and global semantic features of the point cloud scene;
[0065] S54. Based on local geometric features and global semantic features, the incomplete point cloud in the point cloud scene is completed to obtain an optimized three-dimensional model.
[0066] Residual point clouds refer to point clouds with missing point cloud data due to occlusion or sensor limitations.
[0067] Specifically, based on local geometric features and global semantic features, the incomplete point cloud in the point cloud scene is completed to obtain an optimized 3D model, including:
[0068] S541. Use the attention mechanism to capture the correlation between local geometric features and global semantic features and determine the distribution area of residual cloud;
[0069] S542, predicting the outline of the incomplete defect cloud to be completed based on the global semantic features;
[0070] S543, determining the detailed information of the incomplete defect cloud to be completed based on the local geometric features,
[0071] S544. Complete the residual defect cloud in the point cloud scene according to the distribution area of the residual defect cloud, the outline of the residual defect cloud to be completed, and the detail information to obtain an optimized three-dimensional model.
[0072] Reference Figure 2The data is processed using pix2gestalt, a neural rendering algorithm framework for repairing 3D point cloud data defects. Pix2gestalt is a point cloud completion algorithm based on neural rendering. Combining the detail generation capabilities of generative adversarial networks (GANs) with the probabilistic modeling advantages of variational autoencoders (VAEs), it uses IMU data for spatiotemporal alignment. Point cloud stitching and multi-view fusion based on pose estimation jointly encode local geometric features and global semantic information in sparse regions of the point cloud. A VAE-GAN hybrid architecture (VAE generates basic shapes, GAN refines surface details) uses an attention mechanism to capture cross-regional spatial correlations and learn the underlying distribution of the defective region, achieving high-fidelity point cloud reconstruction. Global contour prediction (80% completion) is followed by local iterative optimization, combining multimodal features such as geometric shape, texture distribution, and normal information to enhance the semantic consistency of the completion result. It gradually adapts to real-world scene data through online adaptive DCL (dynamic curriculum learning). In this embodiment, the entire completion process for the defective point cloud to be repaired is: skeleton generation - topology reconstruction - detail carving - texture mapping. This one-stop output from the original point cloud to a semantic mesh, a full-stack restoration method, supports a hybrid mode of human-machine collaboration. This technology can solve the problem of feature loss in sparse areas, thereby improving data quality while reducing processing costs in adaptive DCL data processing.
[0073] Based on the same inventive concept, an embodiment of the present invention provides a multimodal data adaptive collection and annotation system.
[0074] The multimodal data adaptive acquisition and annotation system described in the present invention can be installed in an electronic device. According to the functions implemented, the multimodal data adaptive acquisition and annotation system includes an acquisition module, a feature extraction module, a data annotation model construction module, and an annotation module. The acquisition module is capable of acquiring driving scene data in an autonomous driving scenario; the driving scene data includes a lidar point cloud, an RGB image sequence, and inertial navigation data; the feature extraction module is capable of extracting features from the driving scene data using a deep belief network to obtain a multi-scale feature representation of the driving scene data; the data annotation model construction module is capable of constructing a network structure for the data annotation model and optimizing the network parameters of the data annotation model based on the multi-scale feature representation to obtain a final data annotation model; and the annotation module is capable of inputting the driving scene data into the final data annotation model. The final data annotation model extracts data features from the driving scene data to obtain three-dimensional point cloud feature information of the driving scene, and determines target objects in the driving scene and corresponding labels of the target objects based on the three-dimensional point cloud feature information to obtain an annotation result.
[0075] The module described in the present invention may also be referred to as a unit, which refers to a series of computer program segments that can be executed by a processor of an electronic device and can complete fixed functions, and is stored in a memory of the electronic device.
[0076] The various variations and specific examples of the multimodal data adaptive acquisition and annotation method provided in the above embodiment are also applicable to the multimodal data adaptive acquisition and annotation system of this embodiment. Through the above detailed description of the multimodal data adaptive acquisition and annotation method, those skilled in the art can clearly understand the implementation method of the multimodal data adaptive acquisition and annotation system of this embodiment. For the sake of brevity of the specification, it will not be described in detail here.
[0077] This application also discloses an electronic device, such as Figure 3 Figure 2 is a schematic diagram of the structure of an electronic device for a method for adaptively collecting and annotating multimodal data according to an embodiment of the present invention. The electronic device may include at least one processor 10, a memory 11 communicatively connected to the at least one processor, a communication bus 12, and a communication interface 13. The electronic device may also include a computer program stored in the memory 11 and executable on the processor 10, such as a program for the method for adaptively collecting and annotating multimodal data.
[0078] Among them, in some embodiments, the processor 10 can be composed of an integrated circuit, for example, it can be composed of a single packaged integrated circuit, or it can be composed of multiple integrated circuits with the same or different functions, including one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and a combination of various control chips. The processor 10 is the control core (Control Unit) of the electronic device, which uses various interfaces and lines to connect the various components of the entire electronic device, and executes or executes the programs or modules stored in the memory 11 (such as the method for executing the adaptive acquisition and annotation of multimodal data, etc.), and calls the data stored in the memory 11 to perform various functions of the electronic device and process data.
[0079] The memory 11 includes at least one type of readable storage medium, and the readable storage medium includes a flash memory, a mobile hard disk, a multimedia card, a card-type memory (for example, an SD or DX memory, etc.), a magnetic memory, a disk, an optical disk, etc. In some embodiments, the memory 11 may be an internal storage unit of an electronic device, such as a mobile hard disk of the electronic device. In other embodiments, the memory 11 may also be an external storage device of an electronic device, such as a plug-in mobile hard disk, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc. equipped on the electronic device. Furthermore, the memory 11 may also include both an internal storage unit and an external storage device of the electronic device. The memory 11 can not only be used to store application software and various types of data installed in the electronic device, such as the code of the method program for adaptive acquisition and annotation of multimodal data, but can also be used to temporarily store data that has been output or is to be output.
[0080] The communication bus 12 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. The bus is configured to enable communication between the memory 11 and at least one processor 10, etc.
[0081] The communication interface 13 is used for communication between the above-mentioned electronic device and other devices, including a network interface and a user interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is generally used to establish a communication connection between the electronic device and other electronic devices. The user interface may be a display (Display), an input unit (such as a keyboard (Keyboard)), optionally, the user interface may also be a standard wired interface, a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, and an OLED (Organic Light-Emitting Diode, organic light-emitting diode) touch device, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, for displaying information processed in the electronic device and for displaying a visual user interface.
[0082] Figure 3 Only the electronic device with components is shown, and it can be understood by those skilled in the art that Figure 3The structure shown does not constitute a limitation of the electronic device, and may include fewer or more components than shown, or combine certain components, or arrange the components differently. For example, although not shown, the electronic device may also include a power supply (such as a battery) to power each component. Preferably, the power supply can be logically connected to at least one processor 10 through a power management device, so that functions such as charging management, discharging management, and power consumption management are implemented through the power management device. The power supply may also include one or more DC or AC power supplies, recharging devices, power failure detection circuits, power converters or inverters, power status indicators, and other arbitrary components. The electronic device may also include a variety of sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be repeated here.
[0083] It should be understood that the embodiment is for illustration only and the scope of the patent application is not limited to this structure.
[0084] Furthermore, if the module / unit integrated into the electronic device is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile.
[0085] The present application provides a computer-readable storage medium, for example, including any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard drive, a magnetic disk, an optical disk, a computer memory, or a read-only memory (ROM). The computer-readable storage medium stores a computer program capable of being loaded by a processor and executing the multimodal data adaptive acquisition and annotation method of the above-described embodiment.
[0086] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "example," "specific example," "one implementation," "a preferred implementation," or "some examples" means that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0087] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to the embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the claims and their equivalents.
Claims
1. A method for adaptively collecting and labeling multimodal data, characterized in that: The method comprises: Collect driving scene data in autonomous driving scenarios; driving scene data includes lidar point cloud, RGB image sequence and inertial navigation data; A deep belief network is used to extract features from driving scene data to obtain a multi-scale feature representation of the driving scene data. The multi-scale feature representation is a multimodal feature that integrates the contents of lidar point cloud, RGB image sequence, and inertial navigation data. Construct the network structure of the data annotation model and optimize the network parameters of the data annotation model based on the multi-scale feature representation to obtain the final data annotation model; The driving scene data is input into the final data annotation model, which extracts the data features of the driving scene data to obtain the three-dimensional point cloud feature information of the driving scene, and determines the target objects in the driving scene and the labels corresponding to the target objects based on the three-dimensional point cloud feature information to obtain the annotation results.
2. The method for adaptively collecting and labeling multimodal data according to claim 1, wherein: The deep belief network includes a multi-layer stack of restricted Boltzmann machines. The deep belief network is used to extract features from driving scene data to obtain a multi-scale feature representation of the driving scene data, including: A deep belief network stacked with multiple layers of restricted Boltzmann machines is trained greedily layer by layer to obtain multi-scale feature representation of driving scene data.
3. The multimodal data adaptive collection and annotation method according to claim 1, characterized in that: After the lidar point cloud is input into the final data annotation model, the final data annotation model processes the driving scene data in the following steps: Extract the features of the lidar point cloud to obtain point cloud feature information; Extract the features of the RGB image sequence in sequence to obtain color feature information; Extract the features of inertial navigation data to obtain camera pose feature information; Determine the geometric shape features of the 3D model corresponding to the driving scene based on the point cloud feature information and the camera pose feature information; Determine the texture distribution characteristics of the three-dimensional model corresponding to the driving scene based on point cloud feature information, color feature information, and camera pose feature information; The target objects in the driving scene are determined based on the geometric shape features and texture distribution features, and the labels corresponding to the target objects are predicted to obtain the labeling results.
4. The method for adaptively collecting and labeling multimodal data according to claim 3, wherein: Before inputting the driving scene data into the data annotation model, the data annotation model extracting data features of the driving scene data, and predicting category labels of the driving scene data based on the data features to obtain the data annotation results, the method further includes: Perform spatiotemporal alignment of multi-view LiDAR point clouds based on inertial navigation data, and reconstruct a 3D model of the driving scene based on the aligned LiDAR point clouds to obtain an optimized 3D model. The target objects and labels corresponding to the target objects in the optimized 3D model are determined based on the 3D point cloud feature information to obtain the labeling results.
5. The multimodal data adaptive collection and annotation method according to claim 4, characterized in that: The method of performing spatiotemporal alignment of the multi-view LiDAR point clouds based on the inertial navigation data and reconstructing a three-dimensional model of the driving scene based on the aligned LiDAR point clouds to obtain an optimized three-dimensional model includes: Perform spatiotemporal calibration on the multi-view LiDAR point cloud based on the inertial navigation data to obtain the aligned LiDAR point cloud; Splice the aligned lidar point clouds to construct a globally consistent point cloud scene; Extract local geometric features and global semantic features of point cloud scenes; Based on local geometric features and global semantic features, the incomplete point cloud in the point cloud scene is completed to obtain an optimized 3D model.
6. The multimodal data adaptive collection and annotation method according to claim 1, characterized in that: The method of completing the incomplete point cloud in the point cloud scene based on the local geometric features and the global semantic features to obtain the optimized 3D model includes: The attention mechanism is used to capture the correlation between local geometric features and global semantic features to determine the distribution area of residual cloud. Predict the outline of the incomplete defect cloud to be completed based on global semantic features; Determine the detailed information of the defect cloud to be completed based on local geometric features. The residual defect clouds in the point cloud scene are completed according to the distribution area of the residual defect clouds, the outline and detail information of the residual defect clouds to be completed, and an optimized three-dimensional model is obtained.
7. A multimodal data adaptive collection and annotation system, used to implement the multimodal data adaptive collection and annotation method according to any one of claims 1 to 6, characterized in that: include: The acquisition module is used to collect driving scene data in autonomous driving scenarios; the driving scene data includes lidar point cloud, RGB image sequence and inertial navigation data; A feature extraction module is used to extract features from driving scene data using a deep belief network to obtain a multi-scale feature representation of the driving scene data; Multi-scale features are represented as multimodal features that fuse lidar point cloud, RGB image sequence and inertial navigation data content. The data annotation model construction module is used to build the network structure of the data annotation model and optimize the network parameters of the data annotation model based on the multi-scale feature representation to obtain the final data annotation model; The annotation module is used to input the driving scene data into the final data annotation model. The final data annotation model extracts the data features of the driving scene data to obtain the three-dimensional point cloud feature information of the driving scene, and determines the target objects in the driving scene and the labels corresponding to the target objects based on the three-dimensional point cloud feature information to obtain the annotation results.
8. An electronic device, characterized in that: The electronic device comprises: at least one processor (10); and, a memory (11) communicatively coupled to the at least one processor (10); The memory (11) stores a computer program executable by the at least one processor (10), and the computer program is executed by the at least one processor (10) so that the at least one processor (10) can execute the multimodal data adaptive acquisition and annotation method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program; when the computer program is executed by a processor, the multimodal data adaptive acquisition and annotation method according to any one of claims 1 to 6 is implemented.
Citation Information
Cited By
Driving scene recognition method and device and storage medium
CN121305496A