A method and system for joint enhancement of deep consistency multimodal data

By using a multimodal data joint augmentation method, augmented samples with consistent density are generated, which solves the long-tail problem in autonomous driving systems, improves the system's generalization ability and target detection accuracy, and enhances the system's safety.

CN119445298BActive Publication Date: 2025-10-31TONGJI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411320613.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-20
Publication Date
2025-10-31
Estimated Expiration
2044-09-20

AI Technical Summary

Technical Problem

During the training process, autonomous driving systems suffer from a long-tail problem due to insufficient extreme case samples, which affects the system's generalization ability and performance.

Method used

By employing a multimodal data joint augmentation method based on depth consistency, we can utilize a pre-collected multimodal dataset to expand extreme cases, generate densely consistent augmented sample point clouds and RGB images, and balance the class distribution of the dataset.

Benefits of technology

It improves the generalization performance and target detection accuracy of the autonomous driving system, enhances the system's safety, and avoids domain offset and high-cost manual data collection issues.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119445298B_ABST
    Figure CN119445298B_ABST
Patent Text Reader

Abstract

This case relates to the field of 3D target detection, specifically a multimodal data joint enhancement method based on depth consistency to address the problem of imbalanced class distribution. The steps include: generating a 3D target object based on the RGB pixel blocks corresponding to the target object; copying the target object point cloud from the 3D target object to the point cloud of the sample to be enhanced; acquiring the RGB sample image to be enhanced corresponding to the point cloud of the sample to be enhanced; projecting the target object point cloud from the point cloud of the sample to be enhanced onto the RGB sample image of the RGB sample image of the sample to be enhanced to obtain the enhanced RGB sample image; determining the current point cloud density based on the current distance between the center point of the target object point cloud and the radar using the distance-density linear relationship; and resampling the target object point cloud based on the current point cloud density to obtain an enhanced sample point cloud with consistent density. This scheme can be used to generate application scenarios belonging to long-tail categories, thus achieving a balanced class distribution.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of three-dimensional target detection, and in particular to a method and system for joint enhancement of multimodal data based on depth consistency. Background Technology

[0002] For multimodal data (point cloud and RGB images) collected from autonomous vehicles, it is necessary to analyze some extreme cases or long-tail cases, such as traffic cones, which are often ignored due to their low probability of occurrence in the samples. These cases are often difficult to collect, but they are often more important during the training of autonomous driving systems. These low-probability cases lead to the long-tail problem, which in turn affects the generalization ability of the autonomous driving system. In other words, if there are too few samples for extreme cases during autonomous driving training, the vehicle will assume that such situations will not occur, thus affecting overall performance. Summary of the Invention

[0003] The purpose of this study is to propose a method and system for joint augmentation of multimodal data based on deep consistency. Utilizing a pre-collected multimodal dataset, it automatically augments the data to obtain multimodal data that conforms to the distribution of real samples, thereby achieving a balanced distribution of categories in the multimodal dataset. This method is applicable to augmenting long-tail cases in autonomous driving scenarios.

[0004] To achieve the above objectives, the specific technical solution of this case is as follows.

[0005] Firstly, this paper proposes a multimodal data joint enhancement method based on depth consistency. The method is used to acquire multimodal enhancement data, which includes enhanced RGB sample images and their corresponding enhanced sample point clouds. The steps include: generating a three-dimensional target object based on the RGB pixel blocks corresponding to the target object; copying the target object point cloud in the three-dimensional target object to the sample point cloud to be enhanced; acquiring the RGB sample image to be enhanced corresponding to the sample point cloud to be enhanced; projecting the target object point cloud in the sample point cloud to be enhanced onto the RGB sample image to be enhanced to obtain the enhanced RGB sample image; determining the current point cloud density using the distance-density linear relationship based on the current distance between the center point of the target object point cloud and the radar; and resampling the target object point cloud based on the current point cloud density to obtain an enhanced sample point cloud with consistent density.

[0006] In one embodiment of the above technical solution, after copying the target object point cloud from the three-dimensional target object to the sample point cloud to be enhanced, the method performs attitude adjustment on the target object point cloud, the attitude adjustment including adjustment of direction and angle.

[0007] In one embodiment of the above technical solution, a three-dimensional target object is generated based on the RGB pixel blocks corresponding to the target object. The steps include: acquiring a first sample point cloud containing the target object and its corresponding first RGB sample image; performing instance segmentation based on the first RGB sample image to obtain several RGB pixel blocks including the target object, and then obtaining the point cloud block corresponding to each RGB pixel block in the first sample point cloud; cropping based on the RGB pixel blocks corresponding to the target object, and generating a three-dimensional target object using the Nerf algorithm.

[0008] In one embodiment of the above technical solution, the steps to obtain the point cloud block corresponding to each RGB pixel block in the first sample point cloud include: projecting the first sample point cloud onto the first RGB sample image to obtain the position of each point in the first sample point cloud in the first RGB sample image; and for each RGB pixel block, determining the point cloud block corresponding to the RGB pixel block in the first sample point cloud based on the position of each point in the first sample point cloud in the first RGB sample image.

[0009] In one embodiment of the above technical solution, the distance-density function determination step includes: randomly selecting two object point cloud blocks from the first sample point cloud, obtaining the distance d between the center of each object point cloud block and the radar, and obtaining the density ρ of each object point cloud block; based on the distance-density linear relationship d=αρ+β, obtaining a system of two linear equations in two variables about α and β, and then determining the distance-density linear relationship.

[0010] In one embodiment of the above technical solution, the step of obtaining the density ρ of the object point cloud block includes: voxelizing the object point cloud block to obtain the point cloud quantity matrix PN in n grids, the distribution of which satisfies PN~N(μ, σ 2 ), N(μ, σ 2 ) is a Gaussian distribution, μ is the mean number of point clouds in each grid cell, and σ is the mean number of point clouds in each grid cell. 2 The variance of the point cloud count; PN i Let be the number of point clouds in the i-th grid.

[0011] In one embodiment of the above technical solution, the resampling method includes uniform sampling, farthest point sampling, and random sampling.

[0012] Secondly, this case proposes an autonomous driving system, which includes a target detection and classification model, and the target detection and classification model is trained using multimodal augmented data obtained by any of the above methods.

[0013] Thirdly, this case proposes a computer-readable storage medium storing a computer program that can be loaded by a processor and executed as in any of the above methods.

[0014] The beneficial technical effects of this method are as follows: Compared with simulation-based data augmentation methods, this method avoids the domain shift problem caused by multimodal data augmentation. Compared with manual collection methods, the collection efficiency is higher, and the quality of samples can be better controlled. Compared with single-modal data augmentation, this method can effectively utilize multimodal data information. This method can generate more extreme cases, thereby balancing the data distribution. When applied to autonomous driving, it can enhance the generalization performance of autonomous driving systems. When used as an autonomous driving dataset for training, it can improve the accuracy of target detection in autonomous driving systems, thereby improving their safety. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 , one A schematic diagram of the method flow in one implementation method.

[0017] Figure 2 , one A schematic diagram of the camera coordinate system, radar coordinate system, and pixel coordinate system in one implementation method.

[0018] Figure 3 , one A schematic diagram illustrating the projection of a point cloud onto an RGB image in one implementation method.

[0019] Figure 4 , one A schematic diagram of instance segmentation results of the pre-trained model on sample RGB images in one implementation method.

[0020] Figure 5 , one A schematic diagram of the point cloud blocks corresponding to the RGB image of the sample in one implementation method.

[0021] Figure 6 , one A high-fidelity 3D object diagram generated from a sample RGB image using the Nerf algorithm in one implementation method.

[0022] Figure 7 , one A schematic diagram illustrating the RGB copying of one object to the RGB of another sample in one implementation method.

[0023] Figure 8 , one A schematic diagram illustrating the copying of an object point cloud to a sample point cloud in one implementation method.

[0024] Figure 9 , one A schematic diagram of the object point cloud resampling results in one implementation method. Detailed Implementation

[0025] To address the long-tail problem in existing autonomous driving datasets, current methods include:

[0026] (1) Driving scenario simulation: By building a simulation platform, extreme scenarios are constructed in the simulation environment settings to enrich the data. However, the data collected by this method has a different distribution from the data in real scenarios, which will cause domain shift.

[0027] (2) Based on large-scale manual collection: Collecting a large number of extreme cases can enrich the data distribution. However, this method consumes a lot of time and resources to collect data. Manual collection requires a lot of manpower and time investment, especially when large-scale data collection is required.

[0028] (3) Single-modal data augmentation: For RGB image data, methods such as random cropping, random flipping, and color dithering are used to augment the image data. For point cloud data: New point cloud data is obtained through copying and pasting, GAN network generation, and CAD models, thereby generating a large number of samples. Single-modal data augmentation methods are not suitable for autonomous driving scenarios and cannot meet the requirements of the development of multimodal data in autonomous driving scenarios.

[0029] Based on this, this paper proposes a multimodal data joint augmentation method based on depth consistency. Utilizing pre-collected data with long-tail categories, the method augments the dataset to expand small sample classes and increase the number of long-tail categories, thereby achieving a balanced distribution of categories in the multimodal dataset. A flowchart of the method is shown below. Figure 1 As shown, the steps include:

[0030] A 3D target object is generated based on the RGB pixel blocks corresponding to the target object, and the point cloud of the target object in the 3D target object is copied to the point cloud of the sample to be enhanced.

[0031] Obtain the RGB sample image corresponding to the point cloud of the sample to be enhanced, and project the point cloud of the target object in the point cloud of the sample to be enhanced onto the RGB sample image to obtain the enhanced RGB sample image;

[0032] The current point cloud density is determined by using the distance-density linear relationship based on the current distance between the center point of the target object's point cloud and the radar.

[0033] Based on the current point cloud density, the point cloud of the sample to be enhanced is resampled to obtain an enhanced point cloud with consistent density.

[0034] Figure 1 Operations in the flowchart can be performed out of order. Instead, operations can be performed in reverse order or simultaneously. Furthermore, one or more additional operations can be added to the flowchart. One or more operations can be removed from the flowchart.

[0035] The following is a clear and complete description of how this case utilizes a small sample size to augment multimodal data that conforms to the distribution of real samples, thereby achieving a balanced class distribution in the multimodal dataset. Obviously, the described implementation methods are only a part of the implementation methods in this case, and not all of them. All other implementation methods obtained by those skilled in the art based on the implementation methods in this case without inventive effort are within the scope of protection of this application.

[0036] In this case, point cloud data refers to spatial coordinate point data, and may also include texture data, intensity data, etc. The dataset includes multimodal data for autonomous driving, and in some implementations, it also includes multimodal data acquired in indoor scenes. The multimodal data includes point cloud data and RGB images, with point cloud data acquired by one radar camera and RGB image data acquired by multiple cameras.

[0037] First, establish the camera coordinate system, LiDAR coordinate system, and pixel coordinate system. Their positional relationships are as follows: Figure 2 As shown, the lidar coordinate system is O. L X L Y L Z L The camera coordinate system is O C X C Y C Z C The camera acquires an image with pixel coordinates O. l UV. The aforementioned lidar models can be 32-line, 64-line, synthetic aperture, etc., but this case does not limit the lidar model.

[0038] 3D point cloud data can be generated using the following methods: 1) image-based generation, 2) radar-based acquisition, and 3) depth camera-based acquisition.

[0039] Acquire multimodal data containing the target object. This multimodal data includes a sample point cloud and corresponding RGB sample images. There can be one or multiple RGB sample images. The following method uses one image as an example; if there are multiple images, they can be processed sequentially. By performing statistical analysis on the acquired dataset beforehand, the categories belonging to the long-tail distribution can be determined. The categories belonging to the long-tail distribution are taken as the target classes, and objects belonging to these classes are the target objects.

[0040] The acquired sample point cloud containing the target object is used as the first sample point cloud, and its corresponding sample image is used as the first RGB sample image.

[0041] When the point cloud position of the target object is denoted as P, and it is projected onto the pixel coordinate system, the coordinates of the projected pixel are the coordinates of the point in the 3D point cloud in the image. Figure 3 The coordinates of the projected pixel are marked as (u, v), and the formula for calculating the corresponding pixel is:

[0042]

[0043] In the formula, sample P Let P be the point cloud data of the object point cloud. rect R represents the intrinsic parameter matrix of the camera. rect This represents the correction matrix for the camera. This is the extrinsic parameter matrix for transforming the radar coordinate system to the camera coordinate system.

[0044] By projecting the first sample point cloud onto the corresponding first RGB sample image, the position of each point in the first sample point cloud within the first RGB sample image can be obtained. (See also...) Figure 3 , Figure 3 The lower half shows the coordinate markings in the projection image of the point cloud data.

[0045] A pre-trained instance segmentation network (such as SAM or Mask-RCNN) is used to predict the first sample RGB image data to obtain multiple instance pixel blocks (see...). Figure 4 (Example), the number of instances depends on the scenario. See also Figure 4 The instance pixel blocks include cars, trees, billboards, and figures on the billboards. Based on the instance pixel blocks and the obtained projection positions, the corresponding point cloud block for each RGB pixel block in the first sample point cloud can be obtained.

[0046] The point cloud data corresponding to each instance pixel block is used as the object point cloud block, and the object point cloud block is bound to the RGB pixel block (see [link]). Figure 5 (Example) A one-to-one correspondence is established, allowing for accurate determination of the object's position and pose in the RGB image based on the projection position of the point cloud in the image. For each RGB pixel block, the corresponding point cloud block is obtained through its projection position. By binding the object's point cloud block with the RGB pixel block, both the object's geometric shape information and visual features are obtained.

[0047] Use an instance segmentation network to obtain the corresponding RGB pixel blocks. Represent the instance segmentation result of the target class as obj. RGB =seg(.sample RGB ),obj RGBThat is, the pixel block corresponding to the target class, sample RGB This represents the sample RGB image data, and seg represents the instance segmentation step.

[0048] When the projected point cloud lies within the segmented instance pixel block, the pairing process between the point cloud and the instance pixel block can be expressed as: {obj P obj RGB}, obj P This represents the point cloud data of the target object, (x, y, z) ∈ obj P obj RGB Represents the instance pixel block corresponding to the target object, (u, v) ∈ obj RGB This indicates the projection position of the point cloud onto the RGB image. A point cloud projected onto an instance's RGB pixel block is a point cloud block of an object.

[0049] Next, two object point cloud patches are randomly selected from the sample point clouds. Based on the distance between the object point cloud patch and the radar, and the point cloud density, a distance-density linear relationship is established in the sample point clouds. Specifically, the following steps are included:

[0050] Let P be a randomly selected object point cloud patch from the sample point cloud. Voxelize the object point cloud patch P to obtain a point cloud quantity matrix PN of n grid cells, where PN = {PN0, PN1, ..., PN}. n-1}. PN i Let be the number of point clouds in the i-th grid, i = 0, 1, ..., n-1. Assume the point clouds follow a Gaussian distribution, i.e., PN ~ N(μ, σ). 2 μ is the mean number of point clouds in each raster. σ 2 Let V be the variance of the number of point clouds.

[0051] The probability density can be expressed as The sum of the probability densities of all points is the point cloud density of the object's point cloud patch, which can be represented by the density evaluation function, and can be expressed as: The function exp(x) represents the natural exponential function.

[0052] Based on the density evaluation function, a distance-density linear function is established as follows:

[0053] d = αρ + β, where d represents the distance from the center of the point cloud patch to the radar, and α and β are the slope and intercept, respectively. α and β are unknowns to be determined.

[0054] Given a point cloud block P containing K points, find the distance between the center of the point cloud block and the radar. Let be the coordinates of the i-th point in the object point cloud P. Let be the density of the object point cloud block. The function exp(x) represents the natural exponential function, which can obtain a bivariate equation for the distance-density linear function d = αρ + β, where α and β are unknowns.

[0055] Similarly, based on another object point cloud patch, another bivariate equation for the distance-density linear function can be obtained. Using these two bivariate equations, the values ​​of α and β can be solved, thus obtaining the distance-density function.

[0056] Then, the RGB pixel blocks corresponding to the target object are cropped, and a high-fidelity 3D target object is generated using the Nerf algorithm (see [link]). Figure 6 (Example) This object can be scaled or rotated to generate objects in different poses. The specific algorithm used in Nerf is not limited; it can be Make-It-3D, SinNeRF, MCC, Stable Diffusion, etc. The generated 3D target object is mesh data, composed of dense point clouds and color, with six dimensions: X, Y, Z, R, G, and B. X, Y, and Z represent positional information, while R, G, and B represent color information. The dense point cloud provides the geometric shape information of the object, while the color adds realistic texture and detail to the object's surface. This high-fidelity 3D object representation can provide more accurate and detailed input data for further computer vision tasks, such as object recognition and scene reconstruction. By using the Nerf algorithm to generate such 3D objects, the representation of object categories in the dataset can be enriched, improving the algorithm's generalization performance. The generated high-fidelity 3D target object is represented as 3Dobj. P,RGB Then: 3Dobj P,RGB =Nerf(obj) RGB ).

[0057] Copying the point cloud of the target 3D object to the point cloud of the sample to be enhanced can be represented as newsample. P =copyandpaste(3Dobj) P,RGB `copyandpaste` represents a copy operation. By rotating and copying the target object's point cloud into the point cloud of the sample to be enhanced, the pose of the object at different angles and directions can be simulated, thus adjusting the pose of the target object's point cloud. The point cloud to be enhanced can be the original first sample point cloud or other sample point clouds.

[0058] Since the point cloud of the target object is a dense point cloud, it can be directly projected as an RGB image, thus providing more comprehensive data information (see [link]). Figure 7 Example). Project the point cloud of the target object onto the corresponding RGB sample image of the point cloud to be enhanced, to obtain the enhanced RGB sample image. This process can be represented as: newsample RGB=proj(3Dobj) P,RGB |newsample P ), 3Dobj P,RGB |newsample P This indicates that the projection is performed only on the target object data in the copied object point cloud, and proj represents the point cloud projection process.

[0059] After the point cloud is copied, the point cloud's position information changes in real time. Using this point cloud position information, the current distance between the radar and the center of the target object's point cloud can be obtained, such as... Figure 8 As shown. Figure 8 The bright patches in the image represent the copied target object's point cloud. The current distance between the new radar and the target object's point cloud center at the new location can be directly obtained from the point cloud data. Finally, based on the current distance d′ between the radar and the target object's point cloud center, and the already determined distance-density linear function of α and β, the current point cloud density can be calculated. Based on the point cloud density ρ′, the point cloud of the sample to be enhanced, which has been copied from the point cloud of the target object, is resampled to obtain an enhanced point cloud with consistent density. See also Figure 9 The disappearance of bright patches in the image indicates that the density of the point cloud is consistent after resampling. This process can be represented as:

[0060] newsample P =resampling score(PN′) (3Dobj P,RGB )

[0061] Among them, 3Dobj P,RGB This indicates that only the point cloud data of newly added objects in the point cloud sample to be enhanced is resampled, without affecting other data of existing objects in the new point cloud sample. Methods such as uniform sampling, farthest point sampling, and random sampling can be used to sample point cloud data of a certain density.

[0062] During resampling, the object point cloud is sampled based on the learned distance-density function. This allows for appropriate supplementation of sparse regions and appropriate simplification of dense regions. In other words, the density should remain consistent with the original data density within the same 3D point cloud depth. This ensures that the distribution of the resampled sample data is similar to the original data, while reducing data volume and computational complexity. Resampling the object point cloud yields a new set of more representative and complete sample data. This data can be used to train and evaluate machine learning models, thereby improving the model's generalization ability and stability. Furthermore, resampling provides more reliable and effective input for subsequent data processing and analysis tasks.

[0063] This case does not impose any restrictions on the settings for single-modal enhancement; enhancement can be performed using only RGB images or point cloud data with only minor modifications.

[0064] This paper proposes a multimodal data joint augmentation method based on deep consistency, starting from the needs of high efficiency and multimodal development in autonomous driving. This method preprocesses imbalanced multimodal data, generating more extreme scenarios for autonomous driving. It is a more effective and scientific technique. Using these augmented datasets for autonomous driving training can improve the accuracy of target detection in the system, thereby enhancing its safety. This method can also be applied to mapping and remote sensing fields.

[0065] Through the above description of the embodiments, those skilled in the art can clearly understand that the method of this disclosure can be implemented by means of software plus necessary general-purpose hardware, and of course, it can also be implemented by special-purpose hardware, including application-specific integrated circuits, dedicated CPUs, dedicated memory, dedicated components, etc. Generally, any function performed by a computer program can be easily implemented by corresponding hardware, and the specific hardware structure used to implement the same function can also be diverse, such as analog circuits, digital circuits, or dedicated circuits. However, for the purposes of this disclosure, software program implementation is more often a preferred implementation method.

[0066] Although embodiments of the present invention have been described above in conjunction with the accompanying drawings, the present invention is not limited to the specific embodiments and application fields described above. The specific embodiments described above are merely illustrative and instructive, and not restrictive. Those skilled in the art can make many other forms based on the guidance of this specification and without departing from the scope of protection of the claims of the present invention, and all of these are within the scope of protection of the present invention.

Claims

1. A multimodal data joint enhancement method based on depth consistency, characterized in that, The method is used to acquire multimodal augmentation data, which includes augmented RGB sample images and their corresponding augmented sample point clouds. The steps include: A 3D target object is generated based on the RGB pixel blocks corresponding to the target object, and the point cloud of the target object in the 3D target object is copied to the point cloud of the sample to be enhanced. Obtain the RGB sample image corresponding to the point cloud of the sample to be enhanced, and project the point cloud of the target object in the point cloud of the sample to be enhanced onto the RGB sample image to obtain the enhanced RGB sample image; The current point cloud density is determined using the distance-density linear relationship based on the current distance between the center point of the target object's point cloud and the radar. Resample the target object's point cloud based on the current point cloud density to obtain an enhanced sample point cloud with consistent density. The steps of generating a 3D target object based on the RGB pixel blocks corresponding to the target object include: acquiring a first sample point cloud containing the target object and its corresponding first RGB sample image; performing instance segmentation based on the first RGB sample image to obtain several RGB pixel blocks including the target object, and then obtaining the point cloud block corresponding to each RGB pixel block in the first sample point cloud; cropping based on the RGB pixel blocks corresponding to the target object, and generating a 3D target object using the Nerf algorithm. The steps of obtaining the point cloud block corresponding to each RGB pixel block in the first sample point cloud include: projecting the first sample point cloud onto the first RGB sample image to obtain the position of each point in the first sample point cloud in the first RGB sample image; and for each RGB pixel block, determining the point cloud block corresponding to the RGB pixel block in the first sample point cloud based on the position of each point in the first sample point cloud in the first RGB sample image. The distance-density linear relationship is determined by the following steps: randomly selecting two object point cloud blocks from the first sample point cloud, and obtaining the distance between the center of each object point cloud block and the radar. Obtain the density of each object point cloud patch. Based on the distance-density linear relationship Get information about and The system of two linear equations in two variables is used to determine the linear relationship between distance and density. The density of the object point cloud patch The acquisition steps include: converting the object's point cloud into voxels to obtain a matrix of the number of point clouds in n grid cells. Its distribution satisfies , The distribution follows a Gaussian pattern, where μ is the mean number of point clouds in each grid cell. The variance of the point cloud count; , Let be the number of point clouds in the i-th grid.

2. The method according to claim 1, characterized in that, The method involves copying the point cloud of the target object from the three-dimensional target object to the point cloud of the sample to be enhanced, and then adjusting the pose of the point cloud of the target object, including adjusting the direction and angle.

3. The method according to claim 1, characterized in that, Resampling methods include uniform sampling, farthest point sampling, and random sampling.

4. An autonomous driving system, characterized in that, The system includes an object detection and classification model, which is trained using multimodal augmented data obtained by any one of the methods in claims 1 to 3.

5. A computer-readable storage medium, characterized in that: The computer program is stored that can be loaded by a processor and executed according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Three-dimensional target detection system and method based on multi-task fusion

    CN114118254A

  • Point cloud data enhancement method and device, electronic device and vehicle

    CN115409083A