Environment-adaptive multi-modal data fusion method and device and vehicle
Through the environmentally adaptive multimodal data fusion method, dynamically adjust and fuse data from multi-view camera arrays and lidars, the robustness and computing efficiency problems of multi-modal data fusion in extreme scenarios are solved, and 3D object detection and vehicle trajectory prediction with high accuracy and high reliability are achieved.
Patent Information
- Application Number
- CN202510714685.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-30
AI Technical Summary
There are significant technical bottlenecks in the robustness and computing efficiency of existing multimodal data fusion algorithms in extreme scenarios, resulting in degradation of detection performance, high demand for computing resources, loss of sparse information in sparse areas and redundant computing in dense areas.
A multimodal data fusion method with environmental adaptability is proposed. By acquiring data from a multi-view camera array and lidar, preprocessing images and point clouds, dynamically adjusting resolution and feature extraction, fusing image features and point cloud features, using a Generative Adversarial Network (GAN) to generate pseudo-images, and adjusting modal weights to achieve cross-modal consistency and efficient computing.
It effectively improves the accuracy and reliability of 3D object detection and vehicle trajectory prediction in extreme scenarios, reduces computing resource requirements, reduces information loss and computing redundancy, and improves the robustness and efficiency of multimodal data fusion.
Smart Images

Figure CN120219904A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of autonomous driving, and particularly to an environment-adaptive multimodal data fusion method, apparatus, and vehicle. Background Art
[0002] Currently, multimodal sensor data fusion technology has been widely applied in fields such as autonomous driving and robot environmental perception. However, there are still significant technical bottlenecks in the robustness and computational efficiency of existing fusion algorithms in extreme scenarios. For example, the detection performance for sparse problems and high-reflection problems will drop sharply, the demand for computing resources is high, and there are also problems such as loss of sparse region information and computational redundancy in dense regions. Summary of the Invention
[0003] In view of this, the present disclosure provides an environment-adaptive multimodal data fusion method, apparatus, and vehicle.
[0004] According to a first aspect of the present disclosure, there is provided an environment-adaptive multimodal data fusion method, which is applied to a vehicle equipped with a multi-view camera array and a lidar; the method includes: Obtaining multi-view images collected by the multi-view camera array and point clouds collected by the lidar; Preprocessing the multi-view images, where the preprocessing includes: for each view image in the multi-view images, adjusting the resolution of the view image according to the illumination condition of the view image; Obtaining image features using the preprocessed multi-view images; Preprocessing the point clouds, where the preprocessing includes: downsampling the point clouds according to the vehicle speed, and performing interpolation and completion on sparse regions and downsampling on dense regions of the downsampled point clouds; Obtaining point cloud features using the preprocessed point clouds; Fusing the image features and the point cloud features to obtain multimodal fusion features; Performing 3D object detection and / or vehicle trajectory prediction using the multimodal fusion features.
[0005] In some embodiments of the first aspect of the present disclosure, the resolution of the view image is adjusted in the following manner: determining the illumination level of the view image through a lightweight convolutional neural network classifier; when the illumination level of the view image is lower than a predetermined level, increasing the resolution of the view image to the image resolution corresponding to the predetermined level; when the illumination level of the view image is higher than the predetermined level, decreasing the resolution of the view image to the image resolution corresponding to the predetermined level.
[0006] In some embodiments of the first aspect of the present disclosure, obtaining image features from the preprocessed multi-view images includes: performing the following processing on each view image in the preprocessed multi-view images to extract features of each view image: dynamically adjusting the RGB channel weights of the basic convolution kernels according to the illumination level of the view image through an image feature extraction module including conditional convolution to generate dynamic convolution kernels adapted to the illumination condition of the view image, and using the dynamic convolution kernels to extract features of the view image; aligning and fusing the features of each view image to obtain multi-scale features.
[0007] In some embodiments of the first aspect of the present disclosure, obtaining image features from the adjusted multi-view images further includes: using a lightweight semantic segmentation network to generate a semantic segmentation mask based on the view image, and applying the semantic segmentation mask as an attention weight in the process of extracting features of the view image using the dynamic convolution kernels.
[0008] In some embodiments of the first aspect of the present disclosure, obtaining point cloud features from the preprocessed point cloud includes: adjusting the voxel grid size according to the local point cloud density to adaptively voxelize the preprocessed point cloud to obtain voxel data; using separable 3D convolution to process the voxel data to obtain the point cloud features.
[0009] In some embodiments of the first aspect of the present disclosure, the method further includes: evaluating whether the cross-modal consistency between the image features and the point cloud features meets the requirements; when the cross-modal consistency between the image features and the point cloud features does not meet the requirements, using a generative adversarial network (GAN) to generate pseudo multi-view images based on the point cloud, obtaining the features of the pseudo multi-view images, and merging the features of the pseudo multi-view images into the image features.
[0010] In some embodiments of the first aspect of the present disclosure, evaluating whether the cross-modal consistency between the image features and the point cloud features meets the requirements includes one or more of the following: Calculating the similarity between the point cloud features and the image features at the same spatial position through BEV space contrast learning, and when the similarity is less than the first predetermined similarity threshold, determining that the cross-modal consistency between the point cloud features and the image features does not meet the requirements; Calculating the cross-modal feature similarity between the forward view image in the multi-view images and the point cloud, and when the cross-modal feature similarity is less than the second predetermined similarity threshold, determining that the cross-modal consistency between the point cloud and the forward view image does not meet the requirements.
[0011] In some embodiments of the first aspect of the present disclosure, the method further includes: dynamically adjusting the modal weights of the image features according to the exposure anomaly conditions of each perspective image in the multi-perspective image; and / or, dynamically adjusting the modal weights of the point cloud features according to the local point cloud density of the point cloud; the fusing the image features and the point cloud features to obtain multi-modal fusion features includes: fusing the image features and the point cloud features based on the modal weights of the image features and / or the modal weights of the point cloud features to obtain the multi-modal fusion features.
[0012] According to a second aspect of the present disclosure, there is provided an environment-adaptive multi-modal data fusion device, which is applied to a vehicle, and a multi-perspective camera array and a lidar are installed on the vehicle; The environment-adaptive multi-modal data fusion device includes: A data acquisition unit, configured to acquire multi-perspective images collected by the multi-perspective camera array and point clouds collected by the lidar; An image preprocessing unit, configured to preprocess the multi-perspective images, and the preprocessing includes: for each perspective image in the multi-perspective images, adjusting the resolution of the perspective image according to the illumination condition of the perspective image; An image feature extraction unit, configured to obtain image features by using the preprocessed multi-perspective images; A point cloud preprocessing unit, configured to preprocess the point clouds, and the preprocessing includes: downsampling the point clouds according to the vehicle speed, and performing sparse region interpolation and dense region downsampling on the downsampled point clouds; A point cloud feature extraction unit, configured to obtain point cloud features by using the preprocessed point clouds; A fusion unit, configured to fuse the image features and the point cloud features to obtain multi-modal fusion features; A task execution unit, configured to perform 3D object detection and / or vehicle trajectory prediction by using the multi-modal fusion features.
[0013] According to a third aspect of the present disclosure, there is provided a vehicle, which is equipped with a multi-perspective camera sequence and a lidar, and the vehicle includes the aforementioned environment-adaptive multi-modal data fusion device.
[0014] The embodiments of the present disclosure can achieve environment-adaptive multi-modal data fusion, and at the same time avoid problems such as loss of information in sparse regions and computational redundancy in dense regions, so as to reduce the impact of fusion deviation while reducing the demand for computing resources, improve the reliability and robustness of modal data fusion, and effectively improve the accuracy and reliability of 3D object detection and / or vehicle trajectory prediction in extreme scenarios. Description of the Drawings
[0015] To more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0016] Figure 1 Schematic diagram of the architecture of the system applicable to the embodiments of the present disclosure; Figure 2 Schematic flowchart of an environment - adaptive multi - modal data fusion method provided by the embodiments of the present disclosure; Figure 3 Schematic flowchart of adjusting the resolution of a perspective image according to the illumination condition of the perspective image involved in the embodiments of the present disclosure; Figure 4 Schematic flowchart of obtaining image features using pre - processed multi - perspective images involved in the embodiments of the present disclosure; Figure 5 Another schematic flowchart of the environment - adaptive multi - modal data fusion method provided by the embodiments of the present disclosure; Figure 6 Schematic diagram of the structure of an environment - adaptive multi - modal data fusion device provided by the embodiments of the present disclosure; Figure 7 Schematic block diagram of the schematic structure of an electronic device provided by the embodiments of the present disclosure. Detailed implementation manners
[0017] The following will clearly and completely describe the technical solutions in the embodiments of the present disclosure in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only some embodiments of the present disclosure, rather than all embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present disclosure.
[0018] The terms used in the embodiments of the present disclosure are only for the purpose of describing specific embodiments, and are not intended to limit the present disclosure. The singular forms of "a", "the" and "said" used in the embodiments of the present disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.
[0019] Depending on the context, as used herein, words such as "if", "when", etc. can be interpreted as "when...", "while...", "in response to determining", or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detecting (stated condition or event)" can be interpreted as "when determined", "in response to determining", "when detecting (stated condition or event)", or "in response to detecting (stated condition or event)".
[0020] First, a detailed description of the related technology will be given.
[0021] Currently, there are mainly the following several solutions for multi-modal data fusion in extreme scenarios: 1) Multi-modal fusion based on fixed weights: The radar and camera features are directly added or spliced through fixed ratio weights, relying on the neural network model to automatically learn cross-modal associations. This solution cannot dynamically adjust the modal weights according to the illumination and point cloud density, resulting in a sharp drop in detection performance in extreme scenarios (such as camera failure under low illumination). At the same time, there is an accumulation of projection deviation. The rigid projection of the radar's BEV space and the camera's front view causes positioning offsets of small targets such as pedestrians and traffic cones (average error > 0.5m).
[0022] 2) Multi-modal data fusion based on static voxelized point cloud processing: The point cloud collected by the radar is evenly divided into voxel grids of a fixed size (such as 0.1m³), and features are extracted through a 3D convolutional neural network. This solution has the problem of information loss in sparse regions. In the far-distance region or low-density region, key point clouds are filtered due to the overly large voxel grids, increasing the miss detection rate by 12% or more. At the same time, there is also the problem of redundant calculation in dense regions. In the near-distance high-density region, the same features are repeatedly calculated, reducing the utilization rate of computing resources by about 40% or more.
[0023] As can be seen from the above, there are still significant technical bottlenecks in the robustness and computational efficiency of multi-modal data fusion in the related technology in extreme scenarios. In view of this, the embodiments of the present disclosure provide the following environment-adaptive multi-modal data fusion methods, devices, and vehicles.
[0024] For ease of understanding, a brief description of the system structure applicable to the embodiments of the present disclosure will be given below.
[0025] Figure 1 The structural schematic diagram of the system applicable to the embodiments of the present disclosure is shown. Refer to Figure 1 , the system applicable to the embodiments of the present disclosure may include: an electronic device and peripheral sensor components connected to the electronic device. The peripheral sensor components include, but are not limited to, a multi-view camera array and a lidar.
[0026] The multi-view camera array can be mounted on a vehicle and can be used to collect multi-view images covering the surrounding environment of the vehicle.
[0027] The multi-view camera array can be implemented as, but not limited to, a six-view camera group, which includes a front-view camera, a rear-view camera, a left-view camera, a right-view camera, an upper-view camera, and a lower-view camera. In a specific application, cameras can be evenly deployed around the vehicle so that these cameras can cover the surroundings of the vehicle without dead angles, and each camera captures an image of the surroundings of the vehicle, which contains a part of the surrounding scene of the vehicle.
[0028] The lidar is mounted on the vehicle and can be used to collect point clouds in real time. In a specific application, the installation position of the lidar on the vehicle can be flexibly adjusted according to the actual application so that the point clouds collected by the lidar cover the front, rear, or surrounding environment of the vehicle.
[0029] Each point in the point cloud collected by the lidar can include three-dimensional coordinates, reflection intensity, and timestamp. Among them, the three-dimensional coordinates can be determined by calculating the distance through the laser flight time and combining the radar rotation angles (azimuth angle, pitch angle). The three-dimensional coordinates of each point in the point cloud can be the three-dimensional coordinates in the lidar coordinate system, and can be converted to the vehicle body coordinate system, the world coordinate system, etc. The reflection intensity (also called reflectivity) depends on the surface material of the object and can be used to distinguish material types (such as road marking recognition). The timestamp can be used to record the time difference between laser emission and reception, and can be used for the analysis of dynamic scenes such as the tracking of moving object trajectories.
[0030] Furthermore, the peripheral sensor assembly can also include: an Inertial Measurement Unit (IMU). The IMU is mounted on the vehicle and connected to the electronic device, and can be used to measure IMU data and provide the IMU data to the electronic device. The IMU data can include, but is not limited to, the three-axis acceleration and three-axis angular velocity of the vehicle, etc. The IMU data can be used to determine the instantaneous motion state of the vehicle, assist in positioning, attitude estimation, etc. Exemplarily, the IMU data can be used to determine attitude information such as the roll angle, pitch angle, and yaw angle of the vehicle.
[0031] Furthermore, the above system can also include an in-vehicle communication module (for example, 4G / 5G, Wi-Fi, vehicle-to-everything communication module, etc.) or an in-vehicle environmental sensor. Among them, the in-vehicle communication module can be used to connect to the cloud to obtain meteorological data. For example, it can receive meteorological data on the driving route pushed in real time by Gaode Map, Google Maps, etc. For another example, it can receive meteorological data provided by devices around the vehicle (such as other vehicles, roadside units, etc.). The in-vehicle environmental sensor can be used to detect the weather conditions of the surroundings of the vehicle in real time and generate corresponding meteorological data. The in-vehicle environmental sensor can include, but is not limited to, an ambient light sensor, a humidity sensor, a temperature sensor, etc.
[0032] The system applicable to the embodiments of the present disclosure may be, but is not limited to, any system that requires multi-modal data fusion. For example, the system may be, but is not limited to, an autonomous driving system, an intelligent assisted driving system, a traffic congestion assistance system, etc. Refer to Figure 1 , the system provided by the embodiments of the present disclosure can be loaded on a vehicle and used as, but is not limited to, an intelligent assisted driving system, a traffic congestion assistance system, an autonomous driving system, etc. of the vehicle.
[0033] It should be noted that the "vehicle" described in the embodiments of the present disclosure can be implemented as, but is not limited to, multiple wheeled mobile robots, wheeled mobile robots, mobile robots, general vehicles, aircraft, ships, intelligent rail rapid transit systems (ART, Autonomous rail Rapid Transit), industrial automation equipment, etc. Among them, general vehicles may be, but are not limited to, passenger vehicles, commercial vehicles (such as trucks, buses, freight vehicles, etc.), special-purpose vehicles (such as ambulances, fire trucks, engineering vehicles, rescue vehicles, etc.), agricultural and industrial vehicles (such as harvesters, forklifts, etc.), transportation and logistics vehicles (such as container trucks, refrigerated trucks, etc.), new energy vehicles (such as electric vehicles, hybrid vehicles), special carriers (such as garbage trucks, sprinkler trucks, etc.).
[0034] The embodiments of the present disclosure can be applied to various scenarios such as urban traffic, highways, ports, mines, farms, closed parks, industrial production, etc., and can be applicable to many aspects such as ride-hailing, public transportation, logistics distribution, unmanned transportation, last-mile delivery, automated agricultural operations, automated environmental sanitation, etc. Of course, the embodiments of the present disclosure can also be applied to any other intelligent control scenarios involving, for example, multi-modal data fusion. The present disclosure does not limit the application scenarios and applicable fields of the embodiments of the present disclosure.
[0035] The following will elaborate on the specific implementation manners of the embodiments of the present disclosure.
[0036] Figure 2 The flowchart of the environment-adaptive multi-modal data fusion method provided by the embodiments of the present disclosure is shown. This environment-adaptive multi-modal data fusion method can be executed by the following electronic device, which can be implemented as, but is not limited to, a domain controller or other similar devices. Refer to Figure 2 , the environment-adaptive multi-modal data fusion method of the embodiments of the present disclosure includes the following steps: Step 201, obtaining multi-view images collected by a multi-view camera array and point clouds collected by a lidar; Step 202, preprocessing the multi-view images. The preprocessing includes: for each view image in the multi-view images, adjusting the resolution of the view image according to the illumination condition of the view image; Step 203: Obtain image features using the preprocessed multi-view images; Step 204: Preprocess the point cloud. The preprocessing includes: downsampling the point cloud according to the vehicle speed, and performing interpolation and completion on the sparse regions and downsampling on the dense regions of the downsampled point cloud; Step 205: Obtain point cloud features using the preprocessed point cloud; Step 206: Fuse the image features and the point cloud features to obtain multi-modal fusion features; Step 207: Perform 3D object detection and / or vehicle trajectory prediction using the multi-modal fusion features.
[0037] In the environmental adaptive multi-modal data fusion method of the embodiments of the present disclosure, after adjusting the resolution of each view image according to the lighting condition, the image features are extracted. After downsampling the point cloud according to the vehicle speed and performing interpolation and completion on the sparse regions and downsampling on the dense regions of the downsampled point cloud, the point cloud features are extracted. Thus, the present disclosure can achieve environmental adaptive multi-modal data fusion, and at the same time avoid problems such as loss of information in sparse regions and redundant calculations in dense regions, thereby reducing the impact of fusion deviation while reducing the demand for computing resources, improving the reliability and robustness of modal data fusion, and effectively improving the accuracy and reliability of 3D object detection and / or vehicle trajectory prediction in extreme scenarios.
[0038] In step 201, the multi-view images and the point cloud are synchronized. Specifically, the synchronization of the multi-view images and the point cloud can be achieved through mechanisms such as synchronous triggering and time alignment of the multi-view camera array and the lidar.
[0039] In step 202, for each view image in the multi-view images, the resolution of the view image is adjusted according to the lighting condition of the view image, which can improve the resolution of the view image in a low-light environment to retain details, and reduce the resolution of the view image in a strong-light environment to reduce redundant calculations.
[0040] Figure 3 The specific implementation process diagram of adjusting the resolution of the view image according to the lighting condition of the view image in step 202 is shown. Refer to Figure 3 This process may include the following steps 301 to step 302: Step 301: Determine the lighting level of each view image in the multi-view images through a lightweight Convolutional Neural Network (CNN) classifier; Multiple lighting levels and their image resolutions can be pre-configured. A high lighting level indicates high lighting intensity, and a low lighting level indicates low lighting intensity.
[0041] Among them, the light level can be, but is not limited to, category information such as low light, normal light, strong light, etc., or can also be a continuous value such as a brightness value. For example, three light levels of high, medium, and low can be configured, and the light intensity intervals of each light level are set. The light intensity of the high light level is the highest, the light intensity of the medium light level is moderate, and the light intensity of the low light level is the lowest. The medium level corresponds to the normal light environment, and the medium level can be used as the preset level in step 302.
[0042] A frame of image can be captured from the camera corresponding to each perspective image, and the light level of the corresponding perspective image can be obtained by processing this frame of image using a lightweight CNN classifier. Specifically, for each perspective image, a lightweight CNN classifier can be used to extract global light features to obtain the light level. Among them, the lightweight CNN classifier can be selected but is not limited to MobileNetV3, etc. By estimating the light level through the lightweight CNN classifier, it is possible to more accurately capture features under complex light conditions such as shadows and mixed light sources, and then accurately evaluate the light levels of each perspective image.
[0043] In specific applications, other methods can also be used to determine the light levels of each perspective image according to needs.
[0044] For example, the light intensity of the area covered by each perspective image can be directly obtained through an ambient light sensor (ALS), and the light levels of each perspective image can be determined by the light intensity and the pre-configured light intensity intervals of each light level.
[0045] Another example is that a frame of RGB format image can be captured from the camera corresponding to each perspective image, the image is converted to a grayscale image or the YUV color space to extract the luminance channel, and the luminance mean or luminance median reflecting the overall light intensity is calculated. This luminance mean or luminance median is the light intensity, and the light levels of each perspective image are determined by the light intensity and the pre-configured light intensity intervals of each light level. Step 302: Adjust the resolution of the perspective image according to the light level of the perspective image.
[0046] Specifically, when the light level of the perspective image is lower than the predetermined level, it is determined that the acquisition environment of the perspective image is a low light environment, and the resolution of the perspective image is increased to the image resolution corresponding to the predetermined level; when the light level of the perspective image is higher than the predetermined light level, it is determined that the acquisition environment of the perspective image is a strong light environment, and the resolution of the perspective image is reduced to the image resolution corresponding to the predetermined light level. Thus, the resolution of multi-perspective images can be increased in a low light environment to retain more details, and the image resolution of multi-perspective images can be reduced in a high light environment to reduce its data volume and reduce redundant calculations.
[0047] Dynamically adjusting the resolution of each perspective image in a multi-perspective image according to the light intensity can better balance image quality, processing efficiency, and energy consumption, thereby dynamically adapting to the changing external environment.
[0048] Figure 4 FIG. shows a schematic flow chart of obtaining image features by using the preprocessed multi-perspective image in step 203. Refer to Figure 4 , this process may include steps 401 to 402: Step 401, perform the following processing on each perspective image in the multi-perspective image to extract its features: use an image feature extraction module including conditional convolution to dynamically adjust the RGB channel weights of the basic convolution kernel according to the light level of the perspective image to generate a dynamic convolution kernel adapted to the light condition of the perspective image, and use the dynamic convolution kernel to extract the features of the perspective image.
[0049] Specifically, the process of using an image feature extraction module including conditional convolution to extract the features of a certain perspective image may include: generating an RGB channel weight scaling factor according to the light level of the current perspective image, using the RGB channel weight scaling factor to adjust the RGB channel weights of the basic convolution kernel to generate a dynamic convolution kernel, and using the dynamic convolution kernel to perform feature extraction on the current perspective image to obtain the features of the current perspective image.
[0050] As described above, in low-light regions, the dynamic convolution kernel can enhance the high-frequency detail extraction ability to restore dark information, and in overexposed regions, the dynamic convolution kernel can suppress noise and balance brightness. Thus, the feature extraction process of each perspective image can be adaptively optimized according to the light conditions of each perspective image, enhancing color invariance in low light and suppressing noise in overexposed regions.
[0051] The features of a single perspective image may include one or more of the following information: color, texture, shape, spatial relationship, semantics, etc. The spatial relationship includes the target position, which is used to describe the absolute or relative position of the segmented target in the image. The semantic information includes the target category information, which can be used to describe the object category and its confidence at the pixel level. In addition, the semantic information may also include scene semantic information, which can be used to describe the overall scene of the image. The features of each perspective image can be pixel-level features.
[0052] Step 402, align and fuse the features of each perspective image to obtain multi-scale features.
[0053] Here, the multi-scale features are the image features of the multi-perspective image. The multi-scale features may include one or more of the following information: color, texture, shape, semantics, spatial constraint. Among them, the semantic information includes the target type information. Similarly, the multi-scale features can be pixel-level features.
[0054] Specifically, multi-scale features can be obtained by geometric alignment and attention weighting. The specific implementation method of obtaining multi-scale features is not limited in the embodiments of the present disclosure.
[0055] Furthermore, in step 401, a lightweight semantic segmentation network can be used to generate a semantic segmentation mask based on the view image, and the semantic segmentation mask can be used as an attention weight in the process of extracting features of the view image using a dynamic convolution kernel. Thus, the image feature extraction model can be dynamically guided by the lightweight semantic segmentation network to focus on key areas such as roads and obstacles, and to suppress irrelevant background interference.
[0056] Exemplarily, the lightweight semantic segmentation network may be, but is not limited to, DeepLabv3+ Mobile. DeepLabv3+Mobile can efficiently generate semantic segmentation masks by combining a lightweight architecture with multi-scale feature enhancement technology. The semantic segmentation mask generated based on the view image may be a probability map, a binary mask, or other applicable forms.
[0057] Through the above method, in multi-view scenes, the convolution kernel can be dynamically adjusted to adapt to lighting changes while accurately focusing on key areas, thereby achieving efficient feature extraction of multi-view images.
[0058] In step 204, the point cloud is adaptively downsampled according to the vehicle speed. The resolution of the point cloud can be dynamically adjusted according to the change of the vehicle speed, so as to strike a balance between high precision and efficient processing, so that the point cloud is sparse in high-speed scenes and maintains a high resolution in low-speed scenes. The vehicle speed can be obtained in various applicable ways. Here, the vehicle speed can be the instantaneous speed of the vehicle or the average speed of the vehicle within the time period corresponding to the point cloud.
[0059] In some examples, vehicle speed can be provided in real time by an IMU mounted on the vehicle.
[0060] In some examples, the vehicle speed can be determined based on an extended Kalman filter (EKF) using IMU data provided by an IMU installed on the vehicle. This can compensate for the vehicle's own motion and eliminate the impact of offsets caused by vehicle bumps, vehicle steering, and other conditions on the accuracy of the vehicle speed, thereby improving the stability and reliability of vehicle speed detection.
[0061] In other examples, the vehicle speed can be read from the vehicle CAN bus or estimated by the change in the position of two consecutive frames of point clouds. The specific detection method of the vehicle speed is not limited in the embodiments of the present disclosure.
[0062] In step 204, the downsampling ratio of the point cloud can be dynamically adjusted based on the piecewise linear interpolation method of the vehicle speed to achieve adaptive downsampling of the point cloud. Specifically, multiple speed intervals can be pre-configured, and within the speed interval to which the vehicle speed belongs, a predetermined linear interpolation formula is used to calculate the retention ratio of downsampling, and the point cloud is downsampled according to the retention ratio of downsampling. Through the piecewise linear interpolation method, the resolution of the point cloud can be dynamically adjusted according to the vehicle speed, achieving a balance between data processing efficiency and perception accuracy.
[0063] The starting speed and ending speed of each speed interval are preset values. Exemplarily, the speed intervals can be configured as the following four types: low-speed interval (0 - 30 km / h), medium-speed interval (30 - 60 km / h), high-speed interval (60 - 90 km / h), and ultra-high-speed interval (> 90 km / h). Among them, the low-speed interval corresponds to high resolution, and 100% of the point cloud is retained; the medium-speed interval corresponds to medium resolution, and 70% - 100% of the point cloud is retained; the high-speed interval corresponds to low resolution, and 50% - 70% of the point cloud is retained; the ultra-high-speed interval corresponds to extremely low resolution, and 50% of the point cloud is retained.
[0064] When downsampling (i.e., reducing the resolution), methods such as voxel filtering or random sampling can be used to reduce the number of point clouds. For example, downsampling the point cloud according to the retention ratio of downsampling can be achieved by one of the following methods: 1) Assuming the retention ratio is r, randomly discard (1 - r) proportion of the points in the point cloud; 2) Divide the point cloud into voxel grids, and each voxel retains a representative point. The voxel size is dynamically adjusted according to the retention ratio r of downsampling. The lower the retention ratio r, the larger the voxel.
[0065] In step 204, interpolating and complementing the sparse regions of the point cloud and further downsampling the dense regions can prevent the loss of information in the sparse regions and avoid computational redundancy in the dense regions. In some examples, the sparse regions and dense regions can be dynamically divided according to the local density of the point cloud. For example, the local density of the point cloud can be estimated based on K-Nearest Neighbors (KNN), and the sparse regions and dense regions can be dynamically distinguished according to the local density of the point cloud.
[0066] Specifically, the process of interpolating and complementing the sparse regions and downsampling the dense regions of the downsampled point cloud in step 204 can include: calculating the average distance of the K nearest neighbors of each point through the KNN algorithm, dividing the point cloud into sub-blocks, and each sub-block performs the following processing: independently calculating a threshold based on the average distance of the K nearest neighbors of all points and separating it into dense regions and sparse regions through the threshold. Further downsample the point cloud in the dense regions, interpolate and complement the point cloud in the sparse regions, and finally merge the point cloud in the dense regions and the point cloud in the sparse regions to obtain the preprocessed point cloud.
[0067] Among them, K (for example, K = 10 - 50) is selected according to the point cloud density characteristics to balance noise sensitivity and local detail retention. The median or a specific percentile (such as 25%) can be selected as the threshold by analyzing the K-nearest neighbor average distance distribution of all points.
[0068] Estimate the local density of the point cloud based on KNN and dynamically distinguish sparse regions and dense regions. It can achieve efficient and flexible compression of the point cloud in complex scenarios by dynamically responding to the local characteristics of the point cloud. It has strong adaptability, retains the details of key regions, and reduces the information loss caused by uniform downsampling at the same time.
[0069] In step 204, the adaptive downsampling based on vehicle speed and the adaptive processing based on local point cloud density (i.e., interpolation and completion in sparse regions, downsampling in dense regions) work together. The vehicle speed determines the "overall resolution tone" of the point cloud, and the local point cloud density determines the "local detail enhancement or compression" of the point cloud. It can achieve "global lightweighting and local refinement" in a dynamic environment, and at the same time respond to dynamic environments (i.e., speed changes) and static characteristics (i.e., local density distribution), and can optimize the processing efficiency and accuracy of the point cloud in multiple dimensions.
[0070] In step 205, obtaining point cloud features from the preprocessed point cloud can include: adjusting the voxel grid size according to the local point cloud density to adaptively voxelize the point cloud to obtain voxel data, and using separable 3D convolution to process the voxel data to obtain point cloud features.
[0071] In some examples, the point cloud can be voxelized by dynamic grid division (Density-based Voxelization) (also known as density-based voxelization) to obtain voxel data of the point cloud. Specifically, the point cloud is discretized into a non-uniform voxel grid and the voxel resolution of the non-uniform voxel grid is dynamically adjusted according to the point cloud density. The dynamic grid division voxelization adjusts the voxel resolution in real time according to the point cloud density, geometric features or physical field changes. While retaining details, it can optimize the calculation efficiency, reduce unnecessary resource consumption, avoid information loss in sparse regions, and at the same time suppress calculation redundancy in dense regions, generate a non-uniform voxel grid, and significantly reduce the calculation amount of subsequent convolution (reduce 30 - 50% of invalid calculations).
[0072] Specifically, the process of voxelizing the point cloud through dynamic mesh division may include: first generating an initial voxel grid based on geometric or data features, usually using uniform division or multi-resolution initialization; calculating the local point density of the point cloud in real time, and when the local point density of the point cloud exceeds a preset threshold, updating the voxels in the corresponding area (i.e., high-density areas such as object edges and complex structures) to smaller voxels (e.g., 0.1 m³) to retain details such as vehicle contours. When the local point density is lower than the above preset threshold, adjacent voxels are merged so that the corresponding area (e.g., low-density areas such as empty backgrounds) uses larger voxels (e.g., increased from the initial 0.4 m³ to the current 0.5 m³) to compress the background point cloud and reduce the computational amount. After the voxel grid is adjusted, bilinear interpolation or covariance matrix calculation can be further used to calculate voxel features (such as average coordinates and surface normals) to ensure data continuity. In addition, efficient management of multi-resolution voxels can be achieved using methods such as Adaptive Octree, Sparse Voxel Octree (SVO), or hash tables.
[0073] In step 205, Separable 3D Convolution can be decomposed into 2D spatial convolution and 1D channel convolution. Specifically, the process of using Separable 3D Convolution to process the voxel data to obtain point cloud features may include: independently performing 2D convolution on each depth slice of the three-dimensional voxel grid to extract spatial features within each channel, and performing 1D convolution along the channel dimension to fuse the spatial features of different channels to obtain point cloud features. Among them, the spatial features may include spatial structure features such as object shape, object position, and object edge. The spatial features are the basis of high-level semantics and can be gradually extracted through multiple network layers.
[0074] In other embodiments, the operation of Separable 3D Convolution can also be performed according to the method of "first 1D depth convolution and then 2D spatial convolution, mixed splitting" according to requirements.
[0075] As can be seen from the above, Separable 3D Convolution decomposes the three-dimensional convolution into two independent low-dimensional operations, which can reduce redundant calculations, avoid repeated calculations in three-dimensional convolution, and at the same time reduce the computational waste of empty voxels, significantly improving the computational efficiency without loss of accuracy.
[0076] By combining density adaptive voxelization and Separable 3D Convolution to extract point cloud features, the overall efficiency and accuracy can be effectively improved.
[0077] In step 206, various applicable methods can be adopted to achieve the fusion of point cloud features and image features. The embodiments of the present disclosure do not limit the specific fusion methods. For example, a cross-modal shared attention mechanism can be adopted to map the point cloud features and image features to a shared space, calculate the attention weights, and perform weighted fusion to obtain multi-modal fusion features.
[0078] Figure 5 Another schematic flowchart of the environment-adaptive multi-modal data fusion method provided by the embodiments of the present disclosure is shown. Refer to Figure 5 , further, before step 206, it may further include: Step 208, evaluating whether the cross-modal consistency of the image features and the point cloud features meets the requirements; Step 209, when the cross-modal consistency of the image features and the point cloud features does not meet the requirements, using a Generative Adversarial Network (GAN) to generate pseudo multi-view images based on the point cloud, obtaining the features of the pseudo multi-view images, and merging the features of the pseudo multi-view images into the image features.
[0079] In some embodiments, in step 208, the cross-modal consistency of the image features and the point cloud features can be evaluated for meeting the requirements by one or both of the following methods: 1) calculating the similarity between the point cloud features and the image features at the same spatial position through BEV space contrast learning, and determining that the cross-modal consistency of the point cloud features and the image features does not meet the requirements when the similarity is less than the first predetermined similarity threshold; 2) calculating the cross-modal feature similarity between the forward-view image in the multi-view image and the point cloud, and determining that the cross-modal consistency of the point cloud and the forward-view image does not meet the requirements when the cross-modal feature similarity is less than the second predetermined similarity threshold.
[0080] Specifically, the average cosine similarity between the point cloud features and the image features at the same spatial position can be calculated. When the average cosine similarity is greater than or equal to the first predetermined similarity threshold, the evaluation passes, and the image features obtained in step 203 and the point cloud features obtained in step 205 can be directly used for fusion. When the average cosine similarity is less than the first predetermined similarity threshold, the image features of the pseudo multi-view images generated by the generative adversarial network can be mixed with the image features obtained in step 203 to obtain new image features, and the new image features are used for fusion with the point cloud features obtained in step 205. Thus, the consistency of the feature distributions of the lidar and the camera can be dynamically evaluated in the BEV space through BEV space contrast learning, the projection deviation problem can be solved, and the fusion deviation caused by the external environment in extreme scenarios (such as thunderstorms, nights, etc.) can be further reduced, improving the reliability and robustness of multi-modal data fusion.
[0081] Furthermore, during the model training phase, by adding the contrastive loss (InfoNCE Loss) function in the BEV space contrastive learning to the loss function, the parameters of the image feature-related model and the point cloud feature-related model can be further optimized, forcing the cross-modal features to align in the BEV space, thereby solving the projection deviation problem and further reducing the fusion deviation caused by the external environment in extreme scenarios (such as thunderstorms, nights, etc.), improving the reliability and robustness of multi-modal data fusion.
[0082] Specifically, the average cosine similarity between the forward-view image features (i.e., the features obtained in step 301) and the point cloud features at the same spatial position can be calculated, and this average cosine similarity is used as the cross-modal feature similarity between the forward-view image and the point cloud in the multi-view image. When this average cosine similarity is greater than or equal to the second predetermined similarity threshold, the evaluation passes, and the image features obtained in step 203 and the point cloud features obtained in step 205 can be directly used for fusion. When this average cosine similarity is less than the second predetermined similarity threshold, the image features of the pseudo multi-view image generated by the generative adversarial network can be mixed with the image features obtained in step 203 to obtain new image features, and these new image features are used for fusion with the point cloud features obtained in step 205. Thus, the consistency of the feature distributions of the lidar and the camera can be evaluated by the forward graph feature matching method, solving the projection deviation problem and further reducing the fusion deviation caused by the external environment in extreme scenarios (such as thunderstorms, nights, etc.), improving the reliability and robustness of multi-modal data fusion.
[0083] In a specific application, in the BEV space coordinate system, after projecting the point cloud and the image features onto a unified BEV grid through pre-calibrated parameters or a learnable transformation matrix (such as an MLP), the grid-level cosine similarity can be calculated.
[0084] Furthermore, during the model training phase, by adding the mutual information (MI) related to forward graph feature matching to the loss function, the parameters of the image feature-related model and the point cloud feature-related model can be further optimized, enhancing the representational consistency of the same object in different modalities, solving the projection deviation problem, and further reducing the fusion deviation caused by the external environment (such as occlusion, lighting changes, etc.) in extreme scenarios (such as thunderstorms, nights, etc.), improving the reliability and robustness of multi-modal data fusion.
[0085] In step 209, the GAN can be, but is not limited to, a pixel-level generative adversarial network. Specifically, the GAN can be, but is not limited to, CycleGAN. The GAN includes a Generator and a Discriminator. The Generator is responsible for generating pseudo-data of another modality based on data of one modality, and the Discriminator is used to distinguish between real and fake to drive the Generator to generate more realistic pseudo-data.
[0086] In step 209, the loss function of the GAN includes: adversarial loss, cycle consistency loss, and cross-modal alignment loss. Among them, the adversarial loss (GAN Loss) is used to train the Generator and the Discriminator to compete, so that the features generated by the Generator are as realistic as possible. The cycle consistency loss can ensure that the generated features are consistent with the original features after being transformed back to the original modality, and the cross-modal alignment loss can be used to measure the alignment degree between the generated features and the target modality features. Thus, the GAN can be kept in cross-modal alignment, and the generated features should be consistent with the target modality features semantically and structurally, so as to improve the robustness of feature alignment between images and point clouds in extreme scenarios such as camera failure.
[0087] In specific applications, the GAN can be jointly optimized with cross-modal contrast loss (such as InfoNCE Loss) through adversarial training to ensure that the generated features are consistent with the point cloud features in the BEV space distribution.
[0088] In some extreme scenarios, a certain modality may fail. For example: at night or in foggy weather, the camera may not work properly due to insufficient light or low visibility. Another example is that in rainy or snowy weather, the camera image may be disturbed by raindrops or snowflakes. Another example is at night or in foggy weather. In this case, the camera may fail completely or mostly. If relying on the visual features of the camera as in sunny days, it will lead to performance degradation or even failure. Similarly, in extreme scenarios, the point cloud collected by lidar can be used to generate visual features. In view of this, step 209 may further include: obtaining meteorological data. When the meteorological data indicates that the current scene is a preset extreme scene (such as night, fog, rain, snow, etc.), the Generator in the pixel-level adversarial network can also be used to generate pseudo multi-view images based on the point cloud, obtain the features of the pseudo multi-view images, and merge the features of the pseudo multi-view images into the image features. Thus, the robustness of feature alignment between multi-view images and point clouds in extreme scenarios (such as single modality failure) can be further improved.
[0089] See Figure 5, Further, before step 206, it may further include: step 210, dynamically adjusting the modal weights of image features according to the exposure anomaly conditions of each perspective image in the multi-perspective image; and / or, step 211, dynamically adjusting the modal weights of point cloud features according to the local point cloud density of the point cloud. In step 206, the image features and the point cloud features may be fused based on the modal weights of the image features and / or the modal weights of the point cloud features to obtain multi-modal fusion features.
[0090] In step 210, for each perspective image in the multi-perspective image, an overexposed area and / or an underexposed area may be detected and a two-dimensional mask for marking the overexposed area and / or the underexposed area may be generated, and the modal weights of the image features may be dynamically adjusted according to the two-dimensional mask. Thus, the influence of low-quality image areas on the fusion can be suppressed.
[0091] In some examples, the overexposed / underexposed area may be detected and the mask may be generated by means of histogram analysis and threshold judgment. Specifically, the following processing may be performed for each perspective image to generate its two-dimensional mask: the luminance information of the perspective image is extracted to generate the luminance histogram of the perspective image, the pixel values of the luminance histogram are normalized to [0, 255], and the overexposure mask and the underexposure mask are generated based on the luminance histogram by using a fixed threshold method (i.e., both the overexposure threshold and the underexposure threshold take preset fixed values) or an adaptive threshold method (i.e., the overexposure threshold and the underexposure threshold are dynamically adjusted according to the luminance histogram, the position where the luminance histogram suddenly drops on the right side is taken as the overexposure threshold, and the position where the luminance histogram suddenly rises on the left side is taken as the underexposure threshold), and the joint mask is obtained by taking the dot product of the overexposure mask and the underexposure mask, and this joint mask is the two-dimensional mask of the perspective image.
[0092] In specific applications, an adaptive threshold segmentation such as dynamically setting a threshold based on the mean and variance of image blocks or a lightweight U-Net model may be used, but is not limited to, to dynamically adjust the modal weights of the image features according to the two-dimensional mask.
[0093] In some examples, dynamically adjusting the modal weights of the image features according to the two-dimensional mask may be implemented in the following manner: the two-dimensional masks of each perspective image are mapped to a unified coordinate system (e.g., BEV or panoramic plane) to obtain the two-dimensional mask of the multi-perspective image, the two-dimensional mask of the multi-perspective image is converted into a confidence score C ∈ [0, 1], and the modal weights of the image features are adjusted by the confidence score C.
[0094] Dynamically adjusting the modal weights of the image features through the overexposure / underexposure mask can effectively suppress the influence of low-quality image areas on the fusion in strong light, backlight, and low-light environments at night, reduce the fusion deviation, and improve the fusion accuracy.
[0095] In step 211, the local point cloud density of the point cloud can be calculated through Kernel Density Estimation (KDE), the local point cloud density of the point cloud is mapped to a reliability score, and the modal weights of the point cloud features are adjusted according to the reliability score.
[0096] Due to the large point spacing in the sparse area, the density value after KDE superposition is low, the reliability score of the sparse area is low, and the weight is reduced during fusion, that is, the sparse area can be automatically downweighted to avoid introducing unreliable point cloud features in multi-modal data fusion, thereby effectively improving the robustness of multi-modal fusion in point cloud sparse or occluded scenarios.
[0097] In specific applications, the fusion features of the image feature modal weights dynamically adjusted based on the two-dimensional mask of the overexposed / underexposed area and the point cloud features dynamically adjusted based on the local point cloud density can be used simultaneously in the fusion of image features and point cloud features. The synergistic effect of the two can significantly improve the robustness of the multi-modal system, especially in scenarios where complex lighting and point cloud sparsity coexist.
[0098] In step 207, an Anchor-free based detection head can be used to perform 3D object detection based on the multi-modal fusion features to obtain 3D object detection results. The 3D object detection results include: the parameters of each 3D bounding box and its object category. The parameters of the 3D bounding box include the center coordinates, size, and heading angle. The heading angle represents the direction angle of the target object, the center point represents the position center of the target object, and the size represents the width, height, and length of the target object. Thus, high-precision 3D object detection results can be obtained to meet the decision-making requirements of application scenarios such as autonomous driving.
[0099] In step 207, models such as Recurrent Neural Network (RNN), Long Short-Term Memory Network (LSTM), Gated Recurrent Unit (GRU), or Transformer can be used to model the sequence of multi-modal fusion features, thereby predicting the motion trajectory of the vehicle within a future predetermined time period (for example, within 5 seconds). The motion trajectory of the vehicle can be represented as a sequence of trajectory points and can also include a sequence of motion parameters such as speed and acceleration. Predicting the vehicle motion trajectory through multi-modal fusion features can make full use of the advantages of both point cloud and image modalities and improve the accuracy and robustness of trajectory prediction.
[0100] Furthermore, after step 207, it can also include: smoothing and interpolating and correcting the vehicle motion trajectory obtained in step 207 through a method combining Kalman filtering and polynomial fitting. Thus, the influence of single-frame detection jitter on the vehicle motion trajectory can be eliminated, and the prediction accuracy of the vehicle motion trajectory can be improved.
[0101] Figure 6The structural schematic diagram of the environment-adaptive multi-modal data fusion device provided by the embodiments of the present disclosure is shown. Refer to Figure 6 , the environment-adaptive multi-modal data fusion device 600 of the embodiments of the present disclosure may include: A data acquisition unit 601, configured to acquire multi-view images collected by a multi-view camera array and point clouds collected by a lidar; An image preprocessing unit 602, configured to preprocess the multi-view images, and the preprocessing includes: for each view image in the multi-view images, adjusting the resolution of the view image according to the illumination condition of the view image; An image feature extraction unit 603, configured to obtain image features by using the preprocessed multi-view images; A point cloud preprocessing unit 604, configured to preprocess the point clouds, and the preprocessing includes: downsampling the point clouds according to the vehicle speed, and performing interpolation and completion on the sparse areas and downsampling on the dense areas of the downsampled point clouds; A point cloud feature extraction unit 605, configured to obtain point cloud features by using the preprocessed point clouds; A fusion unit 606, configured to fuse the image features and the point cloud features to obtain multi-modal fusion features; A task execution unit 607, configured to perform 3D object detection and / or vehicle trajectory prediction by using the multi-modal fusion features.
[0102] Further, the image preprocessing unit 602 may specifically be configured to adjust the resolution of the view image in the following manner: determining the illumination level of the view image through a lightweight convolutional neural network classifier; when the illumination level of the view image is lower than a predetermined gear, increasing the resolution of the view image to the image resolution corresponding to the predetermined gear; when the illumination level of the view image is higher than the predetermined gear, decreasing the resolution of the view image to the image resolution corresponding to the predetermined gear.
[0103] Further, the image feature extraction unit 603 may specifically be configured to: perform the following processing on each view image in the preprocessed multi-view images to extract the features of each view image: dynamically adjusting the RGB channel weights of the basic convolutional kernels according to the illumination level of the view image through an image feature extraction module including conditional convolution to generate dynamic convolutional kernels adapted to the illumination condition of the view image and using the dynamic convolutional kernels to extract the features of the view image; and aligning and fusing the features of each view image to obtain multi-scale features.
[0104] Further, the image feature extraction unit 603 may also be configured to: generate a semantic segmentation mask based on the view image by using a lightweight semantic segmentation network, and apply the semantic segmentation mask as an attention weight in the process of extracting the features of the view image by using the dynamic convolutional kernels.
[0105] Further, the point cloud preprocessing unit 604 can specifically be used for: adjusting the voxel grid size according to the local point cloud density to adaptively voxelize the preprocessed point cloud to obtain voxel data; and using separable 3D convolution to process the voxel data to obtain point cloud features.
[0106] Further, the environment-adaptive multimodal data fusion device 600 may further include: a consistency evaluation unit 608 and a forgery unit 609; wherein, the consistency evaluation unit 608 can be used to evaluate whether the cross-modal consistency of the image features and the point cloud features meets the requirements; the forgery unit 609 can be used to generate a pseudo multi-view image based on the point cloud using a generative adversarial network GAN when the cross-modal consistency of the image features and the point cloud features does not meet the requirements, obtain the features of the pseudo multi-view image, and merge the features of the pseudo multi-view image into the image features.
[0107] Further, the consistency evaluation unit 608 can specifically be used to evaluate whether the cross-modal consistency of the image features and the point cloud features meets the requirements by one or both of the following: 1) calculating the similarity between the point cloud features and the image features at the same spatial position through BEV space contrast learning, and determining that the cross-modal consistency of the point cloud features and the image features does not meet the requirements when the similarity is less than the first predetermined similarity threshold; 2) calculating the cross-modal feature similarity between the forward view image and the point cloud in the multi-view image, and determining that the cross-modal consistency of the point cloud and the forward view image does not meet the requirements when the cross-modal feature similarity is less than the second predetermined similarity threshold.
[0108] Further, the environment-adaptive multimodal data fusion device 600 may further include: an image feature modality weight adjustment unit 610 and / or a point cloud feature modality weight adjustment unit 611. The image feature modality weight adjustment unit 610 can be used to dynamically adjust the modality weight of the image features according to the exposure anomaly conditions of each view image in the multi-view image, and the point cloud feature modality weight adjustment unit 611 can be used to dynamically adjust the modality weight of the point cloud features according to the local point cloud density of the point cloud. The fusion unit 606 can specifically be used for: fusing the image features and the point cloud features based on the modality weight of the image features and / or the modality weight of the point cloud features to obtain multi-modal fusion features.
[0109] In a specific application, the environment-adaptive multimodal data fusion device 600 can be implemented by software, hardware, or a combination of both. Exemplarily, the environment-adaptive multimodal data fusion device 600 can be implemented as the following electronic device 700 or can be implemented as software in the following electronic device 700.
[0110] In addition, an embodiment of the present disclosure also provides a computer-readable storage medium, on which a computer program is stored. The program includes instructions that, when executed by one or more processors of a computing device, perform the steps of the aforementioned environment-adaptive multi-modal data fusion method.
[0111] Figure 7 FIG. shows a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. Refer to Figure 7 , the electronic device 700 may include: one or more processors 701, and further includes a memory 702 storing one or more programs, which are executed by the one or more processors 701 to implement the method flow shown in the above embodiments of the present disclosure and / or the program units corresponding to each unit in the device.
[0112] The processor 701 may include one or more single-core processors or multi-core processors. The processor 701 may include a combination of any general-purpose processor or special-purpose processor (such as a CPU, GPU, etc.).
[0113] The memory 702 is the computer-readable storage medium provided by the present disclosure, and can be used to store non-transitory software programs, non-transitory computer-executable programs, and units, such as the program instructions / units corresponding to the environment-adaptive multi-modal data fusion method shown in the embodiments of the present disclosure as Figure 2 shown. The processor 701 executes the non-transitory software programs, instructions, and units stored in the memory 702, thereby executing the programs, instructions, and units corresponding to the environment-adaptive multi-modal data fusion method shown in the above method embodiments as Figure 2 shown.
[0114] The electronic device 700 may further include: an input device 703, an output device 704, a communication component 705, etc. The processor 701, the memory 702, the input device 703, the output device 704, and the communication component 705 may be connected by a bus or other means, Figure 7 taking connection by bus as an example in
[0115] The above program (also referred to as software, software application, or code) includes machine instructions of a programmable processor, and these computing programs can be implemented using a vehicle-oriented programming language, assembly, or machine language.
[0116] With the development of time and technology, the meaning of the medium has become increasingly broad. The dissemination channels of computer programs are no longer limited to tangible media and can also be directly downloaded from the network, etc. Any combination of one or more computer-readable storage media can be adopted. The computer-readable storage medium can be, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device.
[0117] In specific applications, the electronic device 700 can be implemented as, but is not limited to, a domain controller or other similar devices.
[0118] The embodiments of the present disclosure also provide a vehicle, which is equipped with a multi-view image sequence and a lidar. The vehicle may include the aforementioned environment-adaptive multi-modal data fusion device 600 and / or the electronic device 700.
[0119] The above has introduced the technical solutions provided by the present disclosure in detail. Specific examples are used in this article to elaborate on the principles and implementation manners of the present disclosure. The description of the above embodiments is only used to help understand the method and its core idea of the present disclosure; at the same time, for those of ordinary skill in the art, according to the idea of the present disclosure, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present disclosure.
[0120] The above is only a preferred embodiment of the present disclosure and is not intended to limit the present disclosure. Any modifications, equivalent replacements, etc. made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. An environment - adaptive multimodal data fusion method, characterized in that, The method is applied to a vehicle, and a multi-view camera array and a lidar are mounted on the vehicle; the method includes: Obtaining multi-view images collected by the multi-view camera array and point clouds collected by the lidar; Preprocessing the multi-view images, where the preprocessing includes: for each view image in the multi-view images, adjusting the resolution of the view image according to the illumination condition of the view image; Obtaining image features using the preprocessed multi-view images; Preprocessing the point clouds, where the preprocessing includes: downsampling the point clouds according to the vehicle speed, and performing sparse region interpolation completion and dense region downsampling on the downsampled point clouds; Obtaining point cloud features using the preprocessed point clouds; Fusing the image features and the point cloud features to obtain multi-modal fusion features; Performing 3D object detection and / or vehicle trajectory prediction using the multi-modal fusion features.
2. The method according to claim 1, wherein Adjusting the resolution of the view image in the following manner: Determining the illumination level of the view image through a lightweight convolutional neural network classifier; When the illumination level of the view image is lower than a predetermined level, increasing the resolution of the view image to the image resolution corresponding to the predetermined level; When the illumination level of the view image is higher than the predetermined level, decreasing the resolution of the view image to the image resolution corresponding to the predetermined level.
3. The method according to claim 2, characterized in that, The obtaining image features using the preprocessed multi-view images includes: Performing the following processing on each view image in the preprocessed multi-view images to extract the features of each view image: dynamically adjusting the RGB channel weights of the basic convolutional kernel according to the illumination level of the view image through an image feature extraction module including conditional convolution to generate a dynamic convolutional kernel adapted to the illumination condition of the view image, and using the dynamic convolutional kernel to extract the features of the view image; Aligning and fusing the features of each view image to obtain multi-scale features.
4. The method according to claim 3, characterized in that, The obtaining image features using the adjusted multi-view images further includes: generating a semantic segmentation mask based on the view image using a lightweight semantic segmentation network, and applying the semantic segmentation mask as an attention weight in the process of using the dynamic convolutional kernel to extract the features of the view image.
5. The method according to claim 1, wherein The obtaining point cloud features using the preprocessed point clouds includes: Adjusting the voxel grid size according to the local point cloud density to adaptively voxelize the preprocessed point clouds to obtain voxel data; Processing the voxel data using separable 3D convolution to obtain the point cloud features.
6. The method according to claim 1, characterized in that, The method further includes: Evaluating whether the cross-modal consistency of the image features and the point cloud features meets the requirements; When the cross-modal consistency of the image features and the point cloud features does not meet the requirements, using a generative adversarial network GAN to generate pseudo multi-view images based on the point clouds, obtaining the features of the pseudo multi-view images, and merging the features of the pseudo multi-view images into the image features.
7. The method according to claim 6, characterized in that The evaluating whether the cross-modal consistency of the image features and the point cloud features meets the requirements includes one or more of the following: Calculate the similarity between the point cloud features and the image features at the same spatial position through BEV space contrast learning. When the similarity is less than the first predetermined similarity threshold, it is determined that the cross-modal consistency of the point cloud features and the image features does not meet the requirements; Calculate the cross-modal feature similarity between the forward-view image in the multi-view image and the point cloud. When the cross-modal feature similarity is less than the second predetermined similarity threshold, it is determined that the cross-modal consistency of the point cloud and the forward-view image does not meet the requirements.
8. The method according to claim 1, wherein The method further includes: dynamically adjusting the modal weight of the image features according to the exposure anomaly conditions of each view image in the multi-view image; and / or, dynamically adjusting the modal weight of the point cloud features according to the local point cloud density of the point cloud; The fusing the image features and the point cloud features to obtain multi-modal fusion features includes: fusing the image features and the point cloud features based on the modal weight of the image features and / or the modal weight of the point cloud features to obtain the multi-modal fusion features.
9. An environment-adaptive multi-modal data fusion device, characterized in that, The device is applied to a vehicle, and a multi-view camera array and a lidar are installed on the vehicle; The environment-adaptive multi-modal data fusion device includes: A data acquisition unit for acquiring multi-view images collected by the multi-view camera array and point clouds collected by the lidar; An image preprocessing unit for preprocessing the multi-view images, and the preprocessing includes: for each view image in the multi-view images, adjusting the resolution of the view image according to the illumination condition of the view image; An image feature extraction unit for obtaining image features by using the preprocessed multi-view images; A point cloud preprocessing unit for preprocessing the point cloud, and the preprocessing includes: downsampling the point cloud according to the vehicle speed, and performing sparse region interpolation and dense region downsampling on the downsampled point cloud; A point cloud feature extraction unit for obtaining point cloud features by using the preprocessed point cloud; A fusion unit for fusing the image features and the point cloud features to obtain multi-modal fusion features; A task execution unit for performing 3D object detection and / or vehicle trajectory prediction by using the multi-modal fusion features.
10. A vehicle, wherein the vehicle is equipped with a multi-view camera sequence and a lidar, characterized in that, The vehicle includes the device according to claim 9.
Citation Information
Patent Citations
Three-dimensional model restoration method based on three-dimensional deep convolutional generative adversarial network
CN112634145A
Multi-modal target detection method and device, equipment and storage medium
CN117911827A
Vision-language pre-training general framework for realizing multi-granularity cross-modal alignment
CN119206697A
Near-ground landslide detection method and system and computer readable storage medium
CN119323727A
Target detection method, system and equipment based on multi-sensor fusion and medium
CN119399716A
Cited By
3D occupancy grid prediction method and device, electronic equipment and storage medium
CN120976543A