Scene perception method and device based on multiple modes, electronic equipment and storage medium
Through the multimodal data fusion of multi-view camera array and 4D imaging radar group, combined with lightweight feature extraction and cross-modal adaptive fusion, the problems of low data fusion efficiency and high computing resource consumption of autonomous driving technology in complex scenarios are solved, and efficient and accurate scene perception and real-time requirements are achieved.
Patent Information
- Application Number
- CN202510714686.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-30
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-30
AI Technical Summary
In complex scenarios such as night or rainy days, existing autonomous driving technology has low efficiency and large errors in multimodal data fusion, and consumes a lot of computing resources, making it difficult to meet real-time requirements.
Multi-view camera array and 4D imaging radar group are used to obtain multi-view image sequences and 4D radar data sequences, and efficient multi-modal data fusion is achieved through lightweight feature extraction, timing-space joint modeling and cross-modal adaptive fusion.
It improves the efficiency and accuracy of data fusion in complex scenarios, reduces computing resource consumption, meets the real-time requirements of autonomous driving scenarios, and improves the robustness and reliability of the system.
Smart Images

Figure CN120219905A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of autonomous driving, and in particular, to a multi-modal based scene perception method, apparatus, electronic device, and storage medium. Background Art
[0002] With the rapid development of autonomous driving technology, the fusion of multi-modal sensor data has become the core technical route to improve the reliability of environmental perception.
[0003] Currently, multi-modal fusion solutions based on cameras, lidar (LiDAR), and millimeter-wave radars mainly fuse camera images, 4D radar data, and point cloud data collected by lidar, extract features using a voxel encoder with a fixed resolution, and rely on an attention mechanism to achieve cross-modal feature fusion. Finally, 3D object detection boxes and scene occupancy prediction results are synchronously output through a multi-task head.
[0004] Although the above solutions can achieve object detection and occupancy prediction through multi-modal data fusion, there are still problems such as low efficiency and large errors in multi-modal data fusion in complex scenarios such as at night and in rainy days, and relatively large consumption of computing resources. Therefore, there is an urgent need for a new multi-modal data fusion solution to improve the efficiency and accuracy of data fusion in complex scenarios, while reducing the consumption of computing resources, so as to meet the real-time requirements of autonomous driving scenarios. Summary of the Invention
[0005] In view of this, the present disclosure provides a multi-modal based scene perception method, apparatus, electronic device, and storage medium.
[0006] According to a first aspect of the present disclosure, there is provided a multi-modal based scene perception method, the method comprising: Obtaining a current multi-view image sequence and a current 4D radar data sequence from a multi-view camera array and a 4D imaging radar group mounted on an object; Obtaining a current image feature using the current multi-view image sequence; Obtaining a current radar feature using the current 4D radar data sequence, the radar feature including a radar voxel feature and a radar BEV feature; Based on a previously cached historical radar feature and the current radar feature, modeling the spatio-temporal evolution of a dynamic scene and a static scene in BEV space and voxel space to obtain a current dynamic scene feature and a current static scene feature; Performing cross-modal interaction fusion on the current image feature, the current dynamic scene feature, and the current static scene feature to obtain a multi-modal fusion feature, the multi-modal fusion feature being used to perform 3D object detection, semantic occupancy prediction, and / or motion state estimation on the environment around the object.
[0007] In some embodiments of the first aspect of the present disclosure, obtaining the current image features by using the current multi-view image sequence includes: obtaining a multi-view feature map sequence based on the current multi-view image sequence by using an image encoder, where the multi-view feature map sequence includes feature maps of each view image in the current multi-view image sequence, and the image encoder includes a lightweight convolutional neural network that removes the last two convolutional layers and inserts a channel attention mechanism SE module into the penultimate layer; obtaining serialized pyramid features based on the multi-view feature map sequence by using a Feature Pyramid Network (FPN), and the serialized pyramid features are the current image features, and the upsampling of the FPN uses a Content-Aware ReAssembly (CARAFE) operator.
[0008] In some embodiments of the first aspect of the present disclosure, obtaining the current radar features by using the current 4D radar data sequence includes: performing dynamic voxel compression on the current 4D radar data sequence to obtain current radar voxel features; obtaining current radar Bird's Eye View (BEV) features based on the current radar voxel features by using a radar encoder, where the radar encoder includes a plurality of consecutive feature extraction modules, each feature extraction module includes a sparse convolutional layer, and the convolutional kernel sizes of the sparse convolutional layers in the plurality of consecutive feature extraction modules are the same and the number of channels increases gradually.
[0009] In some embodiments of the first aspect of the present disclosure, modeling the spatio-temporal evolution of dynamic and static scenes in BEV space and voxel space based on the previously cached historical radar features and the current radar features to obtain the current dynamic scene features and the current static scene features includes: Reading the previously cached historical radar features from the buffer, where the historical radar features include historical radar voxel features and historical radar BEV features; Obtaining dynamic scene features and static scene features by using a multi-scale dilated temporal fusion network based on the historical radar features and the current radar features, where the multi-scale dilated temporal fusion network includes a dilated temporal convolutional network, and the dilated temporal convolutional network includes a plurality of consecutive dilated convolutional layers, and the dilation rates of the plurality of consecutive dilated convolutional layers increase exponentially by level and the convolutional kernel sizes are the same; Performing dynamic target compensation on the dynamic scene features in a manner of multi-target Kalman filtering tracking.
[0010] In some embodiments of the first aspect of the present disclosure, performing cross-modal interaction fusion on the current image features, the current dynamic scene features, and the current static scene features to obtain multi-modal fusion features includes: Fusing the current image features and the current static scene features to obtain image voxel features; Dynamically fuse the dynamic scene features and the image voxel features using a gated cross-attention module to obtain cross-modal interaction features; Fuse the dynamic scene features and the image voxel features under geometric consistency constraints to obtain geometrically aligned features; Integrate the geometrically aligned features and the cross-modal interaction features through multi-scale pyramid pooling to obtain the multi-modal fusion features.
[0011] In some embodiments of the first aspect of the present disclosure, the dynamically fusing the dynamic scene features and the image voxel features using a gated cross-attention module to obtain cross-modal interaction features includes: Calculate a gated matrix of the image voxel features and the dynamic scene features, where the gated matrix represents the fusion weights of the dynamic scene features; Dynamically adjust the gated matrix according to the current weather conditions; Fuse the image voxel features and the dynamic scene features based on the dynamically adjusted gated matrix to obtain the cross-modal interaction features.
[0012] In some embodiments of the first aspect of the present disclosure, the method further includes: obtaining 3D object detection results, semantic occupancy prediction results, and / or motion estimation results of the environment around the object based on the multi-modal fusion features.
[0013] According to a second aspect of the present disclosure, there is provided a multi-modal based scene perception device, including: A data acquisition unit for acquiring a current multi-view image sequence and a current 4D radar data sequence from a multi-view camera array and a 4D imaging radar group mounted on an object; An image feature extraction unit for obtaining current image features using the current multi-view image sequence; A radar feature extraction unit for obtaining current radar features using the current 4D radar data sequence, where the radar features include radar voxel features and radar BEV features; A temporal fusion unit for modeling the spatio-temporal evolution of dynamic and static scenes in BEV space and voxel space based on previously cached historical radar features and current radar features to obtain current dynamic scene features and current static scene features; A cross-modal interaction fusion unit for performing cross-modal interaction fusion on the current image features, current dynamic scene features, and current static scene features to obtain multi-modal fusion features, where the multi-modal fusion features are used to perform 3D object detection, semantic occupancy prediction, and / or motion state estimation on the environment around the object.
[0014] According to a third aspect of the present disclosure, there is provided an electronic device, the electronic device includes: At least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above-mentioned method.
[0015] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the above-mentioned method.
[0016] As can be seen from the above, the embodiments of the present disclosure provide a novel multi-modal fusion method for multi-view image sequences and 4D radar data sequences, which can achieve high-precision and high-efficiency 3D object detection, semantic occupancy prediction and scene understanding in complex environments, and can meet the high real-time requirements, low computing resource requirements, and relatively high robustness and reliability requirements of scenarios such as autonomous driving. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0018] Figure 1 is a schematic diagram of the architecture of the system applicable to the embodiments of the present disclosure; Figure 2 is a schematic flowchart of the multi-modal based scene perception method provided by the embodiments of the present disclosure; Figure 3 is a schematic diagram of the structure of the scene perception model involved in the embodiments of the present disclosure; Figure 4 is a schematic flowchart of the specific implementation process of modeling the spatio-temporal evolution of dynamic and static scenes in the BEV space and the voxel space involved in the embodiments of the present disclosure; Figure 5 is a schematic flowchart of the specific implementation process of cross-modal interaction and fusion involved in the embodiments of the present disclosure; Figure 6 is a schematic diagram of the structure of the multi-modal based scene perception device provided by the embodiments of the present disclosure; Figure 7 is a schematic structural block diagram of the electronic device provided by the embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] Next, the technical solutions in the embodiments of the present disclosure will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present disclosure.
[0020] The terms used in the embodiments of the present disclosure are only for the purpose of describing specific embodiments, and are not intended to limit the present disclosure. The singular forms "a", "the" and "said" used in the embodiments of the present disclosure and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise.
[0021] Depending on the context, words such as "if" and "when" used herein can be interpreted as "when...", "when...", "in response to determining", or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" can be interpreted as "when determined", "in response to determining", "when detecting (stated condition or event)", or "in response to detecting (stated condition or event)".
[0022] As mentioned above, in related technologies, in extreme scenarios such as thunderstorms and nights, there are still problems in multi-modal data fusion, such as low fusion efficiency, low resource utilization rate, sharp drop in accuracy and reliability, difficulty in reconstructing complex-shaped objects, and high missed detection rate in static scenarios. In view of this, the embodiments of the present disclosure provide the following multi-modal based scene perception methods, devices, electronic devices and storage media, which use cameras and 4D radars as core sensors, and achieve efficient multi-modal based scene perception through lightweight feature extraction, temporal-spatial joint modeling, cross-modal adaptive fusion, etc., so as to highly reliably implement various tasks such as 3D object detection, semantic occupancy prediction, and motion state estimation.
[0023] For ease of understanding, a brief description of the system structure applicable to the embodiments of the present disclosure will be given first.
[0024] Figure 1 The structural schematic diagram of the system applicable to the embodiments of the present disclosure is shown. Refer to Figure 1 , the system applicable to the embodiments of the present disclosure may include: an electronic device and a peripheral sensor component connected to the electronic device, and the peripheral sensor component includes a multi-view camera array and a 4D imaging radar group.
[0025] The multi-view camera array can be used to collect multi-view images of the environment where the vehicle is located, and the multi-view images can cover the surroundings of the environment where the vehicle is located.
[0026] Exemplarily, the multi-view camera array can be implemented as, but not limited to, a six-view camera array, which includes a front-view camera, a rear-view camera, a left-view camera, a right-view camera, an upper-view camera, and a lower-view camera. Taking the Traffic Jam Assistant (TJA) as an example, the cameras in the multi-view camera array can adopt, but not limited to, a 1920×1080@30FPS, 120° wide-angle lens.
[0027] In specific applications, cameras can be evenly deployed around the vehicle so that these cameras can cover the surrounding of the vehicle without dead angles, and each camera captures an image of the surrounding environment of the vehicle, which contains a part of the surrounding scene of the vehicle. The specific implementation manner of the multi-view camera array is not limited in the embodiments of the present disclosure.
[0028] The 4D imaging radar group can be used to collect 4D radar data of the environment where the vehicle is located, and the 4D radar data can cover the surrounding of the environment where the vehicle is located. For example, the 4D imaging radar group can include 4 groups of 4D imaging radars, namely a forward 4D imaging radar group, left and right lateral 4D imaging radar groups, and a rearward 4D imaging radar group. Another example is that the 4D imaging radar group can include 6 groups of 4D imaging radars, namely dual forward 4D imaging radar groups, left and right lateral 4D imaging radar groups, and dual rearward 4D imaging radar groups. The number and deployment positions of the 4D imaging radars in the 4D imaging radar group need to comprehensively consider function coverage, computing power allocation, cost and reliability, etc., and can be flexibly deployed according to actual needs in specific applications.
[0029] The 4D radar data can include, but not limited to, the spatial position, speed, reflection intensity, signal-to-noise ratio (SNR) of the target, etc. Here, "target" refers to the physical entities detected by the 4D imaging radar group, and these physical entities can be, but not limited to, dynamic objects and static objects. Mobile physical entities such as cars, trucks, motorcycles, bicycles, pedestrians, etc. are dynamic objects, and static physical entities such as roadblocks, curbs, stationary vehicles, traffic signs, street lights, bridges, tunnels, ground fixed facilities, manhole covers, potholes, road cracks, cables, tree branches, etc. are static objects.
[0030] The spatial position of the target includes: the distance, azimuth angle, and elevation angle of the target in three-dimensional space. The distance refers to the straight-line distance between the radar and the target, which is determined by calculating the time difference between signal transmission and reception. The azimuth angle represents the horizontal direction angle (left and right direction) of the target relative to the radar, and the elevation angle represents the vertical direction angle (up and down direction) of the target relative to the radar.
[0031] Velocity refers to the radial velocity of the target relative to the radar (i.e., the velocity component along the radar beam direction), which can be measured through the Doppler effect. Stationary and moving objects (such as vehicles and pedestrians) can be distinguished by velocity.
[0032] Reflection Intensity refers to the intensity of the electromagnetic wave energy reflected by the target back to the radar, which is related to the target's material, surface roughness, geometric shape, and radar cross-section (RCS). Metal objects such as vehicles have high reflection intensity, while non-metal objects such as pedestrians have low reflection intensity. Reflection intensity can assist in target classification and enhance the accuracy of environmental perception.
[0033] Signal-to-Noise Ratio (SNR) refers to the ratio of the useful signal to the background noise in the radar signal, reflecting the signal quality. A high SNR indicates reliable target detection and high data confidence; a low SNR may indicate false detection or missed detection due to noise interference. SNR can be used for data filtering to improve the system's robustness.
[0034] In specific applications, the systems applicable to the embodiments of the present disclosure can be applied to, but are not limited to, intelligent control of devices such as multiple wheeled mobile robots, wheeled mobile robots, mobile robots, vehicles, aircraft, ships, Autonomous rail Rapid Transit (ART), industrial automation equipment, etc. Vehicles can be, but are not limited to, passenger vehicles, commercial vehicles (e.g., trucks, buses, freight vehicles, etc.), special-purpose vehicles (e.g., ambulances, fire trucks, engineering vehicles, rescue vehicles, etc.), agricultural and industrial vehicles (e.g., harvesters, forklifts, etc.), transportation and logistics vehicles (e.g., container trucks, refrigerated trucks, etc.), new energy vehicles (e.g., electric vehicles, hybrid vehicles), special carriers (e.g., garbage trucks, sprinkler trucks, etc.). In other words, "vehicles" in the embodiments of the present disclosure are equivalent to the aforementioned various devices.
[0035] The embodiments of the present disclosure can be applied to scenarios such as urban transportation, highways, ports, mines, farms, closed parks, industrial production, etc., and are applicable to many aspects such as ride-hailing, public transportation, logistics distribution, unmanned transportation, last-mile delivery, automated agricultural operations, automated environmental sanitation, etc. Of course, the embodiments of the present disclosure can also be applied to any other intelligent control scenarios involving devices such as vehicles. The present disclosure does not limit the application scenarios and applicable fields of the embodiments of the present disclosure.
[0036] The systems applicable to the embodiments of the present disclosure can be, but are not limited to, any systems that require multi-modal data fusion. For example, the system can be, but is not limited to, an autonomous driving system, an intelligent assisted driving system, a traffic congestion assistance system, etc. See Figure 1, the system provided by the embodiments of the present disclosure can be loaded on a vehicle and used as, but not limited to, an intelligent assisted driving system, a traffic congestion assistance system, an autonomous driving system, etc. of the vehicle.
[0037] In addition, those skilled in the art should understand that the system applicable to the embodiments of the present disclosure is not limited to Figure 1 the architecture shown, and its application scenarios are not limited to the above several types either.
[0038] The following will make a detailed description of the specific implementation manners of the embodiments of the present disclosure.
[0039] Figure 2 shows a schematic flowchart of a multi-modal based scene perception method provided by the embodiments of the present disclosure. This multi-modal based scene perception method can be executed by an electronic device in the Figure 1 system shown. Referring to Figure 2 , the multi-modal based scene perception method of the embodiments of the present disclosure includes the following steps: Step 201, obtain the current multi-view image sequence and the current 4D radar data sequence from the multi-view camera array and the 4D imaging radar group loaded on the object; Step 202, obtain the current image features using the current multi-view image sequence; Step 203, obtain the current radar features using the current 4D radar data sequence. The radar features include radar voxel features and radar bird's eye view (BEV) features; Step 204, based on the previously cached historical radar features and the current radar features, model the spatio-temporal evolution of the dynamic scene and the static scene in the BEV space and the voxel space to obtain the current dynamic scene features and the current static scene features; Among them, the dynamic scene features can be used to characterize the spatio-temporal evolution of dynamic objects such as vehicles and pedestrians in the environment around the object, and the static scene features can be used to characterize the spatio-temporal evolution of static objects such as roads, traffic signs, and signal lights in the environment around the vehicle.
[0040] Exemplarily, the dynamic features can be in various forms such as BEV format, trajectory sequence, etc., and the static features can be in various forms such as voxel format, BEV format, or point cloud. Taking a vehicle as an example, the dynamic features can be real-time local perception implemented in the
[0041] vehicle body coordinate system, and the static features can be in the world coordinate system to ensure global consistency.
[0042] Here, the object can be, but is not limited to, the aforementioned such as multiple wheeled mobile robots, wheeled mobile robots, mobile robots, vehicles, aircraft, ships, intelligent rail rapid transit systems (ART, Autonomous rail Rapid Transit), industrial automation equipment, etc. For details, please refer to the previous description and will not be elaborated here.
[0043] In the embodiments of the present disclosure, multi-modal data is collected through a multi-view camera array and a 4D imaging radar group. After image feature extraction and radar feature extraction, the spatio-temporal evolution and cross-modal interaction fusion of dynamic and static scenes are modeled in the BEV space and the voxel space to obtain multi-modal fusion features obtained by fusing multi-view images and 4D radar data. The multi-modal fusion features can be used to obtain 3D object detection, semantic occupancy prediction, and / or motion of the surrounding environment of the object. Thus, the embodiments of the present disclosure provide a brand-new multi-modal fusion method, which can reduce the impact of fusion deviation and reduce the consumption of computing resources, improve the efficiency and accuracy of data fusion in complex scenarios while reducing the computing resource requirements, and can achieve high-precision and high-efficiency scene perception and understanding in complex environments while taking into account the system robustness and computing efficiency, so as to meet the real-time requirements, low computing resource requirements, high robustness and reliability requirements of various scenarios such as autonomous driving.
[0044] The multi-modal based scene perception method of the embodiments of the present disclosure can be implemented through a pre-constructed scene perception model. Figure 3 The architecture diagram of the scene perception model is shown. Refer to Figure 3 , the scene perception model may include a sensor input layer, an image feature extraction layer, a radar feature extraction layer, a temporal fusion layer, a cross-modal interaction layer, and a multi-task output layer. The sensor input layer is used to implement the aforementioned step 201, the image feature extraction layer is used to implement step 202, the radar feature extraction layer can be used to implement step 203, the temporal fusion layer is used to implement the aforementioned step 204, the cross-modal interaction layer is used to implement the aforementioned step 205, and the multi-task output layer is used to implement the following step 206.
[0045] In step 201, the multi-view image sequence and the 4D radar data sequence are time-synchronized. Specifically, the time synchronization of the multi-view image sequence and the 4D radar data sequence can be achieved through mechanisms such as synchronous triggering and time alignment of the multi-view camera array and the 4D imaging radar group.
[0046] Before step 202, the multi-view image sequence can be pre-processed, and the pre-processing can include, but is not limited to, image signal processor (ISP) processing for the original image (e.g., RAW image), de-distortion / white balance, etc.
[0047] Before step 203, the 4D radar data sequence can also be preprocessed, which may include but is not limited to multi-frame accumulation, Doppler compensation, etc.
[0048] There are various specific implementation manners for obtaining the image features of the multi-view image sequence in step 202. In some embodiments, the image features of the multi-view image sequence can be obtained through a pre-constructed image encoder and a Feature Pyramid Network (FPN).
[0049] Further, step 202 may include the following steps a1 and a2: Step a1, using the image encoder to process the current multi-view image sequence to obtain a multi-view feature map sequence, where the multi-view feature map sequence includes the feature maps of each view image in the current multi-view image sequence. Step a2, using a Feature Pyramid Network (FPN) to process the multi-view feature map sequence to obtain serialized pyramid features, and the serialized pyramid features are the image features of the multi-view image. The upsampling of the FPN can adopt a Content-Aware Reassembly of Features (CARAFE) operator.
[0050] As described above, the feature extraction of the multi-view image sequence is realized through a lightweight image editor, an SE module, and Content-Aware Reassembly of Features in FPN (CARAFE-FPN). The number of model parameters is reduced by about 9.3%, the inference speed is increased by 29%, and the dynamic upsampling of CARAFE improves the small target AP by 3.6%. The independent processing and enhancement strategy of CARAFE-FPN ensures the spatial alignment accuracy of cross-view features (+5.5%), can balance the computational efficiency and the multi-scale feature expression ability, and at the same time retains the multi-view details and spatial consistency. Moreover, this method is especially suitable for scenarios that require real-time processing of multi-view image data on mobile devices and can achieve a better balance between accuracy and efficiency.
[0051] In step a1, the image encoder may include a lightweight convolutional neural network that removes the last two convolutional layers and inserts a Squeeze-and-Excitation (SE) module into the penultimate layer.
[0052] Exemplarily, the above lightweight convolutional neural network can be but is not limited to MobileNetV3-Large. That is, the backbone network of the image encoder can be MobileNetV3-Large with the last two convolutional layers removed and an SE module inserted into the penultimate layer. For example, the backbone network of the image encoder can be configured to: after removing the last two convolutional layers of the MobileNetV3-Large model, insert an SE module with a compression ratio of 16 into the penultimate layer. In practical applications, the compression ratio of the SE module can be adjusted according to requirements to further optimize performance.
[0053] The SE module is a lightweight channel attention mechanism that enhances the feature response of key channels by dynamically adjusting the weights of the channel dimension of the two-dimensional feature map. For example, the penultimate layer of lightweight convolutional neural networks such as MobileNetV3-Large is a key position for high-level semantic feature extraction. Inserting the SE module into the penultimate layer of the lightweight convolutional neural network can strengthen key information in the high-level semantic features, prevent information loss, and optimize computational efficiency, significantly improving accuracy and robustness while maintaining low computational overhead, so as to better apply to tasks that require fine-grained spatial information while retaining the lightweight advantage.
[0054] Specifically, the image encoder can include: an initial convolutional layer, a plurality of inverted residual blocks, and an SE module connected in sequence. Specifically, in step a1, the following steps a11 to a13 can be performed on the images of each view in the multi-view image sequence to obtain the feature maps of each view image: Step a11, using the initial convolutional layer to extract the first feature map of the image, and the size of the first feature map in height and width is half of the corresponding dimension size of the image; Step a12, using a plurality of inverted residual blocks to process the first feature map to obtain a second feature map, and the size of the second feature map in height and width is 1 / 16 of the corresponding dimension size of the image, and the high-resolution spatial information of the image is retained in the second feature map; Step a13, using the SE module to dynamically adjust the weights of each channel in the second feature map to capture key information, so as to obtain a third feature map, and the size of the third feature map in height and width is 1 / 16 of the corresponding dimension size of the image, and the channel dimension size can be fixed to a predetermined value (for example, 1). For example, the channel dimension size of the third feature map can be 256.
[0055] For example, if the size C×H×W of the images of each view in the multi-view image sequence is 1×544×960, the size of the third feature map can be: 256×34×60, where C represents the channel size, H represents the height size, and W represents the width size.
[0056] Among them, the specific process of using the SE module to dynamically adjust the weights of each channel in the second feature map to capture key information may include: First, Squeeze: perform global average pooling on each channel of the second feature map to compress the two-dimensional spatial information of each channel into a scalar to obtain a channel description vector; Second, Excitation: learn the inter-channel dependencies through a fully connected layer to generate channel weights; Finally, multiply the weight vector with the second feature map channel by channel to obtain the third feature map to enhance the response of important channels.
[0057] As can be seen from the above, by combining the SE module and the lightweight convolutional neural network to form an image encoder, it is possible to retain the global context while extracting the high-resolution spatial information of the image, strengthen the capture of key features through channel attention, optimize the computational efficiency, and achieve an effective balance between efficiency and accuracy.
[0058] The checkerboard effect is caused by the "unequal overlap" of deconvolution. When the convolution kernel size cannot be divided evenly by the stride, the result of the deconvolution output will have unequal overlap, resulting in a checkerboard-like block structure or jagged edges in the feature map. In step a2, when using FPN to fuse the multi-view feature map sequence, the upsampling of FPN adopts the CARAFE operator, which can reduce this checkerboard effect, thereby further improving the accuracy and robustness while maintaining low computational overhead.
[0059] Specifically, CARAFE can be used as the upsampling method in the top-down path of FPN to upsample the high-level features and fuse them with the low-level features. CARAFE is an efficient and lightweight feature upsampling module that dynamically adjusts the upsampling process through a content-aware mechanism, avoiding the unequal overlap problem that may occur in traditional upsampling methods, significantly improving the accuracy and efficiency of upsampling, and aggregating context information within a large receptive field, thereby better utilizing the surrounding information and reducing the checkerboard effect. Moreover, the computational overhead of CARAFE is very small and can be designed in a lightweight manner.
[0060] In step a2, after the FPN feature enhancement, the feature map of each view is multi-scale enhanced to generate a set of feature pyramid features, thereby enhancing the perception ability for targets of different sizes. That is, the multi-view feature map sequence can obtain serialized pyramid features after being processed by FPN, and each view corresponds to a set of feature pyramids in the serialized multi-group pyramid features. This serialized pyramid feature is the image feature of the multi-view image sequence obtained in step 202.
[0061] In step a2, CARAFE restores details and improves spatial resolution, while FPN optimizes feature representation and enhances semantic information. The combination of CARAFE and FPN can significantly improve the representation ability of multi-scale features, enhance the accuracy of small object detection and edge segmentation, achieve significant performance gains at a small computational cost without complex structural adjustments, and perform outstandingly in terms of detail retention, small object detection, and spatial consistency. At the same time, it can solve the bottlenecks of traditional methods in detail loss and spatial misalignment.
[0062] Taking a six-view camera array as an example, the multi-view image sequence is a six-view image sequence [6, 1, 544, 960], where "6" represents the number of views. The processing process of step 202 may include: the six-view image sequence [6, 1, 544, 960] is processed by an image encoder to obtain a six-view feature map sequence [6, 256, 34, 60], where "6" represents the number of views. The six-view feature map sequence can also be represented as {F1, F2,..., F6}, Fi ∈ R 256×34×60 , i = 1, 2, 3,..., 6; each view's feature map in the six-view feature map sequence is independently processed by FPN to obtain a sequence composed of 6 groups of feature pyramids, with each group of feature pyramids corresponding to one view. A group of feature pyramids is a hierarchical structure composed of multiple feature maps with different resolutions. For example, a group of feature pyramids corresponding to view 1 can be represented as {P3_1, P4_1, P5_1}, a group of feature pyramids corresponding to view 2 can be represented as {P3_2, P4_2, P5_2}, a group of feature pyramids corresponding to view 3 can be represented as {P3_3, P4_3, P5_3}, and the other views are similar and will not be elaborated here. Among them, P3 represents a low-level feature map, which has a high resolution, weak semantic information but rich detail information and is suitable for detecting small objects; P4 represents a mid-level feature map, which has a lower resolution and stronger semantic information and is suitable for detecting medium-sized objects; P5 represents a high-level feature map, which has a low resolution and strong semantic information and is suitable for detecting large objects.
[0063] Furthermore, step 203 may include: performing dynamic voxel compression on the current 4D radar data sequence to obtain the current radar voxel feature; and processing the radar voxel feature through a radar encoder to obtain a radar BEV feature map. The radar encoder includes multiple consecutive feature extraction modules, and each feature extraction module includes a sparse convolutional layer. The convolutional kernel sizes of the sparse convolutional layers in the multiple consecutive feature extraction modules are the same and the number of channels increases gradually.
[0064] Specifically, the process of dynamic voxel compression can include: dividing the three-dimensional space into a non-uniform voxel grid by adaptively adjusting the voxel resolution, mapping the point cloud coordinates corresponding to the 4D radar data sequence to the corresponding voxel position in the voxel grid, extracting statistical features of the points in each voxel, retaining only non-empty voxels, and using octrees to hierarchically encode the spatial position relationship of these voxels to achieve efficient neighborhood query and dynamic update storage of non-empty voxels. Dynamic voxel compression can achieve efficient and accurate voxelization of 4D radar data sequences, providing an efficient and accurate voxelization data foundation for 3D perception tasks (such as target detection, semantic occupancy prediction, etc.), reducing memory usage by 30%, and improving the detection AP of dynamic targets such as pedestrians by 3.5%.
[0065] The point cloud coordinates corresponding to the 4D radar data sequence can be obtained by converting the spatial position of the target in the 4D radar data sequence through a coordinate system. The point cloud coordinates can be expressed as but not limited to three-dimensional rectangular coordinate system coordinates.
[0066] For example, the point cloud coordinates corresponding to the 4D radar data sequence may be mapped to corresponding voxel positions in the voxel grid by rounding operations, etc. The number of points in each voxel does not exceed a second preset threshold (e.g., 32 points). If the number of points in a voxel exceeds the second preset threshold, some points may be randomly reserved to avoid computational overload. Statistical feature extraction of points in each voxel includes: calculating the position mean (x, y, z), maximum reflection intensity and velocity variance for each point in each voxel. The position mean is obtained by calculating the average of the spatial positions (i.e., x, y, z coordinates) of all points in the voxel, and the position mean represents the center position of the voxel. The maximum reflection intensity is obtained by taking the maximum reflection intensity of all points in the voxel, and the maximum reflection intensity is used to distinguish the target material. The velocity variance can be calculated based on the velocity of all points in the voxel, and the velocity variance can be used to identify dynamic targets such as vehicles and pedestrians.
[0067] Furthermore, in dynamic voxel compression, the voxel resolution (0.1m³~0.4m³) of each area in the point cloud space of the 4D radar data sequence can be adaptively adjusted through confidence-driven subdivision rules so that the target dense area adopts fine granularity and the open area adopts coarse granularity, thereby realizing adaptive adjustment of voxel resolution.
[0068] Specifically, the confidence-driven subdivision rule can be configured as follows: when the detection confidence of a certain area is greater than a first preset threshold (for example, 0.5), the probability of the existence of a target in the area is determined to be high, and it belongs to a high-confidence area. The voxels in the high-confidence area adopt a fine-grained first predetermined resolution (for example, 0.1m×0.1m×0.2m) to capture target details such as vehicle edges and pedestrian postures with higher accuracy.
[0069] Suppose the initial voxel size of dynamic voxel compression is set to 0.4m × 0.4m × 0.8m. The subdivision rule of dynamic voxel compression can also include: when the detection confidence of a certain area is less than or equal to the first preset threshold (e.g., 0.5), it is determined that the target existence probability of this area is relatively high, belonging to a low-confidence area, and the second predetermined resolution with a coarse granularity (e.g., the initial voxel size of 0.4m × 0.4m × 0.8m) is maintained to reduce computational redundancy.
[0070] Specifically, the process of adaptively adjusting the voxel resolution through the confidence-driven subdivision rule can include: traversing the point cloud space of the 4D radar data sequence, judging the voxel size of each area according to the confidence heat map, applying the first predetermined resolution to the area where the confidence is greater than the preset threshold, and retaining the default size, that is, the second predetermined resolution, for the voxels in other areas. Thus, "high precision in key areas and high efficiency in non-key areas" of adaptive dynamic voxel compression can be achieved.
[0071] Furthermore, in dynamic voxel compression, the statistical feature extraction of points within each voxel can be performed in a block calculation manner. Specifically, all points within the voxel can be divided into multiple sub-blocks of a predetermined size (e.g., 8×8). Each sub-block calculates the attention weight independently, and a predetermined number of key points (e.g., 10 points) within the sub-block are selected through the attention weight and retained as sampling points (e.g., the top 10% of key points sorted from high to low according to the attention weight), and then the statistical feature extraction is performed on the sampling points to reduce the computational amount and improve the efficiency.
[0072] Furthermore, in dynamic voxel compression, a Compute Unified Device Architecture (CUDA) can be used to parallel construct an Octree to update the voxel index table in real time. That is, applying CUDA to parallel construct the voxel index table can quickly map each point in the 4D radar data sequence to the corresponding voxel position, record the voxel ID, the number of points, statistical features, etc. Applying CUDA to parallel construct the voxel index table can accelerate and optimize the dynamic voxel compression of the 4D radar data sequence, compress the time consumption of dynamic voxel compression to the 10ms level, meet the real-time requirement, and the memory occupancy can be further reduced by 30%.
[0073] In step 203, the radar encoder can adopt a convolutional kernel with a fixed size and a channel increment strategy to balance the computational efficiency and the feature expression ability. Furthermore, the feature extraction module in the radar encoder can adopt the same hierarchical structure. Each feature extraction module can include a sparse convolutional layer with a convolutional kernel size of a predetermined value (e.g., 3×3×3), a batch normalization layer (BN), and a ReLU activation function layer connected in sequence.
[0074] In some examples, the radar encoder may include multiple consecutive feature extraction modules, each feature extraction module including a SparseConv layer. The convolutional kernel sizes of the SparseConv layers in these multiple consecutive feature extraction modules are the same and the number of channels increases step by step to gradually extract higher-dimensional abstract features, forming an encoding process from local details to global semantics.
[0075] For example, the radar encoder may include multiple consecutive feature extraction modules. The convolutional kernel sizes of the SparseConv layers in these 3 consecutive feature extraction modules are the same and the number of channels doubles step by step, specifically 64→128→256. That is, the number of output data channels of the SparseConv layer in the first feature extraction module is 64, the input data of the SparseConv layer in the second feature extraction module is 64 channels and the output data is 128 channels, and the input data of the SparseConv layer in the third feature extraction module is 128 channels and the output data is 256 channels. The size of the radar BEV feature obtained by this radar encoder is 256×160×240.
[0076] Further, after obtaining the radar BEV feature through the radar encoder, step 203 may further include steps of correcting the spatio-temporal misalignment in the radar BEV feature through DBSCAN clustering and Hungarian matching, combined with Q matrix adaptive Kalman filtering.
[0077] In some embodiments, the process of correcting the spatio-temporal misalignment in the radar BEV feature may include the following steps b1 to b3: Step b1, segment the point cloud in the radar BEV feature into independent targets through DBSCAN clustering; Step b2, associate the clustering result of the current frame with the historical trajectory through Hungarian Matching to maintain the consistency of the target ID; Step b3, perform state prediction and update on each target through Kalman filtering to obtain the prediction result and compensate for the position of the target in the radar BEV feature according to the prediction result of Kalman filtering. The state vector of Kalman filtering tracks the position and speed of the target and the Q matrix is adaptively adjusted.
[0078] Among them, the neighborhood radius eps of DBSCAN clustering can be set to 0.5m.
[0079] High-precision data association is achieved through DBSCAN clustering and Hungarian matching, and the robustness in complex scenarios is improved by combining Q matrix adaptive Kalman filtering, which can effectively cope with the challenges of motion compensation and multi-target tracking in the radar BEV feature.
[0080] Figure 4The figure shows a schematic diagram of the specific implementation process of modeling the spatio-temporal evolution of dynamic and static scenes in the BEV space and the voxel space in step 204. Refer to Figure 4 , step 204 may include the following steps 401 to 403: Step 401, read the previously cached historical radar features from the buffer, where the historical radar features include historical radar voxel features and historical radar BEV features; Among them, the storage depth of the buffer can be 4 frames, that is, the current frame + 3 historical frames, and the data structure can be a circular buffer, covering a 1.2-second time series window. The circular buffer is a first-in-first-out data structure for efficiently managing continuous data, with a fixed capacity and supporting a cyclic overwrite mechanism.
[0081] Step 402, use a multi-scale dilated temporal fusion network to obtain dynamic scene features and static scene features based on the historical radar features and the current radar features; The multi-scale dilated temporal fusion network includes a dilated temporal convolutional network (Dilated Temporal Convolutional Network, Dilated TCN). The dilated temporal convolutional network includes a plurality of consecutive dilated convolutional layers, and the dilation rates of the plurality of consecutive dilated convolutional layers increase exponentially by level and the convolutional kernel sizes are the same. Thus, the receptive field can be exponentially expanded by the dilated TCN, and the global evolution law of dynamic and static scenes can be captured.
[0082] The multi-scale dilated temporal fusion network may also include a two-branch decoder to separately display the dynamic scene features and the static scene features.
[0083] For example, the dilated temporal convolutional network may include 3 consecutive dilated convolutional layers, and the dilation rates of the 3 consecutive dilated convolutional layers are 1, 2, and 4 respectively. The convolutional kernel size can be set to 3×3, and the number of channels of the output data can be set to 256.
[0084] In the dual-modal modeling of the BEV space and the voxel space, when the multi-scale dilated temporal fusion network models the global scene evolution of static and dynamic features, it uses an exponentially increasing dilation rate and a multi-layer stacked structure that can cover multiple frames of data to capture the long-term evolution law. Through low-dilation-rate convolution (dilation rate = 1), the long-term stability of static elements such as road structures and guardrails is captured to extract static scene features. Through the stacking of dilated convolutions with exponentially increasing dilation rates (dilation rate greater than 1), the acceleration changes and trajectory inflection points of the target are associated and the future motion trend is predicted to model dynamic scene features.
[0085] Furthermore, the multi-scale dilated temporal fusion network can obtain static scene features by stacking multiple layers with a low dilation rate, obtain dynamic scene features by using a mixed dilation rate, and enhance the global scene modeling ability through modality fusion.
[0086] In the embodiments of the present disclosure, the multi-scale dilated temporal fusion network is used to model the spatio-temporal evolution of dynamic and static scenes in the BEV and voxel spaces. Compared with the related technology using LSTM, the number of parameters is reduced from 3.8M to about 1.2M, and the inference speed can be increased by about 30%, which can significantly improve the efficiency and reduce the consumption of computing resources at the same time.
[0087] Step 403, perform dynamic target compensation on the dynamic scene features by using the multi-object Kalman filter tracking method.
[0088] Specifically, the process of performing dynamic target compensation on the dynamic scene features by using the multi-object Kalman filter tracking method may include the following steps c1 to c5: Step c1, convert the point cloud corresponding to the dynamic scene features to the global coordinate system to eliminate the observation distortion caused by the movement of the ego vehicle; Specifically, based on the ego vehicle positioning system (such as the GNSS / IMU fused pose), convert the dynamic scene feature point cloud to the global coordinate system (such as the UTM or world coordinate system), and correct the residual distortion caused by positioning delay or error through a motion compensation algorithm (such as eliminating rotational distortion based on the integration of the IMU angular velocity).
[0089] Step c2, segment the point cloud of the dynamic scene features into independent targets through DBSCAN clustering; Specifically, segment the point cloud of the dynamic scene features based on adaptive DBSCAN clustering.
[0090] Step c3, associate the detection clusters of the current frame with the historical trajectories through the Hungarian matching association trajectory to maintain the ID consistency; Specifically, associate the current detection clusters with the historical trajectories through multi-modal Hungarian matching.
[0091] Step c4, perform state estimation on each target through independent Kalman filter tracking to obtain the prediction result of each target, such as the position of the next frame; Specifically, assign an independent Kalman filter (KalmanFilter, KF) to each target, initialize the Kalman filter for each target, and perform Kalman filter prediction and update to predict the position of the next frame of each target. For example, an interactive multi-model Kalman filter (IMM-KF) can be assigned to each target to support the adaptive switching of multiple motion models (such as constant velocity CV, constant turn rate CTRV). Step c5: Match the prediction results obtained by independent Kalman filter tracking with the measured point cloud. For the successfully matched dynamic targets, calculate their motion offsets relative to the host vehicle to correct the coordinates of the dynamic scene features.
[0092] Specifically, after matching the prediction results with the measured point cloud, the coordinates of the dynamic scene features can be corrected based on the association confidence.
[0093] Considering that the multi-scale hollow temporal fusion network may have problems such as insufficient dynamic target motion compensation and ignoring the independent motion of targets in the global pose transformation, these problems are likely to lead to large speed estimation errors. In view of this, the embodiments of the present disclosure perform dynamic target compensation on the dynamic scene features through multi-target independent Kalman filter tracking and compensating the cooperative motion of the host vehicle and the target. It has been verified that this method can reduce the speed error by about 18%.
[0094] Figure 5 The specific implementation process diagram of cross-modal interaction fusion in step 205 is shown. Refer to Figure 5 , step 205 may include the following steps 501 to 504: Step 501: Fuse the image features and static scene features of the multi-view image sequence to obtain image voxel features; Step 502: Use the gated cross-attention module to dynamically fuse the dynamic scene features and image voxel features to obtain cross-modal interaction features; Step 503: Under the geometric consistency constraint, fuse the dynamic scene features and image voxel features to obtain geometric alignment features; Step 504: Integrate the geometric alignment features and cross-modal interaction features through multi-scale pyramid pooling to obtain multi-modal fusion features.
[0095] As described above, heterogeneous features can be fused through a geometric-semantic dual-driven strategy to obtain multi-modal fusion features.
[0096] Furthermore, the exemplary process of using the gated cross-attention module in step 502 to obtain cross-modal interaction features may include the following steps d1 to d2: Step d1: Calculate the gated matrix of the image voxel features and dynamic scene features and dynamically adjust the gated matrix according to the current weather condition. The gated matrix represents the fusion weight of the dynamic scene features; Step d2: Based on the gated matrix, fuse the image voxel features and dynamic scene features to obtain cross-modal interaction features.
[0097] Thus, the gated mechanism can be used to adaptively fuse radar and camera features.
[0098] Furthermore, in step d1, the weight generation of Gated Cross-Attention is achieved through the following steps: Step d11, feature concatenation, that is, concatenating the radar BEV feature and the image voxel feature along the channel dimension to obtain a concatenated feature; Through feature concatenation, the complete information of the two modalities can be retained, avoiding early information loss, and at the same time allowing the model to autonomously learn the interaction relationships between different modalities.
[0099] Step d12, using a 1×1 convolution to compress the number of channels of the concatenated feature obtained in step b21 to the target dimension to obtain a compressed feature; Among them, the target dimension can be a single channel or aligned with the original feature. Compression through 1×1 convolution can reduce the dimension, reduce the computational amount, focus on key interaction features, and learn the cross-modal weight distribution through the convolution kernel parameters.
[0100] Step d13, applying the Sigmoid function to map the feature values of the compressed feature to the interval [0,1] to generate a basic gating matrix G base ; Specifically, the basic gating matrix can be calculated through the following formula (1) G base : (1) where σ represents the Sigmoid function, ConV([F R ;F I ) represents the compressed feature obtained in step b22, F R represents the radar BEV feature, and F I represents the image voxel feature.
[0101] The basic gating matrix G base ∈[0,1], and each gating value in it represents the feature importance weight. A gating value of 0 means suppression, and a gating value of 1 means retention.
[0102] Step d14, dynamically adjusting the gating matrix G according to the weather conditions so that the gating value under complex weather conditions is greater than 0.5, biased towards the 4D radar; Specifically, under complex weather conditions such as thunderstorms and nights, the G value is biased towards the radar (G∈[0.6,0.8]), and under weather conditions such as clear and daytime, G is biased towards the camera (G∈[0.3,0.5]).
[0103] To achieve the dynamic adjustment of the output interval of the gating matrix G under different weather conditions for scene adaptation, it can be achieved by introducing a weather condition encoding and dynamic range mapping module.
[0104] First, obtain the weather label W indicating the current weather conditions, and encode the weather label W into a weather vector w by means such as One-Hot encoding, learnable embedding, etc. Secondly, map the weather vector w to weather condition parameters L and H, where 0 ≤ L < H ≤ 1, through a small neural network such as an MLP. For example, the Sigmoid function can be used to ensure that L, H ∈ [0, 1], and L is guaranteed to be less than H by sorting.
[0105] Then, according to the weather condition parameters L and H, perform a linear transformation on G base to obtain a dynamically adjusted gating matrix. (2) Thus, under complex weather conditions such as thunderstorms and at night, the G value can be adjusted to the first preset interval, i.e., G ∈ [0.6, 0.8], biased towards the radar; under weather conditions such as sunny and during the day, the G value is adjusted to the second preset interval G ∈ [0.3, 0.5], biased towards the camera.
[0106] From the above, by introducing weather condition encoding and dynamic range mapping, the gated cross-attention can adaptively adjust the weight interval, significantly improving the model robustness in complex environments.
[0107] Step d15, use the dynamically adjusted gating matrix to perform dynamic fusion on the dynamic scene features and image voxel features to obtain cross-modal interaction features.
[0108] Specifically, perform dynamic fusion through the following formula (3) to obtain cross-modal interaction features.
[0109] (3) where F fusion represents the cross-modal interaction feature, F R represents the radar BEV feature, and F I represents the image voxel feature.
[0110] As described above, gated cross-attention realizes the dynamic fusion of dynamic scene features and image voxel features through feature concatenation, 1×1 convolution compression, Sigmoid activation, and dynamically allocating radar and camera fusion weights based on weather conditions. Thus, through a learnable gating mechanism, key information is adaptively enhanced, noise is suppressed, and problems such as reduced accuracy caused by low camera imaging quality and fixed multi-modal fusion weights in bad weather are solved. The mean Average Precision (mAP) in rainy scenarios is increased by 14.4%, the speed error is reduced by 15%, and cross-modal feature optimization can be achieved in a lightweight manner, balancing computational efficiency and model performance, and effectively reducing the missed detection rate of object detection.
[0111] Further, the exemplary implementation process of fusing dynamic scene features and image voxel features in step 503 under geometric consistency constraints may include the following steps e1 to e4: Step e1, perform key point sampling on the dynamic scene features and image voxel features; Specifically, preferentially select areas such as the target center and intersection signs to randomly sample a predetermined number (e.g., 1000) of key points in the dynamic scene features and image voxel features.
[0112] Step e2, perform a BEV-to-voxel projection calculation on the key points in the dynamic scene features to obtain the voxel space coordinates of the key points, and calculate the key point alignment loss L align and the feature IoU loss L IoU using the voxel space projection coordinates of the key points and the coordinates of the corresponding key points in the image voxel features. Sum the key point alignment loss L align and the feature IoU loss L IoU with equal weights to obtain the joint optimization target L total . Adjust the geometric parameter p of the BEV-to-voxel projection calculation by minimizing the joint optimization target L total to obtain the optimized geometric parameter p; The BEV-to-voxel projection calculation can be expressed as the following formula (4).
[0113] (4) Formula (4) represents the projection process from the bird's-eye view space to the voxel space. p BEV =(x bev , y bev ) represents a point in the bird's-eye view (BEV), usually corresponding to the two-dimensional plane coordinates (X, Y) of the real world; p voxel =(v x , v y , v z) is the three-dimensional coordinate in the voxel space, representing the position of a point in the three-dimensional voxel grid.
[0114] The key point alignment loss L align That is, the position alignment loss, which is used to measure the deviation between the projected position of the key point and the actual coordinate, and can be calculated by the Euclidean Distance.
[0115] The feature IoU loss L IoU That is, the feature similarity loss, which is used to measure the difference between the BEV feature and the voxel feature in the vector space.
[0116] The joint optimization objective L total Can be calculated by the following formula (5).
[0117] (5) By assigning λ = 0.5 to the feature loss and the position loss, the position and feature similarity can be balanced, and the geometric consistency loss function can evenly optimize the feature alignment in cross-modal data and the position alignment in the physical space.
[0118] Step e3, apply the optimized geometric parameter p to project the dynamic scene feature into the voxel space to obtain the aligned dynamic scene feature.
[0119] Step e4, according to the feature IoU loss L IoU Dynamically weight and fuse the aligned dynamic scene feature and the image voxel feature to obtain the geometric semantic multi-modal fusion feature; Here, the dynamic weighted fusion can be expressed as the following formula (6).
[0120] (6) Among them, F BEV Represents the aligned dynamic scene feature, F voxel Represents the image voxel feature, F fused Represents the geometric semantic multi-modal fusion feature.
[0121] In specific applications, the joint optimization objective Ltotal can be fed back as the GeometricConsistency Loss to the subsequent training of the scene perception model to update the parameters of the scene perception model and improve the geometric consistency of subsequent feature extraction.
[0122] Joint optimization of the key-point alignment loss and the Intersection over Union (IoU) can achieve high-precision geometric consistency in complex scenarios, ensuring the consistency of multi-modal data in the geometric space. By feature alignment, the feature expressions of the same object in different modalities can be made as similar as possible, and by position alignment, the projected coordinates of the key points can be accurately corresponding in the target space. Combining with the dynamic weight mechanism can also effectively improve the model robustness in complex scenarios.
[0123] Through the fusion under geometric consistency constraints, problems such as spatial misalignment between BEV features and voxel features and high missed detection rates of small targets can be solved, reducing the key-point alignment error by about 31.3% and increasing the IoU of small targets by about 2.3%.
[0124] Further, in step 404, the multi-scale pyramid pooling can perform multi-scale pooling operations, concatenation operations, and 1×1 convolution fusion on the geometric alignment features and cross-modal interaction features in sequence by adopting an adaptive pooling strategy, so as to obtain multi-modal fusion features.
[0125] The adaptive pooling strategy can dynamically adjust the stride or padding strategy of the pooling window according to the resolution of the input features (i.e., geometric alignment features or cross-modal interaction features, etc.) to ensure that the output feature sizes are aligned. Specifically, the adaptive pooling strategy can include: adopting a first predetermined pooling size (e.g., 1×1) in high-resolution regions such as dense target regions to retain details, and adopting a second predetermined pooling size (e.g., 4×4) in low-resolution regions such as empty regions to reduce the computational amount, where the first predetermined pooling size is smaller than the second predetermined pooling size. Among them, the high-resolution region and the low-resolution region can be dynamically distinguished by presetting a resolution threshold. Thus, fine-grained pooling is adopted in the high-resolution region and coarse-grained pooling is adopted in the low-resolution region, which can increase the Average Precision (AP) of small dynamic targets such as pedestrians by 3.5%, while reducing the redundant calculation in the empty region and reducing the consumption of computing resources.
[0126] The multi-scale pooling operations can be performed on the geometric alignment features and cross-modal interaction features in parallel. Specifically, the multi-scale pooling operations of the multi-scale pyramid pooling can be implemented through multi-level pooling, and the pooling sizes of the multi-level pooling can increase exponentially with the level. For example, the multi-level pooling of the multi-scale pyramid pooling can be 3-level pooling. The pooling size of the first level is 1×1, which is equivalent to global average pooling or global maximum pooling and can extract global statistical features; the pooling size of the second level is 2×2, which can capture medium-scale regional features; the pooling size of the third level is 4×4, which can retain finer-grained local details.
[0127] The Concatenate operation includes concatenating the pooling results of different scales along the channel dimension to form multi-scale multi-modal fusion features. Thus, the independent information of each scale can be retained to avoid feature coverage or loss.
[0128] Use a 1×1 convolution to reduce the dimension of the concatenated multi-scale multi-modal fusion features, compress redundant channels, and enhance the interaction of cross-scale features, thereby obtaining multi-modal fusion features.
[0129] Through multi-scale pyramid pooling, problems such as insufficient multi-scale feature expression ability and low classification accuracy in complex scenes can be solved. For example, the AP of small targets such as pedestrians can be increased by 3.5%, and the memory occupancy can be reduced by 15%.
[0130] Further, referring to Figure 2 , the method of the embodiment of the present disclosure may further include: step 206, obtaining 3D object detection results, semantic occupancy prediction results, and / or motion estimation results of the environment around the object based on the multi-modal fusion features.
[0131] In step 206, a 3D object detection result of the environment around the vehicle can be obtained based on the multi-modal fusion features through an end-to-end object detection Transformer (DETR) detection head.
[0132] The DETR detection head may include multiple layers of Transformer decoders, and adopt a sparse query mechanism (for example, initializing 900 learnable query vectors, and each query is associated with 3D reference points (x, y, z), etc.) to process the multi-modal fusion features to obtain 3D object detection results. The 3D object detection results may include, but are not limited to, the categories, spatial positions, and speeds of each target in the environment around the object. The targets here may include dynamic targets and static targets, and the spatial position may be represented as, but not limited to, for example, three-dimensional coordinates in an object coordinate system (such as a vehicle body coordinate system) and heading angles, pitch angles, roll angles, etc. used to describe the direction or pose of an object.
[0133] For example, the 3D object detection results may include information such as the categories, coordinates, and current speeds of multiple 3D bounding boxes. Each 3D bounding box represents a target and its confidence is higher than a preset confidence threshold.
[0134] The DETR detection head may include 6 layers of Transformer decoders, and its loss function includes classification loss and regression loss.
[0135] In step 206, a semantic occupancy prediction result can be obtained based on the multi-modal fusion features through a multi-layer perceptron (MLP) decoder. The semantic occupancy prediction result is represented as a semantic occupancy probability map of the environment around the object.
[0136] The MLP decoder may include a plurality of consecutive decoding units, each decoding unit including a plurality of consecutive 3D convolutional layers, normalization layers, and activation function layers, and the number of output data channels of the plurality of consecutive 3D convolutional layers decreases layer by layer. For example, the MLP decoder may be configured as: 3×[Conv3D(3×3×3) + BN + ReLU], that is, the decoding network includes 3 consecutive decoding units, each decoding unit including 3 consecutive 3D convolutional layers with a convolutional kernel size of 3×3×3, a batch normalization layer (BN), and a ReLU activation function layer, and the number of output data channels of these 3 3D convolutional layers is 128, 64, and 11 in sequence.
[0137] Semantic Occupancy Prediction is used to indicate the semantic category and occupancy status of each voxel in the three-dimensional space. The occupancy status indicates whether the voxel is occupied by an object. Compared with 3D bounding boxes, semantic occupancy prediction can provide a finer-grained understanding of the environment and can identify objects of arbitrary shapes such as vegetation and building debris, as well as the continuous changes in dynamic scenes.
[0138] In some examples of the present disclosure, the semantic occupancy prediction result can be represented as a semantic occupancy probability map, and this semantic occupancy probability map supports multi-class semantic prediction. The semantic occupancy probability map characterizes the probability distribution of the semantic categories of each voxel. It has been verified that under a specific dataset, the semantic occupancy probability map of the above MLP decoder supports 11-class semantic prediction, and its overall mIoU reaches 23.2%.
[0139] In step 206, a motion estimation result can be obtained using a velocity regression model based on the 3D object detection result and the multi-modal fusion feature. The motion estimation result may include, but is not limited to, the velocity of dynamic objects in the surrounding environment of the object within a predetermined future time duration.
[0140] Specifically, the 3D object detection result and the multi-modal fusion feature can be input into the velocity regression model and processed. The velocity regression model outputs a two-dimensional matrix, and this two-dimensional matrix contains two-dimensional components (v x , v y ) of each dynamic object, and each two-dimensional component characterizes the velocity component of a dynamic object in the horizontal direction.
[0141] Exemplarily, the velocity regression module can be implemented as, but not limited to, a two-layer fully-connected neural network. Each fully-connected neural network includes an input layer, a hidden layer, and an output layer. The structure is 256→64→2. The input layer is used to receive the 3D object detection prediction results and multi-modal fusion features, and the number of input data channels can be 256. The hidden layer can use the ReLU activation function or other non-linear activation functions such as Sigmoid. The number of output data channels of the hidden layer is 64. The output layer is used to output the velocity components, and the number of output data channels is 2.
[0142] In some examples, the velocity regression model can be trained based on the 4D radar data sequence collected by the 4D imaging radar group. Specifically, the velocity regression model can be trained in the following way: Filter the noise data in the 4D radar data sequence through a preset signal-to-noise ratio threshold (for example, SNR≥10dB), and use the velocity in the filtered 4D radar data sequence as the supervision signal to train the velocity regression model. The supervision signal of the velocity regression model comes from the Doppler velocity measurement value of the 4D radar, which can effectively suppress the influence of noise such as rain and fog reflections and sidelobe interference on velocity estimation, and further improve the accuracy and reliability of velocity estimation in complex weather.
[0143] In some examples, step 206 can be: Obtain the 3D object detection result, semantic occupancy prediction result, and motion estimation result based on the multi-modal fusion features through a multi-task processing layer including a DETR detection head, an MLP decoder, and a velocity regression model. Thus, through sharing encoder features, independent task branch design, and dynamic loss weighting, the collaborative optimization of 3D object detection, three-dimensional semantic understanding, and motion prediction can be achieved.
[0144] Figure 6 Shows the structural schematic diagram of the multi-modal based scene perception device provided by the embodiments of the present disclosure. Refer to Figure 6 , the multi-modal based scene perception device 600 of the embodiments of the present disclosure may include: A data acquisition unit 601, configured to obtain the current multi-view image sequence and the current 4D radar data sequence from the multi-view camera array and the 4D imaging radar group mounted on the object; An image feature extraction unit 602, configured to obtain the current image features by using the current multi-view image sequence; A radar feature extraction unit 603, configured to obtain the current radar features by using the current 4D radar data sequence, where the radar features include radar voxel features and radar BEV features; A temporal fusion unit 604, configured to model the spatio-temporal evolution of the dynamic scene and the static scene in the BEV space and the voxel space based on the previously cached historical radar features and the current radar features to obtain the current dynamic scene features and the current static scene features; The cross-modal interaction fusion unit 605 is used to perform cross-modal interaction fusion on the current image features, current dynamic scene features, and current static scene features to obtain multi-modal fusion features, which are used to perform 3D object detection, semantic occupancy prediction, and / or motion state estimation on the surrounding environment of the object.
[0145] Furthermore, the image feature extraction unit 602 can specifically be used to: obtain a multi-view feature map sequence based on the current multi-view image sequence using an image encoder, where the multi-view feature map sequence includes the feature maps of each view image in the current multi-view image sequence, and the image encoder includes a lightweight convolutional neural network that removes the last two convolutional layers and inserts an SE module in the penultimate layer; and, obtain serialized pyramid features based on the multi-view feature map sequence using FPN, and the serialized pyramid features are the current image features, and the upsampling of FPN uses the CARAFE operator.
[0146] Furthermore, the radar feature extraction unit 603 can specifically be used to: perform dynamic voxel compression on the current 4D radar data sequence to obtain the current radar voxel features; and, obtain the current radar BEV features based on the current radar voxel features through a radar encoder, where the radar encoder includes a plurality of consecutive feature extraction modules, each feature extraction module includes a sparse convolutional layer, and the convolutional kernel sizes of the sparse convolutional layers in the plurality of consecutive feature extraction modules are the same and the number of channels increases step by step.
[0147] Furthermore, the temporal fusion unit 604 can specifically be used to: Read the previously cached historical radar features from the buffer, where the historical radar features include historical radar voxel features and historical radar BEV features; Use a multi-scale dilated temporal fusion network to obtain dynamic scene features and static scene features based on the historical radar features and the current radar features, where the multi-scale dilated temporal fusion network includes a dilated temporal convolutional network, and the dilated temporal convolutional network includes a plurality of consecutive dilated convolutional layers, and the dilation rates of the plurality of consecutive dilated convolutional layers increase exponentially by level and the convolutional kernel sizes are the same; Perform dynamic target compensation on the dynamic scene features using the method of multi-target Kalman filter tracking.
[0148] Furthermore, the cross-modal interaction fusion unit 605 can specifically be used to: Fuse the current image features and the current static scene features to obtain image voxel features; Use a gated cross-attention module to dynamically fuse the dynamic scene features and the image voxel features to obtain cross-modal interaction features; Fuse the dynamic scene features and the image voxel features under the geometric consistency constraint to obtain geometrically aligned features; Integrate geometric alignment features and cross-modal interaction features through multi-scale pyramid pooling to obtain multi-modal fusion features.
[0149] Further, the multi-modal scene perception device 600 may further include: a multi-task unit 606, configured to obtain 3D object detection results, semantic occupancy prediction results, and / or motion estimation results of the environment around the object based on the multi-modal fusion features.
[0150] In a specific application, the multi-modal scene perception device 600 may be implemented by software, hardware, or a combination of both. Exemplarily, the multi-modal scene perception device 600 may be implemented as software running in the following electronic device 700.
[0151] In addition, an embodiment of the present disclosure further provides a computer-readable storage medium, on which a computer program is stored, the program includes instructions, and when the instructions are executed by one or more processors of the computer, the steps of the foregoing multi-modal scene perception method are executed.
[0152] Figure 7 The structural schematic diagram of the electronic device provided by the embodiment of the present disclosure is shown. Refer to Figure 7 , the electronic device 700 may include: one or more processors 701, and further includes a memory 702 storing one or more programs, which are executed by the one or more processors 701 to implement the method flow shown in the foregoing embodiments of the present disclosure and / or the program units corresponding to the respective units in the device.
[0153] Each component is interconnected using different buses and may be installed on a common motherboard or otherwise installed as needed. The processor 701 may process instructions executed within the electronic device, including instructions stored in the memory or on the memory to display graphical information of a user interface on an external input / output device (such as a display device coupled to the interface). In other embodiments, if necessary, multiple processors and / or multiple buses may be used together with multiple memories and multiple memories.
[0154] The processor 701 may include one or more single-core processors or multi-core processors. The processor 701 may include any combination of general-purpose processors or dedicated processors (such as image processors, application processors, baseband processors, etc.).
[0155] The memory 702 is the computer-readable storage medium provided by the present disclosure, and may be used to store non-transitory software programs, non-transitory computer-executable programs, and units, such as those in the embodiments of the present disclosure, such as Figure 2The program instructions / units corresponding to the multi-modal based scene perception method shown. The processor 701 executes, by running the non-transitory software programs, instructions, and units stored in the memory 702, such as those in the above method embodiments as Figure 2 The programs, instructions, and units corresponding to the multi-modal based scene perception method shown.
[0156] The electronic device 700 may further include: an input device 703 and an output device 704. The processor 701, the memory 702, the input device 703, and the output device 704 may be connected through a bus or other means, Figure 7 Taking the connection through the bus as an example.
[0157] The above program (also referred to as software, software application, or code) includes machine instructions for a programmable processor, and these computing programs may be implemented using an object-oriented programming language, assembly, or machine language.
[0158] With the development of time and technology, the meaning of the medium has become more and more extensive. The dissemination path of computer programs is no longer limited to tangible media, and can also be directly downloaded from the network, etc. Any combination of one or more computer-readable storage media may be adopted. The computer-readable storage medium may be, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this document, the computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0159] In specific applications, the electronic device 700 may be implemented as, but is not limited to, a domain controller, a vehicle-mounted device, or other similar devices.
[0160] The embodiments of the present disclosure further provide a vehicle, on which a multi-view camera array and a 4D imaging radar group are loaded. The vehicle includes the aforementioned multi-modal based scene perception device 600, the electronic device 700, and / or a computer-readable storage medium.
[0161] The above has introduced the technical solutions provided by the present disclosure in detail. Specific examples are used herein to elaborate on the principles and implementation manners of the present disclosure. The description of the above embodiments is only used to help understand the method and its core idea of the present disclosure; at the same time, for those of ordinary skill in the art, according to the idea of the present disclosure, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present disclosure.
[0162] The foregoing is only a preferred embodiment of the present disclosure and is not intended to limit the present disclosure. Any modifications, equivalent substitutions, etc. made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A multi-modal based scene perception method, characterized in that, The method includes: Obtaining a current multi-view image sequence and a current 4D radar data sequence from a multi-view camera array and a 4D imaging radar group mounted on an object; Obtaining current image features by using the current multi-view image sequence; Obtaining current radar features by using the current 4D radar data sequence, where the radar features include radar voxel features and radar BEV features; Based on previously cached historical radar features and current radar features, modeling the spatio-temporal evolution of dynamic and static scenes in BEV space and voxel space to obtain current dynamic scene features and current static scene features; Performing cross-modal interaction fusion on the current image features, current dynamic scene features, and current static scene features to obtain multi-modal fusion features, where the multi-modal fusion features are used to perform 3D object detection, semantic occupancy prediction, and / or motion state estimation on the surrounding environment of the object.
2. The method according to claim 1, characterized in that, The obtaining current image features by using the current multi-view image sequence includes: Using an image encoder to obtain a multi-view feature map sequence based on the current multi-view image sequence, where the multi-view feature map sequence contains feature maps of each view image in the current multi-view image sequence, and the image encoder includes a lightweight convolutional neural network that removes the last two convolutional layers and inserts a channel attention mechanism SE module in the penultimate layer; Using a Feature Pyramid Network (FPN) to obtain serialized pyramid features based on the multi-view feature map sequence, where the serialized pyramid features are the current image features, and the upsampling of the FPN uses a Content-Aware ReAssembly (CARAFE) operator.
3. The method according to claim 1, wherein The obtaining current radar features by using the current 4D radar data sequence includes: Performing dynamic voxel compression on the current 4D radar data sequence to obtain current radar voxel features; Obtaining current radar BEV features by using a radar encoder based on the current radar voxel features, where the radar encoder includes a plurality of consecutive feature extraction modules, each of the feature extraction modules includes a sparse convolutional layer, and the convolutional kernel sizes of the sparse convolutional layers in the plurality of consecutive feature extraction modules are the same and the number of channels increases gradually.
4. The method according to claim 1, wherein The modeling the spatio-temporal evolution of dynamic and static scenes in BEV space and voxel space based on previously cached historical radar features and current radar features to obtain current dynamic scene features and current static scene features includes: Reading previously cached historical radar features from a buffer, where the historical radar features include historical radar voxel features and historical radar BEV features; Using a multi-scale dilated temporal fusion network to obtain dynamic scene features and static scene features based on the historical radar features and the current radar features, where the multi-scale dilated temporal fusion network includes a dilated temporal convolutional network, and the dilated temporal convolutional network includes a plurality of consecutive dilated convolutional layers, and the dilation rates of the plurality of consecutive dilated convolutional layers increase exponentially by level and the convolutional kernel sizes are the same; Performing dynamic target compensation on the dynamic scene features in a manner of multi-target Kalman filter tracking.
5. The method according to claim 1, wherein Performing cross-modal interaction fusion on the current image features, current dynamic scene features, and current static scene features to obtain multi-modal fusion features includes: Fusing the current image features and the current static scene features to obtain image voxel features; Using a gated cross-attention module to dynamically fuse the dynamic scene features and the image voxel features to obtain cross-modal interaction features; Fusing the dynamic scene features and the image voxel features under geometric consistency constraints to obtain geometric alignment features; Integrating the geometric alignment features and the cross-modal interaction features through multi-scale pyramid pooling to obtain the multi-modal fusion features.
6. The method according to claim 5, wherein The step of using a gated cross-attention module to dynamically fuse the dynamic scene features and the image voxel features to obtain cross-modal interaction features includes: Calculating a gated matrix of the image voxel features and the dynamic scene features, where the gated matrix represents the fusion weight of the dynamic scene features; Dynamically adjusting the gated matrix according to the current weather conditions; Fusing the image voxel features and the dynamic scene features based on the dynamically adjusted gated matrix to obtain the cross-modal interaction features.
7. The method according to claim 1, wherein The method further includes: obtaining 3D object detection results, semantic occupancy prediction results, and / or motion estimation results of the environment around the object based on the multi-modal fusion features.
8. A multimodal-based scene perception device, characterized in that, Including: A data acquisition unit for obtaining a current multi-view image sequence and a current 4D radar data sequence from a multi-view camera array and a 4D imaging radar group mounted on the object; An image feature extraction unit for obtaining current image features using the current multi-view image sequence; A radar feature extraction unit for obtaining current radar features using the current 4D radar data sequence, where the radar features include radar voxel features and radar BEV features; A temporal fusion unit for modeling the spatio-temporal evolution of the dynamic scene and the static scene in the BEV space and the voxel space based on the previously cached historical radar features and the current radar features to obtain the current dynamic scene features and the current static scene features; A cross-modal interaction fusion unit for performing cross-modal interaction fusion on the current image features, current dynamic scene features, and current static scene features to obtain multi-modal fusion features, where the multi-modal fusion features are used to perform 3D object detection, semantic occupancy prediction, and / or motion state estimation of the environment around the object.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; Wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Multi-mode-based automatic driving perception method and device, equipment and medium
CN115879060A
Cited By
Road traffic potential safety hazard detection method and system based on 4D millimeter wave radar
CN120808289A
A method and system for detecting road traffic safety hazards based on 4D millimeter-wave radar
CN120808289B
Dynamic adaptive BEV perception multi-scale feature fusion method
CN120997790A
Semantic environment perception method and system based on multi-modal dynamic fusion and storage medium
CN121121746A
Multi-view 3D perception method based on space-time modeling and context enhancement
CN121354061A