A multispectral camera and radar feature level data fusion method and system

By using a multispectral camera and radar feature-level data fusion method, the problem of poor data fusion effect of holographic intersection sensors was solved, achieving more efficient data fusion and holographic intersection management, and improving detection capabilities under adverse weather conditions.

CN116883802BActive Publication Date: 2026-01-02AIPARK TECHNOLOGY CO LTD +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310912672.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-25
Publication Date
2026-01-02
Estimated Expiration
2043-07-25

AI Technical Summary

Technical Problem

Existing holographic intersection sensor data fusion methods suffer from poor fusion results and high implementation difficulty, resulting in low accuracy and efficiency of holographic intersection management. In particular, visible light cameras are greatly affected by lighting conditions, and radar detection data is ineffective under adverse weather conditions.

Method used

A feature-level data fusion method combining multispectral camera and radar is adopted. Through joint calibration, feature extraction, bird's-eye view perspective transformation, multimodal feature aggregation and filtering algorithms, feature-level fusion of multispectral camera and radar point cloud data is achieved. Information fusion is carried out by radar-assisted camera perspective transformation and improved cross-attention mechanism.

Benefits of technology

It improves the accuracy and reliability of data fusion, and can provide better assistance in scenarios with poor visibility such as nighttime, rain, snow, and fog, thereby improving the management efficiency and accuracy of holographic intersections and avoiding the failure of traditional camera detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116883802B_ABST
    Figure CN116883802B_ABST
Patent Text Reader

Abstract

The application discloses a multispectral camera and radar feature level data fusion method and system, and relates to the field of intelligent traffic management, which comprises the following steps: feature extraction is performed on multispectral camera images and radar point cloud data; a radar and camera feature level fusion network and a multi-modal feature aggregation are used to perform feature level fusion on the multispectral camera data and the radar point cloud data; and an improved cross attention mechanism is used to perform fusion on camera information and radar multi-modal feature information instead of feature multi-channel series summation, so that the data fusion accuracy is further improved; meanwhile, since the application provides the fusion of multispectral camera and radar feature level data, the multispectral camera can be used to replace a traditional monocular color camera, the radar can be better assisted in scenes with poor visibility conditions such as night, rain, snow and fog, and some detection failure phenomena of the traditional camera in poor visibility conditions can be avoided, so that the management efficiency and accuracy of the holographic intersection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of intelligent traffic management, and particularly relates to a multispectral camera and radar feature level data fusion method and system. BACKGROUND

[0002] The holographic intersection is a modern traffic intersection management system based on sensor and data fusion technology, which can realize real-time sensing of traffic flow, vehicle behavior and pedestrian dynamics, realize on-demand scheduling and intelligent management of road right through intelligent signal control and optimization algorithm, and improve the traffic efficiency of the intersection. In each module of the holographic intersection system, sensing is an important and key step. The configuration scheme of the sensor directly affects the subsequent data processing and algorithm process. A single sensor cannot meet the accurate sensing needs of complex scenes, and often needs multiple sensors for sensing and fusion to improve the sensing ability of the holographic intersection system. Compared with traditional monocular color cameras, multispectral cameras can introduce rich spectral information, capture spectral data in multiple wavebands, and provide more comprehensive visual features and scene understanding. Multispectral cameras usually capture images in the visible spectrum range and some selective infrared and near-infrared wavebands, with high resolution and sensitivity. It can capture the details and texture information of the target and distinguish different target categories through spectral information. This makes multispectral cameras have advantages in target detection, recognition and classification. Radar sensors can sense objects in the surrounding environment by emitting and receiving electromagnetic waves, provide accurate position and motion information by measuring the distance, speed and direction between the target and the vehicle. Compared with optical sensors, radar has better robustness and reliability in bad weather conditions such as rain, snow and fog, and can realize long-distance detection.

[0003] However, limited by factors such as hardware cost and data fusion algorithm applicability and complexity, the current holographic intersection sensor fusion scheme is mainly an integrated machine scheme of visible light camera and millimeter wave radar fusion. Visible light camera and millimeter wave fusion is a technology that combines image collected by a visible light camera and millimeter wave radar data for target detection and tracking. However, the fusion based on visible light camera and millimeter wave still has the following problems and difficulties: 1) Due to the different measurement principles and data characteristics of visible light camera and millimeter wave radar, the calibration and registration of radar are difficult, and accurate calibration is required to align the time and space of the two kinds of data; 2) The different forms of radar and video data also make it impossible to perform information fusion in the feature dimension; 3) During the radar and video fusion process, the camera has poor depth estimation accuracy, and the advantages of radar are not well utilized for assistance. In addition, due to the great influence of traditional cameras on light, they are prone to failure in poor visibility conditions such as night, rain, snow and fog, resulting in only radar detection data without data fusion, causing the loss of some dimensional information of the target. SUMMARY

[0004] To solve the above technical problems, the present application provides a multispectral camera and radar feature level data fusion method and system, which can solve the problems of poor fusion effect, high implementation difficulty, low accuracy and efficiency of holographic intersection management based on the existing holographic intersection sensor data fusion method.

[0005] To achieve the above purpose, in one aspect, the present application provides a multispectral camera and radar feature level data fusion method, which comprises:

[0006] The matching index matches the trajectory information of the queuing stationary target and the trajectory of the starting moving target. The image collected by the multispectral camera and the radar point cloud data collected by the radar sensor are aligned through joint calibration;

[0007] The image and the radar point cloud data after joint calibration are subjected to feature extraction to obtain image feature encoding and depth estimation value corresponding to the image, and distribution data of the radar point cloud data in the observation space;

[0008] The image feature encoding, depth estimation value and distribution data of the radar point cloud data in the observation space are subjected to bird's eye view perspective conversion through a preset radar-assisted camera perspective conversion algorithm to obtain bird's eye view feature data;

[0009] The bird's eye view feature data is subjected to multi-modal feature aggregation through a preset multi-modal feature aggregation algorithm.

[0010] filtering the multi-modal feature aggregation data and outputting filtered data.

[0011] Further, the step of performing feature extraction on the jointly calibrated image and radar point cloud data to obtain image feature encodings and depth estimates corresponding to the image and distribution data of the radar point cloud in the observation space comprises:

[0012] extracting different waveband feature images collected by the multispectral camera at different preset scales through a residual network with a feature pyramid;

[0013] extracting image context feature encodings and depth estimates from the different waveband feature images through a preset LSS algorithm and an additional convolutional layer:

[0014] converting the radar point cloud data into quantized data indicators;

[0015] According to the corresponding quantized data indicators of the radar point cloud data, the statistical features of the radar point cloud data in the observation space are obtained, and the relationship between each point and its neighborhood points in the radar point cloud is obtained through local geometric features.

[0016] Further, the step of performing bird's eye view perspective conversion on the image feature encodings, depth estimates, and distribution data of the radar point cloud in the observation space through a preset radar-assisted camera perspective transformation algorithm to obtain bird's eye view feature data comprises:

[0017] projecting each target point in the radar point cloud data into a multispectral camera image;

[0018] voxelizing the projection points of each target point in the multispectral camera image into image frustum voxels, and extracting a radar context feature map and a radar occupancy map in the frustum view:

[0019] Converting the image context feature map to a frustum view according to the image context feature encodings, depth estimates, and the radar occupancy map;

[0020] Convert the context feature maps of the multispectral camera and the radar in the frustum view to the bird's eye view space.

[0021] Further, the step of performing multi-modal feature aggregation on the bird's eye view feature data through a preset multi-modal feature aggregation algorithm comprises:

[0022] According to the image context feature encoding sequence and the radar point cloud feature sequence collected by the multispectral camera, an intersection attention weight matrix and a value sequence between the multispectral camera and the radar sensor are obtained;

[0023] The multi-modal feature aggregation is performed on the context feature map of the multi-spectral camera and radar in the view frustum view to the feature data of the bird's eye view according to the cross attention weight matrix and the value sequence.

[0024] Further, the step of filtering the multi-modal feature aggregation data and outputting filtered data comprises:

[0025] The multi-modal feature aggregation data is filtered by a preset Kalman filtering algorithm.

[0026] The filtered multi-modal feature aggregation data is outputted, and track management is performed according to the filtered multi-modal feature aggregation data.

[0027] In another aspect, the present application provides a multi-spectral camera and radar feature level data fusion system, which comprises: a calibration unit for data alignment of images collected by a multi-spectral camera and radar point cloud data collected by a radar sensor through joint calibration;

[0028] A obtaining unit is configured to extract features from the joint calibrated images and radar point cloud data, obtain image feature encodings and depth estimation values corresponding to the images, and distribution data of the radar point cloud data in an observation space.

[0029] A conversion unit is configured to perform bird's eye view perspective conversion on the image feature encodings, depth estimation values, and distribution data of the radar point cloud data in the observation space by a preset radar-assisted camera view transformation algorithm to obtain bird's eye view feature data.

[0030] An aggregation unit is configured to perform multi-modal feature aggregation on the bird's eye view feature data by a preset multi-modal feature aggregation algorithm.

[0031] A filtering unit is configured to filter the multi-modal feature aggregation data and output filtered data.

[0032] Further, the obtaining unit is specifically configured to extract different waveband feature images collected by the multi-spectral camera at different preset scales by a residual network with a feature pyramid; extract image context feature encodings and depth estimation values from the different waveband feature images by a preset LSS algorithm and an additional convolution layer; convert the radar point cloud data into quantized data indicators; obtain statistical features of the radar point cloud data in the observation space according to the quantized data indicators corresponding to the radar point cloud data, and obtain the relationship between each point in the radar point cloud and its neighborhood points by local geometric features.

[0033] Further, the conversion unit is specifically used to project each target point in the radar point cloud data onto the multispectral camera image; to voxelize the projection points of each target point in the multispectral camera image into image frustum voxels, and to extract the radar context feature map and radar occupancy map in the frustum view; to convert the image context feature map into a frustum view based on the image context feature encoding, depth estimation value, and radar occupancy map; and to convert the context feature maps of the multispectral camera and radar in the frustum view into the bird's-eye view space.

[0034] Furthermore, the aggregation unit is specifically used to obtain the cross-attention weight matrix and value sequence between the multispectral camera and the radar sensor based on the image context feature encoding sequence and radar point cloud feature sequence acquired by the multispectral camera in different bands; and to perform multimodal feature aggregation on the feature data of the multispectral camera and radar in the view frustum view, which are converted into a bird's-eye view, based on the cross-attention weight matrix and value sequence.

[0035] Furthermore, the filtering unit is specifically used to filter the multimodal feature aggregation data using a preset Kalman filtering algorithm; output the filtered multimodal feature aggregation data; and perform track management based on the filtered multimodal feature aggregation data.

[0036] This invention provides a method and system for fusing multispectral camera and radar feature-level data. It extracts features from images acquired by a multispectral camera and radar point cloud data, and uses a radar and camera feature-level fusion network and multimodal feature aggregation to fuse the multispectral camera data and radar point cloud data at the feature level. Compared to existing post-target fusion methods, this approach offers better reliability. Furthermore, it employs an improved cross-attention mechanism to fuse camera information and radar multimodal feature information, rather than simply summing multiple feature channels, further improving data fusion accuracy. Simultaneously, because this invention achieves the fusion of multispectral camera and radar feature-level data, a multispectral camera can replace a traditional monocular color camera. This provides better radar assistance in low-visibility conditions such as nighttime, rain, snow, and fog, avoiding detection failures common with traditional cameras in poor visibility conditions, thereby improving the management efficiency and accuracy of holographic intersections. Attached Figure Description

[0037] Figure 1 This is a flowchart of a method for fusing multispectral camera and radar feature-level data provided by the present invention;

[0038] Figure 2 This is a schematic diagram of the structure of a multispectral camera and radar feature-level data fusion system provided by the present invention. Detailed Implementation

[0039] The technical solutions of the present application are described in further detail below with reference to the drawings and examples.

[0040] As shown in the drawings and examples, Figure 1 The multispectral camera and radar feature-level data fusion method provided by the embodiment of the present application comprises the following steps:

[0041] 101. Aligning the image collected by the multispectral camera and the radar point cloud data collected by the radar sensor through joint calibration.

[0042] Specifically, the multispectral camera and radar integrated machine is usually installed at a lamp pole with a height of about 5-7 meters on both sides of the intersection. 101.1, Time calibration: aligning the NTP time stamps of the two sensors. In general, the frame rates of the radar sensor and the camera sensor data collection are different. In order to reduce the matching error, the sensor data with higher frame rate can be down-sampled or the sensor data with lower frame rate can be interpolated. 101.2, Spatial calibration and registration: the point cloud data detected by the radar is generally in polar coordinate system, while the commonly used coordinate system of the camera data is pixel coordinate system and world coordinate system. Therefore, the coordinate systems of the radar and the camera and their relative position and attitude relationship need to be determined to obtain the coordinate transformation matrix between them. Secondly, the spatial registration of the point cloud and the image is carried out, including projecting the radar point cloud onto the camera image plane, and using feature matching algorithm or optimization method to obtain the corresponding relationship between the point cloud and the image. The feature points, edges, colors and other information of the point cloud and the image can be used for matching to achieve accurate registration.

[0043] 102. Extracting features from the joint calibrated image and radar point cloud data to obtain image feature encoding and depth estimation value corresponding to the image, and distribution data of the radar point cloud data in the observation space.

[0044] For the embodiment of the present application, step 102 can specifically include: extracting different waveband feature images collected by the multispectral camera at different preset scales through a residual network with a feature pyramid; extracting image context feature encoding and depth estimation value from the different waveband feature images through a preset LSS algorithm and an additional convolution layer; converting the radar point cloud data into quantized data indicators; obtaining statistical features corresponding to the radar point cloud data in the observation space according to the corresponding quantized data indicators of the radar point cloud data, and obtaining the relationship between each point in the radar point cloud and its neighborhood points through local geometric features.

[0045] Specifically, for example, 102.1, acquire different waveband images collected by a multispectral camera, and extract image features at different scales by using a residual network ResNet with a feature pyramid (FP). In the feature pyramid, the bottom layer contains extracted high-resolution but less semantic information feature maps, and the top layer contains low-resolution but more semantic information feature maps. By sampling the feature maps, the intersection target can be better detected and segmented at different scales. Then, the LSS algorithm is used and an additional convolutional layer is used to extract the image context features and depth distribution of the pixels in the perspective view of the multispectral camera, as shown in the following formula: , wherein, and depth distribution (u, v) represents the coordinates of the image plane. 102.2, represent the millimeter wave radar point cloud data as spatial coordinates, reflection intensity, Doppler frequency and other information, denoise, filter and other operations are performed on the original millimeter wave radar point cloud data to reduce noise and improve data quality. Calculate the statistical characteristics of the point cloud distribution histogram, distribution density, maximum reflection intensity, and average reflection intensity. In addition, local geometric features can also be used to describe the relationship between each point and its neighborhood points.

[0046] 103, perform bird's eye view perspective conversion on the image feature encoding, depth estimation value, and radar point cloud data distribution data in the observation space by using a preset radar-assisted camera view transformation algorithm to obtain bird's eye view feature data.

[0047] For the embodiment of the present application, step 103 can specifically include: projecting each target point in the radar point cloud data into a multispectral camera picture; voxelizing the projection point of each target point in the multispectral camera picture into an image view cone voxel, and extracting a radar context feature map and a radar occupancy map in the view cone view; converting the image context feature map to a view cone view according to the image context feature encoding, the depth estimation value, and the radar occupancy map; and converting the context feature maps of the multispectral camera and the radar in the view cone view to a bird's eye view space.

[0048] Specifically, for example, 103.1, for the embodiment of the present application, unlike the LSS algorithm which directly converts the image features to a bird's eye view (BEV) space when estimating the depth distribution, the embodiment of the present application uses a radar-assisted camera view transformation (RVT) to perform view conversion: first, project the radar points onto the multispectral camera picture to find the corresponding image pixels while maintaining their depths, and then voxelize them into image view cone voxels Since radar does not contain altitude information, columnar encoding is used to represent radar features. PointNet and sparse convolution are used to encode non-empty radar columns into features. Extracting radar context features from the view frustum view and radar occupancy map As shown in the formula below: . 103.2 Utilizing Image Context Features The radar occupancy map obtained from the previous sub-step Image context feature map Convert to frustum view As shown in the formula below: Since radar lacks a height dimension, to save memory, image context features are compressed by summing along the height axis. 103.3. Context feature maps of the multispectral camera and radar within the frustum view. Convert to BEV space, use voxel pooling, and modify the summation within each BEV grid to take the mean, so that the measured distance in the BEV feature map is more accurate, as shown in the following formula: .

[0049] 104. Perform multimodal feature aggregation on the bird's-eye view feature data using a pre-set multimodal feature aggregation algorithm.

[0050] In this embodiment of the invention, step 104 may specifically include: obtaining the cross-attention weight matrix and value sequence between the multispectral camera and the radar sensor based on the image context feature encoding sequence and radar point cloud feature sequence acquired by the multispectral camera in different bands; and performing multimodal feature aggregation on the feature data of the multispectral camera and radar in the view frustum view converted into a bird's-eye view based on the cross-attention weight matrix and value sequence.

[0051] Specifically, for example, it should first be noted that directly concatenating or superimposing the feature data from two sensors makes it difficult to handle spatial misalignment and modal ambiguity between the two sensors. Cross-attention mechanisms can help models better understand the relationships and semantic information between different feature sequences, allowing elements from different sequences to interact. In a cross-attention mechanism, each multispectral camera sequence interacts with a radar sequence, and attention weights are calculated to determine the importance and correlation between different sequences.

[0052] Let the band of the multispectral camera be defined. i The feature sequences are X i (i=1, ..., T), the radar point cloud feature sequence is Y, and n and m are sequences X respectively. iand the length of sequence Y. Representations of the query sequence, key sequence, value sequence are Q = X i W Q , K = YW K , V = YW V , where W Q , W V and W K are the query weight matrix, value weight matrix, key weight matrix. First, compute the attention score matrix S = QKᵀ. Second, then compute the attention weight, row softmax normalization is performed on the attention score matrix S, to get the attention weight matrix A = softmax(S). Finally, compute the output of cross attention. According to the attention weight matrix and the value sequence, weighted sum is performed to get the output sequence of cross attention Z = WV T .

[0053] It should be noted that, since the computational complexity of cross attention is quadratic with respect to the length of the sequence, when the length of the sequence is long and the algorithm complexity is high, the cross attention mechanism is deformed to make the computational complexity linear with respect to the length of the sequence, and the influence of distance on the algorithm complexity is weakened by further reducing the input query number. Given the query z q and the multi-modal feature map x m , let q the index query element, p q be the normalized reference point coordinates of q , the feature is aggregated by multi-modal variable cross attention:

[0054] where, h, m, k the attention head, modality and sampling point are indexed, is the output projection matrix of the h th, is the input value projection matrix of the h th head, modality m . In order to better utilize multi-modal information, the attention weight A hmqk and the sampling offset ∆ p hmqk are applied to the multi-modal feature map respectively.

[0055] 105、Filter the multi-modal feature aggregation data and output the filtered data.

[0056] For the embodiment of the application, step 105 can specifically include filtering the multi-modal feature aggregation data by a preset Kalman filtering algorithm; and outputting the filtered multi-modal feature aggregation data and performing track management according to the filtered multi-modal feature aggregation data.

[0057] Specifically, for example, when the target fusion is completed, the matching result is still insufficient to verify the motion state of the target, and further determination of the continuity of the motion state of the target is required, and in this case, the extended Kalman filtering algorithm is selected to track the target. First, in the measurement update stage, there is, , In the prediction update stage, there is: . Where x is the position state vector of the target, ^ represents the predicted value of the current frame, and ~ represents the optimal estimated value of the current frame. z is the observed position vector. A is the state transition matrix, B is the input control matrix, P is the prediction error matrix, Q is the process noise matrix, K is the Kalman gain, R is the measurement error matrix, H is the observation matrix, and I is the unit matrix. Further, the data filtered by the Kalman filter is subjected to track management. The number of tracks is initialized, and the target in the subsequent frame data is subjected to track management, such as creating a new track when a new target appears, deleting a track when a target disappears, and the like

[0058] The multi-spectral camera and radar feature-level data fusion method provided by the embodiment of the application extracts features from the image collected by the multi-spectral camera and the radar point cloud data, uses a radar and camera feature-level fusion network and multi-modal feature aggregation to fuse the multi-spectral camera data and the radar point cloud data at the feature level, has better reliability compared to the existing post-target fusion method, and uses an improved cross-attention mechanism to fuse camera information and radar multi-modal feature information, instead of simply summing up the features in multiple channels in series, thereby further improving the data fusion accuracy. At the same time, since the multi-spectral camera and radar feature-level data fusion is realized, the multi-spectral camera can be used to replace the traditional monocular color camera, and better auxiliary effect can be provided for the radar in poor visual conditions such as night, rain, snow, and fog, thereby avoiding some detection failures of the traditional camera in poor visual conditions, and improving the management efficiency and accuracy of the holographic intersection.

[0059] To implement the method provided by the embodiment of the application, the embodiment of the application provides a multi-spectral camera and radar feature-level data fusion system, as shown in Figure 2 The system includes a calibration unit 21, an acquisition unit 22, a conversion unit 23, an aggregation unit 24, and a filtering unit 25.

[0060] The calibration unit 21 is configured to perform data alignment on the image collected by the multispectral camera and the radar point cloud data collected by the radar sensor through joint calibration.

[0061] The acquisition unit 22 is configured to perform feature extraction on the image and the radar point cloud data after joint calibration, and acquire image feature encoding and depth estimation value corresponding to the image, and distribution data of the radar point cloud data in the observation space.

[0062] The conversion unit 23 is configured to perform bird's-eye view perspective conversion on the image feature encoding, the depth estimation value, and the distribution data of the radar point cloud data in the observation space by using a preset radar-assisted camera perspective transformation algorithm, and obtain bird's-eye view feature data.

[0063] The aggregation unit 24 is configured to perform multi-modal feature aggregation on the bird's-eye view feature data by using a preset multi-modal feature aggregation algorithm.

[0064] The filtering unit 25 is configured to filter the multi-modal feature aggregation data and output filtered data.

[0065] Further, the acquisition unit 22 is specifically configured to extract different waveband feature images collected by the multispectral camera at different preset scales by using a residual network with a feature pyramid; extract image context feature encoding and depth estimation value from the different waveband feature images by using a preset LSS algorithm and an additional convolution layer; convert the radar point cloud data into quantized data indicators; acquire statistical features corresponding to the radar point cloud data in the observation space according to the quantized data indicators corresponding to the radar point cloud data, and acquire the relationship between each point in the radar point cloud and its neighborhood points by using local geometric features.

[0066] Further, the conversion unit 23 is specifically configured to project each target point in the radar point cloud data into a multispectral camera picture; voxelize the projection points of the target points in the multispectral camera picture into image frustum voxels, and extract a radar context feature map and a radar occupancy map in a frustum view; convert the image context feature map into a frustum view according to the image context feature encoding, the depth estimation value, and the radar occupancy map; and convert the context feature maps of the multispectral camera and the radar in the frustum view to a bird's-eye view space.

[0067] Further, the aggregation unit 24 is specifically configured to acquire a cross-attention weight matrix and a value sequence between the multispectral camera and the radar sensor according to a sequence of image context feature encodings of different wavebands collected by the multispectral camera and a sequence of radar point cloud features; and perform multi-modal feature aggregation on the bird's-eye view feature data converted from the context feature maps of the multispectral camera and the radar in the frustum view according to the cross-attention weight matrix and the value sequence.

[0068] Further, the filtering unit 25 is specifically configured to filter the multi-modal feature aggregation data by a preset Kalman filtering algorithm; and output filtered multi-modal feature aggregation data and perform track management according to the filtered multi-modal feature aggregation data.

[0069] The multispectral camera and radar feature-level data fusion system provided by the embodiment of the application performs feature extraction on image and radar point cloud data collected by a multispectral camera, performs feature-level fusion on multispectral camera data and radar point cloud data by using a radar and camera feature-level fusion network and multi-modal feature aggregation, has better reliability compared with an existing target post-fusion mode, and performs fusion of camera information and radar multi-modal feature information based on an improved cross-attention mechanism instead of simple feature multi-channel series summation, thereby further improving data fusion accuracy. At the same time, since the multispectral camera and radar feature-level data are fused, the multispectral camera can be used to replace a traditional monocular color camera, and better auxiliary effect can be provided for the radar in a poor visual condition such as night, rain, snow and fog, thereby avoiding some detection failure phenomena of the traditional camera in a poor visual condition, and improving the management efficiency and accuracy of the holographic intersection.

[0070] It should be understood that the particular order or hierarchy of steps in the processes disclosed is an example. Based upon design preferences, it should be understood that specific order or hierarchy of steps in the processes could be rearranged while remaining within the scope of the present disclosure. The accompanying method claims present elements of the various steps in a sample order, and as such claims are not meant to be limited to the particular order or hierarchy presented.

[0071] In the above detailed description, various features are grouped together in single embodiments for the purpose of streamlining the disclosure. Such disclosed approaches should not be interpreted as reflecting an intention that the claimed embodiments require more features than are explicitly recited in each claim. On the contrary, as indicated previously, the inventiveness lies in less than all features of the disclosed subject matter used in a single embodiment. Accordingly, the claims are hereby expressly incorporated into this detailed description, with each claim standing on its own as a separate preferred embodiment of the application.

[0072] In order for any person skilled in the art to implement or use the present application, the above discloses the disclosed embodiments. Various modifications of these embodiments are obvious to those skilled in the art, and the general principles defined herein can also be applied to other embodiments without departing from the spirit and protection scope of the present disclosure. Therefore, the present disclosure is not limited to the embodiments given herein, but is consistent with the broadest scope of the principles and novel features disclosed in the present disclosure.

[0073] The above description includes examples of one or more embodiments. Of course, not all possible combinations of components or methods described above will be employed to make or use the embodiments nor will all of the following described examples necessarily be realized. One of ordinary skill in the art, however, having the benefit of the present description, can understand how to make and use variations of the embodiments under the teachings and concepts described herein. Thus, the embodiments described herein are intended to embrace all such alterations, modifications, and variations that fall within the scope of the appended claims. Additionally, the term "comprising" as used in the specification and in the following claims is to be construed in the broadest sense as meaning "including" or "including but not limited to," and is not intended to exclude any other elements. Furthermore, the use of any step or embodiment "to" or "for" any purpose is intended to encompass the step or embodiment "to" or "for" that purpose.

[0074] Those of skill would further appreciate that the various illustrative logical blocks, modules, and steps described in connection with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans can implement the described functionality in varying ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present embodiments.

[0075] The various illustrative logical blocks, modules, and steps described in connection with the embodiments disclosed herein can be implemented or performed by a general purpose processor, a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general purpose processor can be a microprocessor, but in the alternative, the general purpose processor can be any conventional processor, controller, microcontroller, or state machine. A processor can also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.

[0076] The steps of a method or algorithm described in connection with the embodiments disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. A software module can reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor such the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium can be integral to the processor. The processor and the storage medium can reside in an ASIC. The ASIC can reside in a user terminal. In the alternative, the processor and the storage medium can reside as discrete components in a user terminal.

[0077] In one or more exemplary designs, the functions described can be implemented in hardware, software, firmware, or any combination thereof. If implemented in software, the functions can be stored on or transmitted over as one or more instructions or code on a computer-readable medium. Computer-readable media include both computer storage media and communication media including any medium that facilitates transfer of a computer program from one place to another. Storage media can be any available media that can be accessed by a general purpose or special purpose computer. By way of example, such computer-readable media can comprise RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code means in the form of instructions or data structures and that can be accessed by a general-purpose or special-purpose computer, or a general-purpose or special-purpose processor. Also, any connection is properly termed a computer-readable medium. For example, if the software is transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or data

[0078] The above detailed description describes the purpose, technical solutions and advantages of the present application. It should be understood that the above description is only a specific embodiment of the present application and is not intended to limit the scope of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the scope of the present application.

Claims

1. A multispectral camera and radar feature level data fusion method, characterized in that, The method comprises: performing data alignment on images collected by a multispectral camera and radar point cloud data collected by a radar sensor through joint calibration; performing feature extraction on the jointly calibrated images and radar point cloud data to obtain image feature encodings and depth estimation values corresponding to the images and distribution data of the radar point cloud data in an observation space; performing bird's eye view perspective conversion on the image feature encodings, depth estimation values, and distribution data of the radar point cloud data in the observation space through a preset radar-assisted camera perspective transformation algorithm to obtain bird's eye view feature data; performing multi-modal feature aggregation on the bird's eye view feature data through a preset multi-modal feature aggregation algorithm; filtering the multi-modal feature aggregation data and outputting filtered data; the step of performing feature extraction on the jointly calibrated images and radar point cloud data to obtain image feature encodings and depth estimation values corresponding to the images and distribution data of the radar point cloud data in an observation space comprises: extracting different waveband feature images collected by a multispectral camera at different preset scales through a residual network with a feature pyramid; extracting image context feature encodings and depth estimation values from the different waveband feature images through a preset LSS algorithm and an additional convolutional layer: converting the radar point cloud data into quantized data indicators; obtaining statistical features of the radar point cloud data in the observation space according to the corresponding quantized data indicators of the radar point cloud data and obtaining the relationship between each point in the radar point cloud and its neighborhood points through local geometric features; the step of performing bird's eye view perspective conversion on the image feature encodings, depth estimation values, and distribution data of the radar point cloud data in the observation space through a preset radar-assisted camera perspective transformation algorithm to obtain bird's eye view feature data comprises: projecting each target point in the radar point cloud data into a multispectral camera picture; voxelizing the projection points of each target point in the multispectral camera picture into image frustum voxels and extracting a radar context feature map and a radar occupancy map in a frustum view: converting the image context feature map to a frustum view according to the image context feature encodings, depth estimation values, and the radar occupancy map; converting the context feature maps of the multispectral camera and the radar in the frustum view to a bird's eye view space.

2. The multispectral camera and radar feature level data fusion method of claim 1, wherein, the step of performing multi-modal feature aggregation on the bird's eye view feature data through a preset multi-modal feature aggregation algorithm comprises: obtaining a cross-attention weight matrix and a value sequence between the multispectral camera and the radar sensor according to a sequence of image context feature encodings of different wavebands collected by the multispectral camera and a sequence of radar point cloud features; performing multi-modal feature aggregation on the bird's eye view feature data converted from the context feature maps of the multispectral camera and the radar in the frustum view according to the cross-attention weight matrix and the value sequence.

3. The multispectral camera and radar feature level data fusion method of claim 1, wherein, the step of filtering the multi-modal feature aggregation data and outputting filtered data comprises: filtering the multi-modal feature aggregation data through a preset Kalman filtering algorithm; outputting the filtered multi-modal feature aggregation data and performing track management according to the filtered multi-modal feature aggregation data.

4. A multispectral camera and radar feature level data fusion system, characterized by, the system comprises: The calibration unit is configured to perform data alignment on images collected by the multispectral camera and radar point cloud data collected by the radar sensor through joint calibration. The acquisition unit is configured to perform feature extraction on the jointly calibrated images and radar point cloud data, to obtain image feature encodings and depth estimation values corresponding to the images, and distribution data of the radar point cloud data in an observation space. The conversion unit is configured to perform bird's-eye view perspective conversion on the image feature encodings, depth estimation values, and distribution data of the radar point cloud data in the observation space by using a preset radar-assisted camera perspective transformation algorithm, to obtain bird's-eye view feature data. The aggregation unit is configured to perform multi-modal feature aggregation on the bird's-eye view feature data by using a preset multi-modal feature aggregation algorithm. The filtering unit is configured to filter the multi-modal feature aggregation data and output filtered data. The acquisition unit is specifically configured to extract different waveband feature images of the multispectral camera collected at different preset scales by using a residual network with a feature pyramid; extract image context feature encodings and depth estimation values from the different waveband feature images by using a preset LSS algorithm and an additional convolution layer; convert the radar point cloud data into quantized data indicators; and obtain statistical features of the radar point cloud data in the observation space according to the quantized data indicators corresponding to the radar point cloud data, and obtain relationships between each point in the radar point cloud and its neighborhood points by using local geometric features. The conversion unit is specifically configured to project each target point in the radar point cloud data into a multispectral camera picture; voxelize the projection points of the target points in the multispectral camera picture into image frustum voxels, and extract a radar context feature map and a radar occupancy map in a frustum view; convert the image context feature map into a frustum view according to the image context feature encodings, the depth estimation values, and the radar occupancy map; and convert the context feature maps of the multispectral camera and the radar in the frustum view into a bird's-eye view space.

5. The multispectral camera and radar feature level data fusion system of claim 4, wherein, The aggregation unit is specifically configured to obtain a cross-attention weight matrix and a value sequence between the multispectral camera and the radar sensor according to a sequence of image context feature encodings of different wavebands collected by the multispectral camera and a sequence of radar point cloud features; and perform multi-modal feature aggregation on the bird's-eye view feature data converted from the context feature maps of the multispectral camera and the radar in the frustum view according to the cross-attention weight matrix and the value sequence.

6. The multispectral camera and radar feature-level data fusion system according to claim 4, wherein The filtering unit is specifically configured to filter the multi-modal feature aggregation data by using a preset Kalman filtering algorithm; output the filtered multi-modal feature aggregation data; and perform track management according to the filtered multi-modal feature aggregation data.

Citation Information

Patent Citations

  • Millimeter wave radar target detection method and system based on fused image features

    CN114218999A

  • Airborne laser point cloud and multispectral image fusion method

    CN115588127A