Obstacle avoidance target recognition method and system based on fusion of infrared image and radar point cloud

CN122473768BActive Publication Date: 2026-09-29GUANGZHOU INSTITUTE OF TECHNOLOY XIDIAN UNIVERSITY +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610975689.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-07-02
Publication Date
2026-09-29
Estimated Expiration
2046-07-02

AI Technical Summary

Technical Problem

[0004]本发明通过提供一种基于红外图像与雷达点云融合的避障目标识别方法及系统,解决了现有技术中多模态融合空间对齐不准与特征交互不足的问题,实现了全天候复杂场景下的高鲁棒性避障目标识别

Benefits of technology

本发明通过点云特征提取网络,能够从稀疏的毫米波雷达点云中提取目标的几何形状和运动信息,其中局部特征聚合单元捕捉点云内部的局部结构关系,全局上下文增强单元通过Transformer捕获长程依赖,从而获得高表征能力的几何运动特征向量,弥补了纯视觉缺乏深度和速度信息的不足。基于预先联合标定的外参矩阵,对纹理语义特征向量与几何运动特征向量进行空间关联,得到关联特征对;利用外参矩阵将雷达点云投影到图像平面,并通过交并比计算和匈牙利算法进行最优匹配,实现了二维图像目标与三维雷达目标的精准空间对齐,解决了异构传感器之间的数据关联问题,确保后续融合的特征在空间上对应同一物理目标,避免误关联。将关联特征对输入基于交叉注意力机制与自注意力机制的融合网络进行特征交互与融合,得到深度融合特征,并将深度融合特征输入预测头,联合预测并输出避障目标的融合辨识结果;采用交叉注意力机制让雷达几何特征与图像语义特征进行跨模态交互,实现不同模态信息的相互补充;自注意力机制进一步增强融合特征的内部关系,使得最终深度融合特征既能保留纹理细节又能包含几何运动信息,预测头同时输出类别、置信度和三维空间参数,显著提升了复杂环境下避障目标识别的准确性和鲁棒性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122473768B_ABST
    Figure CN122473768B_ABST
Patent Text Reader

Abstract

The application discloses an obstacle avoidance target recognition method and system based on infrared image and radar point cloud fusion, relates to the technical field of automatic driving and intelligent transportation, solves the problems of inaccurate space alignment and insufficient feature interaction in the prior art, and realizes high-robustness obstacle avoidance target recognition in all-weather complex scenes. The method comprises the following steps: preprocessing near-infrared images and millimeter wave radar point cloud data collected synchronously, and obtaining two-dimensional boundary boxes of potential obstacles, texture semantic feature vectors and geometric motion feature vectors of candidate target point cloud clusters through a target detection network and a point cloud feature extraction network; based on an external parameter matrix pre-jointly calibrated, the texture semantic feature vectors and the geometric motion feature vectors are spatially associated to obtain associated feature pairs; based on the associated feature pairs, deep fusion features are obtained, and the deep fusion features are input into a prediction head to jointly predict and output fusion recognition results of obstacle avoidance targets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of autonomous driving and intelligent transportation technologies, and in particular to a method and system for obstacle avoidance target recognition based on the fusion of infrared images and radar point clouds. Background Technology

[0002] The core objective of obstacle avoidance target recognition is to continuously and accurately perceive and classify obstacles ahead in complex scenarios such as autonomous driving. Early target perception methods mainly relied on single sensors, such as visible light cameras or millimeter-wave radar. While pure vision methods can provide rich texture and semantic information, they are not robust enough to handle challenging all-weather conditions such as nighttime, rain, fog, or sudden changes in lighting. On the other hand, while pure millimeter-wave radar has all-weather detection capabilities and can directly obtain the distance and velocity of targets, its point cloud data is sparse and lacks visual texture features of targets, making it difficult to accurately identify target categories.

[0003] Multi-sensor data fusion, as an important research direction in the field of target perception, offers a new solution to the perception challenges faced by single sensors by introducing the synergistic effect of near-infrared images and millimeter-wave radar. However, in terms of cross-modal feature processing, most existing vision and radar fusion methods employ simple post-fusion at the result level or crude channel stitching at the feature level, failing to effectively address the spatial heterogeneity problem between the two-dimensional image plane and the three-dimensional radar space. Furthermore, existing networks often employ shallow attention mechanisms, failing to fully explore the deep correlation between visual texture semantics and radar geometric motion modalities, making it difficult to adapt to the adaptive dynamic extraction of key features in complex and changing scenarios, thus limiting the overall accuracy and reliability of obstacle avoidance target recognition. Summary of the Invention

[0004] This invention provides an obstacle avoidance target recognition method and system based on the fusion of infrared images and radar point clouds, which solves the problems of inaccurate spatial alignment and insufficient feature interaction in the existing technology, and realizes highly robust obstacle avoidance target recognition in all-weather complex scenarios.

[0005] In a first aspect, the present invention provides an obstacle avoidance target recognition method based on the fusion of infrared images and radar point clouds, comprising: Preprocessing is performed on the synchronously acquired near-infrared images and millimeter-wave radar point cloud data to obtain enhanced images and multiple candidate target point cloud clusters; Image features of the enhanced image are extracted using an object detection network to obtain two-dimensional bounding boxes and texture semantic feature vectors for each potential obstacle. The three-dimensional center coordinates of each candidate target point cloud cluster are obtained, and the point cloud features of each candidate target point cloud cluster are extracted through a point cloud feature extraction network to obtain the geometric motion feature vector of each candidate target point cloud cluster. The point cloud feature extraction network includes an embedding layer, a local feature aggregation unit, and a global context enhancement unit. The embedding layer maps each candidate target point cloud cluster to a high-dimensional embedding feature. The local feature aggregation unit performs convolutional local feature extraction on the high-dimensional embedding feature to obtain local aggregated features. The global context enhancement unit performs global interaction on the local aggregated features using multiple cascaded Transformer layers to obtain the geometric motion feature vector. Based on the pre-jointly calibrated extrinsic matrix, the texture semantic feature vector and the geometric motion feature vector are spatially associated using the two-dimensional bounding box to obtain associated feature pairs; The associated features are input to a fusion network based on cross-attention and self-attention mechanisms to perform feature interaction and fusion, resulting in deep fusion features. These deep fusion features are then input into a prediction head to jointly predict and output the fusion identification result of the obstacle avoidance target.

[0006] Secondly, the present invention provides an obstacle avoidance target recognition system based on the fusion of infrared images and radar point clouds, comprising: The preprocessing module is used to preprocess the synchronously acquired near-infrared images and millimeter-wave radar point cloud data to obtain enhanced images and multiple candidate target point cloud clusters. The feature extraction module is used to extract image features of the enhanced image through an object detection network to obtain two-dimensional bounding boxes and texture semantic features of each potential obstacle; and to obtain the three-dimensional center coordinates of each candidate target point cloud cluster, and extract the point cloud features of each candidate target point cloud cluster through a point cloud feature extraction network to obtain the geometric motion feature vector of each point cloud cluster; wherein, the point cloud feature extraction network includes: an embedding layer, a local feature aggregation unit, and a global context enhancement unit; the embedding layer is used to map each candidate target point cloud cluster into a high-dimensional embedding feature, the local feature aggregation unit is used to perform convolutional local feature extraction on the high-dimensional embedding feature to obtain local aggregated features, and the global context enhancement unit is used to perform global interaction of multiple cascaded Transformer layers on the local aggregated features to obtain the geometric motion feature vector; The spatial association module is used to spatially associate the texture semantic features and the geometric motion feature vectors based on the pre-jointly calibrated extrinsic parameter matrix using the two-dimensional bounding box to obtain associated feature pairs. The fusion module is used to perform feature interaction and fusion on the input fusion network based on cross-attention mechanism and self-attention mechanism to obtain deep fused features; The prediction output module is used to input the deep fusion features into the prediction head, jointly predict and output the fusion identification result of the obstacle avoidance target.

[0007] One or more technical solutions provided in this invention have at least the following technical effects or advantages: This invention utilizes a point cloud feature extraction network to extract the geometric shape and motion information of targets from sparse millimeter-wave radar point clouds. The local feature aggregation unit captures the local structural relationships within the point cloud, while the global context enhancement unit captures long-range dependencies through a Transformer, thereby obtaining highly representative geometric motion feature vectors that compensate for the lack of depth and velocity information in pure vision. Based on a pre-calibrated extrinsic parameter matrix, the texture semantic feature vector and the geometric motion feature vector are spatially correlated to obtain correlated feature pairs. The radar point cloud is projected onto the image plane using the extrinsic parameter matrix, and optimal matching is achieved through intersection-union calculation and the Hungarian algorithm, realizing precise spatial alignment between 2D image targets and 3D radar targets. This solves the data correlation problem between heterogeneous sensors, ensuring that subsequently fused features spatially correspond to the same physical target and avoiding miscorrelation. The associated features are input to a fusion network based on cross-attention and self-attention mechanisms to perform feature interaction and fusion, resulting in deep fusion features. These deep fusion features are then input into a prediction head to jointly predict and output the fused identification result of the obstacle avoidance target. The cross-attention mechanism enables cross-modal interaction between radar geometric features and image semantic features, achieving mutual complementarity of information from different modalities. The self-attention mechanism further enhances the internal relationships of the fused features, so that the final deep fusion features can retain both texture details and geometric motion information. The prediction head simultaneously outputs the category, confidence score, and three-dimensional spatial parameters, significantly improving the accuracy and robustness of obstacle avoidance target identification in complex environments. Attached Figure Description

[0008] Figure 1 A flowchart illustrating the steps of the obstacle avoidance target recognition method based on the fusion of infrared images and radar point clouds provided in an embodiment of the present invention. Figure 2 This is a structural diagram of the obstacle avoidance target recognition model based on the fusion of infrared images and radar point clouds provided in an embodiment of the present invention; Figure 3 This is a structural diagram of the point cloud feature extraction network provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of the coordinate system provided in an embodiment of the present invention; Figure 5 This is a structural diagram of the fusion network provided in an embodiment of the present invention. Detailed Implementation

[0009] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0010] In a first aspect, the present invention provides an obstacle avoidance target recognition method based on the fusion of infrared images and radar point clouds, see [link to relevant documentation]. Figure 1 The process includes the following steps S101 to S105.

[0011] S101, preprocesses the synchronously acquired near-infrared images and millimeter-wave radar point cloud data to obtain enhanced images and multiple candidate target point cloud clusters; Specifically, in step S101, the synchronously acquired near-infrared image and millimeter-wave radar point cloud data are preprocessed to obtain an enhanced image and multiple candidate target point cloud clusters, including the following steps S1011 to S1012.

[0012] S1011, Perform image enhancement preprocessing on the near-infrared image frame to obtain the enhanced image; S1012 sequentially performs moving target display filtering, constant false alarm rate detection, and density clustering on the millimeter-wave radar point cloud frames to obtain multiple candidate target point cloud clusters.

[0013] For example, near-infrared image frames and millimeter-wave radar point cloud frames are acquired synchronously triggered by hardware; image enhancement preprocessing is performed on the near-infrared image frames to obtain standardized enhanced images; moving target display filtering, constant false alarm rate detection, and density clustering are performed on the millimeter-wave radar point cloud frames to obtain multiple candidate target point cloud clusters.

[0014] S102, the image features of the enhanced image are extracted through the object detection network to obtain the two-dimensional bounding boxes and texture semantic feature vectors of each potential obstacle; Specifically, in step S102, image features of the enhanced image are extracted through the object detection network to obtain the two-dimensional bounding boxes and texture semantic feature vectors of each potential obstacle, including the following steps S1021 to S1023.

[0015] S1021, enhance the image input target detection network, and extract the shallow texture and deep semantic information of the image through the backbone feature extraction structure of the target detection network; S1022 utilizes a feature pyramid network to deeply fuse shallow texture and deep semantic information to obtain multi-scale feature maps; S1023 decodes and predicts multi-scale feature maps using a detection head to obtain two-dimensional bounding box coordinates, and extracts the corresponding texture semantic feature vector from the enhanced image based on the two-dimensional bounding box.

[0016] For example, in this embodiment, a pre-trained YOLO v11 is used as the base target detection network. The enhanced image is input into the backbone feature extraction structure of the network. The shallow texture and deep semantic information of the image are extracted layer by layer through multi-layer convolution and pooling operations. The multi-scale features are deeply fused using a neck network structure such as a feature pyramid. The detection head decodes and predicts the fused feature map to obtain the coordinates of the two-dimensional bounding box of the potential obstacle in the image.

[0017] S103, obtain the three-dimensional center coordinates of each candidate target point cloud cluster, and extract the point cloud features of each candidate target point cloud cluster through the point cloud feature extraction network to obtain the geometric motion feature vector of each candidate target point cloud cluster.

[0018] The point cloud feature extraction network includes an embedding layer, a local feature aggregation unit, and a global context enhancement unit.

[0019] (a) Embedding layer.

[0020] The embedding layer maps each candidate target point cloud cluster into a high-dimensional embedding feature; Here, the embedding layer includes: a multilayer perceptron (MLP); Each candidate target point cloud cluster is mapped to a high-dimensional embedding feature, including: (1) Input the low-dimensional physical features of the point cloud data in each candidate target point cloud cluster into the multilayer perceptron. Through the layer-by-layer nonlinear transformation of the multilayer perceptron, the low-dimensional physical features are mapped to the high-dimensional embedding space, and the high-dimensional embedding features corresponding to each candidate target point cloud cluster are output. Among them, the low-dimensional physical features include the three-dimensional spatial coordinates, radial velocity and reflection intensity of the point.

[0021] For example, such as Figure 3 As shown, the low-dimensional physical features included in the point cloud data of the candidate target point cloud cluster are input into the point cloud feature extraction network. Original input data After being mapped to a high-dimensional feature space through the embedding layer, the initial feature representation is obtained. : ; in, It is an embedded layer composed of multi-layer sensing mechanisms.

[0022] (ii) Local feature aggregation unit.

[0023] The local feature aggregation unit performs convolution on the high-dimensional embedded features to extract local aggregated features; This includes multiple parallel convolutional branches, adders, activation function layers, global average pooling layers, and global max pooling layers; Local aggregated features are obtained by performing convolutional local feature extraction on high-dimensional embedded features, including: The high-dimensional embedding features are input into multiple parallel convolutional branches, and each convolutional branch performs convolution operations independently to obtain the convolutional output features of each convolutional branch. The convolutional output features of each convolutional branch are added element-wise at the same spatial location and channel dimension using an adder to obtain the added features. The summed features are input into an activation function layer for non-linear activation to obtain the activated features. The activated features are then input into the global average pooling layer and the global max pooling layer, respectively, to obtain the average pooling features and the max pooling features. The average pooling features and max pooling features are input into the fusion layer and concatenated to output local aggregated features.

[0024] For example, initial features Enter the multi-branch convolution module. This module contains three parallel convolution branches: the first branch is a 1×1 convolution, used to extract channel information; the second branch is a 3×3 convolution followed by a 1×1 convolution, which extracts local spatial information and integrates channel features; the third branch is a 5×5 convolution followed by a 1×1 convolution, which captures features with a larger receptive field.

[0025] The outputs of the three branches are fused by element-wise addition, and the nonlinear expressive power of the network is further enhanced by the ReLU activation function to obtain a multi-scale feature representation. .

[0026] Multi-scale feature representation after ReLU activation The input branches for global average pooling and global max pooling are represented as follows: ; ; The pooled features are concatenated along the channel dimension to obtain the global aggregated features. : ; (III) Global Context Enhancement Unit.

[0027] The global context enhancement unit performs global interaction on the local aggregated features through multiple cascaded Transformer layers to obtain the geometric motion feature vector; Here, the global context enhancement unit includes: the first normalization layer, the Transformer module, and the second normalization layer; The geometric motion feature vector is obtained by performing global interactions on multiple cascaded Transformer layers on the local aggregated features, including: (1) Input the local aggregated features into the first normalization layer for normalization processing to obtain the first normalized features; (2) Map the first normalized features to query matrix, key matrix and value matrix respectively, and perform global attention interaction through the Transformer module to obtain attention interaction features; (3) Input the attention interaction features into the second normalization layer for normalization processing, and output the geometric motion feature vector.

[0028] For example, layer normalization is used to normalize the global aggregated features to obtain normalized features. Its expression is: ; Normalized features These are mapped to Query, Key, and Value vectors, respectively. Based on multiple concatenated Transformer layers, global feature interaction and enhancement are performed on locally aggregated features. A self-attention mechanism is used to calculate the correlation between features, resulting in context-aware features. The calculation formula is as follows: ; in, Represents the feature dimension of the Key. This represents the activation function.

[0029] Finally, based on the normalization layer, the context-aware features are stabilized to obtain normalized geometric motion feature vectors of the target's three-dimensional position, shape, and radial velocity.

[0030] S104, based on the pre-jointly calibrated extrinsic matrix, uses a two-dimensional bounding box to spatially associate the texture semantic feature vector and the geometric motion feature vector to obtain associated feature pairs; Specifically, in step S104, based on the pre-jointly calibrated extrinsic parameter matrix, the texture semantic feature vector and the geometric motion feature vector are spatially associated using a two-dimensional bounding box to obtain associated feature pairs, including the following steps S1041 to S1044.

[0031] S1041, the three-dimensional center coordinates of each candidate target point cloud cluster are projected from the radar coordinate system to the corresponding image plane through a pre-calibrated extrinsic parameter matrix, to obtain the two-dimensional projection point coordinates of each candidate target point cloud cluster on the image plane; S1042, For each two-dimensional projection point coordinate, construct an associated candidate box with the two-dimensional projection point coordinate as the center and a preset side length, calculate the cross-union ratio between each associated candidate box and the two-dimensional detection bounding box of each potential obstacle, and obtain the cross-union ratio matrix; S1043, based on the intersection-union matrix, uses the Hungarian algorithm to solve the optimal bipartite graph matching, with the optimization objective of maximizing the total association weight of successful matching, and matches candidate target point cloud clusters with two-dimensional bounding boxes one by one; S1044 establishes an association between each successfully matched candidate target point cloud cluster and a two-dimensional bounding box pair, and pairs the texture semantic feature vector corresponding to the two-dimensional bounding box with the geometric motion feature vector corresponding to the candidate target point cloud cluster to form an associated feature pair.

[0032] For example, such as Figure 4 As shown, the 3D center coordinates of each candidate target point cloud cluster output by the point cloud feature extraction network are projected from the radar coordinate system to the corresponding image plane based on the pre-calibrated extrinsic parameter matrix, unifying the coordinates of the radar point cloud and the camera image to obtain the corresponding 2D projection point coordinates. The projection transformation formula is as follows: ; in, The coordinates of the three-dimensional center of the candidate target point cloud cluster in the radar coordinate system. and For the physical length and width of a single pixel, The principal point coordinates of the camera. The physical focal length of the camera. and These are the rotation matrix and the translation matrix, respectively. These are the pixel coordinates of the projected image plane. The depth value of the target point in the camera coordinate system; based on all the 2D detection bounding boxes output by the target detection network, calculate the cross-union ratio (CUI) between each 2D projection point and the 2D detection bounding box of each obstacle. The calculation formula is: ; in, Represents the image detection bounding box, that is, the two-dimensional bounding box of each potential obstacle output by the object detection network; This represents a candidate bounding box with a preset side length centered at the coordinates of a two-dimensional projection point. This represents the area calculation function.

[0033] Based on the calculated intersection-union ratio matrix, the Hungarian algorithm is used to solve for the optimal bipartite graph matching, with the optimization objective being to maximize the total association weight. ; ; in, This refers to the number of point cloud clusters, i.e., the number of candidate target point cloud clusters. This refers to the number of image detection boxes, i.e., the number of two-dimensional bounding boxes. Indicates the first The associated candidate bounding boxes corresponding to the candidate target point cloud cluster and the first candidate cluster are related to the first candidate cluster. Cross-union ratio between two-dimensional bounding boxes; This represents a binary decision variable.

[0034] Indicates the first The candidate target point cloud cluster and the first Two-dimensional bounding boxes were successfully matched; an association was established for each successfully matched "image detection box-radar point cloud cluster" pair, and the texture semantic feature vector corresponding to the image detection box was paired with the geometric motion feature vector corresponding to the radar point cloud cluster to form an associated feature pair.

[0035] S105, the associated features are input to the fusion network based on cross-attention mechanism and self-attention mechanism to perform feature interaction and fusion to obtain deep fusion features, and the deep fusion features are input to the prediction head to jointly predict and output the fusion identification result of the obstacle avoidance target.

[0036] Specifically, in step S105, the associated features are interacted and fused with the input fusion network based on cross-attention mechanism and self-attention mechanism to obtain deep fusion features, including the following S10511 to S10514.

[0037] S10511, input the texture semantic feature vector and geometric motion feature vector from the associated feature pair into the feature fusion network; S10512 utilizes the cross-attention mechanism of the feature fusion network, taking the geometric motion feature vector as the query benchmark and the texture semantic feature vector as the key and value, to perform cross-modal cross-attention calculation and obtain the feature representation after cross-modal interaction; S10513, based on the self-attention mechanism of the feature fusion network, performs self-attention calculation on the feature representation after cross-modal interaction and the position encoding information to obtain the feature representation enhanced by internal relations; S10514 takes the feature representation enhanced by internal relations, inputs it into the feedforward network, performs a nonlinear transformation, and then performs layer normalization to obtain the final deep fusion feature.

[0038] For example, in this embodiment, such as Figure 5 As shown, the geometric motion feature vector is input As the query benchmark, multiplied by the first learnable weight matrix Mapping to generate query vectors ;Texture semantic feature vector As the global context, multiply by the second learnable weight matrix respectively. With the third learnable weight matrix Mapping generates key vectors AND value vector Perform cross-modal correlation calculations to obtain the cross-modal attention score matrix. The calculation formula is as follows: ; in, Represents the key vector Feature dimensions; using the Softmax activation function to evaluate the cross-modal attention score matrix. Normalization is performed according to the feature dimensions, followed by adaptive weighted fusion of features to obtain the feature representation after cross-modal interaction. The overall calculation formula is as follows: ; Feature representation after cross-modal interaction With position encoding matrix Perform element-wise addition and multiply by the fourth learnable weight matrix. Mapping generates self-attention query vectors ; Representing features after cross-modal interaction Directly multiply by the fifth learnable weight matrix. With the sixth learnable weight matrix Mapping generates self-attention key vectors With self-attention value vector Perform self-attention correlation calculations to obtain the self-attention score matrix. The calculation formula is as follows: ; in, Represents the key vector Feature dimensions; using the Softmax activation function on the self-attention score matrix Normalization is performed according to the feature dimension, followed by adaptive weighting of internal features to obtain a feature representation enhanced by internal relationships. The overall calculation formula is as follows: ; The feedforward network performs a nonlinear transformation on the feature representation enhanced by internal relations and stabilizes it through layer normalization to obtain the final deep fusion features.

[0039] Specifically, in step S105, the deep fusion features are input into the prediction head, and the fusion identification results of the obstacle avoidance target are jointly predicted and output, including the following S10521 to S10524.

[0040] S10521, The deep fusion features are fed into the prediction head network; the prediction head network includes a classification branch and a regression branch; S10522 utilizes the deep fusion features received by the classification branch, and after processing by the fully connected layer, batch normalization layer and Softmax activation function, outputs a probability distribution vector representing the target category, and extracts the value with the highest probability in the probability distribution vector as the prediction confidence of the obstacle avoidance target. S10523, the regression branch receives the deep fusion features, performs nonlinear mapping through the deep regression network, and outputs the position coordinates and geometric scale information of the obstacle avoidance target in three-dimensional physical space. S10524, combining the predicted confidence and probability distribution of the target category output from the joint classification branch, and the location coordinates and geometric scale information output from the regression branch, yields a fusion identification result containing the target category, confidence, and spatial parameters.

[0041] For example, the prediction head, as the final decoding and output module of the entire network model, maps the complex high-dimensional multimodal fusion features that have undergone deep interaction to a specific task space. This structure usually contains mutually decoupled classification and regression branches. The former is specifically used to predict the class probability distribution of the target within the perception field, while the latter focuses on predicting the coordinate offset and geometric scale of the target's spatial location. Through a series of customized convolution, feature dimensionality reduction, and nonlinear activation operations, the simultaneous and efficient decoding of the semantic attributes and spatial geometric features of the multimodal target is achieved, resulting in the final deep fusion features.

[0042] In this embodiment, the deep fusion features of texture details from near-infrared images and geometric motion information from millimeter-wave radar point clouds are fed into the prediction head network, which includes a classification branch and a regression branch. The classification branch accurately calculates and outputs the confidence score of the presence of potential obstacles of a specific category within each feature map grid or prior box. In the classification prediction branch, the decoupled features pass through a classification network composed of fully connected layers, batch normalization, and a Softmax activation function, ultimately outputting a probability distribution vector representing the target category, and extracting the value with the highest probability as the prediction confidence of the target. In the regression branch, the decoupled features undergo nonlinear mapping through a deep regression network to directly regress the continuous state parameters of the obstacle avoidance target in three-dimensional physical space, calculates and outputs the precise parameter information of the obstacle, and finally jointly outputs a comprehensive target fusion identification result with high positioning accuracy and high classification robustness.

[0043] In summary, the embodiments of the present invention first preprocess the input near-infrared image and multimodal sensor data such as millimeter-wave radar point cloud; in the feature extraction stage, the texture semantic features of the preprocessed image are extracted based on the target detection network to obtain the two-dimensional bounding box and category probability information of the potential target in the image, and the geometric motion feature vector of the preprocessed point cloud is extracted based on the point cloud feature extraction network to obtain the geometric motion feature vector of the target; in the feature interaction stage, the extracted texture semantic features and geometric motion feature vector are spatially associated based on the target association module to accurately obtain the associated feature pairs; in the feature fusion stage, the associated feature pairs are input into a fusion network based on cross-attention mechanism and self-attention mechanism for efficient interaction and fusion of features to obtain a deep fusion feature with stronger representation ability; the deep fusion feature is input into the prediction head to jointly predict and output the comprehensive fusion identification result of the obstacle avoidance target. The embodiments of the present invention can fully leverage the complementary advantages of high-resolution textures in near-infrared images and three-dimensional geometry and motion information from millimeter-wave radar by accurately spatially associating multimodal data and fusing deep interactive information based on cross-attention. This significantly improves the accuracy of obstacle recognition and the robustness of the system in complex environments and adverse weather conditions, and is suitable for all-weather multimodal cooperative obstacle avoidance and environmental perception tasks in autonomous driving or advanced driver assistance systems.

[0044] Secondly, this invention provides an obstacle avoidance target recognition system based on the fusion of infrared images and radar point clouds, the system comprising: The preprocessing module is used to preprocess the synchronously acquired near-infrared images and millimeter-wave radar point cloud data to obtain enhanced images and multiple candidate target point cloud clusters. The feature extraction module is used to extract image features of the enhanced image through the object detection network to obtain the two-dimensional bounding boxes and texture semantic features of each potential obstacle; and to obtain the three-dimensional center coordinates of each candidate target point cloud cluster, and to extract the point cloud features of each candidate target point cloud cluster through the point cloud feature extraction network to obtain the geometric motion feature vector of each point cloud cluster; wherein, the point cloud feature extraction network includes: an embedding layer, a local feature aggregation unit, and a global context enhancement unit; the embedding layer is used to map each candidate target point cloud cluster into a high-dimensional embedding feature, the local feature aggregation unit is used to perform convolution on the high-dimensional embedding feature to extract local features, and the global context enhancement unit is used to perform global interaction of multiple cascaded Transformer layers on the local aggregated feature to obtain the geometric motion feature vector; The spatial association module is used to spatially associate texture semantic features and geometric motion feature vectors based on a pre-jointly calibrated extrinsic parameter matrix to obtain associated feature pairs. The fusion module is used to perform feature interaction and fusion on the input fusion network based on cross-attention and self-attention mechanisms to obtain deep fused features; The prediction output module is used to input deep fusion features into the prediction head, jointly predict and output the fusion identification result of the obstacle avoidance target.

[0045] For example, see Figure 2 In autonomous vehicles, near-infrared cameras and millimeter-wave radars are installed, and both are synchronously acquired to collect road scene data ahead via hardware triggering. During vehicle operation, the system receives a frame of near-infrared image and a corresponding frame of millimeter-wave radar point cloud data in real time. The near-infrared image is enhanced to obtain a clear enhanced image; the radar point cloud is sequentially subjected to moving target display filtering, constant false alarm rate detection, and density clustering to obtain multiple candidate target point cloud clusters. The enhanced image is input into a pre-trained YOLO v11 target detection network, which outputs multiple two-dimensional bounding boxes and the corresponding texture semantic feature vector for each box. Simultaneously, the three-dimensional center coordinates and point cloud data of each candidate target point cloud cluster are input into a point cloud feature extraction network. The embedding layer (multilayer perceptron) of this network maps the point cloud to high-dimensional embedded features. The local feature aggregation unit obtains local aggregated features through multiple parallel convolutions, additions, activations, global pooling, and fusions. The global context enhancement unit obtains geometric motion feature vectors through layer normalization and a Transformer module. Using a pre-calibrated extrinsic matrix, the three-dimensional center coordinates of each point cloud cluster are projected onto the image plane to obtain two-dimensional projection point coordinates. A candidate bounding box is constructed centered on each projection point. The intersection-union ratio (IU) with all 2D bounding boxes is calculated to form an IU matrix. The Hungarian algorithm is used to find the optimal match. Successfully matched point cloud clusters and bounding boxes constitute associated feature pairs. The texture semantic feature vectors and geometric motion feature vectors from the associated feature pairs are input into a fusion network. First, cross-attention allows geometric features to query image features, and then self-attention enhances the internal relationships. After passing through a feedforward network and layer normalization, a deep fusion feature is output. Finally, the classification branch of the prediction head network outputs the target category (e.g., vehicle, pedestrian, cyclist) and confidence score, while the regression branch outputs the target's 3D position and size. The system sends the fused identification results to the vehicle control unit for decision-making and execution of obstacle avoidance operations, such as automatic deceleration, steering, or braking. The entire process operates in real time at a frequency above 10Hz, ensuring all-weather obstacle avoidance perception capabilities.

[0046] The various embodiments described in this specification are presented in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on its differences from other embodiments. All or part of this invention can be used in numerous general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, mobile communication terminals, multiprocessor systems, microprocessor-based systems, programmable electronic devices, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices, etc.

[0047] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the present invention.

Claims

1. A method for obstacle avoidance target recognition based on the fusion of infrared images and radar point clouds, characterized in that, include: Preprocessing is performed on the synchronously acquired near-infrared images and millimeter-wave radar point cloud data to obtain enhanced images and multiple candidate target point cloud clusters; Image features of the enhanced image are extracted using an object detection network to obtain two-dimensional bounding boxes and texture semantic feature vectors for each potential obstacle. The three-dimensional center coordinates of each candidate target point cloud cluster are obtained, and the point cloud features of each candidate target point cloud cluster are extracted through a point cloud feature extraction network to obtain the geometric motion feature vector of each candidate target point cloud cluster. The point cloud feature extraction network includes an embedding layer, a local feature aggregation unit, and a global context enhancement unit. The embedding layer maps each candidate target point cloud cluster to a high-dimensional embedding feature. The local feature aggregation unit performs convolutional local feature extraction on the high-dimensional embedding feature to obtain local aggregated features. The global context enhancement unit performs global interaction on the local aggregated features using multiple cascaded Transformer layers to obtain the geometric motion feature vector. Based on the pre-jointly calibrated extrinsic matrix, the texture semantic feature vector and the geometric motion feature vector are spatially associated using the two-dimensional bounding box to obtain associated feature pairs; The associated features are input to a fusion network based on cross-attention and self-attention mechanisms to perform feature interaction and fusion, resulting in deep fusion features. These deep fusion features are then input into a prediction head to jointly predict and output the fusion identification result of the obstacle avoidance target.

2. The obstacle avoidance target recognition method based on the fusion of infrared images and radar point clouds according to claim 1, characterized in that, The preprocessing of synchronously acquired near-infrared images and millimeter-wave radar point cloud data yields enhanced images and multiple candidate target point cloud clusters, including: The enhanced image is obtained by performing image enhancement preprocessing on the near-infrared image frame; Moving target display filtering, constant false alarm rate detection, and density clustering are sequentially performed on the millimeter-wave radar point cloud frames to obtain the multiple candidate target point cloud clusters.

3. The obstacle avoidance target recognition method based on the fusion of infrared images and radar point clouds according to claim 1, characterized in that, The step of extracting image features from the enhanced image through an object detection network to obtain two-dimensional bounding boxes and texture semantic feature vectors for each potential obstacle includes: The enhanced image is input into the target detection network, and the shallow texture and deep semantic information of the image are extracted through the backbone feature extraction structure of the target detection network. The shallow texture and the deep semantic information are deeply fused using a feature pyramid network to obtain a multi-scale feature map; The multi-scale feature map is decoded and predicted by the detection head to obtain the coordinates of the two-dimensional bounding box, and the corresponding texture semantic feature vector is extracted from the enhanced image based on the two-dimensional bounding box.

4. The obstacle avoidance target recognition method based on the fusion of infrared images and radar point clouds according to claim 1, characterized in that, The embedded layer includes a multilayer perceptron (MLP); The step of mapping each of the candidate target point cloud clusters into high-dimensional embedding features includes: The low-dimensional physical features included in the point cloud data of each candidate target point cloud cluster are input into the multilayer perceptron. Through the layer-by-layer nonlinear transformation of the multilayer perceptron, the low-dimensional physical features are mapped to a high-dimensional embedding space, and the high-dimensional embedding features corresponding to each candidate target point cloud cluster are output. The low-dimensional physical features include the three-dimensional spatial coordinates, radial velocity, and reflection intensity of the point.

5. The obstacle avoidance target recognition method based on the fusion of infrared images and radar point clouds according to claim 1, characterized in that, The local feature aggregation unit includes: multiple parallel convolutional branches, adders, activation function layers, global average pooling layers, and global max pooling layers; The step of extracting local aggregated features through convolution of the high-dimensional embedded features includes: The high-dimensional embedding features are input into the multiple parallel convolutional branches, and each convolutional branch performs convolution operations independently to obtain the convolutional output features of each convolutional branch. The convolution output features of each convolution branch are added element-wise through the adder at the same spatial location and channel dimension to obtain the added features; The summed features are input into the activation function layer for non-linear activation to obtain the activated features. The activated features are input into the global average pooling layer and the global max pooling layer respectively to obtain average pooling features and max pooling features; The average pooling feature and the max pooling feature are input into the fusion layer and concatenated to output the local aggregated feature.

6. The obstacle avoidance target recognition method based on the fusion of infrared image and radar point cloud as described in claim 1, characterized in that, The global context enhancement unit includes: a first normalization layer, a Transformer module, and a second normalization layer; The process of obtaining a geometric motion feature vector by performing global interaction on the local aggregated features through multiple cascaded Transformer layers includes: The local aggregated features are input into the first normalization layer for normalization processing to obtain the first normalized features; The first normalized feature is mapped to a query matrix, a key matrix, and a value matrix, respectively. Global attention interaction is performed through the Transformer module to obtain attention interaction features. The attention interaction features are input into the second normalization layer for normalization processing, and the geometric motion feature vector is output.

7. The obstacle avoidance target recognition method based on the fusion of infrared image and radar point cloud as described in claim 1, characterized in that, The pre-jointly calibrated extrinsic parameter matrix is ​​used to spatially correlate the texture semantic feature vector and the geometric motion feature vector using the two-dimensional bounding box to obtain correlated feature pairs, including: The three-dimensional center coordinates of each candidate target point cloud cluster are projected from the radar coordinate system onto the corresponding image plane through a pre-calibrated extrinsic parameter matrix to obtain the two-dimensional projection point coordinates of each candidate target point cloud cluster on the image plane. For each two-dimensional projection point coordinate, an associated candidate box is constructed with the two-dimensional projection point coordinate as the center and a preset side length. The cross-union ratio (CUP) between each associated candidate box and the two-dimensional detection bounding box of each potential obstacle is calculated to obtain the CUP matrix. Based on the intersection-union ratio matrix, the optimal bipartite graph matching is solved using the Hungarian algorithm. The optimization objective is to maximize the total association weight of successful matching. The candidate target point cloud clusters are matched one by one with the two-dimensional bounding boxes. Each successfully matched candidate target point cloud cluster is associated with the two-dimensional bounding box pair, and the texture semantic feature vector corresponding to the two-dimensional bounding box is paired with the geometric motion feature vector corresponding to the candidate target point cloud cluster to form the associated feature pair.

8. The obstacle avoidance target recognition method based on the fusion of infrared image and radar point cloud as described in claim 1, characterized in that, The step of performing feature interaction and fusion on the input of the associated features to a fusion network based on cross-attention and self-attention mechanisms to obtain deep fused features includes: The texture semantic feature vector and the geometric motion feature vector in the associated feature pair are input into the feature fusion network; By utilizing the cross-attention mechanism of the feature fusion network, the geometric motion feature vector is used as the query reference, and the texture semantic feature vector is used as the key and value, to perform cross-modal cross-attention calculation and obtain the feature representation after cross-modal interaction; Based on the self-attention mechanism of the feature fusion network, the feature representation after cross-modal interaction and the position encoding information are subjected to self-attention calculation to obtain the feature representation enhanced by internal relations. The feature representation enhanced by internal relations is input into the feedforward network for nonlinear transformation, and then subjected to layer normalization to obtain the final deep fusion feature.

9. The obstacle avoidance target recognition method based on the fusion of infrared image and radar point cloud as described in claim 1, characterized in that, The step of inputting the deep fusion features into the prediction head, jointly predicting and outputting the fusion identification result of the obstacle avoidance target includes: The deep fusion features are fed into the prediction head network; wherein the prediction head network includes a classification branch and a regression branch; The deep fusion features are received using the classification branch, processed by a fully connected layer, a batch normalization layer, and a Softmax activation function, and output as a probability distribution vector representing the target category. The highest probability value in the probability distribution vector is extracted as the prediction confidence of the obstacle avoidance target. The regression branch receives the deep fusion features, performs nonlinear mapping through a deep regression network, and outputs the position coordinates and geometric scale information of the obstacle avoidance target in three-dimensional physical space. By combining the predicted confidence level and the probability distribution of the target category output by the classification branch, and the location coordinates and geometric scale information output by the regression branch, a fusion identification result containing the target category, confidence level, and spatial parameters is obtained.

10. An obstacle avoidance target recognition system based on the fusion of infrared images and radar point clouds, characterized in that, include: The preprocessing module is used to preprocess the synchronously acquired near-infrared images and millimeter-wave radar point cloud data to obtain enhanced images and multiple candidate target point cloud clusters. The feature extraction module is used to extract image features of the enhanced image through the object detection network to obtain two-dimensional bounding boxes and texture semantic features of each potential obstacle; The method includes obtaining the three-dimensional center coordinates of each candidate target point cloud cluster and extracting point cloud features of each candidate target point cloud cluster through a point cloud feature extraction network to obtain the geometric motion feature vector of each point cloud cluster. The point cloud feature extraction network includes an embedding layer, a local feature aggregation unit, and a global context enhancement unit. The embedding layer maps each candidate target point cloud cluster to a high-dimensional embedding feature. The local feature aggregation unit performs convolutional local feature extraction on the high-dimensional embedding feature to obtain local aggregated features. The global context enhancement unit performs global interaction on the local aggregated features using multiple cascaded Transformer layers to obtain the geometric motion feature vector. The spatial association module is used to spatially associate the texture semantic features and the geometric motion feature vectors based on the pre-jointly calibrated extrinsic parameter matrix using the two-dimensional bounding box to obtain associated feature pairs. The fusion module is used to perform feature interaction and fusion on the input fusion network based on cross-attention mechanism and self-attention mechanism to obtain deep fused features; The prediction output module is used to input the deep fusion features into the prediction head, jointly predict and output the fusion identification result of the obstacle avoidance target.

Citation Information

Patent Citations

  • Action recognition method and system based on combination of electrostatic induction and image detection

    CN118747304A

  • Method and system for detecting obstacles around vehicle based on multi-modal sensor

    CN120451936A